Skip to content

Read deletion vectors by content offset without parsing the Puffin footer - #4085

Draft
anoopj wants to merge 1 commit into
apache:mainfrom
anoopj:dv-read-by-content-offset
Draft

anoopj wants to merge 1 commit into
apache:mainfrom
anoopj:dv-read-by-content-offset

Conversation

@anoopj

@anoopj anoopj commented Oct 7, 2026

Copy link
Copy Markdown
Member

Rationale for this change

When a scan applies deletes, we loads the deletion vector that applies to each data file. For Puffin deletion vectors it read the entire file into memory and parsed the footer to locate and deserialize every blob, then returned the one for the referenced data file.

Instead, read only the referenced blob with a single ranged read using content_offset and content_size_in_bytes from the manifest, and take the referenced data file from the manifest as well, matching the Java and Rust readers. Validate the blob's DV_MAGIC and CRC-32 while stripping the framing.

Details

  • Performance: a deletion vector read is now a single ranged read of one blob rather than loading the whole Puffin file and deserializing every blob it contains. That cuts I/O (notably against object storage, where only the blob's byte range is fetched), CPU, and memory, and scales with the referenced vector rather than the size of the shared container.
  • Compatibility: deletion vectors that are not fully-formed Puffin files for example Delta-compatible vectors that omit the footer become readable, since the footer is never consulted.

Note: Reading one blob per manifest entry surfaces gaps that reading every blob previously masked, so match and preserve each deletion vector by its target:

  • Route deletion vectors by referenced_data_file in DeleteFileIndex. Deletion vectors need not carry path bounds, so without this they fall into a shared partition bucket and collapse by file_path, giving every data file in the partition the same vector.
  • Deduplicate delete files on (file_path, content_offset) in _read_all_delete_files. DataFile equality keys only on file_path, so multiple deletion vectors packed into one Puffin file would otherwise collapse into a single read.
  • Fill referenced_data_file from the scan task's data file when converting REST position deletes. The field is optional in the REST schema, but the offset read requires it.

Are these changes tested?

Added unit tests

Are there any user-facing changes?

No

…oter

When a scan applies deletes, _read_deletes loads the deletion vector that
applies to each data file. For Puffin deletion vectors it read the entire file
into memory and parsed the footer to locate and deserialize every blob, then
returned the one for the referenced data file.

Instead, read only the referenced blob with a single ranged read using
content_offset and content_size_in_bytes from the manifest, and take the
referenced data file from the manifest as well, matching the Java and Rust
readers. Validate the blob's DV_MAGIC and CRC-32 while stripping the framing,
so a bad pointer or corrupted body is caught now that the Puffin footer is no
longer validated.

This is both faster and more permissive:

- Performance: a deletion vector read is now a single ranged read of one blob
  rather than loading the whole Puffin file and deserializing every blob it
  contains. That cuts I/O (notably against object storage, where only the
  blob's byte range is fetched), CPU, and memory, and scales with the
  referenced vector rather than the size of the shared container.
- Compatibility: deletion vectors that are not fully-formed Puffin files -- for
  example Delta-compatible vectors that omit the footer -- become readable,
  since the footer is never consulted.

Reading one blob per manifest entry surfaces gaps that reading every blob
previously masked, so match and preserve each deletion vector by its target:

- Route deletion vectors by referenced_data_file in DeleteFileIndex. Deletion
  vectors need not carry path bounds, so without this they fall into a shared
  partition bucket and collapse by file_path, giving every data file in the
  partition the same vector.
- Deduplicate delete files on (file_path, content_offset) in
  _read_all_delete_files. DataFile equality keys only on file_path, so multiple
  deletion vectors packed into one Puffin file would otherwise collapse into a
  single read.
- Fill referenced_data_file from the scan task's data file when converting REST
  position deletes. The field is optional in the REST schema, but the offset
  read requires it.
@anoopj
anoopj marked this pull request as draft October 7, 2026 00:13
@anoopj

anoopj commented Oct 7, 2026

Copy link
Copy Markdown
Member Author

Looks like this is already covered by #3478

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant