Repository navigation
Conversation
|
Thanks for opening a pull request! This pull request has been automatically converted to a draft because its title doesn't match Arrow's required format. If this is not a minor PR. Could you open an issue for this pull request on GitHub? https://github.com/apache/arrow/issues/new/choose Opening GitHub issues ahead of time contributes to the Openness of the Apache Arrow project. Then could you also rename the pull request title in the following format? or After updating the title, you can mark the pull request as ready for review. See also: |
prtkgaur
force-pushed
the
pfor-delta-encoding
branch
from
September 28, 2026 05:32
36395ec to
9174cff
Compare
prtkgaur
force-pushed
the
pfor-delta-encoding
branch
from
October 2, 2026 15:17
3b7bef6 to
af6bbd6
Compare
prtkgaur
force-pushed
the
pfor-delta-encoding
branch
4 times, most recently
from
October 3, 2026 01:25
ee608e6 to
28cff0b
Compare
A frame-of-reference decoder adds the frame to every value right after unpacking it. Doing that in a second pass reads and writes the whole output again, so the scalar and SIMD unpackers now take an optional bias and add it before each store. With no bias the existing paths are unchanged, including the plain memcpy used when the bit width equals the output width. The tests run the biased form across widths, offsets, epilogues and every dispatched instruction set.
PFOR packs each vector of 1024 values at a width chosen for most of them and stores the rest as patched exceptions. This adds frame and width selection, the exception section, an endian-independent page layout, and decoding that checks every length against the page before reading it. int32 and int64 are supported. The codec builds under both CMake and Meson, with unit tests and a microbenchmark.
Registers PFOR for INT32 and INT64 and wires it into the encoder and decoder factories. Reads work for dense and nullable columns,, and the decoder can also return one vector at a time.
The tests write and read whole files, including nulls, batched reads, malformed pages and the factory paths. The benchmark compares PFOR with the existing integer encodings on synthetic columns modeled on common analytic data. It skips any compression codec the build does not include rather than aborting.
The PFOR templates are instantiated in libarrow and used from libparquet and its tests. Without export markers this only links on ELF platforms, where symbols are visible by default.
Lists the supported integer types and the Preview status. Writers never pick PFOR on their own; a column has to ask for it.
Some columns pack far better as differences between neighbors than as raw values, so the planner now costs both per vector and keeps the cheaper one. The frame no longer has to be the minimum either. A histogram pass finds a window that leaves outliers on both sides as exceptions, and the frame then moves down to the smallest value the window covers. If the search does not beat the minimum-frame plan, that plan is kept.
The existing benchmark columns are either unordered or perfectly regular, and neither tells raw packing apart from delta packing. Adds timestamp, sawtooth, bounded-rate, monotonic and clustered columns at both integer widths, some of which also need a frame above the minimum.
The delta frame search costs as much as the raw one and is wasted on uncorrelated data. A strided sample of the differences now gives an estimate first, and if that estimate cannot beat the raw plan, the full search does not run.
PFOR is a Preview encoding, so a writer has to opt in before it is emitted. Delta planning has its own per-column switch on top of that. Readers accept either form regardless.
A delta vector carries a flag, a start value and a frame, and a corrupt page can make any of them inconsistent. The decoder now rejects a mismatched flag, a truncated start value and a vector that runs past the end of the page. How far the unpacker may read is now computed from the full delta header, not the raw one.
Comparing the raw and delta payloads of one input needs both to exist, and the planner usually rejects one. A force_delta option, off by default, skips that rejection while still writing an ordinary delta vector that any reader decodes. The round-trip tests cover it.
prtkgaur
force-pushed
the
pfor-delta-encoding
branch
from
October 3, 2026 22:09
28cff0b to
7d14f50
Compare
prtkgaur
force-pushed
the
pfor-delta-encoding
branch
2 times, most recently
from
October 7, 2026 00:01
af5e415 to
6e15dc9
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rationale for this change
Patched Frame of Reference (PFOR) stores most integers at a common bit width
and records values that do not fit as exceptions. This branch explores PFOR
for Parquet integer columns, including an optional delta mode for locally
correlated values.
The PR also compares the patched-delta payload with Parquet's existing
DELTA_BINARY_PACKEDlayout. The results separate format-level differencesfrom decoder implementation costs.
What changes are included in this PR?
This experimental PR adds:
INT32andINT64.DELTA_BINARY_PACKEDpayloads.The layout study and its measurements are documented in
pfor_delta_layout_report.md.Are these changes tested?
Yes. Tests cover
INT32andINT64round trips, raw and delta modes,exceptions, boundary values, corrupt metadata, writer-property selection,
and Parquet integration.
The benchmark corpus includes timestamp-like values, counters, random walks,
outliers, and generated TPC-style column shapes. The checked-in report records
the method, results, and limitations of the delta-layout comparison.
Are there any user-facing changes?
This is an experimental implementation. It adds writer properties for opting
into PFOR and its delta mode, plus reader support for the corresponding pages.
Neither mode is selected by default.
The encoding is not yet part of the Parquet format specification and should
not be treated as a stable interchange format.