Skip to content

[Experimental][WIP] Pfor delta encoding - #51150

Draft
prtkgaur wants to merge 12 commits into
apache:mainfrom
prtkgaur:pfor-delta-encoding
Draft

prtkgaur wants to merge 12 commits into
apache:mainfrom
prtkgaur:pfor-delta-encoding

Conversation

@prtkgaur

@prtkgaur prtkgaur commented Sep 3, 2026 •

Copy link
Copy Markdown

Rationale for this change

Patched Frame of Reference (PFOR) stores most integers at a common bit width
and records values that do not fit as exceptions. This branch explores PFOR
for Parquet integer columns, including an optional delta mode for locally
correlated values.

The PR also compares the patched-delta payload with Parquet's existing
DELTA_BINARY_PACKED layout. The results separate format-level differences
from decoder implementation costs.

What changes are included in this PR?

This experimental PR adds:

  • PFOR encoding and decoding for INT32 and INT64.
  • Page-level metadata, vector offsets, validation, and exception handling.
  • An optional delta mode selected through writer properties.
  • A sampled estimate that declines delta mode when it is not beneficial.
  • Parquet reader and writer integration.
  • Unit, round-trip, malformed-input, and property tests.
  • Benchmarks covering generated integer-column shapes.
  • A comparison of patched delta and DELTA_BINARY_PACKED payloads.
  • Decoder improvements for equal-width DBP miniblocks and prefix scans.

The layout study and its measurements are documented in
pfor_delta_layout_report.md.

Are these changes tested?

Yes. Tests cover INT32 and INT64 round trips, raw and delta modes,
exceptions, boundary values, corrupt metadata, writer-property selection,
and Parquet integration.

The benchmark corpus includes timestamp-like values, counters, random walks,
outliers, and generated TPC-style column shapes. The checked-in report records
the method, results, and limitations of the delta-layout comparison.

Are there any user-facing changes?

This is an experimental implementation. It adds writer properties for opting
into PFOR and its delta mode, plus reader support for the corresponding pages.
Neither mode is selected by default.

The encoding is not yet part of the Parquet format specification and should
not be treated as a stable interchange format.

@prtkgaur prtkgaur changed the title Pfor delta encoding [Experimental] Pfor delta encoding Sep 3, 2026
@github-actions github-actions Bot added the awaiting review Awaiting review label Sep 3, 2026
@github-actions

github-actions Bot commented Sep 3, 2026

Copy link
Copy Markdown

Thanks for opening a pull request!

This pull request has been automatically converted to a draft because its title doesn't match Arrow's required format.

If this is not a minor PR. Could you open an issue for this pull request on GitHub? https://github.com/apache/arrow/issues/new/choose

Opening GitHub issues ahead of time contributes to the Openness of the Apache Arrow project.

Then could you also rename the pull request title in the following format?

GH-${GITHUB_ISSUE_ID}: [${COMPONENT}] ${SUMMARY}

or

MINOR: [${COMPONENT}] ${SUMMARY}

After updating the title, you can mark the pull request as ready for review.

See also:

@prtkgaur prtkgaur changed the title [Experimental] Pfor delta encoding [Experimental][WIP] Pfor delta encoding Sep 3, 2026
@prtkgaur
prtkgaur force-pushed the pfor-delta-encoding branch from 3b7bef6 to af6bbd6 Compare October 2, 2026 15:17
@prtkgaur
prtkgaur force-pushed the pfor-delta-encoding branch 4 times, most recently from ee608e6 to 28cff0b Compare October 3, 2026 01:25
A frame-of-reference decoder adds the frame to every value right after
unpacking it. Doing that in a second pass reads and writes the whole
output again, so the scalar and SIMD unpackers now take an optional bias
and add it before each store. With no bias the existing paths are
unchanged, including the plain memcpy used when the bit width equals the
output width. The tests run the biased form across widths, offsets,
epilogues and every dispatched instruction set.
PFOR packs each vector of 1024 values at a width chosen for most of them
and stores the rest as patched exceptions. This adds frame and width
selection, the exception section, an endian-independent page layout, and
decoding that checks every length against the page before reading it.
int32 and int64 are supported. The codec builds under both CMake and
Meson, with unit tests and a microbenchmark.
Registers PFOR for INT32 and INT64 and wires it into the encoder and
decoder factories. Reads work for dense and nullable columns,, and the
decoder can also return one vector at a time.
The tests write and read whole files, including nulls, batched reads,
malformed pages and the factory paths. The benchmark compares PFOR with
the existing integer encodings on synthetic columns modeled on common
analytic data. It skips any compression codec the build does not include
rather than aborting.
The PFOR templates are instantiated in libarrow and used from libparquet
and its tests. Without export markers this only links on ELF platforms,
where symbols are visible by default.
Lists the supported integer types and the Preview status. Writers never
pick PFOR on their own; a column has to ask for it.
Some columns pack far better as differences between neighbors than as
raw values, so the planner now costs both per vector and keeps the
cheaper one. The frame no longer has to be the minimum either. A
histogram pass finds a window that leaves outliers on both sides as
exceptions, and the frame then moves down to the smallest value the
window covers. If the search does not beat the minimum-frame plan, that
plan is kept.
The existing benchmark columns are either unordered or perfectly
regular, and neither tells raw packing apart from delta packing. Adds
timestamp, sawtooth, bounded-rate, monotonic and clustered columns at
both integer widths, some of which also need a frame above the minimum.
The delta frame search costs as much as the raw one and is wasted on
uncorrelated data. A strided sample of the differences now gives an
estimate first, and if that estimate cannot beat the raw plan, the full
search does not run.
PFOR is a Preview encoding, so a writer has to opt in before it is
emitted. Delta planning has its own per-column switch on top of that.
Readers accept either form regardless.
A delta vector carries a flag, a start value and a frame, and a corrupt
page can make any of them inconsistent. The decoder now rejects a
mismatched flag, a truncated start value and a vector that runs past the
end of the page. How far the unpacker may read is now computed from the
full delta header, not the raw one.
Comparing the raw and delta payloads of one input needs both to exist,
and the planner usually rejects one. A force_delta option, off by
default, skips that rejection while still writing an ordinary delta
vector that any reader decodes. The round-trip tests cover it.
@prtkgaur
prtkgaur force-pushed the pfor-delta-encoding branch from 28cff0b to 7d14f50 Compare October 3, 2026 22:09
@prtkgaur
prtkgaur force-pushed the pfor-delta-encoding branch 2 times, most recently from af5e415 to 6e15dc9 Compare October 7, 2026 00:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants