Skip to content

[WIP][POC][BenchmarkingOnly] PFOR interleaved and fastlanes numbers - #51296

Draft
prtkgaur wants to merge 27 commits into
apache:mainfrom
prtkgaur:pgaur_interleavedPlusFastLanesDelta
Draft

prtkgaur wants to merge 27 commits into
apache:mainfrom
prtkgaur:pgaur_interleavedPlusFastLanesDelta

Conversation

@prtkgaur

@prtkgaur prtkgaur commented Sep 10, 2026 •

Copy link
Copy Markdown

Rationale for this change

Parquet uses a continuous bit-packed stream, while FastLanes uses a
lane-interleaved grid. This PR measures whether the grid's cheaper unpacking
can offset the cost of returning values in Parquet file order.

The answer depends on packed width, output element width, decode call shape,
memory footprint, instruction set, and the permutation back to file order.
This branch keeps those variables separate instead of reducing the comparison
to one throughput number.

What changes are included in this PR?

This benchmarking-only proof of concept adds:

  • Portable lane-interleaved and FastLanes-style bit-packing kernels.
  • AVX2 and NEON paths that perform the file-order permutation in registers.
  • Runtime dispatch for the interleaved kernels.
  • PFOR experiments using interleaved, lane-delta, and transposed-delta layouts.
  • Benchmarks for packed width, output width, corpus shape, and memory footprint.
  • Controls for call shape, output address, bytes moved, lane order, and stores.
  • Reproducible build and analysis scripts for x86-64 and AArch64.
  • Checked-in benchmark results and documentation of their limitations.

Run ./layouts in fl5_corpus without arguments to see the measurement index.
Each study names the variable it changes and the assumptions required to
interpret its output.

Are these changes tested?

Yes. The codecs and experimental layouts are checked for bit-exact round trips
before they are timed. Unit tests cover packing widths, output order, delta
paths, tails, exceptions, and boundary values.

The benchmark harness includes timing controls and rejects results that do not
meet its validity checks. Results are recorded for multiple compilers,
instruction sets, and working-set sizes.

Are there any user-facing changes?

No production encoding or default writer behavior is proposed by this PR. It
is a benchmarking and research branch used to evaluate layout choices.

The experimental APIs and on-disk layouts in this branch should not be treated
as stable or interoperable formats.

@github-actions

Copy link
Copy Markdown

Thanks for opening a pull request!

This pull request has been automatically converted to a draft because its title doesn't match Arrow's required format.

If this is not a minor PR. Could you open an issue for this pull request on GitHub? https://github.com/apache/arrow/issues/new/choose

Opening GitHub issues ahead of time contributes to the Openness of the Apache Arrow project.

Then could you also rename the pull request title in the following format?

GH-${GITHUB_ISSUE_ID}: [${COMPONENT}] ${SUMMARY}

or

MINOR: [${COMPONENT}] ${SUMMARY}

After updating the title, you can mark the pull request as ready for review.

See also:

@prtkgaur prtkgaur changed the title [WIP][POC][BenchmarkingOnly] PFOR interleaved and flanes numbers [WIP][POC][BenchmarkingOnly] PFOR interleaved and fastlanes numbers Sep 11, 2026
@prtkgaur
prtkgaur force-pushed the pgaur_interleavedPlusFastLanesDelta branch from e737aaf to da4ffbd Compare October 2, 2026 15:18
@prtkgaur
prtkgaur force-pushed the pgaur_interleavedPlusFastLanesDelta branch 4 times, most recently from 633ba32 to 199407c Compare October 3, 2026 01:25
A frame-of-reference decoder adds the frame to every value right after
unpacking it. Doing that in a second pass reads and writes the whole
output again, so the scalar and SIMD unpackers now take an optional bias
and add it before each store. With no bias the existing paths are
unchanged, including the plain memcpy used when the bit width equals the
output width. The tests run the biased form across widths, offsets,
epilogues and every dispatched instruction set.
PFOR packs each vector of 1024 values at a width chosen for most of them
and stores the rest as patched exceptions. This adds frame and width
selection, the exception section, an endian-independent page layout, and
decoding that checks every length against the page before reading it.
int32 and int64 are supported. The codec builds under both CMake and
Meson, with unit tests and a microbenchmark.
Registers PFOR for INT32 and INT64 and wires it into the encoder and
decoder factories. Reads work for dense and nullable columns,, and the
decoder can also return one vector at a time.
The tests write and read whole files, including nulls, batched reads,
malformed pages and the factory paths. The benchmark compares PFOR with
the existing integer encodings on synthetic columns modeled on common
analytic data. It skips any compression codec the build does not include
rather than aborting.
The PFOR templates are instantiated in libarrow and used from libparquet
and its tests. Without export markers this only links on ELF platforms,
where symbols are visible by default.
Lists the supported integer types and the Preview status. Writers never
pick PFOR on their own; a column has to ask for it.
Some columns pack far better as differences between neighbors than as
raw values, so the planner now costs both per vector and keeps the
cheaper one. The frame no longer has to be the minimum either. A
histogram pass finds a window that leaves outliers on both sides as
exceptions, and the frame then moves down to the smallest value the
window covers. If the search does not beat the minimum-frame plan, that
plan is kept.
The existing benchmark columns are either unordered or perfectly
regular, and neither tells raw packing apart from delta packing. Adds
timestamp, sawtooth, bounded-rate, monotonic and clustered columns at
both integer widths, some of which also need a frame above the minimum.
The delta frame search costs as much as the raw one and is wasted on
uncorrelated data. A strided sample of the differences now gives an
estimate first, and if that estimate cannot beat the raw plan, the full
search does not run.
PFOR is a Preview encoding, so a writer has to opt in before it is
emitted. Delta planning has its own per-column switch on top of that.
Readers accept either form regardless.
A delta vector carries a flag, a start value and a frame, and a corrupt
page can make any of them inconsistent. The decoder now rejects a
mismatched flag, a truncated start value and a vector that runs past the
end of the page. How far the unpacker may read is now computed from the
full delta header, not the raw one.
Comparing the raw and delta payloads of one input needs both to exist,
and the planner usually rejects one. A force_delta option, off by
default, skips that rejection while still writing an ordinary delta
vector that any reader decodes. The round-trip tests cover it.
Packing 1024 values as 32 lanes that advance together lets a decoder
unpack them with plain vector shifts at any width. This adds a portable
32-lane kernel and records the layout in the page header. Only complete
int32 vectors use it; int64 and partial tail vectors stay sequential.
Exposes the layout as an experimental writer option. The tests cover
every bit width, exception patching, the sequential tail, equal encoded
size for both layouts, and a round trip through a Parquet file.
Vector-at-a-time reads used to lose the packing mode and decode an
interleaved vector as sequential; they now carry it through.
Registers each decode benchmark in both layouts, generated from the same
values so the encoded size matches and only the layout differs. A
page-sized destination is added next to the full-column one, and a
header comment says which comparisons are valid.
Sequential delta decoding is one long dependency chain. The FastLanes
paper splits a block into lanes of 32 consecutive values, each with its
own base, so 32 prefix sums run side by side. This implements that
assignment with a choice of how the bases are coded, and a decoder that
does the prefix sum and the transpose back to file order in one pass.
The benchmarks separate the cost of base coding, of the lane layout and
of the permutation.
Adds an int32 page decoder for the lane layout that checks the layout
marker and stream length and stops at the declared value count. A
benchmark runs it end to end next to Arrow's existing delta decoder.
Gives both lane-delta and transposed-delta payloads a page frame: a
little-endian value count, followed by sections that are each checked
against the page before decoding. A payload that does not start 4-byte
aligned is copied to an aligned buffer first. The tests only run with
the PFOR Preview option on, like the rest of the encoding.
Encodes each corpus column once per layout with delta selection held
fixed, so the sequential and interleaved rows decode the same values at
the same encoded size.
The generated AVX-512 unpacker builds its vectors from scalar loads and
calls out of line. Runtime dispatch stops at AVX2 until the AVX-512
kernels load vectors directly.
Decoding FL_ORDER to file order used to unpack into a 4 KiB scratch grid
and then transpose it, which is 128 extra stores and loads per block.
The transpose tiles are now built from the unpacked registers directly.
Lane-parallel delta uses the same structure. A benchmark across
working-set sizes and a note on which output orders can be compared are
included.
The baseline and vectorized kernel tables are compiled in separate files
and picked at run time through Arrow's dispatch. The kernels take the
architecture as a template parameter, so the two compiled copies have
different symbols and do not collide at link time.
Moves the synthetic column generators out of the comparison benchmark
into a header that the other benchmarks include, so every benchmark
reads the same columns. The moved code drops boxed headings and
benchmark slang in its comments. No codec behavior changes.
Arm had only the portable scratch-grid path for FL_ORDER. A four-lane
NEON transpose now writes file order directly. Every packed width is
checked with and without a frame bias.
@prtkgaur
prtkgaur force-pushed the pgaur_interleavedPlusFastLanesDelta branch from 199407c to 59a378b Compare October 3, 2026 22:09
@prtkgaur
prtkgaur force-pushed the pgaur_interleavedPlusFastLanesDelta branch from 59a378b to 3dd10b0 Compare October 4, 2026 19:35
The interleaved block kernels only wrote uint32_t, so three-bit values
could not decode into a byte even though nothing about the layout needs
32 bits. BlockGeometry<T> derives the grid from the element: a
1024-value block is 8 rows of 128 lanes for uint8_t and 16 rows of 64
lanes for uint16_t. The packed payload stays 128 * w bytes at every
element width, the same as the sequential layout. kLanes and
kRowsPerBlock still name the 32-bit grid, so the delta and PFOR formats
built on it are unchanged.

The round-trip tests cover every bit width for all three element types,
with and without a bias. They pack into a buffer that already holds
another block, because a page encoder relies on each kernel fully
defining its own payload, and they check that a lane reads as an
ordinary sequentially packed stream.
Times the interleaved block decoder against Arrow's sequential
unpackers: the shipped dispatch entry point, the scalar kernel, and on
AVX2 builds the vector kernel pinned to AVX2. One sweep decodes 16 KiB
and moves the bit width from 1 to 32, decoding into the narrowest whole
byte. The other holds eight widths and moves the decoded size from 16
KiB to 4 MiB.

On 32-bit output the interleaved decoder also runs with the permutation
back to file order, through the fused kernel where there is one and a
scratch grid otherwise. A new test checks that the fused kernel and the
grid path produce the same values. The target links xsimd for the kernel
it names.
The large kernel shifts the high part of each value left by a per-lane
amount. AVX2 has that shift for 32- and 64-bit lanes only, so for byte
and short output xsimd widened, shifted and narrowed again.
right_shift_by_excess already avoids this for the right shift by
shifting pairs of narrow lanes as one wider lane and masking them apart.
left_shift_by_excess does the same for the left shift. Byte output on
AVX2 now also goes through the medium kernel on shorts, as it already
did on SSE4.2, since a byte needs two rounds of widening where a short
needs one.

This affects packed widths whose remainder modulo eight is 3, 5, 6 or 7.
In L1, those widths now run 4.23x to 4.28x as fast as before at short
output and 2.02x to 2.13x as fast at byte output. Other widths stay
within 0.94x to 1.06x, and 32-bit output is unchanged. AVX-512 keeps the
old routing, because there the route through shorts runs 0.16x to 0.23x
as fast. Checked against the scalar reference for every kernel x86
builds, all four output sizes and widths 1 to 63, on SSE4.2, AVX2 and
AVX-512, with and without a bias.
@prtkgaur
prtkgaur force-pushed the pgaur_interleavedPlusFastLanesDelta branch from eb8785d to 37104b8 Compare October 7, 2026 00:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants