Repository navigation
Conversation
|
Thanks for opening a pull request! This pull request has been automatically converted to a draft because its title doesn't match Arrow's required format. If this is not a minor PR. Could you open an issue for this pull request on GitHub? https://github.com/apache/arrow/issues/new/choose Opening GitHub issues ahead of time contributes to the Openness of the Apache Arrow project. Then could you also rename the pull request title in the following format? or After updating the title, you can mark the pull request as ready for review. See also: |
prtkgaur
force-pushed
the
pgaur_interleavedPlusFastLanesDelta
branch
from
October 2, 2026 15:18
e737aaf to
da4ffbd
Compare
prtkgaur
force-pushed
the
pgaur_interleavedPlusFastLanesDelta
branch
4 times, most recently
from
October 3, 2026 01:25
633ba32 to
199407c
Compare
A frame-of-reference decoder adds the frame to every value right after unpacking it. Doing that in a second pass reads and writes the whole output again, so the scalar and SIMD unpackers now take an optional bias and add it before each store. With no bias the existing paths are unchanged, including the plain memcpy used when the bit width equals the output width. The tests run the biased form across widths, offsets, epilogues and every dispatched instruction set.
PFOR packs each vector of 1024 values at a width chosen for most of them and stores the rest as patched exceptions. This adds frame and width selection, the exception section, an endian-independent page layout, and decoding that checks every length against the page before reading it. int32 and int64 are supported. The codec builds under both CMake and Meson, with unit tests and a microbenchmark.
Registers PFOR for INT32 and INT64 and wires it into the encoder and decoder factories. Reads work for dense and nullable columns,, and the decoder can also return one vector at a time.
The tests write and read whole files, including nulls, batched reads, malformed pages and the factory paths. The benchmark compares PFOR with the existing integer encodings on synthetic columns modeled on common analytic data. It skips any compression codec the build does not include rather than aborting.
The PFOR templates are instantiated in libarrow and used from libparquet and its tests. Without export markers this only links on ELF platforms, where symbols are visible by default.
Lists the supported integer types and the Preview status. Writers never pick PFOR on their own; a column has to ask for it.
Some columns pack far better as differences between neighbors than as raw values, so the planner now costs both per vector and keeps the cheaper one. The frame no longer has to be the minimum either. A histogram pass finds a window that leaves outliers on both sides as exceptions, and the frame then moves down to the smallest value the window covers. If the search does not beat the minimum-frame plan, that plan is kept.
The existing benchmark columns are either unordered or perfectly regular, and neither tells raw packing apart from delta packing. Adds timestamp, sawtooth, bounded-rate, monotonic and clustered columns at both integer widths, some of which also need a frame above the minimum.
The delta frame search costs as much as the raw one and is wasted on uncorrelated data. A strided sample of the differences now gives an estimate first, and if that estimate cannot beat the raw plan, the full search does not run.
PFOR is a Preview encoding, so a writer has to opt in before it is emitted. Delta planning has its own per-column switch on top of that. Readers accept either form regardless.
A delta vector carries a flag, a start value and a frame, and a corrupt page can make any of them inconsistent. The decoder now rejects a mismatched flag, a truncated start value and a vector that runs past the end of the page. How far the unpacker may read is now computed from the full delta header, not the raw one.
Comparing the raw and delta payloads of one input needs both to exist, and the planner usually rejects one. A force_delta option, off by default, skips that rejection while still writing an ordinary delta vector that any reader decodes. The round-trip tests cover it.
Packing 1024 values as 32 lanes that advance together lets a decoder unpack them with plain vector shifts at any width. This adds a portable 32-lane kernel and records the layout in the page header. Only complete int32 vectors use it; int64 and partial tail vectors stay sequential.
Exposes the layout as an experimental writer option. The tests cover every bit width, exception patching, the sequential tail, equal encoded size for both layouts, and a round trip through a Parquet file. Vector-at-a-time reads used to lose the packing mode and decode an interleaved vector as sequential; they now carry it through.
Registers each decode benchmark in both layouts, generated from the same values so the encoded size matches and only the layout differs. A page-sized destination is added next to the full-column one, and a header comment says which comparisons are valid.
Sequential delta decoding is one long dependency chain. The FastLanes paper splits a block into lanes of 32 consecutive values, each with its own base, so 32 prefix sums run side by side. This implements that assignment with a choice of how the bases are coded, and a decoder that does the prefix sum and the transpose back to file order in one pass. The benchmarks separate the cost of base coding, of the lane layout and of the permutation.
Adds an int32 page decoder for the lane layout that checks the layout marker and stream length and stops at the declared value count. A benchmark runs it end to end next to Arrow's existing delta decoder.
Gives both lane-delta and transposed-delta payloads a page frame: a little-endian value count, followed by sections that are each checked against the page before decoding. A payload that does not start 4-byte aligned is copied to an aligned buffer first. The tests only run with the PFOR Preview option on, like the rest of the encoding.
Encodes each corpus column once per layout with delta selection held fixed, so the sequential and interleaved rows decode the same values at the same encoded size.
The generated AVX-512 unpacker builds its vectors from scalar loads and calls out of line. Runtime dispatch stops at AVX2 until the AVX-512 kernels load vectors directly.
Decoding FL_ORDER to file order used to unpack into a 4 KiB scratch grid and then transpose it, which is 128 extra stores and loads per block. The transpose tiles are now built from the unpacked registers directly. Lane-parallel delta uses the same structure. A benchmark across working-set sizes and a note on which output orders can be compared are included.
The baseline and vectorized kernel tables are compiled in separate files and picked at run time through Arrow's dispatch. The kernels take the architecture as a template parameter, so the two compiled copies have different symbols and do not collide at link time.
Moves the synthetic column generators out of the comparison benchmark into a header that the other benchmarks include, so every benchmark reads the same columns. The moved code drops boxed headings and benchmark slang in its comments. No codec behavior changes.
Arm had only the portable scratch-grid path for FL_ORDER. A four-lane NEON transpose now writes file order directly. Every packed width is checked with and without a frame bias.
prtkgaur
force-pushed
the
pgaur_interleavedPlusFastLanesDelta
branch
from
October 3, 2026 22:09
199407c to
59a378b
Compare
prtkgaur
force-pushed
the
pgaur_interleavedPlusFastLanesDelta
branch
from
October 4, 2026 19:35
59a378b to
3dd10b0
Compare
The interleaved block kernels only wrote uint32_t, so three-bit values could not decode into a byte even though nothing about the layout needs 32 bits. BlockGeometry<T> derives the grid from the element: a 1024-value block is 8 rows of 128 lanes for uint8_t and 16 rows of 64 lanes for uint16_t. The packed payload stays 128 * w bytes at every element width, the same as the sequential layout. kLanes and kRowsPerBlock still name the 32-bit grid, so the delta and PFOR formats built on it are unchanged. The round-trip tests cover every bit width for all three element types, with and without a bias. They pack into a buffer that already holds another block, because a page encoder relies on each kernel fully defining its own payload, and they check that a lane reads as an ordinary sequentially packed stream.
Times the interleaved block decoder against Arrow's sequential unpackers: the shipped dispatch entry point, the scalar kernel, and on AVX2 builds the vector kernel pinned to AVX2. One sweep decodes 16 KiB and moves the bit width from 1 to 32, decoding into the narrowest whole byte. The other holds eight widths and moves the decoded size from 16 KiB to 4 MiB. On 32-bit output the interleaved decoder also runs with the permutation back to file order, through the fused kernel where there is one and a scratch grid otherwise. A new test checks that the fused kernel and the grid path produce the same values. The target links xsimd for the kernel it names.
The large kernel shifts the high part of each value left by a per-lane amount. AVX2 has that shift for 32- and 64-bit lanes only, so for byte and short output xsimd widened, shifted and narrowed again. right_shift_by_excess already avoids this for the right shift by shifting pairs of narrow lanes as one wider lane and masking them apart. left_shift_by_excess does the same for the left shift. Byte output on AVX2 now also goes through the medium kernel on shorts, as it already did on SSE4.2, since a byte needs two rounds of widening where a short needs one. This affects packed widths whose remainder modulo eight is 3, 5, 6 or 7. In L1, those widths now run 4.23x to 4.28x as fast as before at short output and 2.02x to 2.13x as fast at byte output. Other widths stay within 0.94x to 1.06x, and 32-bit output is unchanged. AVX-512 keeps the old routing, because there the route through shorts runs 0.16x to 0.23x as fast. Checked against the scalar reference for every kernel x86 builds, all four output sizes and widths 1 to 63, on SSE4.2, AVX2 and AVX-512, with and without a bias.
prtkgaur
force-pushed
the
pgaur_interleavedPlusFastLanesDelta
branch
from
October 7, 2026 00:01
eb8785d to
37104b8
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rationale for this change
Parquet uses a continuous bit-packed stream, while FastLanes uses a
lane-interleaved grid. This PR measures whether the grid's cheaper unpacking
can offset the cost of returning values in Parquet file order.
The answer depends on packed width, output element width, decode call shape,
memory footprint, instruction set, and the permutation back to file order.
This branch keeps those variables separate instead of reducing the comparison
to one throughput number.
What changes are included in this PR?
This benchmarking-only proof of concept adds:
Run
./layoutsinfl5_corpuswithout arguments to see the measurement index.Each study names the variable it changes and the assumptions required to
interpret its output.
Are these changes tested?
Yes. The codecs and experimental layouts are checked for bit-exact round trips
before they are timed. Unit tests cover packing widths, output order, delta
paths, tails, exceptions, and boundary values.
The benchmark harness includes timing controls and rejects results that do not
meet its validity checks. Results are recorded for multiple compilers,
instruction sets, and working-set sizes.
Are there any user-facing changes?
No production encoding or default writer behavior is proposed by this PR. It
is a benchmarking and research branch used to evaluate layout choices.
The experimental APIs and on-disk layouts in this branch should not be treated
as stable or interoperable formats.