perf(parquet): SIMD bit-unpack for DeltaBpDecoder dense and sparse visitor paths#5
Closed
jaylisde wants to merge 1 commit into
Closed
perf(parquet): SIMD bit-unpack for DeltaBpDecoder dense and sparse visitor paths#5jaylisde wants to merge 1 commit into
jaylisde wants to merge 1 commit into
Conversation
jaylisde
force-pushed
the
pr/delta-bytearray-perf
branch
from
June 3, 2026 00:01
38359d6 to
7967878
Compare
jaylisde
force-pushed
the
perf/delta-bp-simd-unpack-sparse
branch
from
June 3, 2026 00:02
fb81bda to
18a769a
Compare
jaylisde
force-pushed
the
pr/delta-bytearray-perf
branch
from
June 3, 2026 06:30
7967878 to
e987d8d
Compare
jaylisde
force-pushed
the
perf/delta-bp-simd-unpack-sparse
branch
from
June 3, 2026 06:35
18a769a to
2312b0a
Compare
jaylisde
force-pushed
the
pr/delta-bytearray-perf
branch
from
June 3, 2026 06:49
e987d8d to
a9359a7
Compare
…sitor paths Builds on #3 (PR A) and #4 (PR B) to close the remaining DELTA-vs-PLAIN/DICT scan gap on TPC-H Q12. 1. Inline SIMD bit-unpack kernel in DeltaBpDecoder::decodeLongs. When the read aligns at a miniblock start and consumes a whole miniblock, dispatch on bit_width 0..32 to a compile-time specialized kernel: - bw 0: arithmetic-sequence fast path (no bit-extract). - bw 1..16: 4-value/iter, single unaligned 64-bit load (4*16 = 64 bits fit in one u64 window). - bw 17..32: 2-value/iter, __uint128_t funnel-shift (lowers to two u64 loads + SHRD on x86_64). The trailing u64 read is safe because intra-page overshoots fall into the next miniblock and the last miniblock has PageReader::kPageReadPadding (8) trailing bytes guaranteed. bit_widths 33..64 fall through to the per-row scalar inner loop, which is unchanged. 2. New readWithVisitorSparseBuffered path for the !Visitor::dense + !hasNulls + deterministic filter + NoHook + integral DataType case (the hot path on Q12 because filter chaining produces sparse row sets after l_shipmode is applied). Decodes kBatch=1024 physical values into a stack buffer via decodeLongs (which now uses the SIMD kernel above), then walks the visitor's sparse rows array using rows[k] - batchPhysStart as buffer index. The existing per-row visitor.process is preserved — only the decode side is batched. n is capped to the visitor's residual physical span so the decoder never advances past what the visitor will consume. 3. readWithVisitorDenseBatched kBatch raised from 256 to 1024 to amortize the chunk-loop overhead now that decodeLongs is much faster per call. 4. Add velox/dwio/parquet/tests/reader/DeltaBpDecoderTest.cpp: - Hand-rolled DeltaEncoder for byte-stream control in tests. - 11 boundary tests: bit_widths 0/8/10/16/24/32, multi-block in a single readValues, mid-miniblock split across two readValues calls, negative minDelta, bit_width > 32 fallback, narrowing to int32_t. - 32 parameterized roundtrip tests covering bit_widths 1..32, each forcing the encoder to pick exactly that width by saturating one residual to (1<<bw)-1. Bench (TPC-H Q12 SF10, DELTA-encoded lineitem, num_drivers=4, 5-run median): wall #3 baseline 897 ms +#4 (PR B) 748 ms +this PR 626 ms That is -16.3% on top of PR B, -30.2% from the PR A baseline. PLAIN/DICT parquet paths are unchanged. Test plan - velox_dwio_parquet_delta_bp_decoder_test: 43 tests pass (11 boundary + 32 parameterized bw 1..32). - velox_dwio_parquet_table_scan_test: 54 tests pass (includes the 7 delta tests added in PR A). - velox_parquet_e2e_filter_test: 34 tests pass. Tracking: #2.
jaylisde
force-pushed
the
perf/delta-bp-simd-unpack-sparse
branch
from
June 3, 2026 06:51
2312b0a to
c01426e
Compare
Owner
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Builds on #3 (PR A) and #4 (PR B). Closes the remaining DELTA-vs-PLAIN/DICT scan gap on TPC-H Q12.
Inline SIMD bit-unpack in
DeltaBpDecoder::decodeLongs. When the read aligns at a miniblock start and consumes a whole miniblock, dispatch onbit_width0..32 to a compile-time-specialized kernel:bw=0is an arithmetic-sequence fast path;bw 1..16does 4-value/iter with a single unaligned 64-bit load (4×16 fits in one u64 window);bw 17..32does 2-value/iter with__uint128_tfunnel-shift (lowers to two u64 loads +SHRDon x86_64). The trailing u64 read is safe — intra-page overshoots fall into the next miniblock and the last miniblock haskPageReadPadding(8) trailing bytes guaranteed.bw 33..64falls through to the existing scalar inner loop.readWithVisitorSparseBufferedpath for!Visitor::dense + !hasNulls + deterministic filter + NoHook + integral DataType(Q12's hot path, since chained filters produce a sparse row set afterl_shipmodeis applied). DecodeskBatch=1024physical values into a stack buffer viadecodeLongs(now SIMD), then walks the visitor's sparserows[]array usingrows[k] - batchPhysStartas buffer index. The existing per-rowvisitor.processis preserved — only the decode side is batched.nis capped to the visitor's residual physical span so the decoder never advances past what the visitor will consume.readWithVisitorDenseBatchedkBatchraised from 256 to 1024 to amortize the chunk-loop overhead now thatdecodeLongsis much faster per call.Bench
TPC-H Q12 SF10, DELTA-encoded lineitem,
num_drivers=4, 5-run median:PLAIN/DICT parquet paths are unchanged within noise. Bench will be re-measured (5-run median) before posting upstream.
Test plan
velox_dwio_parquet_delta_bp_decoder_test(new): 43 tests pass — 11 boundary cases (bit_widths0/8/10/16/24/32, multi-block in a singlereadValues, mid-miniblock split across two calls, negativeminDelta,bit_width > 32fallback, narrowing toint32_t) plus 32 parameterized roundtrip tests coveringbit_widths1..32.velox_dwio_parquet_table_scan_test: 54 tests pass.velox_parquet_e2e_filter_test: 34 tests pass.Part of #2.