Skip to content

perf(parquet): SIMD bit-unpack for DeltaBpDecoder dense and sparse visitor paths#5

Closed
jaylisde wants to merge 1 commit into
pr/delta-bytearray-perffrom
perf/delta-bp-simd-unpack-sparse
Closed

perf(parquet): SIMD bit-unpack for DeltaBpDecoder dense and sparse visitor paths#5
jaylisde wants to merge 1 commit into
pr/delta-bytearray-perffrom
perf/delta-bp-simd-unpack-sparse

Conversation

@jaylisde

@jaylisde jaylisde commented Jun 1, 2026

Copy link
Copy Markdown
Owner

Summary

Builds on #3 (PR A) and #4 (PR B). Closes the remaining DELTA-vs-PLAIN/DICT scan gap on TPC-H Q12.

  1. Inline SIMD bit-unpack in DeltaBpDecoder::decodeLongs. When the read aligns at a miniblock start and consumes a whole miniblock, dispatch on bit_width 0..32 to a compile-time-specialized kernel: bw=0 is an arithmetic-sequence fast path; bw 1..16 does 4-value/iter with a single unaligned 64-bit load (4×16 fits in one u64 window); bw 17..32 does 2-value/iter with __uint128_t funnel-shift (lowers to two u64 loads + SHRD on x86_64). The trailing u64 read is safe — intra-page overshoots fall into the next miniblock and the last miniblock has kPageReadPadding (8) trailing bytes guaranteed. bw 33..64 falls through to the existing scalar inner loop.

  2. readWithVisitorSparseBuffered path for !Visitor::dense + !hasNulls + deterministic filter + NoHook + integral DataType (Q12's hot path, since chained filters produce a sparse row set after l_shipmode is applied). Decodes kBatch=1024 physical values into a stack buffer via decodeLongs (now SIMD), then walks the visitor's sparse rows[] array using rows[k] - batchPhysStart as buffer index. The existing per-row visitor.process is preserved — only the decode side is batched. n is capped to the visitor's residual physical span so the decoder never advances past what the visitor will consume.

  3. readWithVisitorDenseBatched kBatch raised from 256 to 1024 to amortize the chunk-loop overhead now that decodeLongs is much faster per call.

Bench

TPC-H Q12 SF10, DELTA-encoded lineitem, num_drivers=4, 5-run median:

Wall lineitem TableScan CPU
#3 baseline (PR A) 0.85 s 2.13 s
+ #4 (PR B) 0.74 s 1.74 s
+ this PR 0.626 s 1.38 s

PLAIN/DICT parquet paths are unchanged within noise. Bench will be re-measured (5-run median) before posting upstream.

Test plan

  • velox_dwio_parquet_delta_bp_decoder_test (new): 43 tests pass — 11 boundary cases (bit_widths 0/8/10/16/24/32, multi-block in a single readValues, mid-miniblock split across two calls, negative minDelta, bit_width > 32 fallback, narrowing to int32_t) plus 32 parameterized roundtrip tests covering bit_widths 1..32.
  • velox_dwio_parquet_table_scan_test: 54 tests pass.
  • velox_parquet_e2e_filter_test: 34 tests pass.

Part of #2.

@jaylisde
jaylisde force-pushed the pr/delta-bytearray-perf branch from 38359d6 to 7967878 Compare June 3, 2026 00:01
@jaylisde
jaylisde force-pushed the perf/delta-bp-simd-unpack-sparse branch from fb81bda to 18a769a Compare June 3, 2026 00:02
@jaylisde
jaylisde force-pushed the pr/delta-bytearray-perf branch from 7967878 to e987d8d Compare June 3, 2026 06:30
@jaylisde
jaylisde force-pushed the perf/delta-bp-simd-unpack-sparse branch from 18a769a to 2312b0a Compare June 3, 2026 06:35
@jaylisde
jaylisde force-pushed the pr/delta-bytearray-perf branch from e987d8d to a9359a7 Compare June 3, 2026 06:49
…sitor paths

Builds on #3 (PR A) and #4 (PR B) to close
the remaining DELTA-vs-PLAIN/DICT scan gap on TPC-H Q12.

1. Inline SIMD bit-unpack kernel in DeltaBpDecoder::decodeLongs.
   When the read aligns at a miniblock start and consumes a whole
   miniblock, dispatch on bit_width 0..32 to a compile-time
   specialized kernel:
   - bw 0: arithmetic-sequence fast path (no bit-extract).
   - bw 1..16: 4-value/iter, single unaligned 64-bit load (4*16 = 64
     bits fit in one u64 window).
   - bw 17..32: 2-value/iter, __uint128_t funnel-shift (lowers to
     two u64 loads + SHRD on x86_64). The trailing u64 read is safe
     because intra-page overshoots fall into the next miniblock and
     the last miniblock has PageReader::kPageReadPadding (8) trailing
     bytes guaranteed.
   bit_widths 33..64 fall through to the per-row scalar inner loop,
   which is unchanged.

2. New readWithVisitorSparseBuffered path for the
     !Visitor::dense + !hasNulls + deterministic filter +
     NoHook + integral DataType
   case (the hot path on Q12 because filter chaining produces sparse
   row sets after l_shipmode is applied). Decodes kBatch=1024
   physical values into a stack buffer via decodeLongs (which now
   uses the SIMD kernel above), then walks the visitor's sparse rows
   array using rows[k] - batchPhysStart as buffer index. The
   existing per-row visitor.process is preserved — only the decode
   side is batched. n is capped to the visitor's residual physical
   span so the decoder never advances past what the visitor will
   consume.

3. readWithVisitorDenseBatched kBatch raised from 256 to 1024 to
   amortize the chunk-loop overhead now that decodeLongs is much
   faster per call.

4. Add velox/dwio/parquet/tests/reader/DeltaBpDecoderTest.cpp:
   - Hand-rolled DeltaEncoder for byte-stream control in tests.
   - 11 boundary tests: bit_widths 0/8/10/16/24/32, multi-block in a
     single readValues, mid-miniblock split across two readValues
     calls, negative minDelta, bit_width > 32 fallback, narrowing to
     int32_t.
   - 32 parameterized roundtrip tests covering bit_widths 1..32, each
     forcing the encoder to pick exactly that width by saturating one
     residual to (1<<bw)-1.

Bench (TPC-H Q12 SF10, DELTA-encoded lineitem, num_drivers=4, 5-run
median):

                                    wall
  #3 baseline         897 ms
  +#4 (PR B)          748 ms
  +this PR                          626 ms

That is -16.3% on top of PR B, -30.2% from the PR A baseline. PLAIN/DICT
parquet paths are unchanged.

Test plan
- velox_dwio_parquet_delta_bp_decoder_test: 43 tests pass (11 boundary
  + 32 parameterized bw 1..32).
- velox_dwio_parquet_table_scan_test: 54 tests pass (includes the 7
  delta tests added in PR A).
- velox_parquet_e2e_filter_test: 34 tests pass.

Tracking: #2.
@jaylisde

jaylisde commented Jun 3, 2026

Copy link
Copy Markdown
Owner Author

Squashed into combined PR #6 (perf/delta-decoder-overhaul). Closing this stacked sub-PR — review continues on #6.

@jaylisde jaylisde closed this Jun 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant