Antalya 26.8: Parquet v3 read concurrency - #2380
Conversation
Parquet v3 read concurrency
…MemoryReservation` The merge of `antalya-26.8` brought a two-argument call to `SharedResourcesExt::getLimitsPerReader`, which now takes three. https://altinity-build-artifacts.s3.amazonaws.com/json.html?PR=2380&sha=01f998a22108bcbe366d1d0286cea441687a0b93&name_0=PR&name_1=Build%20%28amd_debug%29 #2380
…balance The rebalance gives the `BloomFilterBlocksOrDictionary` stage 0.10 of `input_format_parquet_memory_high_watermark` instead of the previous uniform 1/5, halving the budget available to dictionary-filter pruning. The "moderate budget" cases of `04616` and `04651` were calibrated for 1/5 and fell back to a full scan. Double their watermarks so the pruning stage gets the same per-stage budget as before (10 MB and 34 MB); the double-counting bounds these tests guard against still exceed it. CI report: https://altinity-build-artifacts.s3.amazonaws.com/json.html?PR=2380&sha=latest&name_0=PR PR: #2380
…erg `time` mapping change Iceberg `time` now maps to `Time64(6)` (port of #2129), but the unit test still expected `Int64`. CI report: https://altinity-build-artifacts.s3.amazonaws.com/json.html?PR=2380&sha=0d3a8ae0c66bdb9bb6a0a67ebb78df26b3c6d247&name_0=PR&name_1=Unit%20tests%20%28asan_ubsan%29 PR: #2380
CI triage for
|
Changelog category (leave one):
Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md):
Parquet reader now fetches compressed chunks from storage far ahead of decoding. Compressed data is small, so many reads stay in flight while decoding catches up at its own pace — cold scans are no longer bottlenecked on storage latency (S3).
Read-ahead and decoding get separate memory and thread budgets, tunable via
input_format_parquet_prefetch_memory_fraction(0.6) andinput_format_parquet_decode_thread_fraction(0.375) (#2235 by @UnamedRus).Cherry-picked from #2235.
Number of changes to bring parquet v3 reader perf closer to arrow based