Commit 0979fc3
Fix Iceberg min/max pruning lost under object_storage_cluster
Under `object_storage_cluster` (distributed object-storage reads), the WHERE
predicate did not reach `ReadFromCluster`, so Iceberg min/max file pruning was
silently skipped (`IcebergMinMaxIndexPrunedFiles=0`) and selective queries
scanned the whole table. Measured on IcebergBench q16: 3.16B vs 78.7M rows,
~7.0s vs ~0.8s (~60x read amplification).
The prune-only `ObjectFilterStep` that carries the predicate to the cluster
task iterator (`getTaskIteratorExtension`) was gated on `use_hive_partitioning`,
so a non-hive Iceberg cluster read got a null filter. Add `ObjectFilterStep`
for any `ReadFromCluster` with a WHERE, not just hive-partitioned tables.
`ObjectFilterStep::updatePipeline` is a no-op, so it never filters rows on the
initiator -- required at `WithMergeableState`, where the filter columns may be
absent from the blocks returned by cluster replicas.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>1 parent f13f6a5 commit 0979fc3
1 file changed
Lines changed: 9 additions & 3 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
170 | 170 | | |
171 | 171 | | |
172 | 172 | | |
173 | | - | |
174 | 173 | | |
175 | 174 | | |
176 | 175 | | |
| |||
2460 | 2459 | | |
2461 | 2460 | | |
2462 | 2461 | | |
2463 | | - | |
2464 | | - | |
| 2462 | + | |
| 2463 | + | |
| 2464 | + | |
| 2465 | + | |
| 2466 | + | |
| 2467 | + | |
| 2468 | + | |
| 2469 | + | |
| 2470 | + | |
2465 | 2471 | | |
2466 | 2472 | | |
2467 | 2473 | | |
| |||
0 commit comments