Skip to content

Commit 0979fc3

Browse files
UnamedRusclaude
andcommitted
Fix Iceberg min/max pruning lost under object_storage_cluster
Under `object_storage_cluster` (distributed object-storage reads), the WHERE predicate did not reach `ReadFromCluster`, so Iceberg min/max file pruning was silently skipped (`IcebergMinMaxIndexPrunedFiles=0`) and selective queries scanned the whole table. Measured on IcebergBench q16: 3.16B vs 78.7M rows, ~7.0s vs ~0.8s (~60x read amplification). The prune-only `ObjectFilterStep` that carries the predicate to the cluster task iterator (`getTaskIteratorExtension`) was gated on `use_hive_partitioning`, so a non-hive Iceberg cluster read got a null filter. Add `ObjectFilterStep` for any `ReadFromCluster` with a WHERE, not just hive-partitioned tables. `ObjectFilterStep::updatePipeline` is a no-op, so it never filters rows on the initiator -- required at `WithMergeableState`, where the filter columns may be absent from the blocks returned by cluster replicas. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
1 parent f13f6a5 commit 0979fc3

1 file changed

Lines changed: 9 additions & 3 deletions

File tree

‎src/Planner/Planner.cpp‎

Lines changed: 9 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -170,7 +170,6 @@ namespace Setting
170170
extern const SettingsBool serialize_string_in_memory_with_zero_byte;
171171
extern const SettingsString temporary_files_codec;
172172
extern const SettingsNonZeroUInt64 temporary_files_buffer_size;
173-
extern const SettingsBool use_hive_partitioning;
174173
}
175174

176175
namespace ServerSetting
@@ -2460,8 +2459,15 @@ void Planner::buildPlanForQueryNode()
24602459

24612460
if (query_processing_info.isSecondStage() || query_processing_info.isFromAggregationState())
24622461
{
2463-
if (settings[Setting::use_hive_partitioning]
2464-
&& !query_processing_info.isFirstStage()
2462+
/// Deliver the WHERE predicate to distributed object-storage reads (ReadFromCluster) for
2463+
/// object/file pruning: Iceberg min/max & partition pruning, Hive partition pruning, etc.
2464+
/// ObjectFilterStep is prune-only (its updatePipeline is a no-op), so it never filters rows
2465+
/// on the initiator -- required here because at WithMergeableState the filter columns may not
2466+
/// exist in the blocks returned by cluster replicas. This must NOT be gated on
2467+
/// use_hive_partitioning: otherwise a non-hive Iceberg cluster read gets a null filter and
2468+
/// min/max pruning is silently skipped (IcebergMinMaxIndexPrunedFiles=0 -> full-table
2469+
/// over-read on selective queries).
2470+
if (!query_processing_info.isFirstStage()
24652471
&& expression_analysis_result.hasWhere())
24662472
{
24672473
if (typeid_cast<ReadFromCluster *>(query_plan.getRootNode()->step.get()))

0 commit comments

Comments
 (0)