diff --git a/benchmarks/README.md b/benchmarks/README.md index 151f1e7df5fb8..a14bfe4c9a125 100644 --- a/benchmarks/README.md +++ b/benchmarks/README.md @@ -966,6 +966,28 @@ Several queries are included to test hash joins under various workloads. ./bench.sh run hj ``` +## Hash Join Dynamic Filter on an Ordered Subset + +This benchmark (`hj_ordered_subset`) measures the dynamic filter that a hash join pushes to its probe-side Parquet scan when the probe table is sorted by the join key and the build side matches an ordered subset of the key range. This is common for time-ordered fact tables: the events of one day, or of a range of days. + +The load SQL writes an `events` table of 100 days x 200,000 rows (20M rows, 100,000-row row groups) sorted by `event_id`, and two build tables with one row for each event. The queries join `events` to 1 day (Q01, Q04), to 10 days (Q02, Q05), or to 1% of the events scattered over the whole range (Q03, Q06, the control). The bounds of the dynamic filter (`event_id >= min AND event_id <= max`) can prune the `events` row groups outside the matched range, and the membership check passes almost every row that remains. In the control, the bounds cannot prune anything. + +The `partitioned` subgroup (Q01-Q03) forces `HashJoinExec` mode `Partitioned`, and the `collect_left` subgroup (Q04-Q06) forces mode `CollectLeft`. + +### Example Run + +```bash +# No need to generate data: the suite's load SQL writes the Parquet files + +./bench.sh run hj_ordered_subset + +# With parquet filter pushdown +DATAFUSION_EXECUTION_PARQUET_PUSHDOWN_FILTERS=true ./bench.sh run hj_ordered_subset + +# Smaller data: 50 days x 100,000 rows +HJOS_DAYS=50 HJOS_ROWS_PER_DAY=100000 ./bench.sh run hj_ordered_subset +``` + ## Null-Aware Join This benchmark focuses on `NOT IN` subqueries, which plan as null-aware joins: an diff --git a/benchmarks/bench.sh b/benchmarks/bench.sh index 66b6dcd1b12b3..aee0cd25a46c3 100755 --- a/benchmarks/bench.sh +++ b/benchmarks/bench.sh @@ -119,6 +119,10 @@ spill_views: Sort and GROUP BY queries that spill StringView/BinaryVi null_aware_join: Null-aware (NOT IN) hash join micro-benchmarks: uncorrelated, non-equality-correlated and equality-correlated NOT IN across NULL fractions, to measure the per-pair join-filter work the correlated cases do (data generated inline by the suite's load SQL from range(); knobs: NAJ_ROWS, NAJ_LARGE_ROWS) +hj_ordered_subset: Hash join dynamic filter on a probe table sorted by the join key, where the build side matches an ordered subset + (1 day, 10 days) of the key range, so the filter bounds can prune probe row groups; plus a scattered 1% control + (subgroups via BENCH_SUBGROUP: partitioned, collect_left) + (data generated inline by the suite's load SQL; knobs: HJOS_DAYS, HJOS_ROWS_PER_DAY, HJOS_RG_SIZE) projection_subquery: IN / NOT IN / EXISTS subqueries in the SELECT list (see https://github.com/apache/datafusion/issues/25341); each query projects the boolean subquery result and aggregates it, so the cost is the decorrelation plan and not the output size (q07 correlates on '<' instead of '=', so it keeps the nested-loop plan and acts as the control) @@ -295,6 +299,10 @@ main() { # Data is generated inline by the suite's load SQL. echo "projection_subquery: no external data to generate" ;; + hj_ordered_subset) + # Data is generated inline by the suite's load SQL (COPY). + echo "hj_ordered_subset: no external data to generate" + ;; asof_join) data_asof_join ;; @@ -552,6 +560,9 @@ main() { projection_subquery) run_projection_subquery ;; + hj_ordered_subset) + run_hj_ordered_subset + ;; asof_join) run_asof_join ;; @@ -1009,6 +1020,29 @@ run_null_aware_join() { bash -c "$SQL_CARGO_COMMAND" } +# Runs the hj_ordered_subset suite: a hash join whose probe table is sorted by +# the join key and whose build side matches an ordered subset of the key range, +# so the bounds of the join's dynamic filter can prune probe row groups. Data is +# generated inline by the load SQL (COPY), so there is no data step. Set +# DATAFUSION_EXECUTION_PARQUET_PUSHDOWN_FILTERS to vary how the scan uses the +# filter. +# Knobs (string-substituted into the load SQL, not engine config): +# BENCH_SUBGROUP run one subgroup (partitioned, collect_left) +# HJOS_DAYS days in the events table (default 100, at least 50) +# HJOS_ROWS_PER_DAY events per day (default 200_000) +# HJOS_RG_SIZE parquet row-group size (default 100_000) +run_hj_ordered_subset() { + echo "Running hj_ordered_subset benchmark (subgroup=${BENCH_SUBGROUP:-all}, days=${HJOS_DAYS:-100}, rows_per_day=${HJOS_ROWS_PER_DAY:-200000}, rg_size=${HJOS_RG_SIZE:-100000})..." + debug_run env BENCH_NAME=hj_ordered_subset \ + BENCH_RESULTS_FILE="$(sql_results_file hj_ordered_subset)" \ + ${BENCH_SUBGROUP:+BENCH_SUBGROUP="${BENCH_SUBGROUP}"} \ + HJOS_DAYS="${HJOS_DAYS:-100}" \ + HJOS_ROWS_PER_DAY="${HJOS_ROWS_PER_DAY:-200000}" \ + HJOS_RG_SIZE="${HJOS_RG_SIZE:-100000}" \ + ${QUERY:+BENCH_QUERY="${QUERY}"} \ + bash -c "$SQL_CARGO_COMMAND" +} + # Runs the projection_subquery suite: IN / NOT IN / EXISTS subqueries that sit # in the SELECT list instead of a filter (see # https://github.com/apache/datafusion/issues/25341). The load SQL builds the diff --git a/benchmarks/sql_benchmarks/README.md b/benchmarks/sql_benchmarks/README.md index 8047c030927cc..3c5b5d35880fb 100644 --- a/benchmarks/sql_benchmarks/README.md +++ b/benchmarks/sql_benchmarks/README.md @@ -35,6 +35,7 @@ in the community: | `clickbench_sorted` | ClickBench benchmark using a pre-sorted hits file. | | `h2o` | The `h2o` benchmark | | `hj` | Hash join benchmark | +| `hj_ordered_subset` | Hash join dynamic filter on a probe table sorted by the join key, where the build side matches an ordered subset of the key range (1 day, 10 days) or a scattered 1% (control). Subgroups (`--subgroup`): `partitioned`, `collect_left`. Size the data with `HJOS_DAYS`, `HJOS_ROWS_PER_DAY` and `HJOS_RG_SIZE`. | | `imdb` | IMDb benchmark | | `nlj` | Nested‑loop join benchmark | | `null_aware_join` | Null-aware (`NOT IN`) hash join micro-benchmarks. Q01-Q03 are uncorrelated `NOT IN` across NULL fractions and are linear in the table size (`NAJ_LARGE_ROWS`, default `1000000`). Q04-Q08 are correlated, so the correlation predicate stays behind as a join filter that the join applies per candidate (build row × probe row) pair while deciding which outer rows are UNKNOWN; without an equality correlation there are no scope keys to narrow those pairs, so their cost grows with the square of `NAJ_ROWS` (default `10000`). Q08 adds an equality correlation, which turns those pairs into a hash lookup. All tables are built inline from `range()`, so there is no data step. | diff --git a/benchmarks/sql_benchmarks/hj_ordered_subset/benchmarks/q01.benchmark b/benchmarks/sql_benchmarks/hj_ordered_subset/benchmarks/q01.benchmark new file mode 100644 index 0000000000000..7fda4c68149d0 --- /dev/null +++ b/benchmarks/sql_benchmarks/hj_ordered_subset/benchmarks/q01.benchmark @@ -0,0 +1,10 @@ +# Partitioned join, build side = 1 day (1% of the event_id range, ordered). +# The bounds of the dynamic filter can prune ~99% of the events row groups. +subgroup partitioned + +template sql_benchmarks/hj_ordered_subset/hj_ordered_subset.benchmark.template +NAME=Q01_partitioned_ordered_1_day +JOIN_MODE=partitioned +PLAN_MODE=Partitioned +BUILD_TABLE=labels_by_day +BUILD_FILTER=l.day = 42 diff --git a/benchmarks/sql_benchmarks/hj_ordered_subset/benchmarks/q02.benchmark b/benchmarks/sql_benchmarks/hj_ordered_subset/benchmarks/q02.benchmark new file mode 100644 index 0000000000000..7ee85bae2e5d3 --- /dev/null +++ b/benchmarks/sql_benchmarks/hj_ordered_subset/benchmarks/q02.benchmark @@ -0,0 +1,10 @@ +# Partitioned join, build side = 10 days (10% of the event_id range, ordered). +# The bounds of the dynamic filter can prune ~90% of the events row groups. +subgroup partitioned + +template sql_benchmarks/hj_ordered_subset/hj_ordered_subset.benchmark.template +NAME=Q02_partitioned_ordered_10_days +JOIN_MODE=partitioned +PLAN_MODE=Partitioned +BUILD_TABLE=labels_by_day +BUILD_FILTER=l.day BETWEEN 40 AND 49 diff --git a/benchmarks/sql_benchmarks/hj_ordered_subset/benchmarks/q03.benchmark b/benchmarks/sql_benchmarks/hj_ordered_subset/benchmarks/q03.benchmark new file mode 100644 index 0000000000000..1ed1b5dd0768c --- /dev/null +++ b/benchmarks/sql_benchmarks/hj_ordered_subset/benchmarks/q03.benchmark @@ -0,0 +1,11 @@ +# Control: partitioned join, build side = 1% of the events, scattered over the +# whole event_id range. The bounds cannot prune anything; only the membership +# check of the dynamic filter rejects rows. +subgroup partitioned + +template sql_benchmarks/hj_ordered_subset/hj_ordered_subset.benchmark.template +NAME=Q03_partitioned_scattered_1_pct +JOIN_MODE=partitioned +PLAN_MODE=Partitioned +BUILD_TABLE=labels_by_bucket +BUILD_FILTER=l.bucket = 42 diff --git a/benchmarks/sql_benchmarks/hj_ordered_subset/benchmarks/q04.benchmark b/benchmarks/sql_benchmarks/hj_ordered_subset/benchmarks/q04.benchmark new file mode 100644 index 0000000000000..dfcbe3b58470f --- /dev/null +++ b/benchmarks/sql_benchmarks/hj_ordered_subset/benchmarks/q04.benchmark @@ -0,0 +1,9 @@ +# Same as Q01 with a CollectLeft join. +subgroup collect_left + +template sql_benchmarks/hj_ordered_subset/hj_ordered_subset.benchmark.template +NAME=Q04_collect_left_ordered_1_day +JOIN_MODE=collect_left +PLAN_MODE=CollectLeft +BUILD_TABLE=labels_by_day +BUILD_FILTER=l.day = 42 diff --git a/benchmarks/sql_benchmarks/hj_ordered_subset/benchmarks/q05.benchmark b/benchmarks/sql_benchmarks/hj_ordered_subset/benchmarks/q05.benchmark new file mode 100644 index 0000000000000..31ce7cf4e0c92 --- /dev/null +++ b/benchmarks/sql_benchmarks/hj_ordered_subset/benchmarks/q05.benchmark @@ -0,0 +1,9 @@ +# Same as Q02 with a CollectLeft join. +subgroup collect_left + +template sql_benchmarks/hj_ordered_subset/hj_ordered_subset.benchmark.template +NAME=Q05_collect_left_ordered_10_days +JOIN_MODE=collect_left +PLAN_MODE=CollectLeft +BUILD_TABLE=labels_by_day +BUILD_FILTER=l.day BETWEEN 40 AND 49 diff --git a/benchmarks/sql_benchmarks/hj_ordered_subset/benchmarks/q06.benchmark b/benchmarks/sql_benchmarks/hj_ordered_subset/benchmarks/q06.benchmark new file mode 100644 index 0000000000000..962a6d35fb5d5 --- /dev/null +++ b/benchmarks/sql_benchmarks/hj_ordered_subset/benchmarks/q06.benchmark @@ -0,0 +1,9 @@ +# Control: same as Q03 with a CollectLeft join. +subgroup collect_left + +template sql_benchmarks/hj_ordered_subset/hj_ordered_subset.benchmark.template +NAME=Q06_collect_left_scattered_1_pct +JOIN_MODE=collect_left +PLAN_MODE=CollectLeft +BUILD_TABLE=labels_by_bucket +BUILD_FILTER=l.bucket = 42 diff --git a/benchmarks/sql_benchmarks/hj_ordered_subset/hj_ordered_subset.benchmark.template b/benchmarks/sql_benchmarks/hj_ordered_subset/hj_ordered_subset.benchmark.template new file mode 100644 index 0000000000000..91a227734243a --- /dev/null +++ b/benchmarks/sql_benchmarks/hj_ordered_subset/hj_ordered_subset.benchmark.template @@ -0,0 +1,41 @@ +# Shared template for the hj_ordered_subset suite. Each qNN.benchmark sets its +# `subgroup` directive and includes this template with parameters: +# NAME criterion display name +# JOIN_MODE init script stem under init/ (partitioned, collect_left) +# PLAN_MODE expected HashJoinExec mode (Partitioned, CollectLeft) +# BUILD_TABLE build table (labels_by_day, labels_by_bucket) +# BUILD_FILTER filter on the build table, which selects the build side +# +# Every query joins the full `events` table (probe side) to a filtered build +# table on `event_id`. Every build row matches exactly one event. + +load sql_benchmarks/hj_ordered_subset/init/load.sql + +init sql_benchmarks/hj_ordered_subset/init/${JOIN_MODE}.sql + +name ${NAME} +group hj_ordered_subset + +# The queries select days 40..49 and bucket 42 of 100. +assert I +SELECT max(day) >= 49 FROM events; +---- +true + +# Correctness canary: every build row matches exactly one event. A dynamic +# filter that drops a matching row group or row trips this assert. +assert I +SELECT (SELECT count(*) FROM events e JOIN ${BUILD_TABLE} l ON e.event_id = l.event_id WHERE ${BUILD_FILTER}) + = (SELECT count(*) FROM ${BUILD_TABLE} l WHERE ${BUILD_FILTER}); +---- +true + +expect_plan HashJoinExec: mode=${PLAN_MODE} + +run +SELECT count(*) AS n, sum(e.amount_cents) AS amount_cents, sum(e.v1) AS v1, max(e.note) AS note +FROM events e +JOIN ${BUILD_TABLE} l ON e.event_id = l.event_id +WHERE ${BUILD_FILTER}; + +cleanup sql_benchmarks/hj_ordered_subset/init/cleanup.sql diff --git a/benchmarks/sql_benchmarks/hj_ordered_subset/hj_ordered_subset.suite b/benchmarks/sql_benchmarks/hj_ordered_subset/hj_ordered_subset.suite new file mode 100644 index 0000000000000..edac226add99b --- /dev/null +++ b/benchmarks/sql_benchmarks/hj_ordered_subset/hj_ordered_subset.suite @@ -0,0 +1,32 @@ +description = "Hash join dynamic filter on a probe table sorted by the join key, where the build side matches an ordered subset of the key range. The load SQL writes a fact table of HJOS_DAYS x HJOS_ROWS_PER_DAY events sorted by event_id, and build tables that select 1 day, 10 days, or 1% of the events scattered over the whole range (control). The bounds of the dynamic filter can prune the probe row groups outside the matched range, and the membership check passes almost every remaining row. Subgroups: partitioned (HashJoinExec mode=Partitioned) and collect_left (mode=CollectLeft). Toggle parquet filter pushdown with its DATAFUSION_* environment variable." + +query_pattern = "q{QUERY_ID_PADDED}.benchmark" + +[[options]] +name = "days" +env = "HJOS_DAYS" +default = "100" +values = ["100", "..."] +help = "Sets the number of days in the generated events table (at least 50; the queries select days 40..49)." + +[[options]] +name = "rows-per-day" +env = "HJOS_ROWS_PER_DAY" +default = "200000" +values = ["200000", "..."] +help = "Sets the number of events per day. The build side of the 1-day queries has this many rows." + +[[options]] +name = "rg-size" +env = "HJOS_RG_SIZE" +default = "100000" +values = ["100000", "..."] +help = "Sets the Parquet max_row_group_size used when writing the generated tables." + +[[examples]] +command = "cargo run --release --bin benchmark_runner -- hj_ordered_subset" +description = "Run every query with 100 days x 200,000 events (20M rows) and 100,000-row row groups." + +[[examples]] +command = "DATAFUSION_EXECUTION_PARQUET_PUSHDOWN_FILTERS=true cargo run --release --bin benchmark_runner -- hj_ordered_subset --subgroup partitioned" +description = "Run only the partitioned joins, with parquet filter pushdown." diff --git a/benchmarks/sql_benchmarks/hj_ordered_subset/init/cleanup.sql b/benchmarks/sql_benchmarks/hj_ordered_subset/init/cleanup.sql new file mode 100644 index 0000000000000..0e9b81412a0fe --- /dev/null +++ b/benchmarks/sql_benchmarks/hj_ordered_subset/init/cleanup.sql @@ -0,0 +1,3 @@ +DROP TABLE IF EXISTS events; +DROP TABLE IF EXISTS labels_by_day; +DROP TABLE IF EXISTS labels_by_bucket; diff --git a/benchmarks/sql_benchmarks/hj_ordered_subset/init/collect_left.sql b/benchmarks/sql_benchmarks/hj_ordered_subset/init/collect_left.sql new file mode 100644 index 0000000000000..bf0ed7a452327 --- /dev/null +++ b/benchmarks/sql_benchmarks/hj_ordered_subset/init/collect_left.sql @@ -0,0 +1,3 @@ +-- Force HashJoinExec mode=CollectLeft, also for the 10-day build side. +set datafusion.optimizer.hash_join_single_partition_threshold = 1073741824; +set datafusion.optimizer.hash_join_single_partition_threshold_rows = 1000000000; diff --git a/benchmarks/sql_benchmarks/hj_ordered_subset/init/load.sql b/benchmarks/sql_benchmarks/hj_ordered_subset/init/load.sql new file mode 100644 index 0000000000000..4748f5d69fc72 --- /dev/null +++ b/benchmarks/sql_benchmarks/hj_ordered_subset/init/load.sql @@ -0,0 +1,71 @@ +-- Data for the hj_ordered_subset suite. +-- +-- `events` is the probe (fact) table: HJOS_DAYS days x HJOS_ROWS_PER_DAY rows. +-- `event_id` increases with time, and the file is written in `event_id` +-- order, so each row group holds a disjoint, sorted range of `event_id` and +-- of `day`. +-- +-- The two build tables hold one row for every event: +-- - `labels_by_day` is sorted by `event_id` (thus by `day`). A filter on +-- `day` selects a contiguous, ordered range of `event_id`. +-- - `labels_by_bucket` is sorted by a pseudo-random `bucket` (0..99). A +-- filter on `bucket` selects every 100th `event_id`, spread over the whole +-- range. +-- Both build filters are pruned by the static row-group statistics of the +-- build table, so the build scan is cheap and the same for every query shape. +-- +-- The ORDER BY clauses are load-bearing: the benchmark measures row-group +-- pruning of `events` by the join's dynamic filter, which needs disjoint +-- row-group ranges. +-- +-- Knobs: HJOS_DAYS (must be at least 50), HJOS_ROWS_PER_DAY, HJOS_RG_SIZE +-- (parquet max row-group size). +COPY ( + SELECT + value AS event_id, + CAST(value / ${HJOS_ROWS_PER_DAY:-200000} AS INT) AS day, + (value * 7919) % 1000003 AS user_id, + (value * 37) % 10000 AS amount_cents, + (value * 13) % 1000000 AS v1, + concat('note-', CAST(value % 997 AS VARCHAR)) AS note + FROM generate_series(0, ${HJOS_DAYS:-100} * ${HJOS_ROWS_PER_DAY:-200000} - 1) + ORDER BY value +) +TO 'sql_benchmarks/hj_ordered_subset/scratch/events.parquet' +STORED AS PARQUET +OPTIONS ('format.max_row_group_size' '${HJOS_RG_SIZE:-100000}'); + +COPY ( + SELECT + value AS event_id, + CAST(value / ${HJOS_ROWS_PER_DAY:-200000} AS INT) AS day + FROM generate_series(0, ${HJOS_DAYS:-100} * ${HJOS_ROWS_PER_DAY:-200000} - 1) + ORDER BY value +) +TO 'sql_benchmarks/hj_ordered_subset/scratch/labels_by_day.parquet' +STORED AS PARQUET +OPTIONS ('format.max_row_group_size' '${HJOS_RG_SIZE:-100000}'); + +-- 7919 is coprime to 100, so `bucket` is a bijection of `value % 100`. +COPY ( + SELECT + value AS event_id, + CAST((value * 7919) % 100 AS INT) AS bucket + FROM generate_series(0, ${HJOS_DAYS:-100} * ${HJOS_ROWS_PER_DAY:-200000} - 1) + ORDER BY bucket, value +) +TO 'sql_benchmarks/hj_ordered_subset/scratch/labels_by_bucket.parquet' +STORED AS PARQUET +OPTIONS ('format.max_row_group_size' '${HJOS_RG_SIZE:-100000}'); + +CREATE EXTERNAL TABLE events +STORED AS PARQUET +LOCATION 'sql_benchmarks/hj_ordered_subset/scratch/events.parquet'; + +CREATE EXTERNAL TABLE labels_by_day +STORED AS PARQUET +LOCATION 'sql_benchmarks/hj_ordered_subset/scratch/labels_by_day.parquet'; + +CREATE EXTERNAL TABLE labels_by_bucket +STORED AS PARQUET +LOCATION 'sql_benchmarks/hj_ordered_subset/scratch/labels_by_bucket.parquet'; diff --git a/benchmarks/sql_benchmarks/hj_ordered_subset/init/partitioned.sql b/benchmarks/sql_benchmarks/hj_ordered_subset/init/partitioned.sql new file mode 100644 index 0000000000000..14f0cf6adffa1 --- /dev/null +++ b/benchmarks/sql_benchmarks/hj_ordered_subset/init/partitioned.sql @@ -0,0 +1,4 @@ +-- Force HashJoinExec mode=Partitioned: the estimated size of a filtered +-- build side is often below the CollectLeft thresholds. +set datafusion.optimizer.hash_join_single_partition_threshold = 0; +set datafusion.optimizer.hash_join_single_partition_threshold_rows = 0; diff --git a/benchmarks/sql_benchmarks/hj_ordered_subset/scratch/.gitignore b/benchmarks/sql_benchmarks/hj_ordered_subset/scratch/.gitignore new file mode 100644 index 0000000000000..4bed5da93fb28 --- /dev/null +++ b/benchmarks/sql_benchmarks/hj_ordered_subset/scratch/.gitignore @@ -0,0 +1 @@ +*.parquet diff --git a/datafusion/common/src/config.rs b/datafusion/common/src/config.rs index 5f45705301473..44b09368870a5 100644 --- a/datafusion/common/src/config.rs +++ b/datafusion/common/src/config.rs @@ -1185,6 +1185,17 @@ config_namespace! { /// /// Disabled by default, set to a number greater than 0 for enabling it. pub hash_join_buffering_capacity: usize, default = 0 + + /// The assumed work, in nanoseconds, that each row removed by an + /// optional filter saves downstream. Optional filters are filters that + /// are not needed for correctness, such as the dynamic filters that + /// hash joins and TopK push down into scans. When an operator evaluates + /// optional filters adaptively, it pauses an optional filter whose + /// evaluation costs more than the work that it saves. Consumers that can + /// measure the saving (the Parquet scan) add their measured decode cost. + /// The default is about the cost of a hash table probe for one row. The + /// best value depends on the hardware. + pub optional_filter_min_saving_ns_per_row: f64, default = 20.0 } } diff --git a/datafusion/core/tests/physical_optimizer/filter_pushdown.rs b/datafusion/core/tests/physical_optimizer/filter_pushdown.rs index ddd4bb4d271aa..789c39e61fe89 100644 --- a/datafusion/core/tests/physical_optimizer/filter_pushdown.rs +++ b/datafusion/core/tests/physical_optimizer/filter_pushdown.rs @@ -51,8 +51,8 @@ use datafusion_functions_aggregate::{ }; use datafusion_physical_expr::{ LexOrdering, PhysicalSortExpr, - expressions::{DynamicFilterPhysicalExpr, IsNullExpr, cast, col}, - utils::conjunction, + expressions::{IsNullExpr, cast, col}, + utils::{as_dynamic_filter, conjunction}, }; use datafusion_physical_expr::{ Partitioning, RangePartitioning, ScalarFunctionExpr, SplitPoint, @@ -941,7 +941,7 @@ async fn test_topk_filter_passes_through_coalesce_partitions() { Ok: - SortExec: TopK(fetch=1), expr=[b@1 DESC NULLS LAST], preserve_partitioning=[false] - CoalescePartitionsExec - - DataSourceExec: file_groups={2 groups: [[test1.parquet], [test2.parquet]]}, projection=[a, b, c], file_type=test, pushdown_supported=true, predicate=DynamicFilter [ empty ] + - DataSourceExec: file_groups={2 groups: [[test1.parquet], [test2.parquet]]}, projection=[a, b, c], file_type=test, pushdown_supported=true, predicate=Optional(DynamicFilter [ empty ]) " ); } @@ -1010,7 +1010,7 @@ async fn optimize_and_collect_pushdown_plan( // Not portable to sqllogictest: this test pins `PartitionMode::Partitioned` // by hand-wiring `RepartitionExec(Hash, 12)` on both join sides. A SQL // INNER JOIN over small parquet inputs plans as `CollectLeft`, so the -// per-partition CASE filter this test exercises is not reachable via SQL. +// partitioned filter this test exercises is not reachable via SQL. #[tokio::test] async fn test_hashjoin_dynamic_filter_pushdown_partitioned() { // Rough sketch of the MRE we're trying to recreate: @@ -1144,7 +1144,7 @@ async fn test_hashjoin_dynamic_filter_pushdown_partitioned() { - RepartitionExec: partitioning=Hash([a@0, b@1], 12), input_partitions=1 - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b, c], file_type=test, pushdown_supported=true - RepartitionExec: partitioning=Hash([a@0, b@1], 12), input_partitions=1 - - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b, e], file_type=test, pushdown_supported=true, predicate=DynamicFilter [ empty ] + - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b, e], file_type=test, pushdown_supported=true, predicate=Optional(DynamicFilter [ empty ]) AND Optional(DynamicFilter [ empty ]) " ); @@ -1165,7 +1165,7 @@ async fn test_hashjoin_dynamic_filter_pushdown_partitioned() { - RepartitionExec: partitioning=Hash([a@0, b@1], 12), input_partitions=1 - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b, c], file_type=test, pushdown_supported=true - RepartitionExec: partitioning=Hash([a@0, b@1], 12), input_partitions=1 - - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b, e], file_type=test, pushdown_supported=true, predicate=DynamicFilter [ CASE hash_repartition % 12 WHEN 5 THEN a@0 >= ab AND a@0 <= ab AND b@1 >= bb AND b@1 <= bb AND struct(a@0, b@1) IN (SET) ([{c0:ab,c1:bb}]) WHEN 8 THEN a@0 >= aa AND a@0 <= aa AND b@1 >= ba AND b@1 <= ba AND struct(a@0, b@1) IN (SET) ([{c0:aa,c1:ba}]) ELSE false END ] + - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b, e], file_type=test, pushdown_supported=true, predicate=Optional(DynamicFilter [ a@0 >= aa AND a@0 <= ab AND b@1 >= ba AND b@1 <= bb ]) AND Optional(DynamicFilter [ struct(a@0, b@1) IN (SET) ([{c0:aa,c1:ba}, {c0:ab,c1:bb}]) ]) " ); @@ -1360,7 +1360,7 @@ async fn test_hashjoin_dynamic_filter_pushdown_range_partitioned() { - RepartitionExec: partitioning=Range([a@0 ASC, b@1 ASC], [(aa, bb)], 2), input_partitions=1 - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b, c], file_type=test, pushdown_supported=true - RepartitionExec: partitioning=Range([a@0 ASC, b@1 ASC], [(aa, bb)], 2), input_partitions=1 - - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b, e], file_type=test, pushdown_supported=true, predicate=DynamicFilter [ empty ] + - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b, e], file_type=test, pushdown_supported=true, predicate=Optional(DynamicFilter [ empty ]) AND Optional(DynamicFilter [ empty ]) " ); @@ -1369,6 +1369,11 @@ async fn test_hashjoin_dynamic_filter_pushdown_range_partitioned() { config.execution.parquet.pushdown_filters = true; config.optimizer.enable_dynamic_filter_pushdown = true; config.optimizer.preserve_file_partitions = 1; + // Push hash table lookups instead of `InList`s so the filter keeps the + // `range_partition` routing: an all-`InList` build collapses into one `InList`. + config + .optimizer + .hash_join_inlist_pushdown_max_distinct_values = 0; let (plan, batches) = optimize_and_collect_pushdown_plan(plan, config).await; // Now check what our filter looks like @@ -1381,7 +1386,7 @@ async fn test_hashjoin_dynamic_filter_pushdown_range_partitioned() { - RepartitionExec: partitioning=Range([a@0 ASC, b@1 ASC], [(aa, bb)], 2), input_partitions=1 - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b, c], file_type=test, pushdown_supported=true - RepartitionExec: partitioning=Range([a@0 ASC, b@1 ASC], [(aa, bb)], 2), input_partitions=1 - - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b, e], file_type=test, pushdown_supported=true, predicate=DynamicFilter [ CASE range_partition WHEN 0 THEN a@0 >= aa AND a@0 <= aa AND b@1 >= ba AND b@1 <= ba AND struct(a@0, b@1) IN (SET) ([{c0:aa,c1:ba}]) WHEN 1 THEN a@0 >= ab AND a@0 <= ab AND b@1 >= bb AND b@1 <= bb AND struct(a@0, b@1) IN (SET) ([{c0:ab,c1:bb}]) ELSE false END ] + - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b, e], file_type=test, pushdown_supported=true, predicate=Optional(DynamicFilter [ (a@0 >= aa AND a@0 <= aa OR a@0 >= ab AND a@0 <= ab) AND (b@1 >= ba AND b@1 <= ba OR b@1 >= bb AND b@1 <= bb) ]) AND Optional(DynamicFilter [ CASE range_partition WHEN 0 THEN hash_lookup WHEN 1 THEN hash_lookup ELSE false END ]) " ); @@ -1490,7 +1495,7 @@ async fn test_hashjoin_dynamic_filter_pushdown_collect_left() { - HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(a@0, a@0), (b@1, b@1)] - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b, c], file_type=test, pushdown_supported=true - RepartitionExec: partitioning=Hash([a@0, b@1], 12), input_partitions=1 - - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b, e], file_type=test, pushdown_supported=true, predicate=DynamicFilter [ empty ] + - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b, e], file_type=test, pushdown_supported=true, predicate=Optional(DynamicFilter [ empty ]) " ); @@ -1509,7 +1514,7 @@ async fn test_hashjoin_dynamic_filter_pushdown_collect_left() { - HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(a@0, a@0), (b@1, b@1)] - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b, c], file_type=test, pushdown_supported=true - RepartitionExec: partitioning=Hash([a@0, b@1], 12), input_partitions=1 - - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b, e], file_type=test, pushdown_supported=true, predicate=DynamicFilter [ a@0 >= aa AND a@0 <= ab AND b@1 >= ba AND b@1 <= bb AND struct(a@0, b@1) IN (SET) ([{c0:aa,c1:ba}, {c0:ab,c1:bb}]) ] + - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b, e], file_type=test, pushdown_supported=true, predicate=Optional(DynamicFilter [ a@0 >= aa AND a@0 <= ab AND b@1 >= ba AND b@1 <= bb AND struct(a@0, b@1) IN (SET) ([{c0:aa,c1:ba}, {c0:ab,c1:bb}]) ]) " ); @@ -2820,7 +2825,7 @@ async fn test_hashjoin_dynamic_filter_transferred_through_nested_join() { - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[t], file_type=test, pushdown_supported=true - HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(m@0, x@0)] - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[m, c], file_type=test, pushdown_supported=false - - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[x, e], file_type=test, pushdown_supported=true, predicate=DynamicFilter [ empty ] AND DynamicFilter [ empty ] + - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[x, e], file_type=test, pushdown_supported=true, predicate=Optional(DynamicFilter [ empty ]) AND Optional(DynamicFilter [ empty ]) " ); @@ -2839,7 +2844,7 @@ async fn test_hashjoin_dynamic_filter_transferred_through_nested_join() { - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[t], file_type=test, pushdown_supported=true - HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(m@0, x@0)] - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[m, c], file_type=test, pushdown_supported=false - - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[x, e], file_type=test, pushdown_supported=true, predicate=DynamicFilter [ x@0 >= aa AND x@0 <= ad AND x@0 IN (SET) ([aa, ab, ac, ad]) ] AND DynamicFilter [ x@0 >= aa AND x@0 <= ab AND x@0 IN (SET) ([aa, ab]) ] + - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[x, e], file_type=test, pushdown_supported=true, predicate=Optional(DynamicFilter [ x@0 >= aa AND x@0 <= ad AND x@0 IN (SET) ([aa, ab, ac, ad]) ]) AND Optional(DynamicFilter [ x@0 >= aa AND x@0 <= ab AND x@0 IN (SET) ([aa, ab]) ]) " ); @@ -3829,7 +3834,7 @@ async fn test_hashjoin_dynamic_filter_all_partitions_empty() { - RepartitionExec: partitioning=Hash([a@0, b@1], 4), input_partitions=1 - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b], file_type=test, pushdown_supported=true - RepartitionExec: partitioning=Hash([a@0, b@1], 4), input_partitions=1 - - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b], file_type=test, pushdown_supported=true, predicate=DynamicFilter [ empty ] + - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b], file_type=test, pushdown_supported=true, predicate=Optional(DynamicFilter [ empty ]) AND Optional(DynamicFilter [ empty ]) " ); @@ -3854,7 +3859,7 @@ async fn test_hashjoin_dynamic_filter_all_partitions_empty() { - RepartitionExec: partitioning=Hash([a@0, b@1], 4), input_partitions=1 - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b], file_type=test, pushdown_supported=true - RepartitionExec: partitioning=Hash([a@0, b@1], 4), input_partitions=1 - - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b], file_type=test, pushdown_supported=true, predicate=DynamicFilter [ false ] + - DataSourceExec: file_groups={1 group: [[test.parquet]]}, projection=[a, b], file_type=test, pushdown_supported=true, predicate=Optional(DynamicFilter [ false ]) AND Optional(DynamicFilter [ false ]) " ); } @@ -4620,7 +4625,7 @@ async fn test_filter_with_projection_pushdown() { /// Test that ExecutionPlan::apply_expressions() can discover dynamic filters across the plan tree. /// /// Not portable to sqllogictest: asserts by walking the plan tree with -/// `apply_expressions` + `downcast_ref::` and +/// `apply_expressions` + `as_dynamic_filter` and /// counting nodes. Neither API is observable from SQL. #[tokio::test] async fn test_discover_dynamic_filters_via_expressions_api() { @@ -4629,7 +4634,9 @@ async fn test_discover_dynamic_filters_via_expressions_api() { // Check expressions from this node using apply_expressions let _ = plan.apply_expressions(&mut |expr| { - if let Some(_df) = expr.downcast_ref::() { + // Producers push their dynamic filters as + // `Optional(DynamicFilter)`, so look through the wrapper. + if let Some(_df) = as_dynamic_filter(expr) { count += 1; } Ok(TreeNodeRecursion::Continue) diff --git a/datafusion/datasource-parquet/src/source.rs b/datafusion/datasource-parquet/src/source.rs index 42e7116dff73c..90b542139a70e 100644 --- a/datafusion/datasource-parquet/src/source.rs +++ b/datafusion/datasource-parquet/src/source.rs @@ -50,7 +50,9 @@ use datafusion_datasource::file_scan_config::FileScanConfig; use datafusion_functions::core::file_row_index::FileRowIndexFunc; use datafusion_physical_expr::expressions::{Column, DynamicFilterTracking}; use datafusion_physical_expr::projection::ProjectionExprs; -use datafusion_physical_expr::utils::split_conjunction; +use datafusion_physical_expr::utils::{ + debug_assert_optional_on_root_chain, split_conjunction, +}; use datafusion_physical_expr::{EquivalenceProperties, conjunction}; use datafusion_physical_expr_adapter::DefaultPhysicalExprAdapterFactory; use datafusion_physical_expr_adapter::rewrite::{ @@ -911,6 +913,10 @@ impl FileSource for ParquetSource { } None => conjunction(allowed_filters), }; + // Optional filters (for example the dynamic filters of hash joins, + // TopK and aggregates) must stay direct conjuncts of the root AND + // chain, so that a later stage can find them with `split_optional`. + debug_assert_optional_on_root_chain(&predicate); source.predicate = Some(predicate); source = source.with_pushdown_filters(pushdown_filters); let source = Arc::new(source); diff --git a/datafusion/physical-expr/src/expressions/dynamic_filters/mod.rs b/datafusion/physical-expr/src/expressions/dynamic_filters/mod.rs index d53522f5de7c8..6d2fec993c7a9 100644 --- a/datafusion/physical-expr/src/expressions/dynamic_filters/mod.rs +++ b/datafusion/physical-expr/src/expressions/dynamic_filters/mod.rs @@ -21,6 +21,7 @@ use std::{fmt::Display, hash::Hash, sync::Arc}; use tokio::sync::watch; use crate::PhysicalExpr; +use crate::filter_stats::RemovedRowWork; use arrow::datatypes::{DataType, Schema}; #[cfg(feature = "proto")] use datafusion_common::internal_datafusion_err; @@ -94,6 +95,10 @@ pub struct DynamicFilterPhysicalExpr { /// But this can have overhead in production, so it's only included in our tests. data_type: Arc>>, nullable: Arc>>, + /// The work that the producer of the filter does for each row that the + /// filter removes, see [`Self::removed_row_work`]. Shared by all + /// derived filters. + removed_row_work: Arc, } impl std::fmt::Debug for DynamicFilterPhysicalExpr { @@ -223,9 +228,20 @@ impl DynamicFilterPhysicalExpr { state_watch, data_type: Arc::new(RwLock::new(None)), nullable: Arc::new(RwLock::new(None)), + removed_row_work: Arc::default(), } } + /// The work that the producer of this filter does for each row that + /// the filter removes before it (for example the probe of a hash join), + /// as the producer measures it. A consumer that decides if the filter + /// is worth its cost uses it as the saving of a removed row (see + /// [`OptionalFilterGate`](crate::optional_filter_gate::OptionalFilterGate)). + /// The producer records its work in it. + pub fn removed_row_work(&self) -> &Arc { + &self.removed_row_work + } + fn remap_children( children: &[Arc], remapped_children: Option<&Vec>>, @@ -490,6 +506,7 @@ impl DynamicFilterPhysicalExpr { state_watch, data_type: Arc::new(RwLock::new(None)), nullable: Arc::new(RwLock::new(None)), + removed_row_work: Arc::default(), } } } @@ -518,6 +535,7 @@ impl PhysicalExpr for DynamicFilterPhysicalExpr { state_watch: self.state_watch.clone(), data_type: Arc::clone(&self.data_type), nullable: Arc::clone(&self.nullable), + removed_row_work: Arc::clone(&self.removed_row_work), })) } @@ -614,6 +632,7 @@ impl PhysicalExpr for DynamicFilterPhysicalExpr { state_watch: _, // Runtime channel, recreated from inner state by from_parts(). data_type: _, // Cached test invariant, recomputed from the expression. nullable: _, // Cached test invariant, recomputed from the expression. + removed_row_work: _, // Runtime measurement of the producer. } = self; let children = children @@ -837,6 +856,7 @@ impl DynamicFilterPhysicalExpr { state_watch: self.state_watch.clone(), data_type: Arc::clone(&self.data_type), nullable: Arc::clone(&self.nullable), + removed_row_work: Arc::clone(&self.removed_row_work), } } } diff --git a/datafusion/physical-expr/src/expressions/dynamic_filters/tracker.rs b/datafusion/physical-expr/src/expressions/dynamic_filters/tracker.rs index fd4c18b07e2cd..6f0a830d07a85 100644 --- a/datafusion/physical-expr/src/expressions/dynamic_filters/tracker.rs +++ b/datafusion/physical-expr/src/expressions/dynamic_filters/tracker.rs @@ -157,7 +157,7 @@ impl DynamicFilterTracker { } /// `true` once every watched filter has completed and been dropped. - fn is_exhausted(&self) -> bool { + pub(crate) fn is_exhausted(&self) -> bool { self.subscriptions.is_empty() } } diff --git a/datafusion/physical-expr/src/expressions/mod.rs b/datafusion/physical-expr/src/expressions/mod.rs index f2f9285de560a..d36f734afef39 100644 --- a/datafusion/physical-expr/src/expressions/mod.rs +++ b/datafusion/physical-expr/src/expressions/mod.rs @@ -33,6 +33,7 @@ mod literal; mod negative; mod no_op; mod not; +mod optional_filter; mod similar_to_pattern; mod try_cast; mod unknown_column; @@ -59,6 +60,7 @@ pub use literal::{Literal, lit}; pub use negative::{NegativeExpr, negative}; pub use no_op::NoOp; pub use not::{NotExpr, not}; +pub use optional_filter::OptionalFilterPhysicalExpr; pub(crate) use similar_to_pattern::translate_scalar; pub use similar_to_pattern::{SqlSimilarToPattern, sql_similar_to_regex}; pub use try_cast::{TryCastExpr, try_cast, try_cast_with_target_field}; diff --git a/datafusion/physical-expr/src/expressions/optional_filter.rs b/datafusion/physical-expr/src/expressions/optional_filter.rs new file mode 100644 index 0000000000000..5561f90e89345 --- /dev/null +++ b/datafusion/physical-expr/src/expressions/optional_filter.rs @@ -0,0 +1,357 @@ +// Licensed to the Apache Software Foundation (ASF) under one +// or more contributor license agreements. See the NOTICE file +// distributed with this work for additional information +// regarding copyright ownership. The ASF licenses this file +// to you under the Apache License, Version 2.0 (the +// "License"); you may not use this file except in compliance +// with the License. You may obtain a copy of the License at +// +// http://www.apache.org/licenses/LICENSE-2.0 +// +// Unless required by applicable law or agreed to in writing, +// software distributed under the License is distributed on an +// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY +// KIND, either express or implied. See the License for the +// specific language governing permissions and limitations +// under the License. + +//! [`OptionalFilterPhysicalExpr`]: a marker for filters that are not needed +//! for correctness. See the type documentation for the contract that +//! producers, consumers and rewriters must obey. + +use std::fmt; +use std::hash::Hash; +use std::sync::Arc; + +use crate::PhysicalExpr; + +use arrow::array::BooleanArray; +use arrow::datatypes::{DataType, FieldRef, Schema}; +use arrow::record_batch::RecordBatch; +use datafusion_common::{Result, assert_eq_or_internal_err}; +use datafusion_expr::ColumnarValue; +use datafusion_expr::interval_arithmetic::Interval; +use datafusion_expr::sort_properties::ExprProperties; + +/// Marks the inner filter as *optional*: it is not needed for correctness. +/// +/// Some filters are only performance hints. For example, a hash join can push +/// a dynamic filter into the probe side scan, but the join itself still +/// removes the rows that do not match. Such a filter is *optional*: a +/// consumer can skip it (for example, when the filter does not remove enough +/// rows to be worth its cost) and the query result stays the same. +/// +/// # Contract +/// +/// * **Skip only on the root AND chain.** A consumer can skip an optional +/// filter only when the `Optional` node is a direct conjunct of the root +/// `AND` chain of its predicate. For example, in `a AND Optional(b)` the +/// consumer can skip `b`. Use [`split_optional`] to find these conjuncts. +/// * **Transparent everywhere else.** [`PhysicalExpr::evaluate`] always +/// evaluates the inner expression. Thus an `Optional` in a different +/// position (for example under `NOT`, `IS NULL`, `CASE` or `OR`) can make a +/// query slower, but it cannot make the result incorrect. +/// * **Rewriters must not move nodes across the wrapper.** A rewrite must not +/// move an expression into or out of an `Optional`. For example, +/// `NOT(Optional(x))` must not become `Optional(NOT(x))`, because that +/// would make a required filter optional. +/// * **Pruning sees through the wrapper.** [`PhysicalExpr::snapshot`] +/// returns the inner expression, so [`snapshot_physical_expr`] removes the +/// wrapper. Thus statistics pruning uses an optional filter the same as a +/// required filter. +/// +/// [`split_optional`]: crate::utils::split_optional +/// [`snapshot_physical_expr`]: datafusion_physical_expr_common::physical_expr::snapshot_physical_expr +#[derive(Debug, Eq)] +pub struct OptionalFilterPhysicalExpr { + inner: Arc, +} + +// Manually derive PartialEq and Hash to work around https://github.com/rust-lang/rust/issues/78808 +impl PartialEq for OptionalFilterPhysicalExpr { + fn eq(&self, other: &Self) -> bool { + self.inner.eq(&other.inner) + } +} + +impl Hash for OptionalFilterPhysicalExpr { + fn hash(&self, state: &mut H) { + self.inner.hash(state); + } +} + +impl OptionalFilterPhysicalExpr { + /// Create a new optional filter that wraps `inner`. + pub fn new(inner: Arc) -> Self { + Self { inner } + } + + /// Get the wrapped filter expression. + pub fn inner(&self) -> &Arc { + &self.inner + } +} + +impl fmt::Display for OptionalFilterPhysicalExpr { + fn fmt(&self, f: &mut fmt::Formatter) -> fmt::Result { + write!(f, "Optional({})", self.inner) + } +} + +impl PhysicalExpr for OptionalFilterPhysicalExpr { + fn data_type(&self, input_schema: &Schema) -> Result { + self.inner.data_type(input_schema) + } + + fn nullable(&self, input_schema: &Schema) -> Result { + self.inner.nullable(input_schema) + } + + fn evaluate(&self, batch: &RecordBatch) -> Result { + self.inner.evaluate(batch) + } + + fn return_field(&self, input_schema: &Schema) -> Result { + self.inner.return_field(input_schema) + } + + fn evaluate_selection( + &self, + batch: &RecordBatch, + selection: &BooleanArray, + ) -> Result { + self.inner.evaluate_selection(batch, selection) + } + + fn children(&self) -> Vec<&Arc> { + vec![&self.inner] + } + + fn with_new_children( + self: Arc, + children: Vec>, + ) -> Result> { + assert_eq_or_internal_err!( + children.len(), + 1, + "OptionalFilterPhysicalExpr: expected 1 child" + ); + Ok(Arc::new(Self::new(Arc::clone(&children[0])))) + } + + // The wrapper is the identity function, so the bounds and properties of + // the child are also the bounds and properties of the wrapper. + fn evaluate_bounds(&self, children: &[&Interval]) -> Result { + Ok(children[0].clone()) + } + + fn propagate_constraints( + &self, + interval: &Interval, + children: &[&Interval], + ) -> Result>> { + Ok(children[0].intersect(interval)?.map(|result| vec![result])) + } + + fn get_properties(&self, children: &[ExprProperties]) -> Result { + Ok(children[0].clone()) + } + + fn fmt_sql(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { + self.inner.fmt_sql(f) + } + + /// Returns the inner expression, so that snapshot consumers (for example + /// pruning) see through the wrapper. + /// + /// [`snapshot_physical_expr`] transforms the tree bottom up, so the inner + /// expression is already a snapshot when this method is called. For + /// example, `Optional(DynamicFilter)` becomes the current expression of + /// the dynamic filter. + /// + /// [`snapshot_physical_expr`]: datafusion_physical_expr_common::physical_expr::snapshot_physical_expr + fn snapshot(&self) -> Result>> { + Ok(Some(Arc::clone(&self.inner))) + } + + fn snapshot_generation(&self) -> u64 { + // The wrapper is not dynamic. `snapshot_generation(expr)` walks the + // tree and adds the generation of the inner expression. + 0 + } + + #[cfg(feature = "proto")] + fn try_to_proto( + &self, + ctx: &datafusion_physical_expr_common::physical_expr::proto_encode::PhysicalExprEncodeCtx<'_>, + ) -> Result> { + use datafusion_proto_models::protobuf; + + Ok(Some(protobuf::PhysicalExprNode { + expr_id: None, + expr_type: Some(protobuf::physical_expr_node::ExprType::OptionalFilter( + Box::new(protobuf::PhysicalOptionalFilterNode { + inner: Some(Box::new(ctx.encode_child(&self.inner)?)), + }), + )), + })) + } +} + +#[cfg(feature = "proto")] +impl OptionalFilterPhysicalExpr { + /// Reconstruct an [`OptionalFilterPhysicalExpr`] from its protobuf + /// representation. + pub fn try_from_proto( + node: &datafusion_proto_models::protobuf::PhysicalExprNode, + ctx: &datafusion_physical_expr_common::physical_expr::proto_decode::PhysicalExprDecodeCtx<'_>, + ) -> Result> { + use datafusion_physical_expr_common::expect_expr_variant; + use datafusion_proto_models::protobuf; + + let optional = expect_expr_variant!( + node, + protobuf::physical_expr_node::ExprType::OptionalFilter, + "OptionalFilter", + ); + let inner = ctx.decode_required_expression( + optional.inner.as_deref(), + "OptionalFilter", + "inner", + )?; + + Ok(Arc::new(Self::new(inner))) + } +} + +#[cfg(test)] +mod tests { + use super::*; + use crate::expressions::{BinaryExpr, DynamicFilterPhysicalExpr, col, lit, not}; + + use arrow::array::{ArrayRef, Int32Array}; + use arrow::datatypes::Field; + use datafusion_common::cast::as_boolean_array; + use datafusion_expr::Operator; + use datafusion_physical_expr_common::physical_expr::{ + fmt_sql, snapshot_generation, snapshot_physical_expr, + }; + + fn schema() -> Arc { + Arc::new(Schema::new(vec![Field::new("a", DataType::Int32, true)])) + } + + /// `a > 2` + fn a_gt_2(schema: &Schema) -> Arc { + Arc::new(BinaryExpr::new( + col("a", schema).unwrap(), + Operator::Gt, + lit(2i32), + )) + } + + fn optional(inner: Arc) -> Arc { + Arc::new(OptionalFilterPhysicalExpr::new(inner)) + } + + #[test] + fn evaluate_equals_inner() -> Result<()> { + let schema = schema(); + let values: ArrayRef = Arc::new(Int32Array::from(vec![Some(1), None, Some(3)])); + let batch = RecordBatch::try_new(Arc::clone(&schema), vec![values])?; + + let inner = a_gt_2(&schema); + let wrapped = optional(Arc::clone(&inner)); + + let expected = inner.evaluate(&batch)?.into_array(batch.num_rows())?; + let actual = wrapped.evaluate(&batch)?.into_array(batch.num_rows())?; + assert_eq!(as_boolean_array(&actual)?, as_boolean_array(&expected)?); + + let selection = BooleanArray::from(vec![true, false, true]); + let expected = inner + .evaluate_selection(&batch, &selection)? + .into_array(batch.num_rows())?; + let actual = wrapped + .evaluate_selection(&batch, &selection)? + .into_array(batch.num_rows())?; + assert_eq!(as_boolean_array(&actual)?, as_boolean_array(&expected)?); + + assert_eq!(wrapped.data_type(&schema)?, DataType::Boolean); + assert!(wrapped.nullable(&schema)?); + Ok(()) + } + + #[test] + fn display_and_fmt_sql() { + let schema = schema(); + let wrapped = optional(a_gt_2(&schema)); + assert_eq!(wrapped.to_string(), "Optional(a@0 > 2)"); + assert_eq!(fmt_sql(wrapped.as_ref()).to_string(), "a > 2"); + + let negated = not(Arc::clone(&wrapped)).unwrap(); + assert_eq!(negated.to_string(), "NOT Optional(a@0 > 2)"); + } + + #[test] + fn children_and_with_new_children() -> Result<()> { + let schema = schema(); + let wrapped = optional(a_gt_2(&schema)); + assert_eq!(wrapped.children().len(), 1); + + let new_inner = lit(true); + let rewrapped = Arc::clone(&wrapped).with_new_children(vec![new_inner])?; + let rewrapped = rewrapped + .downcast_ref::() + .expect("wrapper is kept"); + assert_eq!(rewrapped.inner().to_string(), "true"); + + assert!(wrapped.with_new_children(vec![]).is_err()); + Ok(()) + } + + #[test] + fn eq_and_hash_use_inner() { + use std::collections::HashSet; + + let schema = schema(); + let a = optional(a_gt_2(&schema)); + let b = optional(a_gt_2(&schema)); + let c = optional(lit(true)); + assert_eq!(&a, &b); + assert_ne!(&a, &c); + // The wrapper is not equal to the inner expression. + assert_ne!(&a, &a_gt_2(&schema)); + + let set: HashSet<_> = [a, b, c].into_iter().collect(); + assert_eq!(set.len(), 2); + } + + #[test] + fn snapshot_sees_through_dynamic_filter() -> Result<()> { + let schema = schema(); + let dynamic = Arc::new(DynamicFilterPhysicalExpr::new( + vec![col("a", &schema)?], + lit(true), + )); + let wrapped = optional(Arc::clone(&dynamic) as Arc); + + assert_eq!( + snapshot_physical_expr(Arc::clone(&wrapped))?.to_string(), + "true" + ); + + let generation = snapshot_generation(&wrapped); + dynamic.update(a_gt_2(&schema))?; + assert_ne!(snapshot_generation(&wrapped), generation); + assert_eq!( + snapshot_physical_expr(Arc::clone(&wrapped))?.to_string(), + "a@0 > 2" + ); + + // A static inner expression is also unwrapped. + let wrapped = optional(a_gt_2(&schema)); + assert_eq!(wrapped.snapshot_generation(), 0); + assert_eq!(snapshot_physical_expr(wrapped)?.to_string(), "a@0 > 2"); + Ok(()) + } +} diff --git a/datafusion/physical-expr/src/filter_stats.rs b/datafusion/physical-expr/src/filter_stats.rs new file mode 100644 index 0000000000000..a837e3bfd30f6 --- /dev/null +++ b/datafusion/physical-expr/src/filter_stats.rs @@ -0,0 +1,265 @@ +// Licensed to the Apache Software Foundation (ASF) under one +// or more contributor license agreements. See the NOTICE file +// distributed with this work for additional information +// regarding copyright ownership. The ASF licenses this file +// to you under the Apache License, Version 2.0 (the +// "License"); you may not use this file except in compliance +// with the License. You may obtain a copy of the License at +// +// http://www.apache.org/licenses/LICENSE-2.0 +// +// Unless required by applicable law or agreed to in writing, +// software distributed under the License is distributed on an +// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY +// KIND, either express or implied. See the License for the +// specific language governing permissions and limitations +// under the License. + +//! Measurements of filter evaluations at runtime. +//! +//! Operators that adapt to the data measure the filters that they evaluate: +//! the rows in, the rows that pass and the evaluation time. This module has +//! the shared parts: +//! +//! * [`Clock`]: a monotonic clock that tests can replace, so that decisions +//! that use time are deterministic in tests. [`SystemClock`] is the real +//! clock and [`ManualClock`] is a clock that only moves when a test moves +//! it. +//! * [`FilterCost`]: the counts and the time of one filter, and the values +//! derived from them (cost for each row, rows removed for each +//! nanosecond). +//! +//! For example, an operator can use them to pause a filter that costs more +//! than it saves, or to change the order of the conjuncts of a predicate. + +use std::fmt::Debug; +use std::sync::Arc; +use std::sync::atomic::{AtomicU64, Ordering}; +use std::time::Duration; + +use datafusion_common::instant::Instant; + +/// A monotonic clock in nanoseconds. +/// +/// Production code uses [`SystemClock`]. Tests use [`ManualClock`] (or their +/// own implementation), thus decisions that use time are deterministic in +/// tests. +pub trait Clock: Debug + Send + Sync { + /// Nanoseconds since an arbitrary fixed point. The value never + /// decreases. + fn now_nanos(&self) -> u64; +} + +/// The real monotonic [`Clock`]. +#[derive(Debug, Clone, Copy)] +pub struct SystemClock { + start: Instant, +} + +impl SystemClock { + /// Creates a clock whose zero is now. + pub fn new() -> Self { + Self { + start: Instant::now(), + } + } + + /// A shared [`SystemClock`], as a trait object. + pub fn shared() -> Arc { + Arc::new(Self::new()) + } +} + +impl Default for SystemClock { + fn default() -> Self { + Self::new() + } +} + +impl Clock for SystemClock { + fn now_nanos(&self) -> u64 { + u64::try_from(self.start.elapsed().as_nanos()).unwrap_or(u64::MAX) + } +} + +/// A [`Clock`] that moves only when [`Self::advance`] is called. For tests. +#[derive(Debug, Default)] +pub struct ManualClock { + nanos: AtomicU64, +} + +impl ManualClock { + /// Creates a clock at zero. + pub fn new() -> Self { + Self::default() + } + + /// Moves the clock forward by `nanos` nanoseconds. + pub fn advance(&self, nanos: u64) { + self.nanos.fetch_add(nanos, Ordering::Relaxed); + } +} + +impl Clock for ManualClock { + fn now_nanos(&self) -> u64 { + self.nanos.load(Ordering::Relaxed) + } +} + +/// Minimum number of evaluated rows before an adaptive decision uses the +/// measurements of a filter. It is one batch of the default +/// `datafusion.execution.batch_size`. With fewer rows, the evaluation time +/// is dominated by the fixed cost of each call (for example 2 to 7 rows of +/// a batch after a selective row filter took 600 to 8000 ns for each row in +/// ClickBench Q23, against 0.4 ns for each row on full batches), and the +/// fraction of removed rows is not reliable. +pub const MIN_OBSERVED_ROWS: u64 = 8192; + +/// Returns the nanoseconds in `elapsed`, saturated to `u64::MAX`. +pub fn duration_nanos(elapsed: Duration) -> u64 { + u64::try_from(elapsed.as_nanos()).unwrap_or(u64::MAX) +} + +/// The measurements of one filter: the rows that it was evaluated on, the +/// rows that passed it, and the evaluation time. +#[derive(Debug, Default, Clone, Copy, PartialEq, Eq)] +pub struct FilterCost { + /// Rows that the filter was evaluated on. + pub rows_in: u64, + /// Rows that passed the filter (`true`; `null` does not pass). + pub rows_out: u64, + /// Evaluation time in nanoseconds. + pub nanos: u64, +} + +impl FilterCost { + /// Adds the result of one evaluation. `rows_out` larger than `rows_in` + /// is used as `rows_in`. + pub fn add(&mut self, rows_in: u64, rows_out: u64, nanos: u64) { + self.rows_in = self.rows_in.saturating_add(rows_in); + self.rows_out = self.rows_out.saturating_add(rows_out.min(rows_in)); + self.nanos = self.nanos.saturating_add(nanos); + } + + /// Rows that the filter removed. + pub fn rows_removed(&self) -> u64 { + self.rows_in.saturating_sub(self.rows_out) + } + + /// Nanoseconds for each evaluated row, or `None` if the filter was not + /// evaluated on any row. + pub fn nanos_per_row(&self) -> Option { + (self.rows_in > 0).then(|| self.nanos as f64 / self.rows_in as f64) + } + + /// Rows removed for each nanosecond, `(1 + rows_in - rows_out) / nanos`, + /// or `None` if the filter was not evaluated on any row. A larger value + /// is a better filter to evaluate first. This is the ranking key of + /// Velox (Pedreira et al., VLDB 2022). The `1 +` ranks a filter that + /// removes no rows by its cost, and a zero time is used as 1 ns. + pub fn rows_removed_per_nano(&self) -> Option { + (self.rows_in > 0) + .then(|| (1 + self.rows_removed()) as f64 / self.nanos.max(1) as f64) + } +} + +/// The work, in nanoseconds for each row, that an operator does on the rows +/// that a filter removes before them, as measured by that operator. +/// +/// For example, a hash join computes the hashes of the join keys of each +/// probe row and looks them up in its hash table. A row that the dynamic +/// filter of the join removes in the scan does not get this work, thus this +/// is the saving of a removed row. The work that depends on a match (the +/// check of the candidates of the lookup and the output of a matched row) +/// is not in it: the filter removes only rows without a match. While the +/// filter is on, most rows that reach the producer are matches, thus work +/// that only matches get would make the saving too large (TPC-DS SF1 Q31: +/// 6 ns for each probe row with the check, 0.1 to 1.2 ns without it). The +/// producer of a dynamic filter measures it and the consumers of the filter +/// read it (see +/// [`DynamicFilterPhysicalExpr::removed_row_work`]). +/// +/// Lock-free. +/// +/// [`DynamicFilterPhysicalExpr::removed_row_work`]: crate::expressions::DynamicFilterPhysicalExpr::removed_row_work +#[derive(Debug, Default)] +pub struct RemovedRowWork { + rows: AtomicU64, + nanos: AtomicU64, +} + +impl RemovedRowWork { + /// Creates an empty measurement. + pub fn new() -> Self { + Self::default() + } + + /// Adds `rows` rows and `nanos` nanoseconds of work. The two can be + /// recorded separately (for example the rows when a batch arrives and + /// the time of each step). + pub fn record(&self, rows: u64, nanos: u64) { + self.rows.fetch_add(rows, Ordering::Relaxed); + self.nanos.fetch_add(nanos, Ordering::Relaxed); + } + + /// The work for each row, or `None` before [`MIN_OBSERVED_ROWS`] rows. + pub fn ns_per_row(&self) -> Option { + let rows = self.rows.load(Ordering::Relaxed); + (rows >= MIN_OBSERVED_ROWS) + .then(|| self.nanos.load(Ordering::Relaxed) as f64 / rows as f64) + } +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn removed_row_work_needs_min_observed_rows() { + let work = RemovedRowWork::new(); + work.record(MIN_OBSERVED_ROWS - 1, 0); + work.record(0, 3 * (MIN_OBSERVED_ROWS - 1)); + assert_eq!(work.ns_per_row(), None); + work.record(1, 3); + assert_eq!(work.ns_per_row(), Some(3.0)); + } + + #[test] + fn filter_cost_derived_values() { + let empty = FilterCost::default(); + assert_eq!(empty.nanos_per_row(), None); + assert_eq!(empty.rows_removed_per_nano(), None); + + let mut cost = FilterCost::default(); + cost.add(100, 25, 1_000); + cost.add(100, 200, 1_000); + assert_eq!(cost.rows_in, 200); + // `rows_out` is at most `rows_in` for each evaluation. + assert_eq!(cost.rows_out, 125); + assert_eq!(cost.rows_removed(), 75); + assert_eq!(cost.nanos_per_row(), Some(10.0)); + assert_eq!(cost.rows_removed_per_nano(), Some(76.0 / 2_000.0)); + + // A zero time is used as 1 ns. + let mut free = FilterCost::default(); + free.add(10, 0, 0); + assert_eq!(free.rows_removed_per_nano(), Some(11.0)); + } + + #[test] + fn manual_clock_moves_only_when_advanced() { + let clock = ManualClock::new(); + assert_eq!(clock.now_nanos(), 0); + clock.advance(5); + clock.advance(7); + assert_eq!(clock.now_nanos(), 12); + } + + #[test] + fn system_clock_is_monotonic() { + let clock = SystemClock::new(); + let first = clock.now_nanos(); + assert!(clock.now_nanos() >= first); + assert_eq!(duration_nanos(Duration::from_micros(3)), 3_000); + } +} diff --git a/datafusion/physical-expr/src/lib.rs b/datafusion/physical-expr/src/lib.rs index 80e9f88b510ed..fba82a96b0c46 100644 --- a/datafusion/physical-expr/src/lib.rs +++ b/datafusion/physical-expr/src/lib.rs @@ -34,8 +34,10 @@ pub mod binary_map { pub mod async_scalar_function; pub mod equivalence; pub mod expressions; +pub mod filter_stats; pub mod higher_order_function; pub mod intervals; +pub mod optional_filter_gate; mod partitioning; mod physical_expr; pub mod planner; diff --git a/datafusion/physical-expr/src/optional_filter_gate.rs b/datafusion/physical-expr/src/optional_filter_gate.rs new file mode 100644 index 0000000000000..15787f101a382 --- /dev/null +++ b/datafusion/physical-expr/src/optional_filter_gate.rs @@ -0,0 +1,1471 @@ +// Licensed to the Apache Software Foundation (ASF) under one +// or more contributor license agreements. See the NOTICE file +// distributed with this work for additional information +// regarding copyright ownership. The ASF licenses this file +// to you under the Apache License, Version 2.0 (the +// "License"); you may not use this file except in compliance +// with the License. You may obtain a copy of the License at +// +// http://www.apache.org/licenses/LICENSE-2.0 +// +// Unless required by applicable law or agreed to in writing, +// software distributed under the License is distributed on an +// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY +// KIND, either express or implied. See the License for the +// specific language governing permissions and limitations +// under the License. + +//! A runtime gate that pauses optional filters that cost more than they +//! save. +//! +//! An *optional filter* is a filter that is not needed for correctness, for +//! example a dynamic filter that a hash join or a TopK pushes down into a +//! scan. An operator can skip such a filter and still produce correct +//! results. When the filter removes few rows, or when it is expensive (for +//! example a hash table lookup with many columns), the cost to evaluate it +//! can be larger than the benefit. +//! +//! [`OptionalFilterGate`] decides, batch by batch, if a stream evaluates the +//! filter or skips it. Each stream has its own gate. The gates of one plan +//! site (for example all files and partitions of one scan) can share their +//! pauses, see "Shared verdict". +//! +//! # State machine +//! +//! ```text +//! keep (see "Decision") +//! (reset backoff, new window) +//! +-------+ +//! | | +//! v | +//! +-------------+-+ pause +---------------------+ +//! start ------>| Evaluate |-------------------------------->| Paused | +//! | (window of | (pause for `backoff` batches, | (skip `remaining` | +//! | sample_batches| then double `backoff`) | batches) | +//! | batches) |<--------------------------------| | +//! +---------------+ pause ends: probe with a +---------------------+ +//! fresh window +//! ``` +//! +//! * In `Evaluate`, the gate collects the rows in, the rows out and the +//! evaluation time of at least `sample_batches` batches and at least +//! [`MIN_OBSERVED_ROWS`] rows (a *window*). Then it +//! decides (see below). To pause, it goes to `Paused` for `backoff` batches +//! and doubles `backoff` (up to `max_pause_batches`). To keep the filter, +//! it stays in `Evaluate`, sets `backoff` to `initial_pause_batches` and +//! starts a new window. +//! * In `Paused`, the gate skips batches. It does not change counters or +//! `backoff` for skipped batches. When the pause ends, the gate evaluates +//! a new window (a *probe*). +//! * Before each batch the gate checks if the filter changed (for example a +//! dynamic filter got new bounds). If so, the gate goes to `Evaluate` with +//! an empty window and sets `backoff` to `initial_pause_batches`. +//! +//! # Decision +//! +//! At the end of each window, the gate pauses the filter if one of these +//! rules is true: +//! +//! 1. The filter removed no rows in the window. +//! 2. The cost of the window (`cost_ns`) is larger than the work that the +//! removed rows save (`saving_ns`): +//! +//! ```text +//! cost_ns = evaluation time + rows_in * measured overhead +//! saving_ns = (rows_in - rows_out) * saving_ns_per_row +//! saving_ns_per_row = producer work + measured saving +//! ``` +//! +//! The *producer work* is the work that a removed row saves after the +//! filter, in the operator that produced the filter, for example the +//! hash and the hash table lookup of a probe row in a hash join. The +//! producer measures it ([`RemovedRowWork`], see +//! [`DynamicFilterPhysicalExpr::removed_row_work`]): a hash join with a +//! small build side does 2 to 8 ns of work for each probe row (TPC-DS +//! SF1 star joins), a join with a large build side much more. Until the +//! producer has measured [`MIN_OBSERVED_ROWS`] rows, the gate uses +//! `min_saving_ns_per_row` from the configuration (a prior). With more +//! than one dynamic filter in the filter, the smallest measured work is +//! used. The *measured saving* and the *measured overhead* are +//! optional: a consumer that can measure more work that a removed row +//! saves (the Parquet scan measures the decode time of the columns that +//! the filter does not read), or a fixed cost for each evaluated row in +//! addition to the evaluation time (the Parquet scan has a cost for each +//! row filter stage), gives them in a shared [`MeasuredRowSaving`] and +//! updates them at any time. +//! +//! To prevent a filter from switching on and off when the cost and the +//! saving are almost equal, this rule has a margin: a running filter is +//! paused only if `cost_ns > saving_ns * 1.1`, and a probe after a pause +//! turns the filter on again only if `cost_ns < saving_ns * 0.9`. +//! +//! Thus a filter that removes most rows but is expensive is paused, and a +//! cheap filter stays on also when it removes only some of the rows. +//! +//! The gate measures time with a [`Clock`]. Tests use a [`ManualClock`], so +//! that the decisions are deterministic. +//! +//! [`ManualClock`]: crate::filter_stats::ManualClock +//! +//! # Change detection +//! +//! The gate walks the filter one time, when it is created, with +//! [`DynamicFilterTracking::classify`]. The walk subscribes to each +//! [`DynamicFilterPhysicalExpr`] in the filter that is not complete. Before +//! each batch, the gate polls these subscriptions with +//! [`DynamicFilterTracker::changed`]. When nothing changed, this is one +//! atomic load for each subscription. The tracker drops a subscription when +//! its filter is complete. A filter without dynamic filters, or with only +//! complete dynamic filters, is never polled and never resets the gate. +//! +//! [`DynamicFilterPhysicalExpr`]: crate::expressions::DynamicFilterPhysicalExpr +//! [`DynamicFilterTracker::changed`]: crate::expressions::DynamicFilterTracker::changed +//! +//! # Shared verdict +//! +//! Each gate pays for at least one window before its first decision. A scan +//! opens many files at the same time, and a filter that removes no rows +//! (for example a hash join filter on a column that has only matching +//! values) then costs one window in each file. Also each gate probes on its +//! own after each pause. +//! +//! Thus the gates of one plan site can share a [`SharedGateVerdict`] (see +//! [`OptionalFilterGate::with_shared_verdict`]). Each gate publishes its +//! pauses and the end of its pauses there. A gate that has no evidence of +//! its own that the filter is worth its cost (before its first decision, +//! after a change of the filter, and after a decision to pause) uses a +//! pause that another gate published after the last shared verdict that it +//! saw: +//! +//! * A new gate starts paused if the shared verdict is a pause. +//! * A gate in its first window, or in a probe window after a pause, stops +//! the window and pauses. Thus after a pause of all gates, usually only +//! the first gate at the end of its pause probes the filter. +//! +//! A gate that keeps the filter does not use the pauses of other gates: +//! with skewed data the filter can be worth its cost for some files only. +//! A gate that sees a change of the filter clears a shared pause, because +//! it was measured on the old filter, and starts again without evidence. + +use std::sync::Arc; +use std::sync::atomic::{AtomicU64, Ordering}; +use std::time::Duration; + +use datafusion_common::config::ExecutionOptions; +use datafusion_physical_expr_common::physical_expr::PhysicalExpr; + +use datafusion_common::tree_node::{TreeNode, TreeNodeRecursion}; + +use crate::expressions::{DynamicFilterPhysicalExpr, DynamicFilterTracking}; +use crate::filter_stats::{ + Clock, FilterCost, MIN_OBSERVED_ROWS, RemovedRowWork, SystemClock, duration_nanos, +}; + +/// A running filter is paused by the cost rule only if its cost is larger +/// than this multiple of its saving. See the [module documentation](self). +const PAUSE_COST_MARGIN: f64 = 1.1; + +/// A probe turns a paused filter on again (by the cost rule) only if its +/// cost is smaller than this multiple of its saving. +const RESUME_COST_MARGIN: f64 = 0.9; + +/// Configuration of an [`OptionalFilterGate`]. +#[derive(Debug, Clone, Copy, PartialEq)] +pub struct OptionalFilterGateConfig { + /// Number of evaluated batches in one window. The gate makes a decision + /// at the end of each window. Values smaller than 1 are used as 1. + pub sample_batches: usize, + /// Number of batches to skip at the first pause, and after the filter + /// was selective again. Values smaller than 1 are used as 1. + pub initial_pause_batches: usize, + /// Maximum number of batches to skip in one pause. Values smaller than + /// `initial_pause_batches` are used as `initial_pause_batches`. + pub max_pause_batches: usize, + /// Work, in nanoseconds, that each row removed by the filter saves after + /// the filter, until the producer of the filter has measured it (see + /// [`RemovedRowWork`]). The gate adds the saving that the consumer + /// measures (see [`MeasuredRowSaving`]). The gate pauses a filter whose + /// evaluation time is larger than the saving of the rows that it + /// removes. Negative values are used as 0. + pub min_saving_ns_per_row: f64, +} + +impl Default for OptionalFilterGateConfig { + fn default() -> Self { + Self { + sample_batches: 2, + initial_pause_batches: 4, + max_pause_batches: 32, + min_saving_ns_per_row: 20.0, + } + } +} + +impl From<&ExecutionOptions> for OptionalFilterGateConfig { + /// Uses `optional_filter_min_saving_ns_per_row` from `options`, and the + /// default values for the other fields. + fn from(options: &ExecutionOptions) -> Self { + Self { + min_saving_ns_per_row: options.optional_filter_min_saving_ns_per_row, + ..Default::default() + } + } +} + +impl OptionalFilterGateConfig { + /// Returns a copy with the documented minimum values applied. + fn normalized(self) -> Self { + let sample_batches = self.sample_batches.max(1); + let initial_pause_batches = self.initial_pause_batches.max(1); + let max_pause_batches = self.max_pause_batches.max(initial_pause_batches); + Self { + sample_batches, + initial_pause_batches, + max_pause_batches, + min_saving_ns_per_row: self.min_saving_ns_per_row.max(0.0), + } + } +} + +/// The decision of an [`OptionalFilterGate`] for one batch. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum GateDecision { + /// Evaluate the filter on this batch, then call + /// [`OptionalFilterGate::record`] with the row counts and the time. + Evaluate, + /// Do not evaluate the filter on this batch. Let all rows pass. + Skip, +} + +/// The terms of the cost rule that the consumer of an optional filter +/// measures: +/// +/// * The *saving*: work, in nanoseconds, that each row removed by the filter +/// saves. The gate adds it to +/// [`OptionalFilterGateConfig::min_saving_ns_per_row`]. For example, the +/// Parquet scan sets it to the time to decode the columns that the filter +/// does not read, for each removed row that the decoder can skip. +/// * The *overhead*: work, in nanoseconds, that the consumer does for each +/// evaluated row because it evaluates the filter, in addition to the +/// evaluation time. The gate adds it to the cost of each window. For +/// example, the Parquet scan sets it to the fixed cost of a row filter +/// stage when the filter is a row filter predicate. +/// +/// The consumer creates one value, gives a clone of the [`Arc`] to the gate +/// with [`OptionalFilterGate::with_measured_saving`], and updates it at any +/// time with [`Self::set_ns_per_row`] and [`Self::set_overhead_ns_per_row`]. +/// The gate reads it at each decision. +/// +/// Each value is an `f64` in an [`AtomicU64`], thus reads and updates are +/// cheap and lock-free. +#[derive(Debug, Default)] +pub struct MeasuredRowSaving { + /// The bits of the `f64` saving for each removed row. + ns_per_row_bits: AtomicU64, + /// The bits of the `f64` overhead for each evaluated row. + overhead_ns_per_row_bits: AtomicU64, +} + +/// `value` if it is finite and not negative, else 0. +fn non_negative(value: f64) -> f64 { + if value.is_finite() { + value.max(0.0) + } else { + 0.0 + } +} + +impl MeasuredRowSaving { + /// Creates a value of 0 ns. + pub fn new() -> Self { + Self::default() + } + + /// Sets the measured saving for each removed row, in nanoseconds. + /// Values that are negative or not finite are used as 0. + pub fn set_ns_per_row(&self, ns_per_row: f64) { + self.ns_per_row_bits + .store(non_negative(ns_per_row).to_bits(), Ordering::Relaxed); + } + + /// The measured saving for each removed row, in nanoseconds. + pub fn ns_per_row(&self) -> f64 { + f64::from_bits(self.ns_per_row_bits.load(Ordering::Relaxed)) + } + + /// Sets the measured overhead for each evaluated row, in nanoseconds. + /// Values that are negative or not finite are used as 0. + pub fn set_overhead_ns_per_row(&self, ns_per_row: f64) { + self.overhead_ns_per_row_bits + .store(non_negative(ns_per_row).to_bits(), Ordering::Relaxed); + } + + /// The measured overhead for each evaluated row, in nanoseconds. + pub fn overhead_ns_per_row(&self) -> f64 { + f64::from_bits(self.overhead_ns_per_row_bits.load(Ordering::Relaxed)) + } +} + +/// The last pause (or end of a pause) that the gates of one plan site +/// published, see "Shared verdict" in the [module documentation](self). +/// +/// Create one value for each plan site (for example each optional filter of +/// one scan), and give a clone of the [`Arc`] to each gate of the site with +/// [`OptionalFilterGate::with_shared_verdict`]. +/// +/// The value is one [`AtomicU64`], thus it is lock-free. Gates write it only +/// at decisions that pause the filter or end a pause, and read it before a +/// batch only while they have no evidence of their own. +#[derive(Debug, Default)] +pub struct SharedGateVerdict { + /// A [`Verdict`] and its sequence number, see [`Verdict::pack`]. + word: AtomicU64, +} + +/// One published verdict of a [`SharedGateVerdict`]. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +enum Verdict { + /// The filter is not paused (or nothing was published yet). + Keep, + /// A gate paused the filter for this number of batches. + Pause(usize), +} + +impl Verdict { + /// Largest pause length that fits in the packed word. + const MAX_BATCHES: usize = u32::MAX as usize; + + /// Packs the verdict with the sequence number `seq`: the sequence + /// number in the low 32 bits, and the pause length (0 for + /// [`Verdict::Keep`]) in the high 32 bits. + fn pack(self, seq: u32) -> u64 { + let batches = match self { + Self::Keep => 0, + Self::Pause(batches) => batches.clamp(1, Self::MAX_BATCHES) as u64, + }; + (batches << 32) | u64::from(seq) + } + + /// The verdict and the sequence number in `word`. + fn unpack(word: u64) -> (Self, u32) { + let verdict = match (word >> 32) as usize { + 0 => Self::Keep, + batches => Self::Pause(batches), + }; + (verdict, word as u32) + } +} + +impl SharedGateVerdict { + /// Creates a shared verdict without any published pause. + pub fn new() -> Self { + Self::default() + } + + /// True if the last published verdict is a pause. + pub fn is_paused(&self) -> bool { + matches!(self.load().0, Verdict::Pause(_)) + } + + /// The current verdict and its sequence number. + fn load(&self) -> (Verdict, u32) { + Verdict::unpack(self.word.load(Ordering::Acquire)) + } + + /// Publishes `verdict` if `replace` returns true for the current + /// verdict. Returns the sequence number of the published verdict. + fn publish_if( + &self, + verdict: Verdict, + replace: impl Fn(Verdict) -> bool, + ) -> Option { + self.word + .fetch_update(Ordering::AcqRel, Ordering::Acquire, |word| { + let (current, seq) = Verdict::unpack(word); + replace(current).then(|| verdict.pack(seq.wrapping_add(1))) + }) + .ok() + .map(|previous| (previous as u32).wrapping_add(1)) + } +} + +/// The state of an [`OptionalFilterGate`]. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +enum GateState { + /// Evaluate the filter and collect counts and time for the current + /// window. + Evaluate { + window: FilterCost, + batches_in_window: usize, + }, + /// Skip the filter for `remaining_batches` more batches. + Paused { remaining_batches: usize }, +} + +impl GateState { + const fn new_window() -> Self { + Self::Evaluate { + window: FilterCost { + rows_in: 0, + rows_out: 0, + nanos: 0, + }, + batches_in_window: 0, + } + } +} + +/// Decides, batch by batch, if one stream evaluates an optional filter. +/// +/// See the [module documentation](self) for the state machine. A gate is +/// for one stream only. Do not share it between streams. +/// +/// Call [`Self::begin_batch`] before each batch. If it returns +/// [`GateDecision::Evaluate`], evaluate [`Self::filter`] on the batch and +/// then call [`Self::record`] with the row counts and the evaluation time. +/// Measure the time with [`Self::clock`], so that tests can replace it. +#[derive(Debug)] +pub struct OptionalFilterGate { + filter: Arc, + /// The dynamic filters in `filter` that can still change. + tracking: DynamicFilterTracking, + config: OptionalFilterGateConfig, + /// The clock that consumers use to measure the evaluation time. + clock: Arc, + /// The saving that the consumer measures, added to the producer work. + measured_saving: Option>, + /// The work that the producers of the dynamic filters in `filter` do + /// for each removed row, see [`RemovedRowWork`]. + producer_work: Vec>, + state: GateState, + /// True while the current window is a probe after a pause. The cost + /// rule then uses [`RESUME_COST_MARGIN`]. + probing: bool, + /// Length of the next pause, in batches. + backoff: usize, + /// True after `begin_batch` returned `Evaluate` and before `record`. + awaiting_record: bool, + pauses: usize, + /// The verdict that this gate shares with the other gates of its plan + /// site, if any. + shared: Option>, + /// Sequence number of the last shared verdict that this gate published + /// or saw. + shared_seq: u32, + /// True while this gate has no evidence of its own that the filter is + /// worth its cost: before its first decision, after a change of the + /// filter, and after a decision that paused the filter. Then it uses + /// the shared pauses of the other gates. + uses_shared_pauses: bool, +} + +impl OptionalFilterGate { + /// Creates a gate for `filter`, a boolean expression. The gate starts to + /// evaluate the filter. + /// + /// The gate uses a [`SystemClock`] and no measured saving. See + /// [`Self::with_clock`] and [`Self::with_measured_saving`]. + /// + /// This walks `filter` one time to find its dynamic filters. + pub fn new(filter: Arc, config: OptionalFilterGateConfig) -> Self { + let config = config.normalized(); + let tracking = DynamicFilterTracking::classify(&filter); + let mut producer_work = vec![]; + filter + .apply(|expr| { + if let Some(dynamic) = expr.downcast_ref::() { + producer_work.push(Arc::clone(dynamic.removed_row_work())); + } + Ok(TreeNodeRecursion::Continue) + }) + .expect("the closure is infallible"); + Self { + filter, + tracking, + config, + clock: SystemClock::shared(), + measured_saving: None, + producer_work, + state: GateState::new_window(), + probing: false, + backoff: config.initial_pause_batches, + awaiting_record: false, + pauses: 0, + shared: None, + shared_seq: 0, + uses_shared_pauses: true, + } + } + + /// Shares the pauses of this gate with the other gates of the same plan + /// site, see "Shared verdict" in the [module documentation](self). If + /// the shared verdict is a pause, the gate starts paused. + pub fn with_shared_verdict(mut self, shared: Arc) -> Self { + let (verdict, seq) = shared.load(); + self.shared_seq = seq; + if let Verdict::Pause(batches) = verdict { + self.start_pause(batches); + } + self.shared = Some(shared); + self + } + + /// Uses `clock` as the clock of this gate, see [`Self::clock`]. + pub fn with_clock(mut self, clock: Arc) -> Self { + self.clock = clock; + self + } + + /// Adds the saving of `saving` to + /// [`OptionalFilterGateConfig::min_saving_ns_per_row`] and its overhead + /// to the cost at each decision. See [`MeasuredRowSaving`]. + pub fn with_measured_saving(mut self, saving: Arc) -> Self { + self.measured_saving = Some(saving); + self + } + + /// The filter of this gate. + pub fn filter(&self) -> &Arc { + &self.filter + } + + /// The clock to measure the evaluation time that is given to + /// [`Self::record`]. + pub fn clock(&self) -> &Arc { + &self.clock + } + + /// The work, in nanoseconds, that the gate assumes each removed row + /// saves now: the work of the producer (measured, or the configured + /// `min_saving_ns_per_row` before the measurement) plus the measured + /// saving of the consumer. + pub fn saving_ns_per_row(&self) -> f64 { + let measured = self + .measured_saving + .as_ref() + .map_or(0.0, |saving| saving.ns_per_row()); + self.producer_work_ns_per_row() + measured + } + + /// The work of the producer for each removed row: the smallest measured + /// [`RemovedRowWork`] of the dynamic filters in the filter, or + /// `min_saving_ns_per_row` if none is measured yet. + fn producer_work_ns_per_row(&self) -> f64 { + self.producer_work + .iter() + .filter_map(|work| work.ns_per_row()) + .reduce(f64::min) + .unwrap_or(self.config.min_saving_ns_per_row) + } + + /// The work, in nanoseconds for each evaluated row, that the gate adds + /// to the evaluation time now: the measured overhead. + pub fn overhead_ns_per_row(&self) -> f64 { + self.measured_saving + .as_ref() + .map_or(0.0, |saving| saving.overhead_ns_per_row()) + } + + /// Call before each batch. Returns if the caller must evaluate the + /// filter on the batch or skip it. + /// + /// If the result is [`GateDecision::Evaluate`], call [`Self::record`] + /// after the evaluation. + pub fn begin_batch(&mut self) -> GateDecision { + // Only a filter with dynamic filters that are not complete can + // change. When nothing changed, this is one atomic load for each + // such dynamic filter. + if let Some(tracker) = self.tracking.watcher() + && tracker.changed() + { + self.state = GateState::new_window(); + self.probing = false; + self.backoff = self.config.initial_pause_batches; + self.uses_shared_pauses = true; + if let Some(shared) = &self.shared { + // A shared pause was measured on the old filter. + let paused = |verdict| matches!(verdict, Verdict::Pause(_)); + if let Some(seq) = shared.publish_if(Verdict::Keep, paused) { + self.shared_seq = seq; + } + } + } + if !self.is_paused() { + self.use_shared_pause(); + } + + match &mut self.state { + GateState::Paused { remaining_batches } => { + // Only count down. Counters and backoff change only at + // decision points. + *remaining_batches = remaining_batches.saturating_sub(1); + if *remaining_batches == 0 { + // The next batch is a probe. + self.state = GateState::new_window(); + self.probing = true; + } + self.awaiting_record = false; + GateDecision::Skip + } + GateState::Evaluate { .. } => { + self.awaiting_record = true; + GateDecision::Evaluate + } + } + } + + /// Records the result of an evaluation that [`Self::begin_batch`] + /// requested. `rows_in` is the number of rows evaluated, `rows_out` + /// the number of rows that passed (a null result does not pass) and + /// `elapsed` the evaluation time. + /// + /// Calls without a matching `begin_batch` that returned + /// [`GateDecision::Evaluate`] are ignored. + pub fn record(&mut self, rows_in: usize, rows_out: usize, elapsed: Duration) { + if !std::mem::take(&mut self.awaiting_record) { + return; + } + let GateState::Evaluate { + window, + batches_in_window, + } = &mut self.state + else { + return; + }; + window.add(rows_in as u64, rows_out as u64, duration_nanos(elapsed)); + *batches_in_window += 1; + // A window has at least `sample_batches` batches and + // `MIN_OBSERVED_ROWS` rows: fewer rows are not enough evidence for a + // decision. + if *batches_in_window >= self.config.sample_batches + && window.rows_in >= MIN_OBSERVED_ROWS + { + let window = *window; + self.decide(window); + } + } + + /// Number of times the gate paused the filter. + pub fn pauses(&self) -> usize { + self.pauses + } + + /// True if the gate skips the next batch, unless the filter changes + /// before it. + pub fn is_paused(&self) -> bool { + matches!(self.state, GateState::Paused { .. }) + } + + /// Makes a decision at the end of a window. + fn decide(&mut self, window: FilterCost) { + if window.rows_in == 0 { + // No rows, thus no information. Start a new window. + self.state = GateState::new_window(); + return; + } + if self.should_pause(&window) { + let batches = self.backoff; + self.start_pause(batches); + self.uses_shared_pauses = true; + self.publish(Verdict::Pause(batches)); + } else { + self.state = GateState::new_window(); + self.probing = false; + self.backoff = self.config.initial_pause_batches; + if std::mem::take(&mut self.uses_shared_pauses) { + // The first decision, or the end of a pause. + self.publish(Verdict::Keep); + } + } + } + + /// Before a batch that the gate would evaluate: if this gate uses the + /// shared pauses and another gate published a pause after the last + /// shared verdict that this gate saw, pauses the filter for the same + /// length. The current window is dropped. + fn use_shared_pause(&mut self) { + if !self.uses_shared_pauses { + return; + } + let Some(shared) = &self.shared else { + return; + }; + let (verdict, seq) = shared.load(); + if seq == self.shared_seq { + return; + } + self.shared_seq = seq; + if let Verdict::Pause(batches) = verdict { + self.start_pause(batches); + } + } + + /// Publishes `verdict` to the shared verdict, if any. + fn publish(&mut self, verdict: Verdict) { + if let Some(shared) = &self.shared { + self.shared_seq = shared + .publish_if(verdict, |_| true) + .expect("an unconditional update always succeeds"); + } + } + + /// The decision rules, see the [module documentation](self). + fn should_pause(&self, window: &FilterCost) -> bool { + let rows_removed = window.rows_removed(); + if rows_removed == 0 { + return true; + } + // The filter costs more than it saves. + let cost_ns = + window.nanos as f64 + window.rows_in as f64 * self.overhead_ns_per_row(); + let saving_ns = rows_removed as f64 * self.saving_ns_per_row(); + if self.probing { + // Turn the filter on again only if it is clearly worth its cost. + cost_ns >= saving_ns * RESUME_COST_MARGIN + } else { + cost_ns > saving_ns * PAUSE_COST_MARGIN + } + } + + /// Goes to `Paused` for `pause_batches` batches and doubles the backoff. + fn start_pause(&mut self, pause_batches: usize) { + let pause_batches = pause_batches.max(1); + self.probing = false; + self.state = GateState::Paused { + remaining_batches: pause_batches, + }; + self.backoff = pause_batches + .saturating_mul(2) + .min(self.config.max_pause_batches); + self.pauses += 1; + } +} + +#[cfg(test)] +mod tests { + use super::*; + use crate::expressions::{BinaryExpr, Column, DynamicFilterPhysicalExpr, col, lit}; + use crate::filter_stats::ManualClock; + use arrow::array::{Array, BooleanArray, Int32Array, RecordBatch}; + use arrow::datatypes::{DataType, Field, Schema}; + use datafusion_common::cast::as_boolean_array; + use datafusion_expr::Operator; + + /// Rows of each test batch: two batches make a window of + /// `MIN_OBSERVED_ROWS` rows. + const ROWS: usize = MIN_OBSERVED_ROWS as usize / 2; + + /// `ROWS` values of `value(i)`. + fn values(value: impl Fn(i32) -> Option) -> Vec> { + (0..ROWS as i32).map(value).collect() + } + + fn static_filter() -> Arc { + lit(true) + } + + /// A gate with the default configuration and a clock that does not + /// move: the evaluation time is 0, thus only a window that removes no + /// rows pauses the filter. + fn gate_with(filter: Arc) -> OptionalFilterGate { + OptionalFilterGate::new(filter, OptionalFilterGateConfig::default()) + .with_clock(Arc::new(ManualClock::new())) + } + + fn new_gate() -> OptionalFilterGate { + gate_with(static_filter()) + } + + /// Feeds one batch with the given pass ratio and an evaluation time of + /// 0. Returns the decision. + fn feed(gate: &mut OptionalFilterGate, pass_ratio: f64) -> GateDecision { + feed_timed(gate, pass_ratio, 0.0) + } + + /// Feeds one batch with the given pass ratio and an evaluation time of + /// `ns_per_row` for each row. Returns the decision. + fn feed_timed( + gate: &mut OptionalFilterGate, + pass_ratio: f64, + ns_per_row: f64, + ) -> GateDecision { + let decision = gate.begin_batch(); + if decision == GateDecision::Evaluate { + let elapsed = Duration::from_nanos((ROWS as f64 * ns_per_row) as u64); + gate.record(ROWS, (ROWS as f64 * pass_ratio) as usize, elapsed); + } + decision + } + + /// Feeds `n` batches and returns how many the gate evaluated. + fn feed_n(gate: &mut OptionalFilterGate, n: usize, pass_ratio: f64) -> usize { + (0..n) + .filter(|_| feed(gate, pass_ratio) == GateDecision::Evaluate) + .count() + } + + /// Feeds batches until the gate evaluates one. Returns the number of + /// skipped batches. + fn skip_until_probe(gate: &mut OptionalFilterGate, pass_ratio: f64) -> usize { + let mut skipped = 0; + while feed(gate, pass_ratio) == GateDecision::Skip { + skipped += 1; + assert!(skipped < 10_000, "gate never probes"); + } + skipped + } + + /// Evaluates the filter of `gate` on `batch` like a consumer does, or + /// skips it. Returns `None` if the gate skipped the filter. + fn evaluate( + gate: &mut OptionalFilterGate, + batch: &RecordBatch, + ) -> Option { + if gate.begin_batch() == GateDecision::Skip { + return None; + } + let num_rows = batch.num_rows(); + let start = gate.clock().now_nanos(); + let result = gate + .filter() + .evaluate(batch) + .unwrap() + .into_array(num_rows) + .unwrap(); + let result = as_boolean_array(&result).unwrap().clone(); + let elapsed = gate.clock().now_nanos().saturating_sub(start); + // `true_count` does not count nulls. + gate.record(num_rows, result.true_count(), Duration::from_nanos(elapsed)); + Some(result) + } + + #[test] + fn pauses_filter_that_removes_no_rows() { + let mut gate = new_gate(); + assert_eq!(feed(&mut gate, 1.0), GateDecision::Evaluate); + assert!(!gate.is_paused()); + assert_eq!(feed(&mut gate, 1.0), GateDecision::Evaluate); + assert!(gate.is_paused()); + assert_eq!(gate.pauses(), 1); + + // Pauses for `initial_pause_batches` batches. + for _ in 0..4 { + assert_eq!(feed(&mut gate, 1.0), GateDecision::Skip); + } + // Then it probes. + assert_eq!(feed(&mut gate, 1.0), GateDecision::Evaluate); + } + + #[test] + fn keeps_free_filter_that_removes_rows() { + let mut gate = new_gate(); + assert_eq!(feed_n(&mut gate, 100, 0.5), 100); + assert_eq!(gate.pauses(), 0); + // A free filter that removes only 1% of the rows also stays on. + assert_eq!(feed_n(&mut gate, 10, 0.99), 10); + assert_eq!(gate.pauses(), 0); + } + + #[test] + fn backoff_doubles_up_to_cap_and_resets() { + let mut gate = new_gate(); + let mut observed = vec![]; + assert_eq!(feed_n(&mut gate, 2, 1.0), 2); + for _ in 0..6 { + observed.push(skip_until_probe(&mut gate, 1.0)); + // `skip_until_probe` evaluated the first batch of the probe + // window. One more batch closes the window. + assert_eq!(feed(&mut gate, 1.0), GateDecision::Evaluate); + } + assert_eq!(observed, vec![4, 8, 16, 32, 32, 32]); + assert_eq!(gate.pauses(), 7); + + // A selective probe resets the backoff. + assert!(gate.is_paused()); + let skipped = skip_until_probe(&mut gate, 0.1); + assert_eq!(skipped, 32); + assert_eq!(feed(&mut gate, 0.1), GateDecision::Evaluate); + assert!(!gate.is_paused()); + assert_eq!(gate.backoff, 4); + + // The next pause is short again. + assert_eq!(feed_n(&mut gate, 2, 1.0), 2); + assert_eq!(skip_until_probe(&mut gate, 1.0), 4); + } + + fn dynamic_filter() -> (Arc, Arc) { + let column: Arc = Arc::new(Column::new("a", 0)); + let dynamic = Arc::new(DynamicFilterPhysicalExpr::new(vec![column], lit(true))); + let filter = Arc::clone(&dynamic) as Arc; + (dynamic, filter) + } + + fn a_gt(value: i32) -> Arc { + Arc::new(BinaryExpr::new( + Arc::new(Column::new("a", 0)), + Operator::Gt, + lit(value), + )) + } + + #[test] + fn generation_change_restarts_evaluation() { + let (dynamic, filter) = dynamic_filter(); + let mut gate = gate_with(filter); + + // Pause two times: the backoff grows to 16. + assert_eq!(feed_n(&mut gate, 2, 1.0), 2); + skip_until_probe(&mut gate, 1.0); + feed(&mut gate, 1.0); + assert!(gate.is_paused()); + assert_eq!(gate.backoff, 16); + + // A new generation of the filter ends the pause at once. + dynamic.update(a_gt(10)).unwrap(); + assert_eq!(feed(&mut gate, 1.0), GateDecision::Evaluate); + assert!(!gate.is_paused()); + assert_eq!(gate.backoff, 4); + + // The new window has an empty history: one more batch decides. + assert_eq!(feed(&mut gate, 1.0), GateDecision::Evaluate); + assert!(gate.is_paused()); + assert_eq!(skip_until_probe(&mut gate, 1.0), 4); + } + + #[test] + fn static_filter_is_never_watched() { + // A filter without dynamic filters, and a filter whose dynamic + // filters are all complete, can not change: the gate never polls + // them and never resets. + let (dynamic, complete) = dynamic_filter(); + dynamic.mark_complete(); + for (filter, expected_complete) in [(static_filter(), false), (complete, true)] { + let mut gate = gate_with(filter); + match (&gate.tracking, expected_complete) { + (DynamicFilterTracking::Static, false) + | (DynamicFilterTracking::AllComplete, true) => {} + (other, _) => panic!("unexpected tracking {other:?}"), + } + assert!(gate.tracking.watcher().is_none()); + + assert_eq!(feed_n(&mut gate, 2, 1.0), 2); + assert!(gate.is_paused()); + assert_eq!(skip_until_probe(&mut gate, 1.0), 4); + assert_eq!(feed(&mut gate, 1.0), GateDecision::Evaluate); + assert_eq!(skip_until_probe(&mut gate, 1.0), 8); + assert_eq!(gate.pauses(), 2); + } + } + + #[test] + fn completed_filter_stops_being_watched() { + let (dynamic, filter) = dynamic_filter(); + let mut gate = gate_with(filter); + assert!(gate.tracking.watcher().is_some()); + assert_eq!(feed_n(&mut gate, 2, 1.0), 2); + assert!(gate.is_paused()); + + // The final update restarts the evaluation one time. + dynamic.update(a_gt(10)).unwrap(); + dynamic.mark_complete(); + assert_eq!(feed(&mut gate, 1.0), GateDecision::Evaluate); + assert!(!gate.is_paused()); + // The tracker dropped the subscription of the complete filter. + assert!(gate.tracking.watcher().unwrap().is_exhausted()); + + // No more resets: the gate pauses and backs off as usual. + assert_eq!(feed(&mut gate, 1.0), GateDecision::Evaluate); + assert!(gate.is_paused()); + assert_eq!(skip_until_probe(&mut gate, 1.0), 4); + assert_eq!(feed(&mut gate, 1.0), GateDecision::Evaluate); + assert_eq!(skip_until_probe(&mut gate, 1.0), 8); + } + + #[test] + fn topk_like_filter_stays_on() { + // The filter gets tighter over time and changes often, like a TopK + // dynamic filter. At the start it removes almost no rows. + let (dynamic, filter) = dynamic_filter(); + let mut gate = gate_with(filter); + for i in 0..100 { + if i % 2 == 0 { + dynamic.update(a_gt(i)).unwrap(); + } + let pass_ratio = 1.0 - (i as f64 / 100.0); + assert_eq!( + feed(&mut gate, pass_ratio), + GateDecision::Evaluate, + "batch {i}" + ); + } + assert_eq!(gate.pauses(), 0); + } + + #[test] + fn skewed_input_probe_re_enables_filter() { + let mut gate = new_gate(); + // Data that the filter does not remove first: the gate pauses with + // growing backoff. + assert_eq!(feed_n(&mut gate, 2, 1.0), 2); + for _ in 0..3 { + skip_until_probe(&mut gate, 1.0); + feed(&mut gate, 1.0); + } + assert!(gate.is_paused()); + + // Then the data becomes selective. A probe finds it. + let skipped = skip_until_probe(&mut gate, 0.05); + assert!(skipped <= 32); + assert_eq!(feed(&mut gate, 0.05), GateDecision::Evaluate); + assert!(!gate.is_paused()); + // The filter stays on for the rest of the input. + assert_eq!(feed_n(&mut gate, 50, 0.05), 50); + } + + /// Regression test: a paused gate must not change its backoff or its + /// counters for each skipped batch. Only decisions change them. + #[test] + fn paused_gate_does_not_grow_backoff_while_skipping() { + let mut gate = new_gate(); + assert_eq!(feed_n(&mut gate, 2, 1.0), 2); + assert!(gate.is_paused()); + let backoff = gate.backoff; + let pauses = gate.pauses(); + + for _ in 0..3 { + assert_eq!(gate.begin_batch(), GateDecision::Skip); + // A stray `record` during a pause is ignored. + gate.record(ROWS, 0, Duration::from_secs(1)); + assert_eq!(gate.backoff, backoff); + assert_eq!(gate.pauses(), pauses); + } + assert_eq!(gate.begin_batch(), GateDecision::Skip); + assert_eq!(gate.backoff, backoff); + // The pause is over: the next batch is evaluated. + assert_eq!(gate.begin_batch(), GateDecision::Evaluate); + } + + /// Small batches (for example after a selective row filter) have a + /// large fixed cost for each row. The window grows until it has + /// `MIN_OBSERVED_ROWS` rows: small batches do not decide early. + #[test] + fn window_needs_min_observed_rows() { + let mut gate = new_gate(); + let small = 3; + let batches = MIN_OBSERVED_ROWS as usize / small; + for _ in 0..batches { + assert_eq!(gate.begin_batch(), GateDecision::Evaluate); + // 8000 ns for each row, and no row removed. + gate.record(small, small, Duration::from_nanos(8000 * small as u64)); + } + assert!(!gate.is_paused()); + assert_eq!(gate.begin_batch(), GateDecision::Evaluate); + gate.record(small, small, Duration::from_nanos(8000 * small as u64)); + assert!(gate.is_paused()); + } + + #[test] + fn empty_window_makes_no_decision() { + let mut gate = new_gate(); + for _ in 0..10 { + assert_eq!(gate.begin_batch(), GateDecision::Evaluate); + gate.record(0, 0, Duration::from_micros(1)); + } + assert_eq!(gate.pauses(), 0); + } + + fn batch(values: Vec>) -> RecordBatch { + let schema = Arc::new(Schema::new(vec![Field::new("a", DataType::Int32, true)])); + RecordBatch::try_new(schema, vec![Arc::new(Int32Array::from(values))]).unwrap() + } + + fn col_gt(value: i32) -> Arc { + let schema = Schema::new(vec![Field::new("a", DataType::Int32, true)]); + Arc::new(BinaryExpr::new( + col("a", &schema).unwrap(), + Operator::Gt, + lit(value), + )) + } + + #[test] + fn evaluate_selective_filter() { + let mut gate = gate_with(col_gt(8)); + let input = batch(values(|i| Some(i % 10))); + for _ in 0..10 { + let result = evaluate(&mut gate, &input).expect("evaluated"); + assert_eq!(result.true_count(), ROWS / 10); + } + assert_eq!(gate.pauses(), 0); + } + + #[test] + fn evaluate_filter_that_removes_no_rows_skips() { + let mut gate = gate_with(col_gt(0)); + let input = batch(values(|i| Some(i % 10 + 1))); + assert!(evaluate(&mut gate, &input).is_some()); + assert!(evaluate(&mut gate, &input).is_some()); + for _ in 0..4 { + assert!(evaluate(&mut gate, &input).is_none()); + } + assert!(evaluate(&mut gate, &input).is_some()); + } + + #[test] + fn evaluate_counts_null_as_not_passing() { + let mut gate = gate_with(col_gt(0)); + // 9 of 10 rows are null: the filter removes them. + let input = batch(values(|i| (i % 10 == 0).then_some(5))); + for _ in 0..10 { + let result = evaluate(&mut gate, &input).expect("evaluated"); + assert_eq!(result.null_count(), ROWS - ROWS.div_ceil(10)); + } + assert_eq!(gate.pauses(), 0); + } + + #[test] + fn config_from_execution_options() { + let mut options = ExecutionOptions::default(); + assert_eq!( + OptionalFilterGateConfig::from(&options), + OptionalFilterGateConfig::default() + ); + options.optional_filter_min_saving_ns_per_row = 7.5; + let config = OptionalFilterGateConfig::from(&options); + assert_eq!(config.min_saving_ns_per_row, 7.5); + assert_eq!(config.sample_batches, 2); + } + + // The tests below use the default `min_saving_ns_per_row` of 20 ns. For + // a window of 1000-row batches with pass ratio `p` and cost `c` ns for + // each row, the gate compares `c` with `(1 - p) * 20` ns for each row. + + /// Like the dynamic filter of a hash join with a multi-column key: it + /// removes 93% of the rows, but costs 70 ns for each row. Each removed + /// row saves only 20 ns, thus 18.6 ns for each evaluated row. + #[test] + fn selective_but_expensive_filter_pauses() { + let mut gate = new_gate(); + assert_eq!(feed_timed(&mut gate, 0.07, 70.0), GateDecision::Evaluate); + assert_eq!(feed_timed(&mut gate, 0.07, 70.0), GateDecision::Evaluate); + assert!(gate.is_paused()); + assert_eq!(gate.pauses(), 1); + + // The probes find the same cost: the pauses get longer. + assert_eq!(skip_until_probe(&mut gate, 0.07), 4); + assert_eq!(feed_timed(&mut gate, 0.07, 70.0), GateDecision::Evaluate); + // `skip_until_probe` evaluated the first batch with no time. The + // window costs 70 ns for each row of one batch and saves 0.93 * 20 + // ns for each row of two batches: still too much. + assert!(gate.is_paused()); + assert_eq!(skip_until_probe(&mut gate, 0.07), 8); + } + + /// A bound check that costs 2 ns for each row and removes 30% of the + /// rows: it saves 6 ns for each row, thus it stays on. + #[test] + fn cheap_weakly_selective_filter_stays_on() { + let mut gate = new_gate(); + for _ in 0..100 { + assert_eq!(feed_timed(&mut gate, 0.7, 2.0), GateDecision::Evaluate); + } + assert_eq!(gate.pauses(), 0); + } + + /// Cheap bounds that cost 1 ns for each row and remove 12% of the rows: + /// they save 2.4 ns for each row, thus they stay on. There is no + /// separate rule for the fraction of rows that pass. + #[test] + fn cheap_bounds_that_remove_few_rows_stay_on() { + let mut gate = new_gate(); + for _ in 0..100 { + assert_eq!(feed_timed(&mut gate, 0.88, 1.0), GateDecision::Evaluate); + } + assert_eq!(gate.pauses(), 0); + } + + /// Like the dynamic filter of a TopK: a cheap comparison that removes + /// almost all rows. It stays on. + #[test] + fn cheap_very_selective_filter_stays_on() { + let mut gate = new_gate(); + for _ in 0..100 { + assert_eq!(feed_timed(&mut gate, 0.001, 3.0), GateDecision::Evaluate); + } + assert_eq!(gate.pauses(), 0); + } + + /// The consumer measures a larger saving (for example the Parquet scan + /// measures the decode time of the columns that the filter does not + /// read). The next probe turns the expensive filter on again. + #[test] + fn measured_saving_re_enables_filter_at_next_probe() { + let saving = Arc::new(MeasuredRowSaving::new()); + let mut gate = new_gate().with_measured_saving(Arc::clone(&saving)); + assert_eq!(gate.saving_ns_per_row(), 20.0); + assert_eq!(feed_timed(&mut gate, 0.07, 70.0), GateDecision::Evaluate); + assert_eq!(feed_timed(&mut gate, 0.07, 70.0), GateDecision::Evaluate); + assert!(gate.is_paused()); + + // Each removed row now saves 20 + 80 = 100 ns: 93 ns for each + // evaluated row, more than the cost of 70 ns with the margin. + saving.set_ns_per_row(80.0); + assert_eq!(gate.saving_ns_per_row(), 100.0); + // The pause does not end early. + for _ in 0..4 { + assert_eq!(feed_timed(&mut gate, 0.07, 70.0), GateDecision::Skip); + } + assert_eq!(feed_timed(&mut gate, 0.07, 70.0), GateDecision::Evaluate); + assert_eq!(feed_timed(&mut gate, 0.07, 70.0), GateDecision::Evaluate); + assert!(!gate.is_paused()); + assert_eq!(gate.backoff, 4); + for _ in 0..20 { + assert_eq!(feed_timed(&mut gate, 0.07, 70.0), GateDecision::Evaluate); + } + + // Invalid measured values are used as 0. + saving.set_ns_per_row(f64::NAN); + assert_eq!(gate.saving_ns_per_row(), 20.0); + saving.set_ns_per_row(-5.0); + assert_eq!(gate.saving_ns_per_row(), 20.0); + } + + /// The consumer measures an overhead for each evaluated row (for example + /// the fixed cost of a Parquet row filter stage). A filter that is worth + /// its evaluation time alone is paused when the overhead is added. + #[test] + fn measured_overhead_adds_to_cost() { + let saving = Arc::new(MeasuredRowSaving::new()); + let mut gate = new_gate().with_measured_saving(Arc::clone(&saving)); + // Removes 50% of the rows: saves 10 ns for each evaluated row. The + // evaluation costs 5 ns for each row. + for _ in 0..10 { + assert_eq!(feed_timed(&mut gate, 0.5, 5.0), GateDecision::Evaluate); + } + assert!(!gate.is_paused()); + + // With 8 ns of overhead, the cost is 13 ns for each row. + saving.set_overhead_ns_per_row(8.0); + assert_eq!(gate.overhead_ns_per_row(), 8.0); + assert_eq!(feed_timed(&mut gate, 0.5, 5.0), GateDecision::Evaluate); + assert_eq!(feed_timed(&mut gate, 0.5, 5.0), GateDecision::Evaluate); + assert!(gate.is_paused()); + + // Invalid measured values are used as 0. + saving.set_overhead_ns_per_row(f64::INFINITY); + assert_eq!(gate.overhead_ns_per_row(), 0.0); + saving.set_overhead_ns_per_row(-1.0); + assert_eq!(gate.overhead_ns_per_row(), 0.0); + } + + /// When the cost is near the saving, the gate keeps its state: a running + /// filter stays on and a paused filter stays paused. + #[test] + fn cost_check_has_hysteresis() { + // The filter removes 50% of the rows: the saving is 10 ns for each + // evaluated row. A cost of 10.5 ns is between 0.9 and 1.1 times the + // saving. + let mut gate = new_gate(); + for _ in 0..20 { + assert_eq!(feed_timed(&mut gate, 0.5, 10.5), GateDecision::Evaluate); + } + assert!(!gate.is_paused()); + + // Pause it with a larger cost. + assert_eq!(feed_timed(&mut gate, 0.5, 30.0), GateDecision::Evaluate); + assert_eq!(feed_timed(&mut gate, 0.5, 30.0), GateDecision::Evaluate); + assert!(gate.is_paused()); + // A probe with the cost near the saving does not turn it on. + for _ in 0..4 { + assert_eq!(feed_timed(&mut gate, 0.5, 10.5), GateDecision::Skip); + } + assert_eq!(feed_timed(&mut gate, 0.5, 10.5), GateDecision::Evaluate); + assert_eq!(feed_timed(&mut gate, 0.5, 10.5), GateDecision::Evaluate); + assert!(gate.is_paused()); + assert_eq!(gate.pauses(), 2); + // A probe with a clearly lower cost turns it on. + for _ in 0..8 { + assert_eq!(feed_timed(&mut gate, 0.5, 8.5), GateDecision::Skip); + } + assert_eq!(feed_timed(&mut gate, 0.5, 8.5), GateDecision::Evaluate); + assert_eq!(feed_timed(&mut gate, 0.5, 8.5), GateDecision::Evaluate); + assert!(!gate.is_paused()); + } + + fn shared_gate( + filter: Arc, + shared: &Arc, + ) -> OptionalFilterGate { + gate_with(filter).with_shared_verdict(Arc::clone(shared)) + } + + #[test] + fn verdict_pack_round_trip() { + for verdict in [Verdict::Keep, Verdict::Pause(1), Verdict::Pause(32)] { + assert_eq!(Verdict::unpack(verdict.pack(7)), (verdict, 7)); + assert_eq!(Verdict::unpack(verdict.pack(u32::MAX)), (verdict, u32::MAX)); + } + // Pause lengths are at least 1 and at most `MAX_BATCHES`. + assert_eq!( + Verdict::unpack(Verdict::Pause(0).pack(1)).0, + Verdict::Pause(1) + ); + assert_eq!( + Verdict::unpack(Verdict::Pause(usize::MAX).pack(1)).0, + Verdict::Pause(Verdict::MAX_BATCHES) + ); + } + + /// A new gate starts from the shared pause of the other gates. + #[test] + fn new_gate_starts_with_shared_pause() { + let shared = Arc::new(SharedGateVerdict::new()); + let mut first = shared_gate(static_filter(), &shared); + assert!(!first.is_paused()); + assert_eq!(feed_n(&mut first, 2, 1.0), 2); + assert!(first.is_paused()); + assert!(shared.is_paused()); + + // A new gate starts paused for the same length, then probes. + let mut second = shared_gate(static_filter(), &shared); + assert!(second.is_paused()); + assert_eq!(second.pauses(), 1); + assert_eq!(skip_until_probe(&mut second, 0.5), 4); + // The probe keeps the filter: new gates evaluate again. + assert_eq!(feed(&mut second, 0.5), GateDecision::Evaluate); + assert!(!second.is_paused()); + assert!(!shared.is_paused()); + assert!(!shared_gate(static_filter(), &shared).is_paused()); + } + + /// Gates that start at the same time: a gate in its first window uses + /// the pause that another gate published, instead of the rest of its + /// own window, and only one gate probes after the pause. + #[test] + fn first_window_and_probe_use_shared_pause() { + let shared = Arc::new(SharedGateVerdict::new()); + let mut gates: Vec<_> = (0..4) + .map(|_| shared_gate(static_filter(), &shared)) + .collect(); + // All gates evaluate their first batch. + for gate in &mut gates { + assert_eq!(feed(gate, 1.0), GateDecision::Evaluate); + } + // The first gate completes its window and pauses. + assert_eq!(feed(&mut gates[0], 1.0), GateDecision::Evaluate); + assert!(gates[0].is_paused()); + // The others do not evaluate the second batch of their window. + for gate in &mut gates[1..] { + assert_eq!(feed(gate, 1.0), GateDecision::Skip); + assert!(gate.is_paused()); + } + + // Only the first gate probes at the end of its pause. The others + // use its verdict: a pause of 8 batches. + assert_eq!(skip_until_probe(&mut gates[0], 1.0), 4); + assert_eq!(feed(&mut gates[0], 1.0), GateDecision::Evaluate); + assert!(gates[0].is_paused()); + for gate in &mut gates[1..] { + let skipped = (0..20) + .take_while(|_| feed(gate, 1.0) == GateDecision::Skip) + .count(); + // 3 more batches of the first pause, then 8. + assert_eq!(skipped, 3 + 8); + assert_eq!(gate.pauses(), 2); + } + } + + /// A gate that keeps the filter does not use the pauses of other gates + /// (skewed data). + #[test] + fn running_gate_ignores_shared_pause() { + let shared = Arc::new(SharedGateVerdict::new()); + let mut selective = shared_gate(static_filter(), &shared); + assert_eq!(feed_n(&mut selective, 2, 0.1), 2); + assert!(!selective.is_paused()); + + let mut other = shared_gate(static_filter(), &shared); + assert_eq!(feed_n(&mut other, 2, 1.0), 2); + assert!(other.is_paused()); + assert!(shared.is_paused()); + + assert_eq!(feed_n(&mut selective, 20, 0.1), 20); + assert_eq!(selective.pauses(), 0); + } + + /// A change of the filter clears the shared pause: it was measured on + /// the old filter. + #[test] + fn change_clears_shared_pause() { + let (dynamic, filter) = dynamic_filter(); + let shared = Arc::new(SharedGateVerdict::new()); + let mut gate = shared_gate(Arc::clone(&filter), &shared); + assert_eq!(feed_n(&mut gate, 2, 1.0), 2); + assert!(shared.is_paused()); + + dynamic.update(a_gt(1)).unwrap(); + // The gate sees the change at its next batch. + assert_eq!(feed(&mut gate, 0.5), GateDecision::Evaluate); + assert!(!shared.is_paused()); + assert!(!shared_gate(filter, &shared).is_paused()); + } + + /// The producer of a dynamic filter measures its work for each row that + /// the filter removes: once measured, it replaces the configured + /// `min_saving_ns_per_row`. + #[test] + fn producer_work_replaces_configured_saving() { + let (dynamic, filter) = dynamic_filter(); + let mut gate = gate_with(filter); + assert_eq!(gate.saving_ns_per_row(), 20.0); + // Removes 80% at 5 ns for each row: 5 < 0.8 * 20 * 1.1, it stays on. + for _ in 0..4 { + assert_eq!(feed_timed(&mut gate, 0.2, 5.0), GateDecision::Evaluate); + } + assert!(!gate.is_paused()); + + // The producer measures 4 ns for each removed row: 0.8 * 4 < 5. + let work = dynamic.removed_row_work(); + work.record(MIN_OBSERVED_ROWS, 4 * MIN_OBSERVED_ROWS); + assert_eq!(gate.saving_ns_per_row(), 4.0); + assert_eq!(feed_timed(&mut gate, 0.2, 5.0), GateDecision::Evaluate); + assert_eq!(feed_timed(&mut gate, 0.2, 5.0), GateDecision::Evaluate); + assert!(gate.is_paused()); + + // A filter without dynamic filters uses the configuration. + assert_eq!(new_gate().saving_ns_per_row(), 20.0); + } + + /// A clock that moves by a fixed step each time it is read. + #[derive(Debug)] + struct SteppingClock { + now: AtomicU64, + step: u64, + } + + impl Clock for SteppingClock { + fn now_nanos(&self) -> u64 { + self.now.fetch_add(self.step, Ordering::Relaxed) + } + } + + /// A consumer measures the time with the clock of the gate. + #[test] + fn consumer_measures_time_with_gate_clock() { + // Each evaluation takes 1000 ns for each row. + let clock = Arc::new(SteppingClock { + now: AtomicU64::new(0), + step: 1000 * ROWS as u64, + }); + let mut gate = + OptionalFilterGate::new(col_gt(8), OptionalFilterGateConfig::default()) + .with_clock(clock); + let input = batch(values(|i| Some(i % 10))); + // The filter removes 90% of the rows, but it costs 1000 ns for each + // row, and each removed row saves 20 ns. + assert!(evaluate(&mut gate, &input).is_some()); + assert!(evaluate(&mut gate, &input).is_some()); + assert!(gate.is_paused()); + assert!(evaluate(&mut gate, &input).is_none()); + } +} diff --git a/datafusion/physical-expr/src/simplifier/mod.rs b/datafusion/physical-expr/src/simplifier/mod.rs index af87ce7f61d35..18fd29709510e 100644 --- a/datafusion/physical-expr/src/simplifier/mod.rs +++ b/datafusion/physical-expr/src/simplifier/mod.rs @@ -96,7 +96,8 @@ mod tests { use super::*; use crate::ScalarFunctionExpr; use crate::expressions::{ - BinaryExpr, CastExpr, Literal, NotExpr, TryCastExpr, col, in_list, lit, + BinaryExpr, CastExpr, Literal, NotExpr, OptionalFilterPhysicalExpr, TryCastExpr, + col, in_list, lit, }; use arrow::datatypes::{DataType, Field}; use datafusion_common::ScalarValue; @@ -238,6 +239,35 @@ mod tests { Ok(()) } + #[test] + fn test_not_optional_filter_unchanged() -> Result<()> { + let schema = not_test_schema(); + let simplifier = PhysicalExprSimplifier::new(&schema); + let optional = |e: Arc| -> Arc { + Arc::new(OptionalFilterPhysicalExpr::new(e)) + }; + + // NOT(Optional(c > 5)) is unchanged: the NOT must not move into the + // Optional, because NOT(c > 5) is a required filter. + let c_gt_5: Arc = Arc::new(BinaryExpr::new( + col("c", &schema)?, + Operator::Gt, + lit(ScalarValue::Int32(Some(5))), + )); + let expr: Arc = + Arc::new(NotExpr::new(optional(Arc::clone(&c_gt_5)))); + assert_not_simplify(&simplifier, Arc::clone(&expr), expr); + + // NOT(Optional(NOT(a))) is unchanged: no double negation elimination + // across the Optional. + let expr: Arc = Arc::new(NotExpr::new(optional(Arc::new( + NotExpr::new(col("a", &schema)?), + )))); + assert_not_simplify(&simplifier, Arc::clone(&expr), expr); + + Ok(()) + } + #[test] fn test_not_literal() -> Result<()> { let schema = not_test_schema(); diff --git a/datafusion/physical-expr/src/utils/mod.rs b/datafusion/physical-expr/src/utils/mod.rs index 1be57c9192626..28543883733bb 100644 --- a/datafusion/physical-expr/src/utils/mod.rs +++ b/datafusion/physical-expr/src/utils/mod.rs @@ -21,7 +21,9 @@ pub use guarantee::{Guarantee, LiteralGuarantee}; use std::borrow::Borrow; use std::sync::Arc; -use crate::expressions::{BinaryExpr, Column, Literal}; +use crate::expressions::{ + BinaryExpr, Column, DynamicFilterPhysicalExpr, Literal, OptionalFilterPhysicalExpr, +}; use crate::tree_node::ExprContext; use crate::{ AcrossPartitions, ConstExpr, EquivalenceProperties, PhysicalExpr, PhysicalSortExpr, @@ -46,6 +48,80 @@ pub fn split_conjunction( split_impl(Operator::And, predicate, vec![]) } +/// Split the root `AND` chain of `predicate` into required and optional +/// filters. +/// +/// Returns `(required, optional)`. `required` holds the conjuncts that must be +/// applied. `optional` holds the *inner* expressions of the conjuncts that are +/// an [`OptionalFilterPhysicalExpr`], which a consumer can skip without +/// affecting correctness. +/// +/// Only the root `AND` chain is examined (the same conjuncts as +/// [`split_conjunction`]). An `Optional` in any other position, for example +/// `NOT(Optional(x))` or `Optional(x) OR y`, stays inside a required conjunct. +/// +/// For example, `a AND Optional(b) AND (c AND Optional(d))` gives +/// `([a, c], [b, d])`. +#[expect(clippy::type_complexity)] +pub fn split_optional( + predicate: &Arc, +) -> (Vec>, Vec>) { + let mut required = vec![]; + let mut optional = vec![]; + for conjunct in split_conjunction(predicate) { + match conjunct.downcast_ref::() { + Some(opt) => optional.push(Arc::clone(opt.inner())), + None => required.push(Arc::clone(conjunct)), + } + } + (required, optional) +} + +/// Returns `true` if `expr` itself is an [`OptionalFilterPhysicalExpr`]. +/// +/// This does not examine the children of `expr`. +pub fn is_optional_filter(expr: &Arc) -> bool { + expr.is::() +} + +/// Returns the [`DynamicFilterPhysicalExpr`] if `expr` is one, or if `expr` is +/// an [`OptionalFilterPhysicalExpr`] that directly wraps one. +pub fn as_dynamic_filter( + expr: &Arc, +) -> Option<&DynamicFilterPhysicalExpr> { + let expr = match expr.downcast_ref::() { + Some(opt) => opt.inner(), + None => expr, + }; + expr.downcast_ref::() +} + +/// In debug builds, panic if `predicate` has an [`OptionalFilterPhysicalExpr`] +/// that is not a direct conjunct of the root `AND` chain, or that is nested +/// in another `Optional`. In release builds, this function does nothing. +/// +/// A misplaced `Optional` does not cause incorrect results (it is always +/// evaluated), but it is not possible to skip it. Thus a misplaced +/// `Optional` usually shows a bug in a producer or a rewriter. +pub fn debug_assert_optional_on_root_chain(predicate: &Arc) { + if !cfg!(debug_assertions) { + return; + } + for conjunct in split_conjunction(predicate) { + let subtree = match conjunct.downcast_ref::() { + Some(opt) => opt.inner(), + None => conjunct, + }; + let misplaced = subtree + .exists(|e| Ok(is_optional_filter(e))) + .expect("infallible closure should not fail"); + debug_assert!( + !misplaced, + "OptionalFilterPhysicalExpr must be a direct conjunct of the root AND chain: {predicate}" + ); + } +} + impl ConstExpr { /// Collects predicate-derived constants from equality conjunctions. /// @@ -644,4 +720,140 @@ pub(crate) mod tests { Ok(()) } + + fn optional_test_schema() -> Schema { + Schema::new(vec![ + Field::new("a", DataType::Boolean, true), + Field::new("b", DataType::Boolean, true), + Field::new("c", DataType::Boolean, true), + Field::new("d", DataType::Boolean, true), + ]) + } + + fn optional(inner: Arc) -> Arc { + Arc::new(OptionalFilterPhysicalExpr::new(inner)) + } + + fn and( + left: Arc, + right: Arc, + ) -> Arc { + Arc::new(BinaryExpr::new(left, Operator::And, right)) + } + + fn or( + left: Arc, + right: Arc, + ) -> Arc { + Arc::new(BinaryExpr::new(left, Operator::Or, right)) + } + + fn to_strings(exprs: &[Arc]) -> Vec { + exprs.iter().map(|e| e.to_string()).collect() + } + + #[test] + fn test_split_optional_root_chain() -> Result<()> { + let schema = optional_test_schema(); + let [a, b, c, d] = ["a", "b", "c", "d"].map(|n| col(n, &schema).unwrap()); + + // a AND Optional(b) AND (c AND Optional(d)) + let predicate = and( + and(Arc::clone(&a), optional(Arc::clone(&b))), + and(Arc::clone(&c), optional(Arc::clone(&d))), + ); + let (required, optional_filters) = split_optional(&predicate); + assert_eq!(to_strings(&required), vec!["a@0", "c@2"]); + assert_eq!(to_strings(&optional_filters), vec!["b@1", "d@3"]); + debug_assert_optional_on_root_chain(&predicate); + + // A single Optional at the root is optional. + let predicate = optional(Arc::clone(&a)); + let (required, optional_filters) = split_optional(&predicate); + assert!(required.is_empty()); + assert_eq!(to_strings(&optional_filters), vec!["a@0"]); + debug_assert_optional_on_root_chain(&predicate); + + // An Optional that wraps an AND is one optional filter. + let predicate = optional(and(Arc::clone(&a), Arc::clone(&b))); + let (required, optional_filters) = split_optional(&predicate); + assert!(required.is_empty()); + assert_eq!(to_strings(&optional_filters), vec!["a@0 AND b@1"]); + debug_assert_optional_on_root_chain(&predicate); + + // No Optional: everything is required. + let predicate = and(Arc::clone(&a), Arc::clone(&b)); + let (required, optional_filters) = split_optional(&predicate); + assert_eq!(to_strings(&required), vec!["a@0", "b@1"]); + assert!(optional_filters.is_empty()); + Ok(()) + } + + #[test] + fn test_split_optional_not_on_root_chain() -> Result<()> { + let schema = optional_test_schema(); + let [x, y] = ["a", "b"].map(|n| col(n, &schema).unwrap()); + + // NOT(Optional(x)) is required + let predicate = crate::expressions::not(optional(Arc::clone(&x)))?; + let (required, optional_filters) = split_optional(&predicate); + assert_eq!(required, vec![Arc::clone(&predicate)]); + assert!(optional_filters.is_empty()); + + // Optional(x) OR y is required + let predicate = or(optional(Arc::clone(&x)), Arc::clone(&y)); + let (required, optional_filters) = split_optional(&predicate); + assert_eq!(required, vec![Arc::clone(&predicate)]); + assert!(optional_filters.is_empty()); + Ok(()) + } + + #[test] + fn test_is_optional_filter_and_as_dynamic_filter() -> Result<()> { + let schema = optional_test_schema(); + let a = col("a", &schema)?; + let dynamic: Arc = Arc::new(DynamicFilterPhysicalExpr::new( + vec![Arc::clone(&a)], + lit(true), + )); + + assert!(is_optional_filter(&optional(Arc::clone(&a)))); + assert!(!is_optional_filter(&a)); + // Only the node itself is examined. + assert!(!is_optional_filter(&crate::expressions::not(optional( + Arc::clone(&a) + ))?)); + + assert!(as_dynamic_filter(&dynamic).is_some()); + assert!(as_dynamic_filter(&optional(Arc::clone(&dynamic))).is_some()); + assert!(as_dynamic_filter(&a).is_none()); + assert!(as_dynamic_filter(&optional(Arc::clone(&a))).is_none()); + // Only one Optional wrapper is removed. + assert!(as_dynamic_filter(&optional(optional(dynamic))).is_none()); + Ok(()) + } + + #[test] + #[cfg_attr( + debug_assertions, + should_panic(expected = "must be a direct conjunct of the root AND chain") + )] + fn test_debug_assert_optional_under_not() { + let schema = optional_test_schema(); + let a = col("a", &schema).unwrap(); + let b = col("b", &schema).unwrap(); + let predicate = and(a, crate::expressions::not(optional(b)).unwrap()); + debug_assert_optional_on_root_chain(&predicate); + } + + #[test] + #[cfg_attr( + debug_assertions, + should_panic(expected = "must be a direct conjunct of the root AND chain") + )] + fn test_debug_assert_nested_optional() { + let schema = optional_test_schema(); + let a = col("a", &schema).unwrap(); + debug_assert_optional_on_root_chain(&optional(optional(a))); + } } diff --git a/datafusion/physical-plan/src/aggregates/mod.rs b/datafusion/physical-plan/src/aggregates/mod.rs index d82d6de61c1aa..ed38bda66a3b3 100644 --- a/datafusion/physical-plan/src/aggregates/mod.rs +++ b/datafusion/physical-plan/src/aggregates/mod.rs @@ -194,7 +194,9 @@ use datafusion_execution::TaskContext; use datafusion_expr::{Accumulator, Aggregate, AggregateMetrics}; use datafusion_physical_expr::aggregate::AggregateFunctionExpr; use datafusion_physical_expr::equivalence::ProjectionMapping; -use datafusion_physical_expr::expressions::{Column, DynamicFilterPhysicalExpr, lit}; +use datafusion_physical_expr::expressions::{ + Column, DynamicFilterPhysicalExpr, OptionalFilterPhysicalExpr, lit, +}; use datafusion_physical_expr::{ ConstExpr, EquivalenceProperties, physical_exprs_contains, }; @@ -2373,8 +2375,12 @@ impl ExecutionPlan for AggregateExec { && config.optimizer.enable_aggregate_dynamic_filter_pushdown && let Some(self_dyn_filter) = &self.dynamic_filter { + // The aggregate itself computes the correct result from all input + // rows, so the filter is not needed for correctness: mark the + // pushed copy as optional. let dyn_filter = Arc::clone(&self_dyn_filter.filter); - child_desc = child_desc.with_self_filter(dyn_filter); + child_desc = child_desc + .with_self_filter(Arc::new(OptionalFilterPhysicalExpr::new(dyn_filter))); } Ok(FilterDescription::new().with_child(child_desc)) diff --git a/datafusion/physical-plan/src/joins/hash_join/bounds_union.rs b/datafusion/physical-plan/src/joins/hash_join/bounds_union.rs new file mode 100644 index 0000000000000..d7985c82dd79d --- /dev/null +++ b/datafusion/physical-plan/src/joins/hash_join/bounds_union.rs @@ -0,0 +1,358 @@ +// Licensed to the Apache Software Foundation (ASF) under one +// or more contributor license agreements. See the NOTICE file +// distributed with this work for additional information +// regarding copyright ownership. The ASF licenses this file +// to you under the Apache License, Version 2.0 (the +// "License"); you may not use this file except in compliance +// with the License. You may obtain a copy of the License at +// +// http://www.apache.org/licenses/LICENSE-2.0 +// +// Unless required by applicable law or agreed to in writing, +// software distributed under the License is distributed on an +// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY +// KIND, either express or implied. See the License for the +// specific language governing permissions and limitations +// under the License. + +//! Merges the per-partition build-side ranges of a partitioned hash join into +//! one bounds predicate that does not need routing. +//! +//! A partitioned hash join knows `[min, max]` for each join key column in each +//! build partition. To test a probe row against the bounds of one partition, +//! the filter must first compute the partition of the row, thus these bounds +//! are inside the routing `CASE`. +//! +//! This module computes the union of the ranges instead. The union does not +//! need routing, so the join can push it as a separate dynamic filter: the +//! pruning code can use it for files, row groups and pages. +//! +//! # This is a relaxation +//! +//! The union accepts more rows than the per-partition bounds: a probe key that +//! is in the range of partition 1 but routes to partition 0 passes the union. +//! This is correct because the join (and the membership filter) remove these +//! rows. +//! +//! With hash partitioning, every partition usually spans almost the full key +//! range, so the union is one range per column, which rejects few rows. It +//! is still useful when the build keys are a narrow slice of a probe side +//! that is clustered by the key. With range partitioning, the +//! partitions hold disjoint key ranges, and the union can keep them as +//! separate ranges (up to [`MAX_RANGES_PER_COLUMN`]), so it also rejects the +//! probe keys in the gaps between them. +//! +//! # Multi-column keys +//! +//! With more than one join key the per-partition bounds describe a box, and a +//! union of boxes is not a box. Each column is merged independently, and the +//! predicate is the product of the merged per-column ranges. This is a +//! superset of the union, thus it is also correct. + +use std::cmp::Ordering; +use std::sync::Arc; + +use super::shared_bounds::PartitionBounds; + +use datafusion_common::ScalarValue; +use datafusion_expr::Operator; +use datafusion_physical_expr::expressions::{BinaryExpr, lit}; +use datafusion_physical_expr::{PhysicalExpr, PhysicalExprRef}; + +/// Maximum number of disjoint ranges for one join key column when the probe +/// side is range partitioned. More ranges are merged into one range from the +/// smallest minimum to the largest maximum: each range adds evaluation cost +/// for each probe batch, and the pruning code gets little from a long `OR` +/// chain. +/// +/// With hash partitioning the gaps between the ranges of the partitions are +/// random, not a property of the data, so the ranges are always merged into +/// one (see [`merge_partition_bounds`]). +pub(super) const MAX_RANGES_PER_COLUMN: usize = 8; + +/// A closed `[min, max]` range of one join key column. +pub(super) type Range = (ScalarValue, ScalarValue); + +/// Merges the bounds of the given partitions into, for each of the +/// `num_columns` key columns, the smallest set of disjoint ranges that contains +/// the bounds of every partition. If a column gets more than +/// `max_ranges_per_column` ranges, they are merged into one range. +/// +/// `partitions` must contain every partition that can hold a build row. If a +/// partition is missing (for example, a canceled partition with unknown +/// contents), the result can reject probe rows that match. +/// +/// An empty set for a column means that the column has no usable bounds and +/// the predicate must not constrain it. This occurs when a partition has no +/// bounds for the column, or when two bounds cannot be compared. NULL bounds +/// are skipped: they occur only when every key of the column in that partition +/// is NULL, and a NULL key cannot satisfy a range check in any case. If every +/// bound of a column is NULL, the column gets an empty set. +pub(super) fn merge_partition_bounds( + num_columns: usize, + partitions: &[&PartitionBounds], + max_ranges_per_column: usize, +) -> Vec> { + (0..num_columns) + .map(|column| merge_column(column, partitions, max_ranges_per_column)) + .collect() +} + +fn merge_column( + column: usize, + partitions: &[&PartitionBounds], + max_ranges: usize, +) -> Vec { + let mut ranges: Vec = Vec::with_capacity(partitions.len()); + for bounds in partitions { + // A partition without bounds for this column can hold any value, so + // the column must stay unconstrained. + let Some(column_bounds) = bounds.get_column_bounds(column) else { + return Vec::new(); + }; + if column_bounds.min.is_null() || column_bounds.max.is_null() { + continue; + } + ranges.push((column_bounds.min.clone(), column_bounds.max.clone())); + } + + // Values of different types cannot be compared. This is not expected (all + // partitions have the same key types), but in that case the column stays + // unconstrained instead of producing an incorrect range. + let Some((first, _)) = ranges.first() else { + return Vec::new(); + }; + let comparable = ranges.iter().all(|(min, max)| { + min.partial_cmp(first).is_some() && max.partial_cmp(first).is_some() + }); + if !comparable { + return Vec::new(); + } + + // `ScalarValue` uses the same order as the `>=` and `<=` comparisons of + // the predicate, so a sort by minimum and a sweep merge exactly the ranges + // that overlap. + ranges.sort_by(|(a, _), (b, _)| a.partial_cmp(b).unwrap_or(Ordering::Equal)); + let mut merged: Vec = Vec::with_capacity(ranges.len()); + for (min, max) in ranges { + match merged.last_mut() { + Some((_, current_max)) + if min.partial_cmp(current_max) != Some(Ordering::Greater) => + { + if max.partial_cmp(current_max) == Some(Ordering::Greater) { + *current_max = max; + } + } + _ => merged.push((min, max)), + } + } + + // The ranges are sorted by minimum and disjoint, so the last one has the + // largest maximum. + if merged.len() > max_ranges { + let min = merged[0].0.clone(); + let max = merged[merged.len() - 1].1.clone(); + merged = vec![(min, max)]; + } + merged +} + +/// Creates the predicate for merged bounds: for each column with ranges, +/// `col >= min AND col <= max`, with the ranges of one column combined with +/// `OR`, and the columns combined with `AND`. +/// +/// Returns `None` if no column has ranges. +pub(super) fn create_merged_bounds_predicate( + on_right: &[PhysicalExprRef], + merged: &[Vec], +) -> Option> { + on_right + .iter() + .zip(merged) + .filter_map(|(right_expr, ranges)| { + ranges + .iter() + .map(|(min, max)| range_predicate(right_expr, min, max)) + .reduce(|acc, range| { + Arc::new(BinaryExpr::new(acc, Operator::Or, range)) + as Arc + }) + }) + .reduce(|acc, predicate| { + Arc::new(BinaryExpr::new(acc, Operator::And, predicate)) + as Arc + }) +} + +/// Creates the predicate `expr >= min AND expr <= max`. +pub(super) fn range_predicate( + expr: &PhysicalExprRef, + min: &ScalarValue, + max: &ScalarValue, +) -> Arc { + let min_expr = Arc::new(BinaryExpr::new( + Arc::clone(expr), + Operator::GtEq, + lit(min.clone()), + )) as Arc; + let max_expr = Arc::new(BinaryExpr::new( + Arc::clone(expr), + Operator::LtEq, + lit(max.clone()), + )) as Arc; + Arc::new(BinaryExpr::new(min_expr, Operator::And, max_expr)) +} + +#[cfg(test)] +mod tests { + use super::*; + + use crate::joins::hash_join::shared_bounds::ColumnBounds; + + use datafusion_physical_expr::expressions::Column; + + fn partition(ranges: &[(i32, i32)]) -> PartitionBounds { + PartitionBounds::new( + ranges + .iter() + .map(|(min, max)| { + ColumnBounds::new( + ScalarValue::Int32(Some(*min)), + ScalarValue::Int32(Some(*max)), + ) + }) + .collect(), + ) + } + + fn merge(num_columns: usize, partitions: &[PartitionBounds]) -> Vec> { + let refs = partitions.iter().collect::>(); + merge_partition_bounds(num_columns, &refs, MAX_RANGES_PER_COLUMN) + } + + fn ranges(merged: &[Vec], column: usize) -> Vec<(i32, i32)> { + merged[column] + .iter() + .map(|(min, max)| match (min, max) { + (ScalarValue::Int32(Some(min)), ScalarValue::Int32(Some(max))) => { + (*min, *max) + } + other => panic!("expected Int32 range, got {other:?}"), + }) + .collect() + } + + #[test] + fn overlapping_ranges_merge_into_one() { + let merged = merge(1, &[partition(&[(0, 10)]), partition(&[(5, 20)])]); + assert_eq!(ranges(&merged, 0), vec![(0, 20)]); + } + + #[test] + fn touching_ranges_merge_into_one() { + let merged = merge(1, &[partition(&[(0, 10)]), partition(&[(10, 20)])]); + assert_eq!(ranges(&merged, 0), vec![(0, 20)]); + } + + #[test] + fn disjoint_ranges_stay_separate() { + let merged = merge(1, &[partition(&[(100, 110)]), partition(&[(0, 10)])]); + assert_eq!(ranges(&merged, 0), vec![(0, 10), (100, 110)]); + } + + #[test] + fn contained_range_is_absorbed() { + let merged = merge(1, &[partition(&[(0, 100)]), partition(&[(10, 20)])]); + assert_eq!(ranges(&merged, 0), vec![(0, 100)]); + } + + #[test] + fn too_many_disjoint_ranges_merge_into_one() { + let partitions = (0..=MAX_RANGES_PER_COLUMN) + .map(|i| partition(&[(i as i32 * 100, i as i32 * 100 + 1)])) + .collect::>(); + let merged = merge(1, &partitions); + assert_eq!( + ranges(&merged, 0), + vec![(0, MAX_RANGES_PER_COLUMN as i32 * 100 + 1)] + ); + } + + #[test] + fn one_range_per_column_merges_disjoint_ranges() { + let partitions = [partition(&[(100, 110)]), partition(&[(0, 10)])]; + let refs = partitions.iter().collect::>(); + let merged = merge_partition_bounds(1, &refs, 1); + assert_eq!(ranges(&merged, 0), vec![(0, 110)]); + } + + #[test] + fn partition_without_bounds_leaves_column_unconstrained() { + // The second partition has no bounds for column 0, so it can hold any + // value. + let merged = merge(1, &[partition(&[(0, 10)]), partition(&[])]); + assert!(merged[0].is_empty()); + } + + #[test] + fn null_bounds_are_skipped() { + let all_null = PartitionBounds::new(vec![ColumnBounds::new( + ScalarValue::Int32(None), + ScalarValue::Int32(None), + )]); + let merged = merge(1, &[partition(&[(0, 10)]), all_null.clone()]); + assert_eq!(ranges(&merged, 0), vec![(0, 10)]); + + let merged = merge(1, &[all_null]); + assert!(merged[0].is_empty()); + } + + #[test] + fn incomparable_bounds_leave_column_unconstrained() { + let utf8 = PartitionBounds::new(vec![ColumnBounds::new( + ScalarValue::from("a"), + ScalarValue::from("b"), + )]); + let merged = merge(1, &[partition(&[(0, 10)]), utf8]); + assert!(merged[0].is_empty()); + } + + #[test] + fn columns_are_merged_independently() { + let merged = merge( + 2, + &[ + partition(&[(0, 10), (100, 110)]), + partition(&[(20, 30), (0, 5)]), + ], + ); + assert_eq!(ranges(&merged, 0), vec![(0, 10), (20, 30)]); + assert_eq!(ranges(&merged, 1), vec![(0, 5), (100, 110)]); + } + + #[test] + fn predicate_ors_disjoint_ranges_and_conjoins_columns() { + let merged = merge( + 2, + &[ + partition(&[(0, 10), (0, 5)]), + partition(&[(100, 110), (0, 5)]), + ], + ); + let on_right: Vec = + vec![Arc::new(Column::new("a", 0)), Arc::new(Column::new("b", 1))]; + let predicate = create_merged_bounds_predicate(&on_right, &merged) + .expect("expected a bounds predicate"); + assert_eq!( + predicate.to_string(), + "(a@0 >= 0 AND a@0 <= 10 OR a@0 >= 100 AND a@0 <= 110) AND b@1 >= 0 AND b@1 <= 5" + ); + } + + #[test] + fn unconstrained_columns_have_no_predicate() { + let merged = merge(1, &[partition(&[(0, 10)]), partition(&[])]); + let on_right: Vec = vec![Arc::new(Column::new("a", 0))]; + assert!(create_merged_bounds_predicate(&on_right, &merged).is_none()); + } +} diff --git a/datafusion/physical-plan/src/joins/hash_join/exec.rs b/datafusion/physical-plan/src/joins/hash_join/exec.rs index 1bffdb1ab2d7c..bbdaecd23ba65 100644 --- a/datafusion/physical-plan/src/joins/hash_join/exec.rs +++ b/datafusion/physical-plan/src/joins/hash_join/exec.rs @@ -92,7 +92,9 @@ use datafusion_functions_aggregate_common::min_max::{MaxAccumulator, MinAccumula use datafusion_physical_expr::equivalence::{ ProjectionMapping, join_equivalence_properties, }; -use datafusion_physical_expr::expressions::{Column, DynamicFilterPhysicalExpr, lit}; +use datafusion_physical_expr::expressions::{ + Column, DynamicFilterPhysicalExpr, OptionalFilterPhysicalExpr, lit, +}; use datafusion_physical_expr::projection::{ProjectionRef, combine_projections}; use datafusion_physical_expr::{PhysicalExpr, PhysicalExprRef}; @@ -949,15 +951,45 @@ pub struct HashJoinExec { fetch: Option, } +/// The dynamic filters that a hash join updates with the results of the build +/// side once that is done. +/// +/// The join pushes two filters to the probe side: one for the build-side +/// bounds and one for the membership check (see +/// [`HashJoinExec::gather_filters_for_pushdown`]). At least one of them is set: +/// the join keeps only the filters that reached a consumer. #[derive(Clone)] struct HashJoinExecDynamicFilter { - /// Dynamic filter that we'll update with the results of the build side once that is done. - filter: Arc, + /// The membership check (`InList` or hash table lookup, routed by + /// partition in partitioned mode). When `bounds` is `None`, this filter + /// also holds the bounds (`bounds AND membership`). + membership: Option>, + /// The build-side bounds (`col >= min AND col <= max`). In partitioned + /// mode, the union of the bounds of all partitions. + bounds: Option>, /// Build accumulator to collect build-side information (hash maps and/or bounds) from each partition. /// It is lazily initialized during execution to make sure we use the actual execution time partition counts. build_accumulator: OnceLock>, } +impl HashJoinExecDynamicFilter { + fn new( + membership: Option>, + bounds: Option>, + ) -> Self { + Self { + membership, + bounds, + build_accumulator: OnceLock::new(), + } + } + + /// The dynamic filters of this join, membership first. + fn filters(&self) -> impl Iterator> { + self.membership.iter().chain(self.bounds.iter()) + } +} + impl fmt::Debug for HashJoinExec { fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { f.debug_struct("HashJoinExec") @@ -1223,7 +1255,9 @@ impl HashJoinExec { note = "Use ExecutionPlan::dynamic_expressions_produced instead" )] pub fn dynamic_filter_expr(&self) -> Option<&Arc> { - self.dynamic_filter.as_ref().map(|df| &df.filter) + self.dynamic_filter + .as_ref() + .and_then(|df| df.membership.as_ref().or(df.bounds.as_ref())) } /// Set the dynamic filter on this hash join. @@ -1248,27 +1282,72 @@ impl HashJoinExec { self.set_dynamic_filter(filter) } - /// Set the dynamic filter on this hash join, resetting any internal state - /// that depends on an existing one and validating that the filter's - /// children reference valid columns in the probe (right) side's schema. + /// Returns the dynamic filter in a self filter that this join pushed to + /// the probe side, if a node in the probe side holds it. + /// + /// Note that we don't check `PushedDownPredicate::discriminant`: a node + /// that replies `PushedDown::No` may still retain the filter for + /// statistics pruning, so the reply does not tell us whether anyone will + /// actually read the filter. Instead we look for a consumer holding the + /// expression in the probe subtree. + /// + /// `self` here is the join with its post-pushdown children, so anything + /// that accepted the filter is already wired into `self.right`. Searching + /// from `self` would always find the producer expression pushed by this + /// join. This is the last chance to make the decision: the Post-phase + /// `FilterPushdown` rule is the final rule that mutates the plan. + fn consumed_dynamic_filter( + &self, + predicate: &Arc, + ) -> Result>> { + // The pushed self filter is `Optional(DynamicFilter)`: look through + // the wrapper to recover the dynamic filter this join must update. + let predicate = match predicate.downcast_ref::() { + Some(optional) => Arc::clone(optional.inner()), + None => Arc::clone(predicate), + }; + let Ok(dynamic_filter) = Arc::downcast::(predicate) + else { + return Ok(None); + }; + let has_consumer = dynamic_filter + .expression_id() + .map(|id| plan_contains_expression_id(&self.right, id)) + .transpose()? + .unwrap_or(false); + Ok(has_consumer.then_some(dynamic_filter)) + } + + /// Set one dynamic filter on this hash join, which holds both the bounds + /// and the membership check. See [`Self::set_dynamic_filters`]. + fn set_dynamic_filter(self, filter: Arc) -> Result { + self.set_dynamic_filters(Some(filter), None) + } + + /// Set the membership and bounds dynamic filters on this hash join (see + /// [`HashJoinExecDynamicFilter`]), resetting any internal state that + /// depends on existing ones and validating that the filters' children + /// reference valid columns in the probe (right) side's schema. Without a + /// bounds filter, the membership filter also holds the bounds. If both are + /// `None`, the join has no dynamic filter. /// - /// Only used to restore the filter when decoding a serialized plan: every - /// other code path installs the filter in + /// Only used to restore the filters when decoding a serialized plan: every + /// other code path installs the filters in /// [`ExecutionPlan::handle_child_pushdown_result`]. - fn set_dynamic_filter( + fn set_dynamic_filters( mut self, - filter: Arc, + membership: Option>, + bounds: Option>, ) -> Result { let probe_schema = self.right.schema(); - for child in filter.children() { - child.data_type(&probe_schema)?; + for filter in membership.iter().chain(bounds.iter()) { + for child in filter.children() { + child.data_type(&probe_schema)?; + } } - self.dynamic_filter = Some(HashJoinExecDynamicFilter { - filter, - // Initialize with an empty accumulator which will be lazily populated - // during execution. - build_accumulator: OnceLock::new(), - }); + // The accumulator is empty and lazily populated during execution. + self.dynamic_filter = (membership.is_some() || bounds.is_some()) + .then(|| HashJoinExecDynamicFilter::new(membership, bounds)); Ok(self) } @@ -1628,19 +1707,16 @@ impl ExecutionPlan for HashJoinExec { .filter .iter() .map(|filter| Arc::clone(filter.expression())); - let dynamic_filter = self.dynamic_filter.iter().map(|dynamic_filter| { - Arc::::clone(&dynamic_filter.filter) - as Arc - }); - crate::apply_expression_roots(join_keys.chain(filter).chain(dynamic_filter), f) + let dynamic_filters = self.dynamic_expressions_produced(); + crate::apply_expression_roots(join_keys.chain(filter).chain(dynamic_filters), f) } fn dynamic_expressions_produced(&self) -> Vec> { self.dynamic_filter .iter() - .map(|dynamic_filter| { - Arc::::clone(&dynamic_filter.filter) - as Arc + .flat_map(HashJoinExecDynamicFilter::filters) + .map(|filter| { + Arc::::clone(filter) as Arc }) .collect() } @@ -1724,7 +1800,8 @@ impl ExecutionPlan for HashJoinExec { let build_accumulator = enable_dynamic_filter_pushdown .then(|| { self.dynamic_filter.as_ref().map(|df| { - let filter = Arc::clone(&df.filter); + let membership_filter = df.membership.as_ref().map(Arc::clone); + let bounds_filter = df.bounds.as_ref().map(Arc::clone); let on_right = self .on .iter() @@ -1735,7 +1812,8 @@ impl ExecutionPlan for HashJoinExec { self.mode, self.left.as_ref(), self.right.as_ref(), - filter, + membership_filter, + bounds_filter, on_right, repartition_random_state, self.null_equality, @@ -1812,6 +1890,21 @@ impl ExecutionPlan for HashJoinExec { let batch_size = context.session_config().batch_size(); + // The join measures the work that its dynamic filters save for each + // row that they remove (see `RemovedRowWork`). + let removed_row_work = self + .dynamic_filter + .as_ref() + .filter(|_| enable_dynamic_filter_pushdown) + .map(|df| { + df.membership + .iter() + .chain(df.bounds.iter()) + .map(|filter| Arc::clone(filter.removed_row_work())) + .collect() + }) + .unwrap_or_default(); + // we have the batches and the hash map with their keys. We can how create a stream // over the right that uses this information to issue new batches. let right_stream = self.right.execute(partition, context)?; @@ -1831,27 +1924,30 @@ impl ExecutionPlan for HashJoinExec { .map(|(_, right_expr)| Arc::clone(right_expr)) .collect::>(); - Ok(Box::pin(HashJoinStream::new( - partition, - self.schema(), - on_right, - self.filter.clone(), - self.join_type, - right_stream, - self.random_state.random_state().clone(), - join_metrics, - column_indices_after_projection, - self.null_equality, - HashJoinStreamState::WaitBuildSide, - BuildSide::Initial(BuildSideInitialState { left_fut }), - batch_size, - vec![], - self.right.output_ordering().is_some(), - build_accumulator, - self.mode, - null_aware, - self.fetch, - ))) + Ok(Box::pin( + HashJoinStream::new( + partition, + self.schema(), + on_right, + self.filter.clone(), + self.join_type, + right_stream, + self.random_state.random_state().clone(), + join_metrics, + column_indices_after_projection, + self.null_equality, + HashJoinStreamState::WaitBuildSide, + BuildSide::Initial(BuildSideInitialState { left_fut }), + batch_size, + vec![], + self.right.output_ordering().is_some(), + build_accumulator, + self.mode, + null_aware, + self.fetch, + ) + .with_removed_row_work(removed_row_work), + )) } fn metrics(&self) -> Option { @@ -2042,9 +2138,36 @@ impl ExecutionPlan for HashJoinExec { && self.dynamic_filter.is_none() && self.allow_join_dynamic_filter_pushdown(config) { - // Add actual dynamic filter to right side (probe side) - let dynamic_filter = Self::create_dynamic_filter(&self.on); - right_child = right_child.with_self_filter(dynamic_filter); + // Add actual dynamic filters to right side (probe side). + // + // A partitioned join pushes two filters: first the build-side + // bounds, then the membership check. Two filters, not one filter + // with an `AND`, let each consumer use them independently. The + // bounds are the union of the bounds of all partitions and do not + // need the routing `CASE` of the membership check, so the pruning + // code can use them. A collect-left join pushes one filter that + // holds both (`bounds AND membership`): it has no routing `CASE`, + // so the pruning code can already use its bounds. See + // `SharedBuildAccumulator` for the filter contents. + // + // `handle_child_pushdown_result` relies on this order. + // + // The join itself removes the rows that do not match, so the + // filters are not needed for correctness: mark the pushed copies + // as optional. + let dynamic_filter = || { + Arc::new(OptionalFilterPhysicalExpr::new( + Self::create_dynamic_filter(&self.on), + )) as Arc + }; + right_child = if self.mode == PartitionMode::Partitioned { + right_child.with_self_filters(vec![ + dynamic_filter(), // bounds + dynamic_filter(), // membership + ]) + } else { + right_child.with_self_filter(dynamic_filter()) + }; } Ok(FilterDescription::new() @@ -2061,43 +2184,38 @@ impl ExecutionPlan for HashJoinExec { let mut result = FilterPushdownPropagation::if_any(child_pushdown_result.clone()); assert_eq!(child_pushdown_result.self_filters.len(), 2); // Should always be 2, we have 2 children let right_child_self_filters = &child_pushdown_result.self_filters[1]; // We only push down filters to the right child - // We expect 0 or 1 self filters - if let Some(filter) = right_child_self_filters.first() { - let predicate = Arc::clone(&filter.predicate); - if let Ok(dynamic_filter) = - Arc::downcast::(predicate) - { - // Note that we don't check `PushedDownPredicate::discriminant`: a node - // that replies `PushedDown::No` may still retain the filter for - // statistics pruning, so the reply does not tell us whether anyone will - // actually read the filter. Instead we look for a consumer holding the - // expression in the probe subtree. - // - // `self` here is the join with its post-pushdown children, so anything - // that accepted the filter is already wired into `self.right`. Searching - // from `self` would always find the producer expression pushed by this - // join. This is the last chance to make the decision: the Post-phase - // `FilterPushdown` rule is the final rule that mutates the plan. - let has_consumer = dynamic_filter - .expression_id() - .map(|id| plan_contains_expression_id(&self.right, id)) - .transpose()? - .unwrap_or(false); - if has_consumer { - // Our self filter reached a consumer: rebuild the node holding onto - // the dynamic filter so that `execute` populates it from the build - // side. If it did not, we leave `dynamic_filter` as `None` and skip - // the (not cheap) bounds accumulation entirely. - let new_node = self - .builder() - .with_dynamic_filter(Some(HashJoinExecDynamicFilter { - filter: dynamic_filter, - build_accumulator: OnceLock::new(), - })) - .build_exec()?; - result = result.with_updated_node(new_node); - } + + // We expect 0 self filters, 1 (collect-left: bounds and membership + // in one filter) or 2 (partitioned: the bounds filter, then the + // membership filter). See `gather_filters_for_pushdown`. Keep each + // one that reached a consumer. + let consumed = right_child_self_filters + .iter() + .map(|filter| self.consumed_dynamic_filter(&filter.predicate)) + .collect::>>()?; + let (bounds, membership) = match consumed.as_slice() { + [] => (None, None), + [membership] => (None, membership.clone()), + [bounds, membership] => (bounds.clone(), membership.clone()), + _ => { + return internal_err!( + "HashJoinExec expected 0, 1 or 2 self filters, got {}", + consumed.len() + ); } + }; + if membership.is_some() || bounds.is_some() { + // Our self filters reached a consumer: rebuild the node holding onto + // the dynamic filters so that `execute` populates them from the build + // side. If none did, we leave `dynamic_filter` as `None` and skip + // the (not cheap) bounds accumulation entirely. + let new_node = self + .builder() + .with_dynamic_filter(Some(HashJoinExecDynamicFilter::new( + membership, bounds, + ))) + .build_exec()?; + result = result.with_updated_node(new_node); } Ok(result) } @@ -2190,14 +2308,21 @@ impl ExecutionPlan for HashJoinExec { .map(|f| crate::joins::proto::join_filter_to_proto(f, ctx)) .transpose()?; - let dynamic_filter = dynamic_filter - .as_ref() - .map(|df| { - let df_expr: Arc = - Arc::clone(&df.filter) as Arc; - ctx.encode_expr(&df_expr) - }) - .transpose()?; + let encode_dynamic_filter = |filter: Option<&Arc>| { + filter + .map(|filter| { + ctx.encode_expr(&(Arc::clone(filter) as Arc)) + }) + .transpose() + }; + let dynamic_filter_bounds = encode_dynamic_filter( + dynamic_filter.as_ref().and_then(|df| df.bounds.as_ref()), + )?; + let dynamic_filter = encode_dynamic_filter( + dynamic_filter + .as_ref() + .and_then(|df| df.membership.as_ref()), + )?; Ok(Some(protobuf::PhysicalPlanNode { physical_plan_type: Some( @@ -2225,6 +2350,7 @@ impl ExecutionPlan for HashJoinExec { null_aware: *null_aware, dynamic_filter, fetch: fetch.map(|f| f as u64), + dynamic_filter_bounds, }, )), ), @@ -2263,6 +2389,7 @@ impl HashJoinExec { null_aware, dynamic_filter, fetch, + dynamic_filter_bounds, } = &**hashjoin; let left = ctx.decode_required_child(left.as_deref(), "HashJoinExec", "left")?; @@ -2328,7 +2455,7 @@ impl HashJoinExec { .map(|fetch| usize_from_wire(fetch, "HashJoinExec", "fetch")) .transpose()?; - let mut hash_join = HashJoinExecBuilder::new(left, right, on, join_type) + let hash_join = HashJoinExecBuilder::new(left, right, on, join_type) .with_filter(filter) .with_projection(projection) .with_partition_mode(partition_mode) @@ -2337,20 +2464,35 @@ impl HashJoinExec { .with_fetch(fetch) .build()?; - if let Some(dynamic_filter_proto) = dynamic_filter { - // The dynamic filter is a `DynamicFilterPhysicalExpr` over the probe - // (right) side; decode against the right schema then downcast. - let dynamic_filter_expr = - ctx.decode_expr(dynamic_filter_proto, right_schema.as_ref())?; - let df = (dynamic_filter_expr as Arc) - .downcast::() - .map_err(|_| { - internal_datafusion_err!( - "HashJoinExec dynamic_filter did not decode to a DynamicFilterPhysicalExpr" - ) - })?; - hash_join = hash_join.set_dynamic_filter(df)?; - } + // The dynamic filters are `DynamicFilterPhysicalExpr`s over the probe + // (right) side; decode against the right schema then downcast. A plan + // without `dynamic_filter_bounds` restores one filter that holds both + // the bounds and the membership check. + let decode_dynamic_filter = |proto: Option<&protobuf::PhysicalExprNode>, + field: &str| + -> Result< + Option>, + > { + proto + .map(|proto| { + let expr = ctx.decode_expr(proto, right_schema.as_ref())?; + (expr as Arc) + .downcast::() + .map_err(|_| { + internal_datafusion_err!( + "HashJoinExec {field} did not decode to a DynamicFilterPhysicalExpr" + ) + }) + }) + .transpose() + }; + let membership = + decode_dynamic_filter(dynamic_filter.as_ref(), "dynamic_filter")?; + let bounds = decode_dynamic_filter( + dynamic_filter_bounds.as_ref(), + "dynamic_filter_bounds", + )?; + let hash_join = hash_join.set_dynamic_filters(membership, bounds)?; Ok(Arc::new(hash_join)) } @@ -2429,6 +2571,7 @@ mod proto_tests { null_aware: false, dynamic_filter: None, fetch: None, + dynamic_filter_bounds: None, } } @@ -3834,10 +3977,10 @@ mod tests { NullEquality::NullEqualsNothing, false, )?; - join.dynamic_filter = Some(HashJoinExecDynamicFilter { - filter: Arc::clone(&dynamic_filter), - build_accumulator: OnceLock::new(), - }); + join.dynamic_filter = Some(HashJoinExecDynamicFilter::new( + Some(Arc::clone(&dynamic_filter)), + None, + )); Ok((join, dynamic_filter)) } @@ -9158,10 +9301,10 @@ mod tests { NullEquality::NullEqualsNull, false, )?; - join.dynamic_filter = Some(HashJoinExecDynamicFilter { - filter: Arc::clone(&dynamic_filter), - build_accumulator: OnceLock::new(), - }); + join.dynamic_filter = Some(HashJoinExecDynamicFilter::new( + Some(Arc::clone(&dynamic_filter)), + None, + )); // (1) Building the hash table must not fail on the dictionary key. let stream = join.execute(0, task_ctx)?; @@ -9191,6 +9334,87 @@ mod tests { Ok(()) } + /// The join measures the work that its dynamic filter saves for each + /// row that the filter removes: its work for each probe row. + #[tokio::test] + async fn test_dynamic_filter_measures_removed_row_work() -> Result<()> { + let task_ctx = Arc::new(TaskContext::default()); + let left = build_table( + ("a1", &vec![1, 2, 3]), + ("b1", &vec![4, 5, 6]), + ("c1", &vec![7, 8, 9]), + ); + // `MIN_OBSERVED_ROWS` probe rows. + let rows = datafusion_physical_expr::filter_stats::MIN_OBSERVED_ROWS as i32; + let values: Vec = (0..rows).collect(); + let right = build_table(("a2", &values), ("b2", &values), ("c2", &values)); + let on = vec![( + Arc::new(Column::new_with_schema("a1", &left.schema())?) as _, + Arc::new(Column::new_with_schema("a2", &right.schema())?) as _, + )]; + let dynamic_filter = HashJoinExec::create_dynamic_filter(&on); + let consumer: Arc = Arc::clone(&dynamic_filter) as _; + // The consumer does not apply the filter here: the join sees all + // probe rows. + let right = Arc::new(FilterExecBuilder::new(consumer, right).build()?); + let mut join = HashJoinExec::try_new( + left, + right, + on, + None, + &JoinType::Inner, + None, + PartitionMode::CollectLeft, + NullEquality::NullEqualsNothing, + false, + )?; + join.dynamic_filter = Some(HashJoinExecDynamicFilter::new( + Some(Arc::clone(&dynamic_filter)), + None, + )); + let work = Arc::clone(dynamic_filter.removed_row_work()); + assert_eq!(work.ns_per_row(), None); + let batches = common::collect(join.execute(0, task_ctx)?).await?; + assert_eq!(batches.iter().map(|b| b.num_rows()).sum::(), 3); + // The filter removed all but 3 probe rows, thus the join saw fewer + // than `MIN_OBSERVED_ROWS` rows: no measurement yet. + assert_eq!(work.ns_per_row(), None); + + // Without a consumer that filters, the join sees all probe rows. + let task_ctx = Arc::new(TaskContext::default()); + let left = build_table( + ("a1", &vec![1, 2, 3]), + ("b1", &vec![4, 5, 6]), + ("c1", &vec![7, 8, 9]), + ); + let right = build_table(("a2", &values), ("b2", &values), ("c2", &values)); + let on = vec![( + Arc::new(Column::new_with_schema("a1", &left.schema())?) as _, + Arc::new(Column::new_with_schema("a2", &right.schema())?) as _, + )]; + let dynamic_filter = HashJoinExec::create_dynamic_filter(&on); + let mut join = HashJoinExec::try_new( + left, + right, + on, + None, + &JoinType::Inner, + None, + PartitionMode::CollectLeft, + NullEquality::NullEqualsNothing, + false, + )?; + join.dynamic_filter = Some(HashJoinExecDynamicFilter::new( + Some(Arc::clone(&dynamic_filter)), + None, + )); + let batches = common::collect(join.execute(0, task_ctx)?).await?; + assert_eq!(batches.iter().map(|b| b.num_rows()).sum::(), 3); + let measured = dynamic_filter.removed_row_work().ns_per_row(); + assert!(measured.is_some_and(|ns| ns > 0.0), "{measured:?}"); + Ok(()) + } + /// The [`PartitionMode::Partitioned`] counterpart of /// [`test_null_equal_dynamic_filter_keeps_probe_nulls_for_build_logical_null`]. /// @@ -9207,6 +9431,12 @@ mod tests { .options_mut() .optimizer .enable_dynamic_filter_pushdown = true; + // Push hash table lookups instead of `InList`s: when every partition pushes + // an `InList`, the routed `CASE` collapses into a single `InList`. + session_config + .options_mut() + .optimizer + .hash_join_inlist_pushdown_max_distinct_values = 0; let task_ctx = Arc::new(TaskContext::default().with_session_config(session_config)); @@ -9268,10 +9498,10 @@ mod tests { NullEquality::NullEqualsNull, false, )?; - join.dynamic_filter = Some(HashJoinExecDynamicFilter { - filter: Arc::clone(&dynamic_filter), - build_accumulator: OnceLock::new(), - }); + join.dynamic_filter = Some(HashJoinExecDynamicFilter::new( + Some(Arc::clone(&dynamic_filter)), + None, + )); let batches = crate::execution_plan::collect(Arc::new(join), task_ctx).await?; @@ -9966,6 +10196,123 @@ mod tests { Ok(()) } + /// A partitioned join pushes two dynamic filters to the probe side (the + /// bounds, then the membership check), each with its own expression id. + /// A collect-left join pushes one dynamic filter that holds both. + #[test] + fn test_pushed_dynamic_filters_by_partition_mode() -> Result<()> { + use datafusion_physical_expr::utils::{as_dynamic_filter, is_optional_filter}; + + let mut config = ConfigOptions::default(); + config.optimizer.enable_join_dynamic_filter_pushdown = true; + + for (mode, expected_filters) in [ + (PartitionMode::Partitioned, 2), + (PartitionMode::CollectLeft, 1), + ] { + let (_, _, on) = build_schema_and_on()?; + let left = build_table(("a1", &vec![1]), ("b1", &vec![1]), ("c1", &vec![1])); + let right = build_table(("a2", &vec![1]), ("b1", &vec![1]), ("c2", &vec![1])); + let join = HashJoinExec::try_new( + left, + right, + on, + None, + &JoinType::Inner, + None, + mode, + NullEquality::NullEqualsNothing, + false, + )?; + let description = join.gather_filters_for_pushdown( + FilterPushdownPhase::Post, + vec![], + &config, + )?; + + // The self filters go to the probe side only. + let self_filters = description.self_filters(); + assert!(self_filters[0].is_empty()); + assert_eq!(self_filters[1].len(), expected_filters, "{mode:?}"); + let ids = self_filters[1] + .iter() + .map(|filter| { + // The pushed self filters are `Optional(DynamicFilter)`. + assert!(is_optional_filter(filter)); + as_dynamic_filter(filter) + .expect("the self filter should be a dynamic filter") + .expression_id() + }) + .collect::>(); + assert_eq!(ids.len(), expected_filters, "{mode:?}"); + } + Ok(()) + } + + /// The join pushes its own dynamic filter as `Optional(DynamicFilter)`. + /// A key transfer keeps the optionality of the parent filter: an optional + /// parent filter stays optional, and a required parent filter stays + /// required. + #[test] + fn test_pushed_dynamic_filters_are_optional() -> Result<()> { + use crate::filter_pushdown::PushedDown; + use datafusion_physical_expr::utils::{as_dynamic_filter, is_optional_filter}; + + let (_, _, on) = build_schema_and_on()?; + let left = build_table(("a1", &vec![1]), ("b1", &vec![1]), ("c1", &vec![1])); + let right = build_table(("a2", &vec![1]), ("b1", &vec![1]), ("c2", &vec![1])); + let join = HashJoinExec::try_new( + left, + right, + on, + None, + &JoinType::Inner, + None, + PartitionMode::CollectLeft, + NullEquality::NullEqualsNothing, + false, + )?; + + // Parent filters over the left join key `b1@1`. + let left_key: Arc = Arc::new(Column::new("b1", 1)); + let parent_dynamic = Arc::new(DynamicFilterPhysicalExpr::new( + vec![Arc::clone(&left_key)], + lit(true), + )); + let optional_parent: Arc = Arc::new( + OptionalFilterPhysicalExpr::new(Arc::clone(&parent_dynamic) as _), + ); + let required_parent: Arc = Arc::clone(&parent_dynamic) as _; + + let mut config = ConfigOptions::default(); + config.optimizer.enable_join_dynamic_filter_pushdown = true; + let description = join.gather_filters_for_pushdown( + FilterPushdownPhase::Post, + vec![optional_parent, required_parent], + &config, + )?; + + // The self filter goes to the probe side only, and it is optional. + let self_filters = description.self_filters(); + assert!(self_filters[0].is_empty()); + assert_eq!(self_filters[1].len(), 1); + assert!(is_optional_filter(&self_filters[1][0])); + assert!(as_dynamic_filter(&self_filters[1][0]).is_some()); + + // Both parent filters are transferred to the probe side. The transfer + // does not add or remove the `Optional` wrapper. + let right_parent_filters = &description.parent_filters()[1]; + for (pushed, expect_optional) in right_parent_filters.iter().zip([true, false]) { + assert!(matches!(pushed.discriminant, PushedDown::Yes)); + assert_eq!(is_optional_filter(&pushed.predicate), expect_optional); + let view = as_dynamic_filter(&pushed.predicate) + .expect("the transferred filter should be a dynamic filter view"); + assert_eq!(view.expression_id(), parent_dynamic.expression_id()); + assert_eq!(view.children()[0].to_string(), "b1@1"); + } + Ok(()) + } + #[test] fn test_swap_inputs_rejects_dynamic_filter() -> Result<()> { let left = build_table( diff --git a/datafusion/physical-plan/src/joins/hash_join/exec/prepared.rs b/datafusion/physical-plan/src/joins/hash_join/exec/prepared.rs index fb46cedf96cf1..b2e3032d5b23b 100644 --- a/datafusion/physical-plan/src/joins/hash_join/exec/prepared.rs +++ b/datafusion/physical-plan/src/joins/hash_join/exec/prepared.rs @@ -91,9 +91,14 @@ impl PreparedHashJoinBuild { "Prepared hash-join build does not match schema, keys or null equality" ); } - if let Some(filter) = &join.dynamic_filter { - let probe_keys = join.on.iter().map(|(_, right)| right); - if !filter.filter.children().into_iter().eq(probe_keys) { + if let Some(dynamic_filter) = &join.dynamic_filter { + // Each dynamic filter of the join (membership and bounds) is on + // the probe keys. + let keys_match = dynamic_filter.filters().all(|filter| { + let probe_keys = join.on.iter().map(|(_, right)| right); + filter.children().into_iter().eq(probe_keys) + }); + if !keys_match { return plan_err!( "Prepared hash-join dynamic filter keys do not match probe keys" ); diff --git a/datafusion/physical-plan/src/joins/hash_join/exec/prepared/tests.rs b/datafusion/physical-plan/src/joins/hash_join/exec/prepared/tests.rs index e83f953a4c9f0..fe3c58ecd7e4e 100644 --- a/datafusion/physical-plan/src/joins/hash_join/exec/prepared/tests.rs +++ b/datafusion/physical-plan/src/joins/hash_join/exec/prepared/tests.rs @@ -132,10 +132,7 @@ fn with_probe_filter(join: HashJoinExec) -> Result { )?); join.builder() .with_new_children(vec![Arc::clone(join.left()), probe])? - .with_dynamic_filter(Some(HashJoinExecDynamicFilter { - filter, - build_accumulator: OnceLock::new(), - })) + .with_dynamic_filter(Some(HashJoinExecDynamicFilter::new(Some(filter), None))) .build() } @@ -378,7 +375,8 @@ async fn prepared_build_reuses_data_with_independent_dynamic_filters() -> Result ); assert_eq!(metrics.output_rows(), Some(1)); let dynamic = plan.dynamic_filter.as_ref().unwrap(); - assert!(futures::poll!(Box::pin(dynamic.filter.wait_complete())).is_ready()); + let filter = dynamic.membership.as_ref().unwrap(); + assert!(futures::poll!(Box::pin(filter.wait_complete())).is_ready()); } assert_eq!(pool.reserved(), bytes); drop(plans); @@ -896,7 +894,13 @@ async fn prepared_byte_keys_use_hash_membership() -> Result<()> { .build()?; let output = run(&task).await?; assert_eq!(output.iter().map(RecordBatch::num_rows).sum::(), 1); - let filter = &task.dynamic_filter.as_ref().unwrap().filter; + let filter = task + .dynamic_filter + .as_ref() + .unwrap() + .membership + .as_ref() + .unwrap(); assert!(futures::poll!(Box::pin(filter.wait_complete())).is_ready()); drop(task); assert_eq!(pool.reserved(), 0); diff --git a/datafusion/physical-plan/src/joins/hash_join/inlist_builder.rs b/datafusion/physical-plan/src/joins/hash_join/inlist_builder.rs index 2fc3201c6363f..83a4a32c73ae4 100644 --- a/datafusion/physical-plan/src/joins/hash_join/inlist_builder.rs +++ b/datafusion/physical-plan/src/joins/hash_join/inlist_builder.rs @@ -17,10 +17,13 @@ //! Utilities for building InList expressions from hash join build side data +use std::collections::HashSet; use std::sync::Arc; -use arrow::array::{ArrayRef, StructArray}; +use arrow::array::{ArrayRef, StructArray, UInt64Array}; +use arrow::compute::take; use arrow::datatypes::{Field, FieldRef, Fields}; +use arrow::row::{Row, RowConverter, SortField}; use arrow_schema::DataType; use datafusion_common::Result; @@ -77,6 +80,48 @@ pub(super) fn build_struct_inlist_values( Ok(Some(source_array)) } +/// Removes duplicate entries from an `IN` list value array and sorts the +/// remaining entries. +/// +/// Equality and order are those of the arrow row format produced by +/// [`RowConverter`] with default [`SortField`] options (ascending, NULLs +/// first). This works for single values, for the struct values of +/// multi-column keys and for dictionaries. NULLs compare equal to each other, +/// so many NULLs collapse into one NULL. That does not change the result of +/// `IN`: one NULL in the list gives the same three-valued result as many. +/// +/// Sorting makes the list independent of the order in which build rows +/// arrive, so the displayed filter is deterministic. +/// +/// Returns the input unchanged when it has fewer than two entries, or when +/// its type cannot be row-encoded (deduplication is an optimization, not a +/// requirement). +pub(super) fn sorted_distinct_inlist_values(values: ArrayRef) -> Result { + if values.len() < 2 { + return Ok(values); + } + + let sort_field = SortField::new(values.data_type().clone()); + if !RowConverter::supports_fields(std::slice::from_ref(&sort_field)) { + return Ok(values); + } + + let converter = RowConverter::new(vec![sort_field])?; + let rows = converter.convert_columns(std::slice::from_ref(&values))?; + + // Select the first entry of each distinct value by index. The entries are + // taken from the input instead of decoded from the rows, so that the type + // of the input (for example a dictionary) is kept. + let mut seen: HashSet = HashSet::with_capacity(values.len()); + let mut indices: Vec = (0..rows.num_rows()) + .filter(|&idx| seen.insert(rows.row(idx))) + .collect(); + indices.sort_unstable_by(|&a, &b| rows.row(a).cmp(&rows.row(b))); + + let indices = UInt64Array::from_iter_values(indices.into_iter().map(|i| i as u64)); + Ok(take(values.as_ref(), &indices, None)?) +} + #[cfg(test)] mod tests { use super::*; @@ -155,4 +200,25 @@ mod tests { assert_eq!(result.len(), 3); assert_eq!(result.data_type(), dict_array.data_type()); } + + #[test] + fn test_sorted_distinct_inlist_values_keeps_dictionary_type() { + let keys = Int8Array::from(vec![1i8, 0, 1, 1, 0]); + let values = Arc::new(StringArray::from(vec!["foo", "bar"])); + let dict_array = Arc::new(DictionaryArray::new(keys, values)) as ArrayRef; + + let result = sorted_distinct_inlist_values(Arc::clone(&dict_array)).unwrap(); + + assert_eq!(result.data_type(), dict_array.data_type()); + assert_eq!(result.len(), 2); + // "bar" sorts before "foo". + assert_eq!( + arrow::util::display::array_value_to_string(&result, 0).unwrap(), + "bar" + ); + assert_eq!( + arrow::util::display::array_value_to_string(&result, 1).unwrap(), + "foo" + ); + } } diff --git a/datafusion/physical-plan/src/joins/hash_join/mod.rs b/datafusion/physical-plan/src/joins/hash_join/mod.rs index 7b50435e18792..2e177a5c734ef 100644 --- a/datafusion/physical-plan/src/joins/hash_join/mod.rs +++ b/datafusion/physical-plan/src/joins/hash_join/mod.rs @@ -20,6 +20,7 @@ pub use exec::{HashJoinExec, HashJoinExecBuilder, PreparedHashJoinBuild}; pub use partitioned_hash_eval::{HashExpr, HashTableLookupExpr, SeededRandomState}; +mod bounds_union; mod exec; mod inlist_builder; mod partitioned_hash_eval; diff --git a/datafusion/physical-plan/src/joins/hash_join/shared_bounds.rs b/datafusion/physical-plan/src/joins/hash_join/shared_bounds.rs index 62087c14c5179..cb41a180d645a 100644 --- a/datafusion/physical-plan/src/joins/hash_join/shared_bounds.rs +++ b/datafusion/physical-plan/src/joins/hash_join/shared_bounds.rs @@ -26,13 +26,20 @@ use crate::ExecutionPlanProperties; use crate::Partitioning; use crate::joins::Map; use crate::joins::PartitionMode; +use crate::joins::hash_join::bounds_union::{ + MAX_RANGES_PER_COLUMN, create_merged_bounds_predicate, merge_partition_bounds, + range_predicate, +}; use crate::joins::hash_join::exec::HASH_JOIN_SEED; -use crate::joins::hash_join::inlist_builder::build_struct_fields; +use crate::joins::hash_join::inlist_builder::{ + build_struct_fields, sorted_distinct_inlist_values, +}; use crate::joins::hash_join::partitioned_hash_eval::{ HashExpr, HashTableLookupExpr, SeededRandomState, }; use crate::repartition::RangeExpr; use arrow::array::ArrayRef; +use arrow::compute::concat; use arrow::datatypes::{DataType, Field, Schema}; use datafusion_common::config::ConfigOptions; use datafusion_common::{ @@ -155,40 +162,20 @@ fn create_bounds_predicate( on_right: &[PhysicalExprRef], bounds: &PartitionBounds, ) -> Option> { - let mut column_predicates = Vec::new(); - - for (col_idx, right_expr) in on_right.iter().enumerate() { - if let Some(column_bounds) = bounds.get_column_bounds(col_idx) { - // Create predicate: col >= min AND col <= max - let min_expr = Arc::new(BinaryExpr::new( - Arc::clone(right_expr), - Operator::GtEq, - lit(column_bounds.min.clone()), - )) as Arc; - let max_expr = Arc::new(BinaryExpr::new( - Arc::clone(right_expr), - Operator::LtEq, - lit(column_bounds.max.clone()), - )) as Arc; - let range_expr = Arc::new(BinaryExpr::new(min_expr, Operator::And, max_expr)) - as Arc; - column_predicates.push(range_expr); - } - } - - if column_predicates.is_empty() { - None - } else { - Some( - column_predicates - .into_iter() - .reduce(|acc, pred| { - Arc::new(BinaryExpr::new(acc, Operator::And, pred)) - as Arc - }) - .unwrap(), - ) - } + on_right + .iter() + .enumerate() + .filter_map(|(col_idx, right_expr)| { + let column_bounds = bounds.get_column_bounds(col_idx)?; + Some(range_predicate( + right_expr, + &column_bounds.min, + &column_bounds.max, + )) + }) + .reduce(|acc, pred| { + Arc::new(BinaryExpr::new(acc, Operator::And, pred)) as Arc + }) } /// Combines a membership predicate and a bounds predicate with logical AND. @@ -253,8 +240,22 @@ pub(crate) struct SharedBuildAccumulator { /// result, then broadcasts), so late subscribers simply re-check the /// state under the mutex and return immediately. completion_notify: Notify, - /// Dynamic filter for pushdown to probe side - dynamic_filter: Arc, + /// Dynamic filter for the membership check, pushed to the probe side. + /// + /// When [`Self::bounds_filter`] is `None`, this filter also holds the + /// build-side bounds (`bounds AND membership`), so that it is complete on + /// its own. `None` when no probe-side node holds it. + membership_filter: Option>, + /// Dynamic filter for the build-side bounds (`col >= min AND col <= max`), + /// pushed to the probe side separately from [`Self::membership_filter`]. + /// + /// Two filters, not one filter with an `AND`, let each consumer use them + /// independently. In partitioned mode the bounds are the union of the + /// bounds of all partitions (see [`merge_partition_bounds`]), which does + /// not need the routing `CASE` of the membership check, so the pruning + /// code can use them. `None` when no probe-side node holds it, and for a + /// collect-left join, which pushes one filter. + bounds_filter: Option>, /// Right side join expressions needed for creating filter expressions on_right: Vec, /// Random state for partitioning (RepartitionExec's hash function with 0,0,0,0 seeds) @@ -272,6 +273,15 @@ pub(crate) struct SharedBuildAccumulator { null_aware: bool, } +/// Ceiling on the size of the deduplicated union `InList` array that +/// [`SharedBuildAccumulator::union_inlist_membership`] will push. +/// +/// Each partition's list is independently capped by +/// `hash_join_inlist_pushdown_max_size`, so without a combined cap the union +/// grows with the partition count. Past this size, keeping the routed `CASE` +/// (where each probe row only probes one list) is the cheaper shape. +const MAX_UNIONED_INLIST_BYTES: usize = 1024 * 1024; + /// Strategy for filter pushdown (decided at collection time) #[derive(Clone)] pub(crate) enum PushdownStrategy { @@ -308,6 +318,17 @@ struct PartitionData { keys_have_null: bool, } +/// The new expressions for the dynamic filters, built from finalized build +/// data. `None` leaves a filter unchanged. +struct BuildFilterExprs { + /// For [`SharedBuildAccumulator::bounds_filter`]. + bounds: Option>, + /// For [`SharedBuildAccumulator::membership_filter`]. + membership: Option>, + /// Whether any build key is NULL (or can be, for a canceled partition). + keys_have_null: bool, +} + /// Build-side data organized by partition mode enum AccumulatedBuildData { Partitioned { @@ -376,7 +397,8 @@ impl SharedBuildAccumulator { partition_mode: PartitionMode, left_child: &dyn ExecutionPlan, right_child: &dyn ExecutionPlan, - dynamic_filter: Arc, + membership_filter: Option>, + bounds_filter: Option>, on_right: Vec, repartition_random_state: SeededRandomState, null_equality: NullEquality, @@ -431,7 +453,8 @@ impl SharedBuildAccumulator { completion: CompletionState::Pending, }), completion_notify: Notify::new(), - dynamic_filter, + membership_filter, + bounds_filter, on_right, repartition_random_state, probe_schema: right_child.schema(), @@ -583,7 +606,9 @@ impl SharedBuildAccumulator { fn finish(&self, finalize_input: FinalizeInput) { let result = self.build_filter(finalize_input).map_err(Arc::new); - self.dynamic_filter.mark_complete(); + for filter in self.filters() { + filter.mark_complete(); + } let mut guard = self.inner.lock(); guard.completion = CompletionState::Ready(result); @@ -609,19 +634,50 @@ impl SharedBuildAccumulator { } } + /// The dynamic filters that this accumulator updates. + fn filters(&self) -> impl Iterator> { + self.membership_filter + .iter() + .chain(self.bounds_filter.iter()) + } + fn build_filter(&self, finalize_input: FinalizeInput) -> Result<()> { - match finalize_input { + let exprs = match finalize_input { FinalizeInput::CollectLeft(partition) => { - self.build_collect_left_filter(partition) + self.collect_left_filter_exprs(partition)? } FinalizeInput::Partitioned(partitions) => { - self.build_partitioned_filter(partitions) + self.partitioned_filter_exprs(&partitions)? } + }; + self.update_filters(exprs) + } + + /// Updates each dynamic filter that has a new expression. + fn update_filters(&self, exprs: BuildFilterExprs) -> Result<()> { + let BuildFilterExprs { + bounds, + membership, + keys_have_null, + } = exprs; + if let (Some(filter), Some(expr)) = (&self.bounds_filter, bounds) { + filter.update(self.preserve_probe_nulls(expr, keys_have_null)?)?; } + if let (Some(filter), Some(expr)) = (&self.membership_filter, membership) { + filter.update(self.preserve_probe_nulls(expr, keys_have_null)?)?; + } + Ok(()) } - /// Builds the single global filter used by a collect-left join. - fn build_collect_left_filter(&self, partition: PartitionStatus) -> Result<()> { + /// Builds the global filters used by a collect-left join. + /// + /// With a bounds filter, the bounds and the membership check go to + /// separate filters. Without one, the membership filter gets + /// `bounds AND membership`. + fn collect_left_filter_exprs( + &self, + partition: PartitionStatus, + ) -> Result { match partition { PartitionStatus::Reported(PartitionData { bounds, @@ -635,15 +691,22 @@ impl SharedBuildAccumulator { self.probe_schema.as_ref(), )?; let bounds_expr = create_bounds_predicate(&self.on_right, &bounds); - - if let Some(filter_expr) = - combine_membership_and_bounds(membership_expr, bounds_expr) - { - self.dynamic_filter.update( - self.preserve_probe_nulls(filter_expr, keys_have_null)?, - )?; - } - Ok(()) + Ok(if self.bounds_filter.is_some() { + BuildFilterExprs { + bounds: bounds_expr, + membership: membership_expr, + keys_have_null, + } + } else { + BuildFilterExprs { + bounds: None, + membership: combine_membership_and_bounds( + membership_expr, + bounds_expr, + ), + keys_have_null, + } + }) } PartitionStatus::Pending => datafusion_common::internal_err!( "attempted to finalize collect-left dynamic filter without reported build data" @@ -654,47 +717,50 @@ impl SharedBuildAccumulator { } } - /// Builds one routed probe-side filter from finalized partitioned build data. - /// Empty partitions reject their routed rows, while canceled partitions stay - /// permissive because their build contents are unknown. - fn build_partitioned_filter(&self, partitions: Vec) -> Result<()> { - let mut partition_filters = Vec::with_capacity(partitions.len()); + /// Builds the probe-side filters from finalized partitioned build data. + /// + /// The bounds filter gets the union of the bounds of all partitions (see + /// [`merge_partition_bounds`]). The membership filter gets a `CASE` that + /// routes each probe row to the membership check of its partition. Empty + /// partitions reject their routed rows, while canceled partitions stay + /// permissive because their build contents are unknown. When every + /// non-empty partition pushes an `InList`, the routed `CASE` is replaced by + /// one `InList` over the union of the lists (see + /// [`Self::union_inlist_membership`]). + /// + /// The per-partition bounds reject no row that the membership check of the + /// partition accepts, so when the bounds filter holds the union, the + /// membership check does not repeat them. The bounds stay in the `CASE` + /// (`bounds_i AND membership_i`) when there is no bounds filter, or when + /// the union cannot describe the build side: a canceled partition can hold + /// any key, and without usable bounds there is nothing to hoist. + fn partitioned_filter_exprs( + &self, + partitions: &[PartitionStatus], + ) -> Result { let mut real_partition_ids = Vec::new(); let mut empty_partition_ids = Vec::new(); + let mut real_partition_bounds = Vec::new(); let mut has_canceled_unknown = false; let mut keys_have_null = false; - for (partition_id, partition) in partitions.into_iter().enumerate() { + for (partition_id, partition) in partitions.iter().enumerate() { match partition { PartitionStatus::Reported(PartitionData { pushdown: PushdownStrategy::Empty, .. - }) => { - empty_partition_ids.push(partition_id); - partition_filters.push(lit(false)); - } + }) => empty_partition_ids.push(partition_id), PartitionStatus::Reported(PartitionData { bounds, - pushdown, keys_have_null: partition_keys_have_null, + .. }) => { real_partition_ids.push(partition_id); + real_partition_bounds.push(bounds); keys_have_null |= partition_keys_have_null; - let membership_expr = create_membership_predicate( - &self.on_right, - pushdown, - &HASH_JOIN_SEED, - self.probe_schema.as_ref(), - )?; - let bounds_expr = create_bounds_predicate(&self.on_right, &bounds); - let then_expr = - combine_membership_and_bounds(membership_expr, bounds_expr) - .unwrap_or_else(|| lit(true)); - partition_filters.push(then_expr); } PartitionStatus::CanceledUnknown => { has_canceled_unknown = true; - partition_filters.push(lit(true)); // A canceled partition's build content is unknown, so it // may hold a NULL key. keys_have_null = true; @@ -707,95 +773,259 @@ impl SharedBuildAccumulator { } } - let all_partitions_canceled = has_canceled_unknown - && real_partition_ids.is_empty() - && empty_partition_ids.is_empty(); - let all_partitions_empty = !has_canceled_unknown && real_partition_ids.is_empty(); - let one_non_empty_partition = - !has_canceled_unknown && real_partition_ids.len() == 1; - - let filter_expr = if all_partitions_canceled { - // No build data is known, so filtering any probe row could discard a match. - lit(true) - } else if all_partitions_empty { + if !has_canceled_unknown && real_partition_ids.is_empty() { // No build row exists, so no probe row can match. - lit(false) - } else if one_non_empty_partition { - // Only one build partition contains rows, so its filter covers every - // possible probe match without routing. - Arc::clone(&partition_filters[real_partition_ids[0]]) + return Ok(BuildFilterExprs { + bounds: Some(lit(false)), + membership: Some(lit(false)), + keys_have_null, + }); + } + if has_canceled_unknown + && real_partition_ids.is_empty() + && empty_partition_ids.is_empty() + { + // No build data is known, so filtering any probe row could discard + // a match. + return Ok(BuildFilterExprs { + bounds: None, + membership: Some(lit(true)), + keys_have_null, + }); + } + + // The union of the bounds describes every build row only when no + // partition is canceled. + let union_bounds = if has_canceled_unknown { + None } else { - // Builds the shared sparse `CASE` for partition filter routing. - // Without cancellation, omitted branches are known empty and safely fall - // through to `ELSE false`. With cancellation, omitted canceled partitions - // have unknown contents and must fall through to `ELSE true`, so known-empty - // partitions are emitted explicitly as false branches. - let mut branches = if has_canceled_unknown { - empty_partition_ids - .iter() - .map(|&partition_id| (lit(partition_id as u64), lit(false))) - .collect::>() + // Range partitions hold disjoint key ranges, so keep them apart. + // With hash partitioning, the gaps between the ranges of the + // partitions are random: one range per column is enough. + let max_ranges_per_column = if self.probe_range_partitioning.is_some() { + MAX_RANGES_PER_COLUMN } else { - vec![] + 1 }; - branches.extend(real_partition_ids.iter().map(|&partition_id| { - ( - lit(partition_id as u64), - Arc::clone(&partition_filters[partition_id]), - ) - })); - - let routing_expr = if let Some(range_partitioning) = - &self.probe_range_partitioning - { - // Routes probe rows with the partition id selected by [`RangeExpr`]. - // CASE range_partition(keys) - // WHEN empty_partition_id THEN false -- only when cancellation exists - // WHEN real_partition_id THEN F(real_partition_id) - // ... - // ELSE has_canceled_unknown - // END - assert_or_internal_err!( - partition_filters.len() == range_partitioning.partition_count(), - "Dynamic filter partition count {} does not match Range partition count {}", - partition_filters.len(), - range_partitioning.partition_count() - ); - Arc::new(RangeExpr::try_new_with_schema( - self.on_right.clone(), - range_partitioning, - &self.probe_schema, - )?) as Arc + create_merged_bounds_predicate( + &self.on_right, + &merge_partition_bounds( + self.on_right.len(), + &real_partition_bounds, + max_ranges_per_column, + ), + ) + }; + let hoist_bounds = self.bounds_filter.is_some() && union_bounds.is_some(); + + let membership = if self.membership_filter.is_none() { + None + } else if let Some(in_list) = self.union_inlist_membership(partitions)? { + // The collapsed `InList` has no per-partition bounds. Without a + // bounds filter, put the union of the bounds before it: the + // pruning code can use the range when the list has more than + // `max_in_list_size` entries. + if hoist_bounds { + Some(in_list) } else { - // Routes probe rows with the same `hash(keys) % partition_count` expression used - // by Hash repartitioning. - // CASE hash(keys) % partition_count - // WHEN empty_partition_id THEN false -- only when cancellation exists - // WHEN real_partition_id THEN F(real_partition_id) - // ... - // ELSE has_canceled_unknown - // END - let routing_hash_expr = Arc::new(HashExpr::new( - self.on_right.clone(), - self.repartition_random_state.clone(), - "hash_repartition".to_string(), - )) as Arc; - Arc::new(BinaryExpr::new( - routing_hash_expr, - Operator::Modulo, - lit(partition_filters.len() as u64), - )) as Arc - }; - - Arc::new(CaseExpr::try_new( - Some(routing_expr), - branches, - Some(lit(has_canceled_unknown)), + combine_membership_and_bounds(Some(in_list), union_bounds.clone()) + } + } else if !has_canceled_unknown && real_partition_ids.len() == 1 { + // Only one build partition contains rows, so its filter covers + // every possible probe match without routing. + Some( + self.partition_filter(&partitions[real_partition_ids[0]], !hoist_bounds)?, + ) + } else { + Some(self.routed_membership( + partitions, + &real_partition_ids, + &empty_partition_ids, + has_canceled_unknown, + !hoist_bounds, )?) }; - self.dynamic_filter - .update(self.preserve_probe_nulls(filter_expr, keys_have_null)?) + Ok(BuildFilterExprs { + bounds: if hoist_bounds { union_bounds } else { None }, + membership, + keys_have_null, + }) + } + + /// The filter for the probe rows that route to one non-empty partition: + /// its membership check, with its bounds before it if `with_bounds`. + fn partition_filter( + &self, + partition: &PartitionStatus, + with_bounds: bool, + ) -> Result> { + let PartitionStatus::Reported(PartitionData { + bounds, pushdown, .. + }) = partition + else { + return datafusion_common::internal_err!( + "attempted to build the dynamic filter of a partition without reported build data" + ); + }; + let membership_expr = create_membership_predicate( + &self.on_right, + pushdown.clone(), + &HASH_JOIN_SEED, + self.probe_schema.as_ref(), + )?; + let bounds_expr = if with_bounds { + create_bounds_predicate(&self.on_right, bounds) + } else { + None + }; + Ok(combine_membership_and_bounds(membership_expr, bounds_expr) + .unwrap_or_else(|| lit(true))) + } + + /// Builds the sparse `CASE` that routes each probe row to the filter of its + /// partition. + /// + /// Without cancellation, omitted branches are known empty and safely fall + /// through to `ELSE false`. With cancellation, omitted canceled partitions + /// have unknown contents and must fall through to `ELSE true`, so + /// known-empty partitions are emitted explicitly as false branches. + fn routed_membership( + &self, + partitions: &[PartitionStatus], + real_partition_ids: &[usize], + empty_partition_ids: &[usize], + has_canceled_unknown: bool, + with_bounds: bool, + ) -> Result> { + let mut branches = if has_canceled_unknown { + empty_partition_ids + .iter() + .map(|&partition_id| (lit(partition_id as u64), lit(false))) + .collect::>() + } else { + vec![] + }; + for &partition_id in real_partition_ids { + branches.push(( + lit(partition_id as u64), + self.partition_filter(&partitions[partition_id], with_bounds)?, + )); + } + + let routing_expr = if let Some(range_partitioning) = + &self.probe_range_partitioning + { + // Routes probe rows with the partition id selected by [`RangeExpr`]. + // CASE range_partition(keys) + // WHEN empty_partition_id THEN false -- only when cancellation exists + // WHEN real_partition_id THEN F(real_partition_id) + // ... + // ELSE has_canceled_unknown + // END + assert_or_internal_err!( + partitions.len() == range_partitioning.partition_count(), + "Dynamic filter partition count {} does not match Range partition count {}", + partitions.len(), + range_partitioning.partition_count() + ); + Arc::new(RangeExpr::try_new_with_schema( + self.on_right.clone(), + range_partitioning, + &self.probe_schema, + )?) as Arc + } else { + // Routes probe rows with the same `hash(keys) % partition_count` + // expression used by Hash repartitioning. + // CASE hash(keys) % partition_count + // WHEN empty_partition_id THEN false -- only when cancellation exists + // WHEN real_partition_id THEN F(real_partition_id) + // ... + // ELSE has_canceled_unknown + // END + let routing_hash_expr = Arc::new(HashExpr::new( + self.on_right.clone(), + self.repartition_random_state.clone(), + "hash_repartition".to_string(), + )) as Arc; + Arc::new(BinaryExpr::new( + routing_hash_expr, + Operator::Modulo, + lit(partitions.len() as u64), + )) as Arc + }; + + Ok(Arc::new(CaseExpr::try_new( + Some(routing_expr), + branches, + Some(lit(has_canceled_unknown)), + )?)) + } + + /// Collapses an all-`InList` partitioned build into one `InList` over the + /// union of the per-partition lists, instead of a `CASE` that routes each + /// probe row to the list of its partition. + /// + /// This is exact, not a relaxation: routing is a deterministic function of + /// the key columns, so every build row with key `K` is in the partition that + /// a probe row with key `K` routes to. A test of `K` against the union thus + /// accepts the same rows as the routed `CASE`. The result does not compute + /// the routing hash for each probe row, and unlike a `CASE`, pruning can use + /// an `InList`. + /// + /// The union is deduplicated and sorted (see + /// [`sorted_distinct_inlist_values`]). The per-partition lists hold one + /// entry per build row, not per distinct key, and the pruning code uses an + /// `InList` only up to `max_in_list_size` entries. + /// + /// Returns `None` when the collapse does not apply: fewer than two + /// partitions have rows, a partition is canceled or pushes a hash table, + /// the lists have different types, or the deduplicated union is larger + /// than [`MAX_UNIONED_INLIST_BYTES`]. + fn union_inlist_membership( + &self, + partitions: &[PartitionStatus], + ) -> Result>> { + let mut arrays: Vec<&ArrayRef> = Vec::with_capacity(partitions.len()); + for partition in partitions { + let PartitionStatus::Reported(PartitionData { pushdown, .. }) = partition + else { + return Ok(None); + }; + let values = match pushdown { + PushdownStrategy::InList(values) => values, + PushdownStrategy::Empty => continue, + PushdownStrategy::Map(_) => return Ok(None), + }; + if arrays + .first() + .is_some_and(|first| first.data_type() != values.data_type()) + { + return Ok(None); + } + arrays.push(values); + } + + // With zero or one non-empty partition, `partitioned_filter_exprs` + // already skips the `CASE`. + if arrays.len() < 2 { + return Ok(None); + } + + // Each partition's list is at most `hash_join_inlist_pushdown_max_size`, + // so the concatenation is at most the partition count times that size. + let union = concat(&arrays.iter().map(|a| a.as_ref()).collect::>())?; + let union = sorted_distinct_inlist_values(union)?; + if union.get_array_memory_size() > MAX_UNIONED_INLIST_BYTES { + return Ok(None); + } + + create_membership_predicate( + &self.on_right, + PushdownStrategy::InList(union), + &HASH_JOIN_SEED, + self.probe_schema.as_ref(), + ) } /// Keeps probe rows with a NULL key when the join semantics need them. @@ -883,7 +1113,8 @@ pub(super) fn make_partitioned_accumulator_for_test( completion: CompletionState::Pending, }), completion_notify: Notify::new(), - dynamic_filter, + membership_filter: Some(dynamic_filter), + bounds_filter: None, on_right: vec![], repartition_random_state: SeededRandomState::with_seed(1), probe_schema, @@ -910,7 +1141,9 @@ pub(super) fn completed_partitions_for_test(acc: &SharedBuildAccumulator) -> usi mod tests { use super::*; - use arrow::array::{ArrayRef, BooleanArray, Float64Array, Int32Array}; + use crate::joins::hash_join::inlist_builder::build_struct_inlist_values; + use crate::joins::join_hash_map::JoinHashMapU32; + use arrow::array::{ArrayRef, BooleanArray, Float64Array, Int32Array, StringArray}; use arrow::compute::SortOptions; use arrow::record_batch::RecordBatch; use datafusion_common::SplitPoint; @@ -948,7 +1181,8 @@ mod tests { completion: CompletionState::Pending, }), completion_notify: Notify::new(), - dynamic_filter, + membership_filter: Some(dynamic_filter), + bounds_filter: None, on_right, repartition_random_state: SeededRandomState::with_seed(1), probe_schema: test_probe_schema(), @@ -985,6 +1219,12 @@ mod tests { PushdownStrategy::InList(Arc::new(Int32Array::from(values.to_vec())) as ArrayRef) } + fn map_pushdown() -> PushdownStrategy { + PushdownStrategy::Map(Arc::new(Map::HashMap(Box::new( + JoinHashMapU32::with_capacity(1), + )))) + } + fn bounds(min: i32, max: i32) -> PartitionBounds { PartitionBounds::new(vec![ColumnBounds::new( ScalarValue::Int32(Some(min)), @@ -1005,7 +1245,9 @@ mod tests { } fn current_expr(acc: &SharedBuildAccumulator) -> PhysicalExprRef { - acc.dynamic_filter + acc.membership_filter + .as_ref() + .unwrap() .current() .expect("dynamic filter current expression should be available") } @@ -1109,7 +1351,11 @@ mod tests { #[test] fn collect_left_empty_build_data_does_not_update_filter() { let acc = make_collect_left_accumulator_for_test(); - let initial_generation = acc.dynamic_filter.snapshot_generation(); + let initial_generation = acc + .membership_filter + .as_ref() + .unwrap() + .snapshot_generation(); acc.build_filter(FinalizeInput::CollectLeft(reported( PushdownStrategy::Empty, @@ -1118,7 +1364,10 @@ mod tests { .unwrap(); assert_eq!( - acc.dynamic_filter.snapshot_generation(), + acc.membership_filter + .as_ref() + .unwrap() + .snapshot_generation(), initial_generation, "empty CollectLeft input must not update with a no-op filter" ); @@ -1142,6 +1391,248 @@ mod tests { assert!(expr.downcast_ref::().is_none()); } + #[test] + fn partitioned_all_inlist_collapses_to_a_single_union_inlist() { + let acc = make_partitioned_expr_accumulator_for_test(3); + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported(in_list(&[1, 4]), bounds(1, 4)), + reported(in_list(&[2, 5]), bounds(2, 5)), + reported(PushdownStrategy::Empty, no_bounds()), + ])) + .unwrap(); + + // Routing is a function of the key, so a probe key can only match the + // list of the partition it routes to: the union is exact, and the `CASE` + // is not necessary. The per-partition bounds are replaced by one range + // that contains all of them. + let expr = current_expr(&acc); + assert_eq!( + expr.to_string(), + "probe_key@0 >= 1 AND probe_key@0 <= 5 AND probe_key@0 IN (SET) ([1, 2, 4, 5])" + ); + let union = binary_expr(&expr).right(); + assert_in_list_column_values(union, "probe_key", 0, &[1, 2, 4, 5]); + } + + #[test] + fn partitioned_union_inlist_drops_duplicates_within_and_across_partitions() { + let acc = make_partitioned_expr_accumulator_for_test(3); + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported(in_list(&[7, 3, 7, 7, 3]), no_bounds()), + reported(in_list(&[5, 5, 3]), no_bounds()), + reported(in_list(&[9, 9]), no_bounds()), + ])) + .unwrap(); + + // The per-partition lists hold one entry per build row. The union holds + // each distinct key once, in sorted order. + let expr = current_expr(&acc); + assert_in_list_column_values(&expr, "probe_key", 0, &[3, 5, 7, 9]); + } + + #[test] + fn partitioned_union_inlist_keeps_one_null_key() { + let acc = null_equal_partitioned_accumulator(2); + let with_nulls = |values: Vec>| { + PushdownStrategy::InList(Arc::new(Int32Array::from(values)) as ArrayRef) + }; + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported_with_null_keys( + with_nulls(vec![Some(2), None, Some(2), None]), + bounds(2, 2), + ), + reported_with_null_keys(with_nulls(vec![None, Some(1)]), bounds(1, 1)), + ])) + .unwrap(); + + // The NULL keys collapse into one NULL, and the filter still keeps the + // probe NULLs that a null-equal join can match. + let expr = current_expr(&acc); + assert_eq!( + expr.to_string(), + "probe_key@0 IS NULL OR probe_key@0 >= 1 AND probe_key@0 <= 2 AND probe_key@0 IN (SET) ([NULL, 1, 2])" + ); + } + + #[test] + fn partitioned_union_inlist_bounds_skip_all_null_partitions() { + let acc = null_equal_partitioned_accumulator(2); + let null_bounds = PartitionBounds::new(vec![ColumnBounds::new( + ScalarValue::Int32(None), + ScalarValue::Int32(None), + )]); + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported(in_list(&[4, 6]), bounds(4, 6)), + reported_with_null_keys( + PushdownStrategy::InList( + Arc::new(Int32Array::from(vec![None, None])) as ArrayRef + ), + null_bounds, + ), + ])) + .unwrap(); + + // A partition with only NULL keys has NULL bounds. It adds nothing to + // the range, because a NULL key cannot pass a range check. + let expr = current_expr(&acc); + assert_eq!( + expr.to_string(), + "probe_key@0 IS NULL OR probe_key@0 >= 4 AND probe_key@0 <= 6 AND probe_key@0 IN (SET) ([NULL, 4, 6])" + ); + } + + #[test] + fn partitioned_union_inlist_without_bounds_in_one_partition_has_no_range() { + let acc = make_partitioned_expr_accumulator_for_test(2); + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported(in_list(&[1, 2]), bounds(1, 2)), + reported(in_list(&[3]), no_bounds()), + ])) + .unwrap(); + + // The bounds of the second partition are unknown, so no range covers + // all partitions. + let expr = current_expr(&acc); + assert_in_list_column_values(&expr, "probe_key", 0, &[1, 2, 3]); + } + + #[test] + fn partitioned_multi_column_union_inlist_is_deduplicated_with_bounds() { + let probe_schema = Arc::new(Schema::new(vec![ + Field::new("a", DataType::Int32, false), + Field::new("b", DataType::Utf8, false), + ])); + let on_right: Vec = + vec![Arc::new(Column::new("a", 0)), Arc::new(Column::new("b", 1))]; + let mut acc = make_accumulator_for_test( + AccumulatedBuildData::Partitioned { + partitions: vec![PartitionStatus::Pending; 2], + completed_partitions: 0, + }, + on_right, + ); + acc.probe_schema = probe_schema; + + let struct_list = |a: Vec, b: Vec<&str>| { + PushdownStrategy::InList( + build_struct_inlist_values(&[ + Arc::new(Int32Array::from(a)) as ArrayRef, + Arc::new(StringArray::from(b)) as ArrayRef, + ]) + .unwrap() + .unwrap(), + ) + }; + let two_column_bounds = |a: (i32, i32), b: (&str, &str)| { + PartitionBounds::new(vec![ + ColumnBounds::new( + ScalarValue::Int32(Some(a.0)), + ScalarValue::Int32(Some(a.1)), + ), + ColumnBounds::new(ScalarValue::from(b.0), ScalarValue::from(b.1)), + ]) + }; + + acc.build_filter(FinalizeInput::Partitioned(vec![ + // `(2, x)` is present twice; `(1, x)` and `(1, y)` differ only in `b`. + reported( + struct_list(vec![2, 1, 2, 1], vec!["x", "y", "x", "x"]), + two_column_bounds((1, 2), ("x", "y")), + ), + reported( + struct_list(vec![3, 3], vec!["w", "w"]), + two_column_bounds((3, 3), ("w", "w")), + ), + ])) + .unwrap(); + + // Deduplication is on the whole tuple, and each column gets its own + // range. + let expr = current_expr(&acc); + assert_eq!( + expr.to_string(), + "a@0 >= 1 AND a@0 <= 3 AND b@1 >= w AND b@1 <= y AND struct(a@0, b@1) IN (SET) ([{c0:1,c1:x}, {c0:1,c1:y}, {c0:2,c1:x}, {c0:3,c1:w}])" + ); + } + + #[test] + fn partitioned_mixed_strategies_keep_the_routing_case() { + let acc = make_partitioned_expr_accumulator_for_test(2); + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported(in_list(&[1, 2]), bounds(1, 2)), + reported(map_pushdown(), bounds(3, 4)), + ])) + .unwrap(); + + // One partition needs a hash table lookup, so routing is necessary. + let expr = current_expr(&acc); + assert_eq!(case_expr(&expr).when_then_expr().len(), 2); + } + + #[test] + fn partitioned_canceled_partition_keeps_the_routing_case() { + let acc = make_partitioned_expr_accumulator_for_test(3); + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported(in_list(&[1]), no_bounds()), + reported(in_list(&[2]), no_bounds()), + PartitionStatus::CanceledUnknown, + ])) + .unwrap(); + + // The canceled partition's keys are unknown, so the union is incomplete + // and the rows routed to that partition must stay permissive. + let expr = current_expr(&acc); + let case = case_expr(&expr); + assert_eq!(case.when_then_expr().len(), 2); + assert_literal_bool( + case.else_expr().expect("expected permissive fallback"), + true, + ); + } + + #[test] + fn partitioned_oversized_inlist_union_keeps_the_routing_case() { + let acc = make_partitioned_expr_accumulator_for_test(2); + let half = (MAX_UNIONED_INLIST_BYTES / size_of::() / 2 + 1) as i32; + let low = (0..half).collect::>(); + let high = (half..2 * half).collect::>(); + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported(in_list(&low), no_bounds()), + reported(in_list(&high), no_bounds()), + ])) + .unwrap(); + + let expr = current_expr(&acc); + assert_eq!(case_expr(&expr).when_then_expr().len(), 2); + } + + #[test] + fn partitioned_inlist_union_cap_applies_after_deduplication() { + let acc = make_partitioned_expr_accumulator_for_test(2); + // Before deduplication the two lists are larger than the cap, but they + // hold only two distinct keys. + let half = MAX_UNIONED_INLIST_BYTES / size_of::() / 2 + 1; + let ones = vec![1; half]; + let twos = vec![2; half]; + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported(in_list(&ones), no_bounds()), + reported(in_list(&twos), no_bounds()), + ])) + .unwrap(); + + let expr = current_expr(&acc); + assert_in_list_column_values(&expr, "probe_key", 0, &[1, 2]); + } + #[test] fn partitioned_canceled_unknown_partitions_keep_unknown_routes_permissive() { let acc = make_partitioned_expr_accumulator_for_test(2); @@ -1436,7 +1927,11 @@ mod tests { completion: CompletionState::Pending, }), completion_notify: Notify::new(), - dynamic_filter: Arc::new(DynamicFilterPhysicalExpr::new(vec![], lit(true))), + membership_filter: Some(Arc::new(DynamicFilterPhysicalExpr::new( + vec![], + lit(true), + ))), + bounds_filter: None, on_right, repartition_random_state: SeededRandomState::with_seed(1), probe_schema, @@ -1497,8 +1992,9 @@ mod tests { fn partitioned_null_keys_in_one_partition_widen_whole_routed_filter() { let acc = null_equal_partitioned_accumulator(2); + // One partition pushes a hash table, so the filter keeps the routed `CASE`. acc.build_filter(FinalizeInput::Partitioned(vec![ - reported(in_list(&[1]), no_bounds()), + reported(map_pushdown(), no_bounds()), reported_with_null_keys(in_list(&[2]), no_bounds()), ])) .unwrap(); @@ -1518,6 +2014,28 @@ mod tests { ); } + /// The collapsed `InList` must be widened by the NULL flag of every partition, + /// not only of the first one. + #[test] + fn partitioned_null_keys_in_one_partition_widen_union_inlist() { + let acc = null_equal_partitioned_accumulator(2); + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported(in_list(&[1]), no_bounds()), + reported_with_null_keys(in_list(&[2]), no_bounds()), + ])) + .unwrap(); + + let expr = current_expr(&acc); + assert_top_binary_op(&expr, Operator::Or); + let widened = binary_expr(&expr); + assert!( + widened.left().downcast_ref::().is_some(), + "expected the IS NULL disjunct first, got: {expr}" + ); + assert_in_list_column_values(widened.right(), "probe_key", 0, &[1, 2]); + } + /// A canceled partition's build content is unknown, so it may hold a NULL key: /// the aggregated flag must be permissive even when no partition reported one. #[test] @@ -1640,7 +2158,8 @@ mod tests { completion: CompletionState::Pending, }), completion_notify: Notify::new(), - dynamic_filter, + membership_filter: Some(dynamic_filter), + bounds_filter: None, on_right, repartition_random_state: SeededRandomState::with_seed(1), probe_schema: Arc::new(Schema::new(vec![ @@ -1676,4 +2195,318 @@ mod tests { .expect("expected column under IS NULL"); assert_eq!(column.index(), 0, "escape must target the NOT IN value key"); } + + // Tests for an accumulator with a separate bounds filter: the bounds and + // the membership check go to two dynamic filters. + + /// Adds a bounds filter to `acc`. + fn with_bounds_filter(mut acc: SharedBuildAccumulator) -> SharedBuildAccumulator { + acc.bounds_filter = Some(test_dynamic_filter(&acc.on_right)); + acc + } + + fn bounds_filter_string(acc: &SharedBuildAccumulator) -> String { + acc.bounds_filter + .as_ref() + .expect("expected a bounds filter") + .current() + .expect("bounds filter current expression should be available") + .to_string() + } + + fn membership_filter_string(acc: &SharedBuildAccumulator) -> String { + current_expr(acc).to_string() + } + + fn range_partitioning_for_test( + acc: &SharedBuildAccumulator, + split_points: &[i32], + ) -> Result { + RangePartitioning::try_new( + [PhysicalSortExpr::new( + Arc::clone(&acc.on_right[0]), + Default::default(), + )] + .into(), + split_points + .iter() + .map(|point| SplitPoint::new(vec![ScalarValue::Int32(Some(*point))])) + .collect(), + ) + } + + fn evaluate_to_bools(expr: &PhysicalExprRef, values: Vec) -> Vec { + let batch = RecordBatch::try_new( + test_probe_schema(), + vec![Arc::new(Int32Array::from(values))], + ) + .unwrap(); + let result = expr + .evaluate(&batch) + .unwrap() + .into_array(batch.num_rows()) + .unwrap(); + result + .as_any() + .downcast_ref::() + .expect("dynamic filter should evaluate to BooleanArray") + .iter() + .map(|value| value.unwrap_or(false)) + .collect() + } + + #[test] + fn split_partitioned_hoists_union_bounds_out_of_routing_case() { + let acc = with_bounds_filter(make_partitioned_expr_accumulator_for_test(2)); + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported(map_pushdown(), bounds(1, 10)), + reported(map_pushdown(), bounds(5, 20)), + ])) + .unwrap(); + + assert_eq!( + bounds_filter_string(&acc), + "probe_key@0 >= 1 AND probe_key@0 <= 20" + ); + // The routed `CASE` does not repeat the per-partition bounds. + assert_eq!( + membership_filter_string(&acc), + "CASE hash_repartition % 2 WHEN 0 THEN hash_lookup WHEN 1 THEN hash_lookup ELSE false END" + ); + } + + #[test] + fn split_partitioned_union_inlist_has_no_bounds() { + let acc = with_bounds_filter(make_partitioned_expr_accumulator_for_test(2)); + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported(in_list(&[3, 1]), bounds(1, 3)), + reported(in_list(&[7]), bounds(7, 7)), + ])) + .unwrap(); + + assert_eq!( + bounds_filter_string(&acc), + "probe_key@0 >= 1 AND probe_key@0 <= 7" + ); + assert_eq!( + membership_filter_string(&acc), + "probe_key@0 IN (SET) ([1, 3, 7])" + ); + } + + #[test] + fn split_partitioned_one_real_partition_splits_its_filter() { + let acc = with_bounds_filter(make_partitioned_expr_accumulator_for_test(2)); + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported(map_pushdown(), bounds(1, 10)), + reported(PushdownStrategy::Empty, no_bounds()), + ])) + .unwrap(); + + assert_eq!( + bounds_filter_string(&acc), + "probe_key@0 >= 1 AND probe_key@0 <= 10" + ); + assert_eq!(membership_filter_string(&acc), "hash_lookup"); + } + + #[test] + fn split_partitioned_all_empty_rejects_in_both_filters() { + let acc = with_bounds_filter(make_partitioned_expr_accumulator_for_test(2)); + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported(PushdownStrategy::Empty, no_bounds()), + reported(PushdownStrategy::Empty, no_bounds()), + ])) + .unwrap(); + + assert_eq!(bounds_filter_string(&acc), "false"); + assert_eq!(membership_filter_string(&acc), "false"); + } + + /// A canceled partition can hold any key, so the union of the known + /// bounds would reject its matches. The bounds filter stays `true`, and + /// the bounds stay in the routed `CASE`. + #[test] + fn split_partitioned_canceled_partition_keeps_bounds_in_case() { + let acc = with_bounds_filter(make_partitioned_expr_accumulator_for_test(2)); + let bounds_generation = acc.bounds_filter.as_ref().unwrap().snapshot_generation(); + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported(map_pushdown(), bounds(1, 10)), + PartitionStatus::CanceledUnknown, + ])) + .unwrap(); + + assert_eq!( + acc.bounds_filter.as_ref().unwrap().snapshot_generation(), + bounds_generation + ); + assert_eq!(bounds_filter_string(&acc), "true"); + assert_eq!( + membership_filter_string(&acc), + "CASE hash_repartition % 2 WHEN 0 THEN probe_key@0 >= 1 AND probe_key@0 <= 10 AND hash_lookup ELSE true END" + ); + } + + #[test] + fn split_partitioned_without_bounds_leaves_bounds_filter_unchanged() { + let acc = with_bounds_filter(make_partitioned_expr_accumulator_for_test(2)); + let bounds_generation = acc.bounds_filter.as_ref().unwrap().snapshot_generation(); + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported(map_pushdown(), no_bounds()), + reported(map_pushdown(), no_bounds()), + ])) + .unwrap(); + + assert_eq!( + acc.bounds_filter.as_ref().unwrap().snapshot_generation(), + bounds_generation + ); + assert_eq!( + membership_filter_string(&acc), + "CASE hash_repartition % 2 WHEN 0 THEN hash_lookup WHEN 1 THEN hash_lookup ELSE false END" + ); + } + + /// Range partitions hold disjoint key ranges, so the bounds filter keeps + /// them apart and rejects the probe keys between them. + #[test] + fn split_range_partitioned_keeps_disjoint_ranges() -> Result<()> { + let mut acc = with_bounds_filter(make_partitioned_expr_accumulator_for_test(4)); + acc.probe_range_partitioning = + Some(range_partitioning_for_test(&acc, &[10, 20, 30])?); + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported(in_list(&[5]), bounds(5, 5)), + reported(PushdownStrategy::Empty, no_bounds()), + reported(map_pushdown(), bounds(20, 25)), + reported(map_pushdown(), bounds(30, 30)), + ]))?; + + assert_eq!( + bounds_filter_string(&acc), + "probe_key@0 >= 5 AND probe_key@0 <= 5 OR probe_key@0 >= 20 AND probe_key@0 <= 25 OR probe_key@0 >= 30 AND probe_key@0 <= 30" + ); + let bounds_expr = acc.bounds_filter.as_ref().unwrap().current()?; + assert_eq!( + evaluate_to_bools(&bounds_expr, vec![5, 6, 19, 20, 25, 26, 30, 31]), + vec![true, false, false, true, true, false, true, false] + ); + assert_eq!( + membership_filter_string(&acc), + "CASE range_partition WHEN 0 THEN probe_key@0 IN (SET) ([5]) WHEN 2 THEN hash_lookup WHEN 3 THEN hash_lookup ELSE false END" + ); + Ok(()) + } + + /// Without a bounds filter (for example, a plan from a version that + /// pushed one filter), a range-partitioned join keeps the bounds in the + /// routed `CASE`. + #[test] + fn range_partitioned_without_bounds_filter_keeps_bounds_in_case() -> Result<()> { + let mut acc = make_partitioned_expr_accumulator_for_test(2); + acc.probe_range_partitioning = Some(range_partitioning_for_test(&acc, &[10])?); + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported(map_pushdown(), bounds(1, 5)), + reported(map_pushdown(), bounds(10, 15)), + ]))?; + + assert_eq!( + membership_filter_string(&acc), + "CASE range_partition WHEN 0 THEN probe_key@0 >= 1 AND probe_key@0 <= 5 AND hash_lookup WHEN 1 THEN probe_key@0 >= 10 AND probe_key@0 <= 15 AND hash_lookup ELSE false END" + ); + Ok(()) + } + + #[test] + fn split_collect_left_updates_bounds_and_membership_separately() { + let acc = with_bounds_filter(make_collect_left_accumulator_for_test()); + + acc.build_filter(FinalizeInput::CollectLeft(reported( + in_list(&[1, 2, 3]), + bounds(1, 3), + ))) + .unwrap(); + + assert_eq!( + bounds_filter_string(&acc), + "probe_key@0 >= 1 AND probe_key@0 <= 3" + ); + assert_in_list_column_values(¤t_expr(&acc), "probe_key", 0, &[1, 2, 3]); + } + + /// Both filters must keep probe NULLs when a null-equal join has a NULL + /// build key: either filter alone would drop the match. + #[test] + fn split_null_equal_widens_both_filters() { + let mut acc = null_equal_partitioned_accumulator(2); + acc.bounds_filter = Some(test_dynamic_filter(&acc.on_right)); + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported(map_pushdown(), bounds(1, 10)), + reported_with_null_keys(map_pushdown(), bounds(5, 20)), + ])) + .unwrap(); + + assert_eq!( + bounds_filter_string(&acc), + "probe_key@0 IS NULL OR probe_key@0 >= 1 AND probe_key@0 <= 20" + ); + assert_eq!( + membership_filter_string(&acc), + "probe_key@0 IS NULL OR CASE hash_repartition % 2 WHEN 0 THEN hash_lookup WHEN 1 THEN hash_lookup ELSE false END" + ); + } + + #[test] + fn split_bounds_filter_without_membership_filter() { + let mut acc = with_bounds_filter(make_partitioned_expr_accumulator_for_test(2)); + acc.membership_filter = None; + + acc.build_filter(FinalizeInput::Partitioned(vec![ + reported(in_list(&[1]), bounds(1, 1)), + reported(in_list(&[4]), bounds(4, 4)), + ])) + .unwrap(); + + assert_eq!( + bounds_filter_string(&acc), + "probe_key@0 >= 1 AND probe_key@0 <= 4" + ); + } + + #[tokio::test] + async fn split_filters_are_both_marked_complete() -> Result<()> { + let acc = with_bounds_filter(make_partitioned_expr_accumulator_for_test(1)); + + acc.report_build_data(PartitionBuildData::Partitioned { + partition_id: 0, + pushdown: in_list(&[1, 2]), + bounds: bounds(1, 2), + keys_have_null: false, + }) + .await?; + + for filter in [&acc.bounds_filter, &acc.membership_filter] { + let filter = filter.as_ref().unwrap(); + tokio::time::timeout( + std::time::Duration::from_secs(5), + filter.wait_complete(), + ) + .await + .expect("both filters should be marked complete"); + } + assert_eq!( + bounds_filter_string(&acc), + "probe_key@0 >= 1 AND probe_key@0 <= 2" + ); + assert_in_list_column_values(¤t_expr(&acc), "probe_key", 0, &[1, 2]); + Ok(()) + } } diff --git a/datafusion/physical-plan/src/joins/hash_join/stream.rs b/datafusion/physical-plan/src/joins/hash_join/stream.rs index 03387ea574e63..2041a3ba6441d 100644 --- a/datafusion/physical-plan/src/joins/hash_join/stream.rs +++ b/datafusion/physical-plan/src/joins/hash_join/stream.rs @@ -49,10 +49,12 @@ use arrow::array::{Array, ArrayRef, UInt32Array, UInt64Array}; use arrow::buffer::{BooleanBuffer, NullBuffer}; use arrow::datatypes::{Schema, SchemaRef}; use arrow::record_batch::RecordBatch; +use datafusion_common::instant::Instant; use datafusion_common::{ JoinSide, JoinType, NullEquality, Result, internal_datafusion_err, internal_err, }; use datafusion_physical_expr::PhysicalExprRef; +use datafusion_physical_expr::filter_stats::{RemovedRowWork, duration_nanos}; use datafusion_common::hash_utils::RandomState; use datafusion_physical_expr_common::utils::evaluate_expressions_to_arrays; @@ -382,6 +384,16 @@ pub(super) struct HashJoinStream { batch_size: usize, /// Scratch space for computing hashes hashes_buffer: Vec, + /// The work measurements of the dynamic filters that this join produces + /// (see [`RemovedRowWork`]). The join records the rows of each probe + /// batch and the time of the work that it does for a probe row whether + /// or not the row matches: the evaluation and the hashes of the join + /// keys and the hash table lookup. A row that the dynamic filter removes + /// before the join does not get this work. The check of the candidates + /// of the lookup and the output are not in it: only matches get them, + /// and while the filter is on, most probe rows are matches. Empty + /// without dynamic filter pushdown. + removed_row_work: Vec>, /// Scratch space for probe indices during hash lookup probe_indices_buffer: Vec, /// Scratch space for build indices during hash lookup @@ -415,6 +427,18 @@ impl RecordBatchStream for HashJoinStream { } } +/// Records the time since `start` in `works` (see the `removed_row_work` +/// field of [`HashJoinStream`]), without rows: the rows of a probe batch are +/// recorded once, when the batch arrives. +fn record_lookup_work(works: &[Arc], start: Option) { + if let Some(start) = start { + let nanos = duration_nanos(start.elapsed()); + for work in works { + work.record(0, nanos); + } + } +} + /// Executes lookups by hash against JoinHashMap and resolves potential /// hash collisions. /// Returns build/probe indices satisfying the equality condition, along with @@ -484,7 +508,27 @@ pub(super) fn lookup_join_hashmap( probe_indices_buffer, build_indices_buffer, ); + let (build_indices, probe_indices) = equal_candidates( + build_side_values, + probe_side_values, + null_equality, + probe_indices_buffer, + build_indices_buffer, + )?; + Ok((build_indices, probe_indices, next_offset)) +} +/// Keeps the candidates of a hash table lookup (the indices in +/// `probe_indices_buffer` and `build_indices_buffer`) whose join keys are +/// equal, and returns their indices. Only the probe rows with a candidate +/// (a match or a hash collision) get this work. +fn equal_candidates( + build_side_values: &[ArrayRef], + probe_side_values: &[ArrayRef], + null_equality: NullEquality, + probe_indices_buffer: &mut Vec, + build_indices_buffer: &mut Vec, +) -> Result<(UInt64Array, UInt32Array)> { let build_indices_unfiltered: UInt64Array = std::mem::take(build_indices_buffer).into(); let probe_indices_unfiltered: UInt32Array = @@ -504,7 +548,7 @@ pub(super) fn lookup_join_hashmap( *build_indices_buffer = build_indices_unfiltered.into_parts().1.into(); *probe_indices_buffer = probe_indices_unfiltered.into_parts().1.into(); - Ok((build_indices, probe_indices, next_offset)) + Ok((build_indices, probe_indices)) } /// Counts the number of distinct elements in the input array. @@ -585,6 +629,7 @@ impl HashJoinStream { null_mark_hashes_buffer: Vec::new(), null_mark_probe_indices_buffer: Vec::new(), null_mark_build_indices_buffer: Vec::new(), + removed_row_work: vec![], right_side_ordered, build_report: BuildReportHandle::new(partition, mode, build_accumulator), mode, @@ -593,6 +638,24 @@ impl HashJoinStream { } } + /// Records the work for each probe row in `works`, see + /// the `removed_row_work` field. + pub(super) fn with_removed_row_work( + mut self, + works: Vec>, + ) -> Self { + self.removed_row_work = works; + self + } + + /// Records `rows` probe rows and `nanos` nanoseconds of work for each + /// row in the dynamic filters of this join. + fn record_removed_row_work(&self, rows: usize, nanos: u64) { + for work in &self.removed_row_work { + work.record(rows as u64, nanos); + } + } + /// Returns the next state after the build side has been fully collected /// and any required build-side coordination has completed. fn state_after_build_ready( @@ -779,6 +842,7 @@ impl HashJoinStream { self.state = HashJoinStreamState::ExhaustedProbeSide; } Some(Ok(batch)) => { + let work_start = (!self.removed_row_work.is_empty()).then(Instant::now); // Precalculate hash values for fetched batch let keys_values = evaluate_expressions_to_arrays(&self.on_right, &batch)?; @@ -797,6 +861,12 @@ impl HashJoinStream { None }; + if let Some(work_start) = work_start { + self.record_removed_row_work( + batch.num_rows(), + duration_nanos(work_start.elapsed()), + ); + } self.join_metrics.input_batches.add(1); self.join_metrics.input_rows.add(batch.num_rows()); @@ -880,21 +950,33 @@ impl HashJoinStream { return Ok(StatefulStreamResult::Continue); } - // get the matched by join keys indices + // get the matched by join keys indices. The lookup of the hashes is + // the work that every probe row gets, and a dynamic filter saves it + // for each row that it removes: its time goes to `removed_row_work`. + // The check and the output of the candidates is work for the matches + // only, which a dynamic filter does not remove. + let work_start = (!self.removed_row_work.is_empty()).then(Instant::now); let (left_indices, right_indices, next_offset) = match build_side.left_data.map() { - Map::HashMap(map) => lookup_join_hashmap( - map.as_ref(), - build_side.left_data.values(), - &state.values, - self.null_equality, - &self.hashes_buffer, - state.valid_keys.as_ref(), - self.batch_size, - state.offset, - &mut self.probe_indices_buffer, - &mut self.build_indices_buffer, - )?, + Map::HashMap(map) => { + let next_offset = map.get_matched_indices_with_limit_offset( + &self.hashes_buffer, + state.valid_keys.as_ref(), + self.batch_size, + state.offset, + &mut self.probe_indices_buffer, + &mut self.build_indices_buffer, + ); + record_lookup_work(&self.removed_row_work, work_start); + let (left_indices, right_indices) = equal_candidates( + build_side.left_data.values(), + &state.values, + self.null_equality, + &mut self.probe_indices_buffer, + &mut self.build_indices_buffer, + )?; + (left_indices, right_indices, next_offset) + } Map::ArrayMap(array_map) => { let next_offset = array_map.get_matched_indices_with_limit_offset( &state.values, @@ -903,6 +985,7 @@ impl HashJoinStream { &mut self.probe_indices_buffer, &mut self.build_indices_buffer, )?; + record_lookup_work(&self.removed_row_work, work_start); ( UInt64Array::from(self.build_indices_buffer.clone()), UInt32Array::from(self.probe_indices_buffer.clone()), @@ -910,7 +993,6 @@ impl HashJoinStream { ) } }; - let matched_probe_rows = state.count_new_matched_probe_rows(&right_indices); self.join_metrics diff --git a/datafusion/physical-plan/src/joins/nested_loop_join.rs b/datafusion/physical-plan/src/joins/nested_loop_join.rs index e15f34ced6e73..f0059ab44f552 100644 --- a/datafusion/physical-plan/src/joins/nested_loop_join.rs +++ b/datafusion/physical-plan/src/joins/nested_loop_join.rs @@ -79,7 +79,7 @@ use datafusion_physical_expr::PhysicalExpr; use datafusion_physical_expr::equivalence::{ ProjectionMapping, join_equivalence_properties, }; -use datafusion_physical_expr::expressions::DynamicFilterPhysicalExpr; +use datafusion_physical_expr::utils::as_dynamic_filter; use datafusion_physical_expr::projection::{ProjectionRef, combine_projections}; use futures::future::{BoxFuture, Shared}; @@ -876,7 +876,9 @@ impl ExecutionPlan for NestedLoopJoinExec { .iter() .zip(&mut child_description.parent_filters) { - if !filter.is::() { + // Producers push their dynamic filters as + // `Optional(DynamicFilter)`, so look through the wrapper. + if as_dynamic_filter(filter).is_none() { *pushed = PushedDownPredicate::unsupported(Arc::clone(filter)); } } @@ -4056,7 +4058,9 @@ pub(crate) mod tests { use datafusion_execution::runtime_env::RuntimeEnvBuilder; use datafusion_execution::spill_file::{SpillFile, SpillWriter, TempFileFactory}; use datafusion_expr::Operator; - use datafusion_physical_expr::expressions::{BinaryExpr, Literal}; + use datafusion_physical_expr::expressions::{ + BinaryExpr, DynamicFilterPhysicalExpr, Literal, OptionalFilterPhysicalExpr, + }; use datafusion_physical_expr::{Partitioning, PhysicalExpr}; use datafusion_physical_expr_common::sort_expr::{LexOrdering, PhysicalSortExpr}; @@ -4203,6 +4207,47 @@ pub(crate) mod tests { Ok(()) } + /// Producers push their dynamic filters as `Optional(DynamicFilter)`. The + /// NLJ must route the wrapped filter like a bare dynamic filter, and keep + /// the wrapper on the pushed copy. + #[test] + fn test_nlj_routes_optional_dynamic_filter() -> Result<()> { + use crate::filter_pushdown::PushedDown; + use datafusion_physical_expr::expressions::lit; + + let join = NestedLoopJoinExec::try_new( + build_left_table(), + build_right_table(), + None, + &JoinType::Inner, + None, + )?; + let column: Arc = + Arc::new(Column::new(join.schema().field(0).name(), 0)); + let source = Arc::new(DynamicFilterPhysicalExpr::new(vec![column], lit(true))); + let optional: Arc = + Arc::new(OptionalFilterPhysicalExpr::new(Arc::clone(&source) as _)); + + let filters = join + .gather_filters_for_pushdown( + FilterPushdownPhase::Post, + vec![optional], + &ConfigOptions::default(), + )? + .parent_filters(); + + // Column 0 is a left column, so only the left child accepts it. + let left_filter = &filters[0][0]; + assert!(matches!(left_filter.discriminant, PushedDown::Yes)); + assert!(left_filter.predicate.is::()); + assert_eq!( + as_dynamic_filter(&left_filter.predicate).and_then(|df| df.expression_id()), + source.expression_id() + ); + assert!(matches!(filters[1][0].discriminant, PushedDown::No)); + Ok(()) + } + fn delayed_stream(batch: RecordBatch, delay: Duration) -> SendableRecordBatchStream { let schema = batch.schema(); Box::pin(crate::stream::RecordBatchStreamAdapter::new( diff --git a/datafusion/physical-plan/src/sorts/sort.rs b/datafusion/physical-plan/src/sorts/sort.rs index 9a149cc7c38c1..11667fadcbfbc 100644 --- a/datafusion/physical-plan/src/sorts/sort.rs +++ b/datafusion/physical-plan/src/sorts/sort.rs @@ -73,7 +73,9 @@ use datafusion_execution::memory_pool::{ use datafusion_execution::runtime_env::RuntimeEnv; use datafusion_physical_expr::LexOrdering; use datafusion_physical_expr::PhysicalExpr; -use datafusion_physical_expr::expressions::{DynamicFilterPhysicalExpr, lit}; +use datafusion_physical_expr::expressions::{ + DynamicFilterPhysicalExpr, OptionalFilterPhysicalExpr, lit, +}; use futures::{StreamExt, TryStreamExt}; use log::{debug, trace}; @@ -1632,7 +1634,11 @@ impl ExecutionPlan for SortExec { if let Some(filter) = &self.filter && config.optimizer.enable_topk_dynamic_filter_pushdown { - child = child.with_self_filter(filter.read().expr()); + // The TopK itself keeps only the top `fetch` rows, so the filter + // is not needed for correctness: mark the pushed copy as optional. + child = child.with_self_filter(Arc::new(OptionalFilterPhysicalExpr::new( + filter.read().expr(), + ))); } Ok(FilterDescription::new().with_child(child)) diff --git a/datafusion/proto-models/proto/datafusion.proto b/datafusion/proto-models/proto/datafusion.proto index 4cb4ddebe4be6..541001cf13c22 100644 --- a/datafusion/proto-models/proto/datafusion.proto +++ b/datafusion/proto-models/proto/datafusion.proto @@ -1093,6 +1093,8 @@ message PhysicalExprNode { PhysicalLambdaVariableExprNode lambda_variable = 26; PhysicalRangeExprNode range_expr = 27; PhysicalSqlSimilarToPatternNode sql_similar_to_pattern = 28; + + PhysicalOptionalFilterNode optional_filter = 29; } } @@ -1104,6 +1106,11 @@ message PhysicalDynamicFilterNode { bool is_complete = 5; } +// Marks the wrapped filter as optional: it is not needed for correctness. +message PhysicalOptionalFilterNode { + PhysicalExprNode inner = 1; +} + message PhysicalSqlSimilarToPatternNode { PhysicalExprNode expr = 1; } @@ -1403,7 +1410,9 @@ message HashJoinExecNode { JoinFilter filter = 8; repeated uint32 projection = 9; bool null_aware = 10; - // Optional dynamic filter expression for pushing down to the probe side. + // Optional dynamic filter expression for pushing down to the probe side: + // the membership check. When `dynamic_filter_bounds` is absent, it also + // holds the build-side bounds. PhysicalExprNode dynamic_filter = 11; // Optional row limit pushed into the join by the `limit_pushdown` rule. // @@ -1413,6 +1422,10 @@ message HashJoinExecNode { // turning old plans into empty results. With `optional`, absent decodes to // `None`, which is the correct reading of an older message. optional uint64 fetch = 12; + // Optional dynamic filter expression for pushing down to the probe side: + // the build-side bounds, separate from the membership check in + // `dynamic_filter`. + PhysicalExprNode dynamic_filter_bounds = 13; } enum StreamPartitionMode { diff --git a/datafusion/proto-models/src/generated/pbjson.rs b/datafusion/proto-models/src/generated/pbjson.rs index 0f42dea134975..a18be07ea6c92 100644 --- a/datafusion/proto-models/src/generated/pbjson.rs +++ b/datafusion/proto-models/src/generated/pbjson.rs @@ -9754,6 +9754,9 @@ impl serde::Serialize for HashJoinExecNode { if self.fetch.is_some() { len += 1; } + if self.dynamic_filter_bounds.is_some() { + len += 1; + } let mut struct_ser = serializer.serialize_struct("datafusion.HashJoinExecNode", len)?; if let Some(v) = self.left.as_ref() { struct_ser.serialize_field("left", v)?; @@ -9796,6 +9799,9 @@ impl serde::Serialize for HashJoinExecNode { #[allow(clippy::needless_borrows_for_generic_args)] struct_ser.serialize_field("fetch", ToString::to_string(&v).as_str())?; } + if let Some(v) = self.dynamic_filter_bounds.as_ref() { + struct_ser.serialize_field("dynamicFilterBounds", v)?; + } struct_ser.end() } } @@ -9822,6 +9828,8 @@ impl<'de> serde::Deserialize<'de> for HashJoinExecNode { "dynamic_filter", "dynamicFilter", "fetch", + "dynamic_filter_bounds", + "dynamicFilterBounds", ]; #[allow(clippy::enum_variant_names)] @@ -9837,6 +9845,7 @@ impl<'de> serde::Deserialize<'de> for HashJoinExecNode { NullAware, DynamicFilter, Fetch, + DynamicFilterBounds, } impl<'de> serde::Deserialize<'de> for GeneratedField { fn deserialize(deserializer: D) -> std::result::Result @@ -9869,6 +9878,7 @@ impl<'de> serde::Deserialize<'de> for HashJoinExecNode { "nullAware" | "null_aware" => Ok(GeneratedField::NullAware), "dynamicFilter" | "dynamic_filter" => Ok(GeneratedField::DynamicFilter), "fetch" => Ok(GeneratedField::Fetch), + "dynamicFilterBounds" | "dynamic_filter_bounds" => Ok(GeneratedField::DynamicFilterBounds), _ => Err(serde::de::Error::unknown_field(value, FIELDS)), } } @@ -9899,6 +9909,7 @@ impl<'de> serde::Deserialize<'de> for HashJoinExecNode { let mut null_aware__ = None; let mut dynamic_filter__ = None; let mut fetch__ = None; + let mut dynamic_filter_bounds__ = None; while let Some(k) = map_.next_key()? { match k { GeneratedField::Left => { @@ -9972,6 +9983,12 @@ impl<'de> serde::Deserialize<'de> for HashJoinExecNode { map_.next_value::<::std::option::Option<::pbjson::private::NumberDeserialize<_>>>()?.map(|x| x.0) ; } + GeneratedField::DynamicFilterBounds => { + if dynamic_filter_bounds__.is_some() { + return Err(serde::de::Error::duplicate_field("dynamicFilterBounds")); + } + dynamic_filter_bounds__ = map_.next_value()?; + } } } Ok(HashJoinExecNode { @@ -9986,6 +10003,7 @@ impl<'de> serde::Deserialize<'de> for HashJoinExecNode { null_aware: null_aware__.unwrap_or_default(), dynamic_filter: dynamic_filter__, fetch: fetch__, + dynamic_filter_bounds: dynamic_filter_bounds__, }) } } @@ -19549,6 +19567,9 @@ impl serde::Serialize for PhysicalExprNode { physical_expr_node::ExprType::SqlSimilarToPattern(v) => { struct_ser.serialize_field("sqlSimilarToPattern", v)?; } + physical_expr_node::ExprType::OptionalFilter(v) => { + struct_ser.serialize_field("optionalFilter", v)?; + } } } struct_ser.end() @@ -19608,6 +19629,8 @@ impl<'de> serde::Deserialize<'de> for PhysicalExprNode { "rangeExpr", "sql_similar_to_pattern", "sqlSimilarToPattern", + "optional_filter", + "optionalFilter", ]; #[allow(clippy::enum_variant_names)] @@ -19639,6 +19662,7 @@ impl<'de> serde::Deserialize<'de> for PhysicalExprNode { LambdaVariable, RangeExpr, SqlSimilarToPattern, + OptionalFilter, } impl<'de> serde::Deserialize<'de> for GeneratedField { fn deserialize(deserializer: D) -> std::result::Result @@ -19687,6 +19711,7 @@ impl<'de> serde::Deserialize<'de> for PhysicalExprNode { "lambdaVariable" | "lambda_variable" => Ok(GeneratedField::LambdaVariable), "rangeExpr" | "range_expr" => Ok(GeneratedField::RangeExpr), "sqlSimilarToPattern" | "sql_similar_to_pattern" => Ok(GeneratedField::SqlSimilarToPattern), + "optionalFilter" | "optional_filter" => Ok(GeneratedField::OptionalFilter), _ => Err(serde::de::Error::unknown_field(value, FIELDS)), } } @@ -19898,6 +19923,13 @@ impl<'de> serde::Deserialize<'de> for PhysicalExprNode { return Err(serde::de::Error::duplicate_field("sqlSimilarToPattern")); } expr_type__ = map_.next_value::<::std::option::Option<_>>()?.map(physical_expr_node::ExprType::SqlSimilarToPattern) +; + } + GeneratedField::OptionalFilter => { + if expr_type__.is_some() { + return Err(serde::de::Error::duplicate_field("optionalFilter")); + } + expr_type__ = map_.next_value::<::std::option::Option<_>>()?.map(physical_expr_node::ExprType::OptionalFilter) ; } } @@ -21359,6 +21391,97 @@ impl<'de> serde::Deserialize<'de> for PhysicalNot { deserializer.deserialize_struct("datafusion.PhysicalNot", FIELDS, GeneratedVisitor) } } +impl serde::Serialize for PhysicalOptionalFilterNode { + #[allow(deprecated)] + fn serialize(&self, serializer: S) -> std::result::Result + where + S: serde::Serializer, + { + use serde::ser::SerializeStruct; + let mut len = 0; + if self.inner.is_some() { + len += 1; + } + let mut struct_ser = serializer.serialize_struct("datafusion.PhysicalOptionalFilterNode", len)?; + if let Some(v) = self.inner.as_ref() { + struct_ser.serialize_field("inner", v)?; + } + struct_ser.end() + } +} +impl<'de> serde::Deserialize<'de> for PhysicalOptionalFilterNode { + #[allow(deprecated)] + fn deserialize(deserializer: D) -> std::result::Result + where + D: serde::Deserializer<'de>, + { + const FIELDS: &[&str] = &[ + "inner", + ]; + + #[allow(clippy::enum_variant_names)] + enum GeneratedField { + Inner, + } + impl<'de> serde::Deserialize<'de> for GeneratedField { + fn deserialize(deserializer: D) -> std::result::Result + where + D: serde::Deserializer<'de>, + { + struct GeneratedVisitor; + + impl serde::de::Visitor<'_> for GeneratedVisitor { + type Value = GeneratedField; + + fn expecting(&self, formatter: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + write!(formatter, "expected one of: {:?}", &FIELDS) + } + + #[allow(unused_variables)] + fn visit_str(self, value: &str) -> std::result::Result + where + E: serde::de::Error, + { + match value { + "inner" => Ok(GeneratedField::Inner), + _ => Err(serde::de::Error::unknown_field(value, FIELDS)), + } + } + } + deserializer.deserialize_identifier(GeneratedVisitor) + } + } + struct GeneratedVisitor; + impl<'de> serde::de::Visitor<'de> for GeneratedVisitor { + type Value = PhysicalOptionalFilterNode; + + fn expecting(&self, formatter: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + formatter.write_str("struct datafusion.PhysicalOptionalFilterNode") + } + + fn visit_map(self, mut map_: V) -> std::result::Result + where + V: serde::de::MapAccess<'de>, + { + let mut inner__ = None; + while let Some(k) = map_.next_key()? { + match k { + GeneratedField::Inner => { + if inner__.is_some() { + return Err(serde::de::Error::duplicate_field("inner")); + } + inner__ = map_.next_value()?; + } + } + } + Ok(PhysicalOptionalFilterNode { + inner: inner__, + }) + } + } + deserializer.deserialize_struct("datafusion.PhysicalOptionalFilterNode", FIELDS, GeneratedVisitor) + } +} impl serde::Serialize for PhysicalPlanNode { #[allow(deprecated)] fn serialize(&self, serializer: S) -> std::result::Result diff --git a/datafusion/proto-models/src/generated/prost.rs b/datafusion/proto-models/src/generated/prost.rs index 31e060707a399..b51696f2e5b15 100644 --- a/datafusion/proto-models/src/generated/prost.rs +++ b/datafusion/proto-models/src/generated/prost.rs @@ -1613,7 +1613,7 @@ pub struct PhysicalExprNode { pub expr_id: ::core::option::Option, #[prost( oneof = "physical_expr_node::ExprType", - tags = "1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 15, 16, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28" + tags = "1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 15, 16, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29" )] pub expr_type: ::core::option::Option, } @@ -1682,6 +1682,8 @@ pub mod physical_expr_node { SqlSimilarToPattern( ::prost::alloc::boxed::Box, ), + #[prost(message, tag = "29")] + OptionalFilter(::prost::alloc::boxed::Box), } } #[derive(Clone, PartialEq, ::prost::Message)] @@ -1697,6 +1699,12 @@ pub struct PhysicalDynamicFilterNode { #[prost(bool, tag = "5")] pub is_complete: bool, } +/// Marks the wrapped filter as optional: it is not needed for correctness. +#[derive(Clone, PartialEq, ::prost::Message)] +pub struct PhysicalOptionalFilterNode { + #[prost(message, optional, boxed, tag = "1")] + pub inner: ::core::option::Option<::prost::alloc::boxed::Box>, +} #[derive(Clone, PartialEq, ::prost::Message)] pub struct PhysicalSqlSimilarToPatternNode { #[prost(message, optional, boxed, tag = "1")] @@ -2131,7 +2139,9 @@ pub struct HashJoinExecNode { pub projection: ::prost::alloc::vec::Vec, #[prost(bool, tag = "10")] pub null_aware: bool, - /// Optional dynamic filter expression for pushing down to the probe side. + /// Optional dynamic filter expression for pushing down to the probe side: + /// the membership check. When `dynamic_filter_bounds` is absent, it also + /// holds the build-side bounds. #[prost(message, optional, tag = "11")] pub dynamic_filter: ::core::option::Option, /// Optional row limit pushed into the join by the `limit_pushdown` rule. @@ -2143,6 +2153,11 @@ pub struct HashJoinExecNode { /// `None`, which is the correct reading of an older message. #[prost(uint64, optional, tag = "12")] pub fetch: ::core::option::Option, + /// Optional dynamic filter expression for pushing down to the probe side: + /// the build-side bounds, separate from the membership check in + /// `dynamic_filter`. + #[prost(message, optional, tag = "13")] + pub dynamic_filter_bounds: ::core::option::Option, } #[derive(Clone, PartialEq, ::prost::Message)] pub struct SymmetricHashJoinExecNode { diff --git a/datafusion/proto/src/physical_plan/from_proto.rs b/datafusion/proto/src/physical_plan/from_proto.rs index bc443149df413..91ec5180c22ed 100644 --- a/datafusion/proto/src/physical_plan/from_proto.rs +++ b/datafusion/proto/src/physical_plan/from_proto.rs @@ -52,7 +52,9 @@ use super::{ }; use crate::protobuf::physical_expr_node::ExprType; use crate::{convert_required, protobuf}; -use datafusion_physical_expr::expressions::DynamicFilterPhysicalExpr; +use datafusion_physical_expr::expressions::{ + DynamicFilterPhysicalExpr, OptionalFilterPhysicalExpr, +}; /// Parses a physical sort expression from a protobuf. /// @@ -362,6 +364,9 @@ pub fn parse_physical_expr_with_converter( ExprType::DynamicFilter(_) => { DynamicFilterPhysicalExpr::try_from_proto(proto, &decode_ctx)? } + ExprType::OptionalFilter(_) => { + OptionalFilterPhysicalExpr::try_from_proto(proto, &decode_ctx)? + } ExprType::SqlSimilarToPattern(_) => { SqlSimilarToPattern::try_from_proto(proto, &decode_ctx)? } diff --git a/datafusion/proto/tests/cases/plans/dynamic_filters.rs b/datafusion/proto/tests/cases/plans/dynamic_filters.rs index 7892ffd9a1ab0..7edc5273b70d5 100644 --- a/datafusion/proto/tests/cases/plans/dynamic_filters.rs +++ b/datafusion/proto/tests/cases/plans/dynamic_filters.rs @@ -40,7 +40,8 @@ use datafusion::physical_plan::aggregates::{ }; use datafusion::physical_plan::empty::EmptyExec; use datafusion::physical_plan::expressions::{ - BinaryExpr, Column, DynamicFilterPhysicalExpr, PhysicalSortExpr, lit, + BinaryExpr, Column, DynamicFilterPhysicalExpr, OptionalFilterPhysicalExpr, + PhysicalSortExpr, lit, }; use datafusion::physical_plan::filter::FilterExec; use datafusion::physical_plan::joins::{HashJoinExec, PartitionMode}; @@ -273,6 +274,42 @@ fn parquet_source_predicate(child: &Arc) -> Arc, +) -> Vec> { + let predicate = parquet_source_predicate(child); + datafusion::physical_expr::split_conjunction(&predicate) + .into_iter() + .map(|conjunct| { + let optional = conjunct + .downcast_ref::() + .unwrap_or_else(|| { + panic!("pushed dynamic filter should be optional, got {conjunct}") + }); + Arc::clone(optional.inner()) + }) + .collect() +} + +/// Extract the dynamic filter that a producer pushed down to the parquet scan +/// at the bottom of the plan tree. Producers mark their dynamic filters as +/// optional, so the predicate must be `Optional(DynamicFilter)` after the +/// roundtrip. Returns the inner dynamic filter. +fn pushed_optional_dynamic_filter( + child: &Arc, +) -> Arc { + let predicate = parquet_source_predicate(child); + let optional = predicate + .downcast_ref::() + .unwrap_or_else(|| { + panic!("pushed dynamic filter should be optional, got {predicate}") + }); + Arc::clone(optional.inner()) +} + /// Assert that two dynamic filters are equal both structurally (Debug output) /// and by identity (`expression_id`). fn assert_dynamic_filters_equal( @@ -326,6 +363,68 @@ fn test_dynamic_filter_roundtrip_dedupe() -> Result<()> { Ok(()) } +// An `Optional` wrapper survives the roundtrip, and a dynamic filter inside it +// is deduped with an unwrapped clone of the same dynamic filter. +#[test] +fn test_optional_dynamic_filter_roundtrip_dedupe() -> Result<()> { + let schema = Arc::new(Schema::new(vec![Field::new("a", DataType::Int64, false)])); + let dynamic_filter = make_dynamic_filter(); + let optional_filter = + Arc::new(OptionalFilterPhysicalExpr::new(Arc::clone(&dynamic_filter))) + as Arc; + + let (optional_after_roundtrip, dynamic_after_roundtrip) = + roundtrip_dynamic_filter_expr_pair( + Arc::clone(&optional_filter), + Arc::clone(&dynamic_filter), + schema, + )?; + + assert_eq!( + optional_filter.to_string(), + optional_after_roundtrip.to_string() + ); + let inner_after_roundtrip = optional_after_roundtrip + .downcast_ref::() + .expect("Expected OptionalFilterPhysicalExpr") + .inner(); + assert_dynamic_filters_equal(&dynamic_filter, inner_after_roundtrip); + assert_dynamic_filters_equal(&dynamic_filter, &dynamic_after_roundtrip); + + // Assert referential integrity through the wrapper. + assert_dynamic_filter_update_is_visible( + inner_after_roundtrip, + &dynamic_after_roundtrip, + )?; + + Ok(()) +} + +// An `Optional` wrapper around a static expression survives the roundtrip. +#[test] +fn test_optional_filter_roundtrip() -> Result<()> { + let schema = Arc::new(Schema::new(vec![Field::new("a", DataType::Int64, false)])); + let expr = Arc::new(OptionalFilterPhysicalExpr::new(Arc::new(BinaryExpr::new( + Arc::new(Column::new("a", 0)), + Operator::Gt, + lit(5_i64), + )))) as Arc; + + let codec = DefaultPhysicalExtensionCodec {}; + let converter = DefaultPhysicalProtoConverter {}; + let proto = converter.physical_expr_to_proto(&expr, &codec)?; + let ctx = SessionContext::new(); + let task_ctx = ctx.task_ctx(); + let decode_ctx = PhysicalPlanDecodeContext::new(task_ctx.as_ref(), &codec); + let roundtrip = converter.proto_to_physical_expr(&proto, &schema, &decode_ctx)?; + + assert!(roundtrip.is::()); + assert_eq!(&expr, &roundtrip); + assert_eq!(roundtrip.to_string(), "Optional(a@0 > 5)"); + + Ok(()) +} + /// Roundtrip test for an execution plan where there are multiple instances of a dynamic filter /// with different children. #[test] @@ -396,58 +495,73 @@ fn datasource_for_dynamic_filter_pushdown( } /// Test that plan containing a HashJoinExec with dynamic filter pushdown -/// can be serialized and deserialized while preserving references to the dynamic filter. +/// can be serialized and deserialized while preserving references to the dynamic +/// filters. A partitioned join pushes two filters (bounds and membership), a +/// collect-left join pushes one. Each must stay linked to its copy in the probe +/// side. #[test] fn test_hash_join_with_dynamic_filter_roundtrip() -> Result<()> { - let schema = Arc::new(Schema::new(vec![Field::new("col", DataType::Int64, false)])); - - let left_child = Arc::new(EmptyExec::new(Arc::clone(&schema))); - let (right_child, config) = datasource_for_dynamic_filter_pushdown(&schema); - - let on: Vec<(Arc, Arc)> = vec![( - Arc::new(Column::new("col", 0)), - Arc::new(Column::new("col", 0)), - )]; - - let hash_join = Arc::new(HashJoinExec::try_new( - left_child, - right_child, - on, - None, - &JoinType::Inner, - None, - PartitionMode::CollectLeft, - NullEquality::NullEqualsNothing, - false, - )?) as Arc; - - // Run the optimizer rule for filter pushdown. - let optimizer = FilterPushdown::new_post_optimization(); - let plan = optimizer.optimize(hash_join, &config)?; - - let ctx = SessionContext::new(); - let codec = DefaultPhysicalExtensionCodec {}; - let converter = DeduplicatingProtoConverter {}; - let deserialized = roundtrip_test_and_return(plan, &ctx, &codec, &converter)?; - - // Extract the deserialized HashJoinExec and its dynamic filter. - let deserialized_join = deserialized - .downcast_ref::() - .expect("Should be HashJoinExec"); - let deserialized_hash_join_df = deserialized_join - .dynamic_expressions_produced() - .into_iter() - .next() - .expect("HashJoinExec should have a dynamic filter after roundtrip"); - - // Extract the dynamic filter pushed down to the probe side's ParquetSource. - let deserialized_predicate = parquet_source_predicate(deserialized_join.right()); - - // The HashJoinExec's dynamic filter and the probe side's predicate should - // refer to the same underlying expression. - let plan_df = deserialized_hash_join_df; - assert_dynamic_filters_equal(&plan_df, &deserialized_predicate); - assert_dynamic_filter_update_is_visible(&plan_df, &deserialized_predicate)?; + for (mode, expected_filters) in [ + (PartitionMode::CollectLeft, 1), + (PartitionMode::Partitioned, 2), + ] { + let schema = + Arc::new(Schema::new(vec![Field::new("col", DataType::Int64, false)])); + + let left_child = Arc::new(EmptyExec::new(Arc::clone(&schema))); + let (right_child, config) = datasource_for_dynamic_filter_pushdown(&schema); + + let on: Vec<(Arc, Arc)> = vec![( + Arc::new(Column::new("col", 0)), + Arc::new(Column::new("col", 0)), + )]; + + let hash_join = Arc::new(HashJoinExec::try_new( + left_child, + right_child, + on, + None, + &JoinType::Inner, + None, + mode, + NullEquality::NullEqualsNothing, + false, + )?) as Arc; + + // Run the optimizer rule for filter pushdown. + let optimizer = FilterPushdown::new_post_optimization(); + let plan = optimizer.optimize(hash_join, &config)?; + + let ctx = SessionContext::new(); + let codec = DefaultPhysicalExtensionCodec {}; + let converter = DeduplicatingProtoConverter {}; + let deserialized = roundtrip_test_and_return(plan, &ctx, &codec, &converter)?; + + // Extract the deserialized HashJoinExec and its dynamic filters + // (membership, then bounds if any). + let deserialized_join = deserialized + .downcast_ref::() + .expect("Should be HashJoinExec"); + let produced = deserialized_join.dynamic_expressions_produced(); + assert_eq!(produced.len(), expected_filters, "{mode:?}"); + + // Extract the dynamic filters pushed down to the probe side's + // ParquetSource (bounds first, then membership). + let pushed = parquet_source_conjuncts(deserialized_join.right()); + assert_eq!(pushed.len(), expected_filters, "{mode:?}"); + + // Each dynamic filter of the HashJoinExec and its copy in the probe + // side's predicate should refer to the same underlying expression. + for (plan_df, pushed_df) in produced.iter().zip(pushed.iter().rev()) { + assert_ne!(plan_df.expression_id(), None); + assert_eq!(plan_df.expression_id(), pushed_df.expression_id()); + assert_dynamic_filters_equal(plan_df, pushed_df); + assert_dynamic_filter_update_is_visible(plan_df, pushed_df)?; + } + if let [membership, bounds] = produced.as_slice() { + assert_ne!(membership.expression_id(), bounds.expression_id()); + } + } Ok(()) } @@ -596,7 +710,7 @@ fn test_aggregate_with_dynamic_filter_roundtrip() -> Result<()> { .expect("AggregateExec should have a dynamic filter after roundtrip"); // Extract the dynamic filter pushed down to the child ParquetSource. - let deserialized_predicate = parquet_source_predicate(deserialized_agg.input()); + let deserialized_predicate = pushed_optional_dynamic_filter(deserialized_agg.input()); // The AggregateExec's dynamic filter and the child's predicate should // refer to the same underlying expression. @@ -721,7 +835,8 @@ fn test_sort_topk_with_dynamic_filter_roundtrip() -> Result<()> { .expect("SortExec should have a dynamic filter after roundtrip"); // Extract the dynamic filter pushed down to the child ParquetSource. - let deserialized_predicate = parquet_source_predicate(deserialized_sort.input()); + let deserialized_predicate = + pushed_optional_dynamic_filter(deserialized_sort.input()); // The SortExec's dynamic filter and the child's predicate should // refer to the same underlying expression. diff --git a/datafusion/pruning/src/pruning_predicate.rs b/datafusion/pruning/src/pruning_predicate.rs index b49b72058e0cd..0332ff903b005 100644 --- a/datafusion/pruning/src/pruning_predicate.rs +++ b/datafusion/pruning/src/pruning_predicate.rs @@ -3448,6 +3448,47 @@ mod tests { assert_eq!(result, expected); } + /// Pruning sees through `OptionalFilterPhysicalExpr`: an optional filter + /// prunes the same as the inner filter. + #[test] + fn prune_optional_filter_same_as_inner() { + let (schema, statistics) = int32_setup(); + let expected = &[true, true, false, true, true]; + + let required = logical2physical(&col("i").gt(lit(0)), &schema); + let optional = Arc::new(phys_expr::OptionalFilterPhysicalExpr::new(Arc::clone( + &required, + ))) as Arc; + let dynamic = Arc::new(DynamicFilterPhysicalExpr::new( + collect_columns(&required) + .into_iter() + .map(|c| Arc::new(c) as Arc) + .collect(), + Arc::clone(&required), + )) as Arc; + let optional_dynamic = + Arc::new(phys_expr::OptionalFilterPhysicalExpr::new(dynamic)) + as Arc; + + let build = |expr: Arc| { + PruningPredicateBuilder::new() + .with_file_schema(Arc::clone(&schema)) + .try_build(expr) + .unwrap() + }; + let baseline = build(required); + assert_eq!(baseline.prune(&statistics).unwrap(), expected); + + for expr in [optional, optional_dynamic] { + let p = build(expr); + assert_eq!( + p.predicate_expr().to_string(), + baseline.predicate_expr().to_string() + ); + assert_eq!(p.prune(&statistics).unwrap(), expected); + } + } + #[test] fn row_group_predicate_lt_bool() -> Result<()> { let schema = Schema::new(vec![Field::new("c1", DataType::Boolean, false)]); diff --git a/datafusion/sqllogictest/test_files/clickbench.slt b/datafusion/sqllogictest/test_files/clickbench.slt index 7cb5547383c38..059397268021f 100644 --- a/datafusion/sqllogictest/test_files/clickbench.slt +++ b/datafusion/sqllogictest/test_files/clickbench.slt @@ -652,7 +652,7 @@ physical_plan 03)----ProjectionExec: expr=[WatchID@0 as WatchID, JavaEnable@1 as JavaEnable, Title@2 as Title, GoodEvent@3 as GoodEvent, EventTime@4 as EventTime, CounterID@6 as CounterID, ClientIP@7 as ClientIP, RegionID@8 as RegionID, UserID@9 as UserID, CounterClass@10 as CounterClass, OS@11 as OS, UserAgent@12 as UserAgent, URL@13 as URL, Referer@14 as Referer, IsRefresh@15 as IsRefresh, RefererCategoryID@16 as RefererCategoryID, RefererRegionID@17 as RefererRegionID, URLCategoryID@18 as URLCategoryID, URLRegionID@19 as URLRegionID, ResolutionWidth@20 as ResolutionWidth, ResolutionHeight@21 as ResolutionHeight, ResolutionDepth@22 as ResolutionDepth, FlashMajor@23 as FlashMajor, FlashMinor@24 as FlashMinor, FlashMinor2@25 as FlashMinor2, NetMajor@26 as NetMajor, NetMinor@27 as NetMinor, UserAgentMajor@28 as UserAgentMajor, UserAgentMinor@29 as UserAgentMinor, CookieEnable@30 as CookieEnable, JavascriptEnable@31 as JavascriptEnable, IsMobile@32 as IsMobile, MobilePhone@33 as MobilePhone, MobilePhoneModel@34 as MobilePhoneModel, Params@35 as Params, IPNetworkID@36 as IPNetworkID, TraficSourceID@37 as TraficSourceID, SearchEngineID@38 as SearchEngineID, SearchPhrase@39 as SearchPhrase, AdvEngineID@40 as AdvEngineID, IsArtifical@41 as IsArtifical, WindowClientWidth@42 as WindowClientWidth, WindowClientHeight@43 as WindowClientHeight, ClientTimeZone@44 as ClientTimeZone, ClientEventTime@45 as ClientEventTime, SilverlightVersion1@46 as SilverlightVersion1, SilverlightVersion2@47 as SilverlightVersion2, SilverlightVersion3@48 as SilverlightVersion3, SilverlightVersion4@49 as SilverlightVersion4, PageCharset@50 as PageCharset, CodeVersion@51 as CodeVersion, IsLink@52 as IsLink, IsDownload@53 as IsDownload, IsNotBounce@54 as IsNotBounce, FUniqID@55 as FUniqID, OriginalURL@56 as OriginalURL, HID@57 as HID, IsOldCounter@58 as IsOldCounter, IsEvent@59 as IsEvent, IsParameter@60 as IsParameter, DontCountHits@61 as DontCountHits, WithHash@62 as WithHash, HitColor@63 as HitColor, LocalEventTime@64 as LocalEventTime, Age@65 as Age, Sex@66 as Sex, Income@67 as Income, Interests@68 as Interests, Robotness@69 as Robotness, RemoteIP@70 as RemoteIP, WindowName@71 as WindowName, OpenerName@72 as OpenerName, HistoryLength@73 as HistoryLength, BrowserLanguage@74 as BrowserLanguage, BrowserCountry@75 as BrowserCountry, SocialNetwork@76 as SocialNetwork, SocialAction@77 as SocialAction, HTTPError@78 as HTTPError, SendTiming@79 as SendTiming, DNSTiming@80 as DNSTiming, ConnectTiming@81 as ConnectTiming, ResponseStartTiming@82 as ResponseStartTiming, ResponseEndTiming@83 as ResponseEndTiming, FetchTiming@84 as FetchTiming, SocialSourceNetworkID@85 as SocialSourceNetworkID, SocialSourcePage@86 as SocialSourcePage, ParamPrice@87 as ParamPrice, ParamOrderID@88 as ParamOrderID, ParamCurrency@89 as ParamCurrency, ParamCurrencyID@90 as ParamCurrencyID, OpenstatServiceName@91 as OpenstatServiceName, OpenstatCampaignID@92 as OpenstatCampaignID, OpenstatAdID@93 as OpenstatAdID, OpenstatSourceID@94 as OpenstatSourceID, UTMSource@95 as UTMSource, UTMMedium@96 as UTMMedium, UTMCampaign@97 as UTMCampaign, UTMContent@98 as UTMContent, UTMTerm@99 as UTMTerm, FromTag@100 as FromTag, HasGCLID@101 as HasGCLID, RefererHash@102 as RefererHash, URLHash@103 as URLHash, CLID@104 as CLID, CAST(CAST(EventDate@5 AS Int32) AS Date32) as EventDate] 04)------FilterExec: URL@13 LIKE %google% 05)--------RepartitionExec: partitioning=RoundRobinBatch(4), input_partitions=1 -06)----------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/core/tests/data/clickbench_hits_10.parquet]]}, projection=[WatchID, JavaEnable, Title, GoodEvent, EventTime, EventDate, CounterID, ClientIP, RegionID, UserID, CounterClass, OS, UserAgent, URL, Referer, IsRefresh, RefererCategoryID, RefererRegionID, URLCategoryID, URLRegionID, ResolutionWidth, ResolutionHeight, ResolutionDepth, FlashMajor, FlashMinor, FlashMinor2, NetMajor, NetMinor, UserAgentMajor, UserAgentMinor, CookieEnable, JavascriptEnable, IsMobile, MobilePhone, MobilePhoneModel, Params, IPNetworkID, TraficSourceID, SearchEngineID, SearchPhrase, AdvEngineID, IsArtifical, WindowClientWidth, WindowClientHeight, ClientTimeZone, ClientEventTime, SilverlightVersion1, SilverlightVersion2, SilverlightVersion3, SilverlightVersion4, PageCharset, CodeVersion, IsLink, IsDownload, IsNotBounce, FUniqID, OriginalURL, HID, IsOldCounter, IsEvent, IsParameter, DontCountHits, WithHash, HitColor, LocalEventTime, Age, Sex, Income, Interests, Robotness, RemoteIP, WindowName, OpenerName, HistoryLength, BrowserLanguage, BrowserCountry, SocialNetwork, SocialAction, HTTPError, SendTiming, DNSTiming, ConnectTiming, ResponseStartTiming, ResponseEndTiming, FetchTiming, SocialSourceNetworkID, SocialSourcePage, ParamPrice, ParamOrderID, ParamCurrency, ParamCurrencyID, OpenstatServiceName, OpenstatCampaignID, OpenstatAdID, OpenstatSourceID, UTMSource, UTMMedium, UTMCampaign, UTMContent, UTMTerm, FromTag, HasGCLID, RefererHash, URLHash, CLID], file_type=parquet, predicate=URL@13 LIKE %google% AND DynamicFilter [ empty ], dynamic_rg_pruning=eligible +06)----------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/core/tests/data/clickbench_hits_10.parquet]]}, projection=[WatchID, JavaEnable, Title, GoodEvent, EventTime, EventDate, CounterID, ClientIP, RegionID, UserID, CounterClass, OS, UserAgent, URL, Referer, IsRefresh, RefererCategoryID, RefererRegionID, URLCategoryID, URLRegionID, ResolutionWidth, ResolutionHeight, ResolutionDepth, FlashMajor, FlashMinor, FlashMinor2, NetMajor, NetMinor, UserAgentMajor, UserAgentMinor, CookieEnable, JavascriptEnable, IsMobile, MobilePhone, MobilePhoneModel, Params, IPNetworkID, TraficSourceID, SearchEngineID, SearchPhrase, AdvEngineID, IsArtifical, WindowClientWidth, WindowClientHeight, ClientTimeZone, ClientEventTime, SilverlightVersion1, SilverlightVersion2, SilverlightVersion3, SilverlightVersion4, PageCharset, CodeVersion, IsLink, IsDownload, IsNotBounce, FUniqID, OriginalURL, HID, IsOldCounter, IsEvent, IsParameter, DontCountHits, WithHash, HitColor, LocalEventTime, Age, Sex, Income, Interests, Robotness, RemoteIP, WindowName, OpenerName, HistoryLength, BrowserLanguage, BrowserCountry, SocialNetwork, SocialAction, HTTPError, SendTiming, DNSTiming, ConnectTiming, ResponseStartTiming, ResponseEndTiming, FetchTiming, SocialSourceNetworkID, SocialSourcePage, ParamPrice, ParamOrderID, ParamCurrency, ParamCurrencyID, OpenstatServiceName, OpenstatCampaignID, OpenstatAdID, OpenstatSourceID, UTMSource, UTMMedium, UTMCampaign, UTMContent, UTMTerm, FromTag, HasGCLID, RefererHash, URLHash, CLID], file_type=parquet, predicate=URL@13 LIKE %google% AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible query IITIIIIIIIIITTIIIIIIIIIITIIITIIIITTIIITIIIIIIIIIITIIIIITIIIIIITIIIIIIIIIITTTTIIIIIIIITITTITTTTTTTTTTIIIID SELECT * FROM hits WHERE "URL" LIKE '%google%' ORDER BY "EventTime" LIMIT 10; @@ -676,7 +676,7 @@ physical_plan 04)------SortExec: TopK(fetch=10), expr=[EventTime@0 ASC NULLS LAST], preserve_partitioning=[true] 05)--------FilterExec: SearchPhrase@1 != 06)----------RepartitionExec: partitioning=RoundRobinBatch(4), input_partitions=1 -07)------------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/core/tests/data/clickbench_hits_10.parquet]]}, projection=[EventTime, SearchPhrase], file_type=parquet, predicate=SearchPhrase@39 != AND DynamicFilter [ empty ], dynamic_rg_pruning=eligible, pruning_predicate=SearchPhrase_null_count@2 != row_count@3 AND (SearchPhrase_min@0 != OR != SearchPhrase_max@1), required_guarantees=[SearchPhrase not in ()] +07)------------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/core/tests/data/clickbench_hits_10.parquet]]}, projection=[EventTime, SearchPhrase], file_type=parquet, predicate=SearchPhrase@39 != AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible, pruning_predicate=SearchPhrase_null_count@2 != row_count@3 AND (SearchPhrase_min@0 != OR != SearchPhrase_max@1), required_guarantees=[SearchPhrase not in ()] query T SELECT "SearchPhrase" FROM hits WHERE "SearchPhrase" <> '' ORDER BY "EventTime" LIMIT 10; @@ -696,7 +696,7 @@ physical_plan 02)--SortExec: TopK(fetch=10), expr=[SearchPhrase@0 ASC NULLS LAST], preserve_partitioning=[true] 03)----FilterExec: SearchPhrase@0 != 04)------RepartitionExec: partitioning=RoundRobinBatch(4), input_partitions=1 -05)--------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/core/tests/data/clickbench_hits_10.parquet]]}, projection=[SearchPhrase], file_type=parquet, predicate=SearchPhrase@39 != AND DynamicFilter [ empty ], dynamic_rg_pruning=eligible, pruning_predicate=SearchPhrase_null_count@2 != row_count@3 AND (SearchPhrase_min@0 != OR != SearchPhrase_max@1), required_guarantees=[SearchPhrase not in ()] +05)--------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/core/tests/data/clickbench_hits_10.parquet]]}, projection=[SearchPhrase], file_type=parquet, predicate=SearchPhrase@39 != AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible, pruning_predicate=SearchPhrase_null_count@2 != row_count@3 AND (SearchPhrase_min@0 != OR != SearchPhrase_max@1), required_guarantees=[SearchPhrase not in ()] query T SELECT "SearchPhrase" FROM hits WHERE "SearchPhrase" <> '' ORDER BY "SearchPhrase" LIMIT 10; @@ -720,7 +720,7 @@ physical_plan 04)------SortExec: TopK(fetch=10), expr=[EventTime@0 ASC NULLS LAST, SearchPhrase@1 ASC NULLS LAST], preserve_partitioning=[true] 05)--------FilterExec: SearchPhrase@1 != 06)----------RepartitionExec: partitioning=RoundRobinBatch(4), input_partitions=1 -07)------------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/core/tests/data/clickbench_hits_10.parquet]]}, projection=[EventTime, SearchPhrase], file_type=parquet, predicate=SearchPhrase@39 != AND DynamicFilter [ empty ], dynamic_rg_pruning=eligible, pruning_predicate=SearchPhrase_null_count@2 != row_count@3 AND (SearchPhrase_min@0 != OR != SearchPhrase_max@1), required_guarantees=[SearchPhrase not in ()] +07)------------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/core/tests/data/clickbench_hits_10.parquet]]}, projection=[EventTime, SearchPhrase], file_type=parquet, predicate=SearchPhrase@39 != AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible, pruning_predicate=SearchPhrase_null_count@2 != row_count@3 AND (SearchPhrase_min@0 != OR != SearchPhrase_max@1), required_guarantees=[SearchPhrase not in ()] query T SELECT "SearchPhrase" FROM hits WHERE "SearchPhrase" <> '' ORDER BY "EventTime", "SearchPhrase" LIMIT 10; diff --git a/datafusion/sqllogictest/test_files/dynamic_filter_pushdown_config.slt b/datafusion/sqllogictest/test_files/dynamic_filter_pushdown_config.slt index 6a6bad99f0840..594c0209a8975 100644 --- a/datafusion/sqllogictest/test_files/dynamic_filter_pushdown_config.slt +++ b/datafusion/sqllogictest/test_files/dynamic_filter_pushdown_config.slt @@ -90,7 +90,7 @@ logical_plan 02)--TableScan: test_parquet projection=[id, value, name] physical_plan 01)SortExec: TopK(fetch=3), expr=[value@1 DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/test_data.parquet]]}, projection=[id, value, name], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[value@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/test_data.parquet]]}, projection=[id, value, name], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[value@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible statement ok set datafusion.explain.analyze_level = summary; @@ -104,7 +104,7 @@ Plan with Metrics 03)----ProjectionExec: expr=[id@0 as id, value@1 as v, value@1 + id@0 as name], metrics=[output_rows=10, ] 04)------FilterExec: value@1 > 3, metrics=[output_rows=10, , selectivity=100% (10/10)] 05)--------RepartitionExec: partitioning=RoundRobinBatch(4), input_partitions=1, metrics=[output_rows=10, ] -06)----------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/test_data.parquet]]}, projection=[id, value], file_type=parquet, predicate=value@1 > 3 AND DynamicFilter [ value@1 IS NULL OR value@1 > 800 ], dynamic_rg_pruning=eligible, pruning_predicate=value_null_count@1 != row_count@2 AND value_max@0 > 3 AND (value_null_count@1 > 0 OR value_null_count@1 != row_count@2 AND value_max@0 > 800), required_guarantees=[], metrics=[output_rows=10, elapsed_compute=, output_bytes=80.0 B, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched -> 1 fully matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, bytes_processed=1147.0 B, bytes_scanned=210.0 B, page_index_load_skipped=1, row_groups_pruned_dynamic_filter=0, metadata_load_time=, scan_efficiency_ratio=18.31% (210/1.15 K)] +06)----------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/test_data.parquet]]}, projection=[id, value], file_type=parquet, predicate=value@1 > 3 AND Optional(DynamicFilter [ value@1 IS NULL OR value@1 > 800 ]), dynamic_rg_pruning=eligible, pruning_predicate=value_null_count@1 != row_count@2 AND value_max@0 > 3 AND (value_null_count@1 > 0 OR value_null_count@1 != row_count@2 AND value_max@0 > 800), required_guarantees=[], metrics=[output_rows=10, elapsed_compute=, output_bytes=80.0 B, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched -> 1 fully matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, bytes_processed=1147.0 B, bytes_scanned=210.0 B, page_index_load_skipped=1, row_groups_pruned_dynamic_filter=0, metadata_load_time=, scan_efficiency_ratio=18.31% (210/1.15 K)] statement ok set datafusion.explain.analyze_level = dev; @@ -157,7 +157,7 @@ logical_plan physical_plan 01)HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(id@0, id@0)], projection=[id@2, data@3, info@1] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_right.parquet]]}, projection=[id, info], file_type=parquet -03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id, data], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id, data], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Disable Join dynamic filter pushdown statement ok @@ -235,7 +235,7 @@ logical_plan physical_plan 01)HashJoinExec: mode=CollectLeft, join_type=Left, on=[(id@0, id@0)], projection=[id@2, data@3, info@1] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_right.parquet]]}, projection=[id, info], file_type=parquet -03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id, data], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id, data], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # RIGHT JOIN correctness: all right rows appear, unmatched left rows produce NULLs query ITT @@ -284,7 +284,7 @@ logical_plan physical_plan 01)HashJoinExec: mode=CollectLeft, join_type=RightSemi, on=[(id@0, id@0)] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_right.parquet]]}, projection=[id], file_type=parquet -03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id, data], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id, data], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # LEFT SEMI JOIN (physical LeftSemi): reverse table roles so optimizer keeps LeftSemi # (right_parquet has 3 rows < left_parquet has 5 rows, so no swap occurs). @@ -304,7 +304,7 @@ logical_plan physical_plan 01)HashJoinExec: mode=CollectLeft, join_type=LeftSemi, on=[(id@0, id@0)] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_right.parquet]]}, projection=[id, info], file_type=parquet -03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # LEFT SEMI (physical LeftSemi) correctness: only right rows with matching left ids query IT rowsort @@ -337,8 +337,8 @@ physical_plan 01)HashJoinExec: mode=CollectLeft, join_type=RightSemi, on=[(id@0, id@0)] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_right.parquet]]}, projection=[id], file_type=parquet 03)--SortExec: expr=[data@1 DESC], preserve_partitioning=[false] -04)----FilterExec: DynamicFilter [ empty ] -05)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id, data], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[data@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible +04)----FilterExec: Optional(DynamicFilter [ empty ]) +05)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id, data], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[data@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible statement count 0 SET datafusion.execution.parquet.pushdown_filters = true; @@ -361,7 +361,7 @@ physical_plan 01)HashJoinExec: mode=CollectLeft, join_type=RightSemi, on=[(id@0, id@0)] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_right.parquet]]}, projection=[id], file_type=parquet 03)--SortExec: expr=[data@1 DESC], preserve_partitioning=[false] -04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id, data], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[data@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible +04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id, data], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[data@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible statement count 0 RESET datafusion.execution.parquet.pushdown_filters; @@ -407,7 +407,7 @@ physical_plan 02)--RepartitionExec: partitioning=RoundRobinBatch(4), input_partitions=1 03)----HashJoinExec: mode=CollectLeft, join_type=LeftMark, on=[(id@0, id@0)] 04)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_right.parquet]]}, projection=[id, info], file_type=parquet -05)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +05)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # LEFT MARK correctness: all right rows match EXISTS, so all 3 appear query IT rowsort @@ -444,8 +444,8 @@ logical_plan physical_plan 01)SortExec: TopK(fetch=2), expr=[id@0 ASC NULLS LAST], preserve_partitioning=[false] 02)--HashJoinExec: mode=CollectLeft, join_type=RightSemi, on=[(id@0, id@0)] -03)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_right.parquet]]}, projection=[id], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible -04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id, data], file_type=parquet, predicate=DynamicFilter [ empty ] AND DynamicFilter [ empty ], dynamic_rg_pruning=eligible +03)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_right.parquet]]}, projection=[id], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible +04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id, data], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]) AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Correctness check query IT @@ -479,7 +479,7 @@ physical_plan 01)SortExec: TopK(fetch=2), expr=[id@0 ASC NULLS LAST], preserve_partitioning=[false] 02)--HashJoinExec: mode=CollectLeft, join_type=RightAnti, on=[(id@0, id@0)], null_aware 03)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_right.parquet]]}, projection=[id], file_type=parquet -04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id, data], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id, data], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Correctness check query IT @@ -550,7 +550,7 @@ logical_plan physical_plan 01)HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(id@0, id@0)], projection=[id@2, data@3, info@1] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_right.parquet]]}, projection=[id, info], file_type=parquet -03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id, data], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id, data], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Enable TopK, disable Join statement ok @@ -622,7 +622,7 @@ physical_plan 02)--CoalescePartitionsExec 03)----AggregateExec: mode=Partial, gby=[], aggr=[max(agg_parquet.score)] 04)------RepartitionExec: partitioning=RoundRobinBatch(4), input_partitions=1 -05)--------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/agg_data.parquet]]}, projection=[score], file_type=parquet, predicate=category@0 = alpha AND DynamicFilter [ empty ], dynamic_rg_pruning=eligible, pruning_predicate=category_null_count@2 != row_count@3 AND category_min@0 <= alpha AND alpha <= category_max@1, required_guarantees=[category in (alpha)] +05)--------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/agg_data.parquet]]}, projection=[score], file_type=parquet, predicate=category@0 = alpha AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible, pruning_predicate=category_null_count@2 != row_count@3 AND category_min@0 <= alpha AND alpha <= category_max@1, required_guarantees=[category in (alpha)] # Test 4b: COUNT + MAX — DynamicFilter should NOT appear here in mixed aggregates @@ -770,7 +770,7 @@ logical_plan physical_plan 01)HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(id@0, id@0)], projection=[id@2, data@3, info@1] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_right.parquet]]}, projection=[id, info], file_type=parquet -03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id, data], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_filter_pushdown_config/join_left.parquet]]}, projection=[id, data], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Test 6: Regression test for issue #20213 - dynamic filter applied to wrong table # when subquery join has same column names on both sides. diff --git a/datafusion/sqllogictest/test_files/dynamic_row_group_pruning.slt b/datafusion/sqllogictest/test_files/dynamic_row_group_pruning.slt index c674dede75706..cdf0e050b3f0a 100644 --- a/datafusion/sqllogictest/test_files/dynamic_row_group_pruning.slt +++ b/datafusion/sqllogictest/test_files/dynamic_row_group_pruning.slt @@ -84,7 +84,7 @@ logical_plan 02)--TableScan: t projection=[v] physical_plan 01)SortExec: TopK(fetch=3), expr=[v@0 DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_row_group_pruning/data.parquet]]}, projection=[v], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[v@0 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_row_group_pruning/data.parquet]]}, projection=[v], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[v@0 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible # `EXPLAIN ANALYZE` must surface the runtime metric # `row_groups_pruned_dynamic_filter` with a non-zero value. Five @@ -98,7 +98,7 @@ explain analyze select v from t order by v desc limit 3; ---- Plan with Metrics 01)SortExec: TopK(fetch=3), expr=[v@0 DESC], preserve_partitioning=[false], filter=[v@0 IS NULL OR v@0 > 12], metrics=[output_rows=3, elapsed_compute=, output_bytes=] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_row_group_pruning/data.parquet]]}, projection=[v], file_type=parquet, predicate=DynamicFilter [ v@0 IS NULL OR v@0 > 12 ], sort_order_for_reorder=[v@0 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible, pruning_predicate=v_null_count@0 > 0 OR v_null_count@0 != row_count@2 AND v_max@1 > 12, required_guarantees=[], metrics=[output_rows=3, elapsed_compute=, output_bytes=, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=5 total → 5 matched, row_groups_pruned_bloom_filter=5 total → 5 matched, page_index_pages_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, bytes_processed=, bytes_scanned=, row_groups_pruned_dynamic_filter=4, metadata_load_time=, scan_efficiency_ratio=] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/dynamic_row_group_pruning/data.parquet]]}, projection=[v], file_type=parquet, predicate=Optional(DynamicFilter [ v@0 IS NULL OR v@0 > 12 ]), sort_order_for_reorder=[v@0 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible, pruning_predicate=v_null_count@0 > 0 OR v_null_count@0 != row_count@2 AND v_max@1 > 12, required_guarantees=[], metrics=[output_rows=3, elapsed_compute=, output_bytes=, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=5 total → 5 matched, row_groups_pruned_bloom_filter=5 total → 5 matched, page_index_pages_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, bytes_processed=, bytes_scanned=, row_groups_pruned_dynamic_filter=4, metadata_load_time=, scan_efficiency_ratio=] # `EXPLAIN ANALYZE` with a *static* predicate must surface the # `row_filter_skipped_fully_matched` metric with a non-zero value — this diff --git a/datafusion/sqllogictest/test_files/explain_analyze.slt b/datafusion/sqllogictest/test_files/explain_analyze.slt index 511ba88ed3345..42a9129d49031 100644 --- a/datafusion/sqllogictest/test_files/explain_analyze.slt +++ b/datafusion/sqllogictest/test_files/explain_analyze.slt @@ -231,7 +231,7 @@ explain analyze select * from cat_tracking where species > 'M' AND s >= 50 order ---- Plan with Metrics 01)SortExec: TopK(fetch=3), expr=[species@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[species@0 < Nlpine Sheep], metrics=[] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND DynamicFilter [ species@0 < Nlpine Sheep ], sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND Optional(DynamicFilter [ species@0 < Nlpine Sheep ]), sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[] statement ok reset datafusion.explain.analyze_categories; @@ -247,7 +247,7 @@ explain analyze select * from cat_tracking where species > 'M' AND s >= 50 order ---- Plan with Metrics 01)SortExec: TopK(fetch=3), expr=[species@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[species@0 < Nlpine Sheep], metrics=[output_rows=3] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND DynamicFilter [ species@0 < Nlpine Sheep ], sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, scan_efficiency_ratio=21.75% (485/2.23 K)] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND Optional(DynamicFilter [ species@0 < Nlpine Sheep ]), sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, scan_efficiency_ratio=21.75% (485/2.23 K)] statement ok reset datafusion.explain.analyze_categories; @@ -262,7 +262,7 @@ explain analyze select * from cat_tracking where species > 'M' AND s >= 50 order ---- Plan with Metrics 01)SortExec: TopK(fetch=3), expr=[species@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[species@0 < Nlpine Sheep], metrics=[output_rows=3, output_bytes=] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND DynamicFilter [ species@0 < Nlpine Sheep ], sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, output_bytes=, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, bytes_processed=, bytes_scanned=, scan_efficiency_ratio=] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND Optional(DynamicFilter [ species@0 < Nlpine Sheep ]), sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, output_bytes=, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, bytes_processed=, bytes_scanned=, scan_efficiency_ratio=] statement ok reset datafusion.explain.analyze_categories; @@ -277,7 +277,7 @@ explain analyze select * from cat_tracking where species > 'M' AND s >= 50 order ---- Plan with Metrics 01)SortExec: TopK(fetch=3), expr=[species@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[species@0 < Nlpine Sheep], metrics=[output_rows=3, output_bytes=] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND DynamicFilter [ species@0 < Nlpine Sheep ], sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, output_bytes=, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, bytes_processed=, bytes_scanned=, scan_efficiency_ratio=] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND Optional(DynamicFilter [ species@0 < Nlpine Sheep ]), sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, output_bytes=, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, bytes_processed=, bytes_scanned=, scan_efficiency_ratio=] statement ok reset datafusion.explain.analyze_categories; @@ -292,7 +292,7 @@ explain analyze select * from cat_tracking where species > 'M' AND s >= 50 order ---- Plan with Metrics 01)SortExec: TopK(fetch=3), expr=[species@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[species@0 < Nlpine Sheep], metrics=[elapsed_compute=] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND DynamicFilter [ species@0 < Nlpine Sheep ], sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[elapsed_compute=, metadata_load_time=] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND Optional(DynamicFilter [ species@0 < Nlpine Sheep ]), sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[elapsed_compute=, metadata_load_time=] statement ok reset datafusion.explain.analyze_categories; @@ -550,7 +550,7 @@ EXPLAIN (ANALYZE, METRICS 'none', LEVEL summary) select * from cat_tracking wher ---- Plan with Metrics 01)SortExec: TopK(fetch=3), expr=[species@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[species@0 < Nlpine Sheep], metrics=[] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND DynamicFilter [ species@0 < Nlpine Sheep ], sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND Optional(DynamicFilter [ species@0 < Nlpine Sheep ]), sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[] # ---- (METRICS 'rows', LEVEL summary) — row-count metrics only ---- @@ -559,7 +559,7 @@ EXPLAIN (ANALYZE, METRICS 'rows', LEVEL summary) select * from cat_tracking wher ---- Plan with Metrics 01)SortExec: TopK(fetch=3), expr=[species@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[species@0 < Nlpine Sheep], metrics=[output_rows=3] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND DynamicFilter [ species@0 < Nlpine Sheep ], sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, scan_efficiency_ratio=21.75% (485/2.23 K)] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND Optional(DynamicFilter [ species@0 < Nlpine Sheep ]), sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, scan_efficiency_ratio=21.75% (485/2.23 K)] # ---- Quoted-string METRICS with multiple categories ---- @@ -568,7 +568,7 @@ EXPLAIN (ANALYZE, METRICS 'rows,bytes', LEVEL summary) select * from cat_trackin ---- Plan with Metrics 01)SortExec: TopK(fetch=3), expr=[species@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[species@0 < Nlpine Sheep], metrics=[output_rows=3, output_bytes=] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND DynamicFilter [ species@0 < Nlpine Sheep ], sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, output_bytes=, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, bytes_processed=, bytes_scanned=, scan_efficiency_ratio=] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND Optional(DynamicFilter [ species@0 < Nlpine Sheep ]), sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, output_bytes=, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, bytes_processed=, bytes_scanned=, scan_efficiency_ratio=] # ---- (METRICS 'timing', LEVEL summary) — timing metrics only ---- @@ -577,7 +577,7 @@ EXPLAIN (ANALYZE, METRICS 'timing', LEVEL summary) select * from cat_tracking wh ---- Plan with Metrics 01)SortExec: TopK(fetch=3), expr=[species@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[species@0 < Nlpine Sheep], metrics=[elapsed_compute=] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND DynamicFilter [ species@0 < Nlpine Sheep ], sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[elapsed_compute=, metadata_load_time=] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND Optional(DynamicFilter [ species@0 < Nlpine Sheep ]), sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[elapsed_compute=, metadata_load_time=] # ---- TIMING sugar: `METRICS 'rows,bytes', TIMING off` ↔ rows+bytes only ---- # Equivalent to METRICS 'rows,bytes' since the sugar removes the timing @@ -588,7 +588,7 @@ EXPLAIN (ANALYZE, METRICS 'rows,bytes', TIMING off, LEVEL summary) select * from ---- Plan with Metrics 01)SortExec: TopK(fetch=3), expr=[species@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[species@0 < Nlpine Sheep], metrics=[output_rows=3, output_bytes=] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND DynamicFilter [ species@0 < Nlpine Sheep ], sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, output_bytes=, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, bytes_processed=, bytes_scanned=, scan_efficiency_ratio=] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND Optional(DynamicFilter [ species@0 < Nlpine Sheep ]), sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, output_bytes=, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, bytes_processed=, bytes_scanned=, scan_efficiency_ratio=] # ---- TIMING sugar: `METRICS 'rows', TIMING on` ↔ rows + timing ---- @@ -597,7 +597,7 @@ EXPLAIN (ANALYZE, METRICS 'rows', TIMING on, LEVEL summary) select * from cat_tr ---- Plan with Metrics 01)SortExec: TopK(fetch=3), expr=[species@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[species@0 < Nlpine Sheep], metrics=[output_rows=3, elapsed_compute=] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND DynamicFilter [ species@0 < Nlpine Sheep ], sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, elapsed_compute=, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, metadata_load_time=, scan_efficiency_ratio=21.75% (485/2.23 K)] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND Optional(DynamicFilter [ species@0 < Nlpine Sheep ]), sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, elapsed_compute=, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, metadata_load_time=, scan_efficiency_ratio=21.75% (485/2.23 K)] # ---- SUMMARY sugar: `SUMMARY on` ↔ `LEVEL summary` ---- # Equivalent to METRICS 'rows', LEVEL summary above. @@ -607,7 +607,7 @@ EXPLAIN (ANALYZE, METRICS 'rows', SUMMARY on) select * from cat_tracking where s ---- Plan with Metrics 01)SortExec: TopK(fetch=3), expr=[species@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[species@0 < Nlpine Sheep], metrics=[output_rows=3] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND DynamicFilter [ species@0 < Nlpine Sheep ], sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, scan_efficiency_ratio=21.75% (485/2.23 K)] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND Optional(DynamicFilter [ species@0 < Nlpine Sheep ]), sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, scan_efficiency_ratio=21.75% (485/2.23 K)] # ---- Statement option overrides session config ---- # Session says 'timing' but statement-level `METRICS 'rows'` wins. @@ -620,7 +620,7 @@ EXPLAIN (ANALYZE, METRICS 'rows', LEVEL summary) select * from cat_tracking wher ---- Plan with Metrics 01)SortExec: TopK(fetch=3), expr=[species@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[species@0 < Nlpine Sheep], metrics=[output_rows=3] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND DynamicFilter [ species@0 < Nlpine Sheep ], sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, scan_efficiency_ratio=21.75% (485/2.23 K)] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND Optional(DynamicFilter [ species@0 < Nlpine Sheep ]), sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, scan_efficiency_ratio=21.75% (485/2.23 K)] # ---- pgjson format: structural golden with no metrics ---- @@ -682,7 +682,7 @@ EXPLAIN (ANALYZE, METRICS rows, LEVEL summary) select * from cat_tracking where ---- Plan with Metrics 01)SortExec: TopK(fetch=3), expr=[species@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[species@0 < Nlpine Sheep], metrics=[output_rows=3] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND DynamicFilter [ species@0 < Nlpine Sheep ], sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, scan_efficiency_ratio=21.75% (485/2.23 K)] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/explain_analyze/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND Optional(DynamicFilter [ species@0 < Nlpine Sheep ]), sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, scan_efficiency_ratio=21.75% (485/2.23 K)] statement ok reset datafusion.sql_parser.dialect; diff --git a/datafusion/sqllogictest/test_files/information_schema.slt b/datafusion/sqllogictest/test_files/information_schema.slt index 5305a7613326a..abf5e251d1d8f 100644 --- a/datafusion/sqllogictest/test_files/information_schema.slt +++ b/datafusion/sqllogictest/test_files/information_schema.slt @@ -231,6 +231,7 @@ datafusion.execution.max_spill_file_size_bytes 134217728 datafusion.execution.meta_fetch_concurrency 32 datafusion.execution.minimum_parallel_output_files 4 datafusion.execution.objectstore_writer_buffer_size 10485760 +datafusion.execution.optional_filter_min_saving_ns_per_row 20 datafusion.execution.parquet.allow_single_file_parallelism true datafusion.execution.parquet.binary_as_string false datafusion.execution.parquet.bloom_filter_fpp NULL @@ -392,6 +393,7 @@ datafusion.execution.max_spill_file_size_bytes 134217728 Maximum size in bytes f datafusion.execution.meta_fetch_concurrency 32 Number of files to read in parallel when inferring schema and statistics datafusion.execution.minimum_parallel_output_files 4 Guarantees a minimum level of output files running in parallel. RecordBatches will be distributed in round robin fashion to each parallel writer. Each writer is closed and a new file opened once soft_max_rows_per_output_file is reached. datafusion.execution.objectstore_writer_buffer_size 10485760 Size (bytes) of data buffer DataFusion uses when writing output files. This affects the size of the data chunks that are uploaded to remote object stores (e.g. AWS S3). If very large (>= 100 GiB) output files are being written, it may be necessary to increase this size to avoid errors from the remote end point. +datafusion.execution.optional_filter_min_saving_ns_per_row 20 The assumed work, in nanoseconds, that each row removed by an optional filter saves downstream. Optional filters are filters that are not needed for correctness, such as the dynamic filters that hash joins and TopK push down into scans. When an operator evaluates optional filters adaptively, it pauses an optional filter whose evaluation costs more than the work that it saves. Consumers that can measure the saving (the Parquet scan) add their measured decode cost. The default is about the cost of a hash table probe for one row. The best value depends on the hardware. datafusion.execution.parquet.allow_single_file_parallelism true (writing) Controls whether DataFusion will attempt to speed up writing parquet files by serializing them in parallel. Each column in each row group in each output file are serialized in parallel leveraging a maximum possible core count of n_files\*n_row_groups\*n_columns. datafusion.execution.parquet.binary_as_string false (reading) If true, parquet reader will read columns of `Binary/LargeBinary` with `Utf8`, and `BinaryView` with `Utf8View`. Parquet files generated by some legacy writers do not correctly set the UTF8 flag for strings, causing string columns to be loaded as BLOB instead. The parquet reader has special optimizations for `Utf8` validation, so reading such columns as strings is significantly faster than reading them as binary and then casting to string. datafusion.execution.parquet.bloom_filter_fpp NULL (writing) Sets bloom filter false positive probability. If NULL, uses default parquet writer setting diff --git a/datafusion/sqllogictest/test_files/join_dynamic_filter_transfer.slt b/datafusion/sqllogictest/test_files/join_dynamic_filter_transfer.slt index 04d3e4374f86a..65064dd09343e 100644 --- a/datafusion/sqllogictest/test_files/join_dynamic_filter_transfer.slt +++ b/datafusion/sqllogictest/test_files/join_dynamic_filter_transfer.slt @@ -104,8 +104,8 @@ physical_plan 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/join_dynamic_filter_transfer/dim.parquet]]}, projection=[d_key, d_val], file_type=parquet 03)--RepartitionExec: partitioning=RoundRobinBatch(4), input_partitions=1 04)----HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(m_key@0, f_key@0)], projection=[m_key@0, m_c@1, f_e@3] -05)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/join_dynamic_filter_transfer/mid.parquet]]}, projection=[m_key, m_c], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible -06)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/join_dynamic_filter_transfer/fact.parquet]]}, projection=[f_key, f_e], file_type=parquet, predicate=DynamicFilter [ empty ] AND DynamicFilter [ empty ], dynamic_rg_pruning=eligible +05)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/join_dynamic_filter_transfer/mid.parquet]]}, projection=[m_key, m_c], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible +06)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/join_dynamic_filter_transfer/fact.parquet]]}, projection=[f_key, f_e], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]) AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible query TII rowsort SELECT d.d_val, m.m_c, f.f_e diff --git a/datafusion/sqllogictest/test_files/joins.slt b/datafusion/sqllogictest/test_files/joins.slt index b594c68ce17d6..a1b7a5c77f7f3 100644 --- a/datafusion/sqllogictest/test_files/joins.slt +++ b/datafusion/sqllogictest/test_files/joins.slt @@ -2980,7 +2980,7 @@ physical_plan 02)--HashJoinExec: mode=CollectLeft, join_type=RightSemi, on=[(t2_id@0, t1_id@0)] 03)----DataSourceExec: partitions=1, partition_sizes=[1] 04)----SortExec: expr=[t1_id@0 ASC NULLS LAST], preserve_partitioning=[true] -05)------FilterExec: DynamicFilter [ empty ] +05)------FilterExec: Optional(DynamicFilter [ empty ]) 06)--------RepartitionExec: partitioning=RoundRobinBatch(2), input_partitions=1 07)----------DataSourceExec: partitions=1, partition_sizes=[1] @@ -3016,7 +3016,7 @@ physical_plan 02)--HashJoinExec: mode=CollectLeft, join_type=RightSemi, on=[(t2_id@0, t1_id@0)] 03)----DataSourceExec: partitions=1, partition_sizes=[1] 04)----SortExec: expr=[t1_id@0 ASC NULLS LAST], preserve_partitioning=[true] -05)------FilterExec: DynamicFilter [ empty ] +05)------FilterExec: Optional(DynamicFilter [ empty ]) 06)--------RepartitionExec: partitioning=RoundRobinBatch(2), input_partitions=1 07)----------DataSourceExec: partitions=1, partition_sizes=[1] @@ -3073,7 +3073,7 @@ physical_plan 02)--HashJoinExec: mode=CollectLeft, join_type=RightSemi, on=[(t2_id@0, t1_id@0)] 03)----DataSourceExec: partitions=1, partition_sizes=[1] 04)----SortExec: expr=[t1_id@0 ASC NULLS LAST], preserve_partitioning=[true] -05)------FilterExec: DynamicFilter [ empty ] +05)------FilterExec: Optional(DynamicFilter [ empty ]) 06)--------RepartitionExec: partitioning=RoundRobinBatch(2), input_partitions=1 07)----------DataSourceExec: partitions=1, partition_sizes=[1] @@ -3109,7 +3109,7 @@ physical_plan 02)--HashJoinExec: mode=CollectLeft, join_type=RightSemi, on=[(t2_id@0, t1_id@0)] 03)----DataSourceExec: partitions=1, partition_sizes=[1] 04)----SortExec: expr=[t1_id@0 ASC NULLS LAST], preserve_partitioning=[true] -05)------FilterExec: DynamicFilter [ empty ] +05)------FilterExec: Optional(DynamicFilter [ empty ]) 06)--------RepartitionExec: partitioning=RoundRobinBatch(2), input_partitions=1 07)----------DataSourceExec: partitions=1, partition_sizes=[1] @@ -3167,7 +3167,7 @@ physical_plan 02)--HashJoinExec: mode=CollectLeft, join_type=RightSemi, on=[(t2_id@0, t1_id@0)], filter=t2_name@1 != t1_name@0 03)----DataSourceExec: partitions=1, partition_sizes=[1] 04)----SortExec: expr=[t1_id@0 ASC NULLS LAST], preserve_partitioning=[true] -05)------FilterExec: DynamicFilter [ empty ] +05)------FilterExec: Optional(DynamicFilter [ empty ]) 06)--------RepartitionExec: partitioning=RoundRobinBatch(2), input_partitions=1 07)----------DataSourceExec: partitions=1, partition_sizes=[1] @@ -3184,7 +3184,7 @@ physical_plan 02)--HashJoinExec: mode=CollectLeft, join_type=RightSemi, on=[(t2_id@0, t1_id@0)], filter=t2_name@0 != t1_name@1 03)----DataSourceExec: partitions=1, partition_sizes=[1] 04)----SortExec: expr=[t1_id@0 ASC NULLS LAST], preserve_partitioning=[true] -05)------FilterExec: DynamicFilter [ empty ] +05)------FilterExec: Optional(DynamicFilter [ empty ]) 06)--------RepartitionExec: partitioning=RoundRobinBatch(2), input_partitions=1 07)----------DataSourceExec: partitions=1, partition_sizes=[1] @@ -3239,7 +3239,7 @@ physical_plan 02)--HashJoinExec: mode=CollectLeft, join_type=RightSemi, on=[(t2_id@0, t1_id@0)], filter=t2_name@1 != t1_name@0 03)----DataSourceExec: partitions=1, partition_sizes=[1] 04)----SortExec: expr=[t1_id@0 ASC NULLS LAST], preserve_partitioning=[true] -05)------FilterExec: DynamicFilter [ empty ] +05)------FilterExec: Optional(DynamicFilter [ empty ]) 06)--------RepartitionExec: partitioning=RoundRobinBatch(2), input_partitions=1 07)----------DataSourceExec: partitions=1, partition_sizes=[1] @@ -3256,7 +3256,7 @@ physical_plan 02)--HashJoinExec: mode=CollectLeft, join_type=RightSemi, on=[(t2_id@0, t1_id@0)], filter=t2_name@0 != t1_name@1 03)----DataSourceExec: partitions=1, partition_sizes=[1] 04)----SortExec: expr=[t1_id@0 ASC NULLS LAST], preserve_partitioning=[true] -05)------FilterExec: DynamicFilter [ empty ] +05)------FilterExec: Optional(DynamicFilter [ empty ]) 06)--------RepartitionExec: partitioning=RoundRobinBatch(2), input_partitions=1 07)----------DataSourceExec: partitions=1, partition_sizes=[1] @@ -4264,7 +4264,7 @@ physical_plan 01)HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(b@1, y@1)], filter=a@0 < x@1 02)--DataSourceExec: partitions=1, partition_sizes=[0] 03)--SortExec: expr=[x@0 ASC NULLS LAST], preserve_partitioning=[false] -04)----FilterExec: DynamicFilter [ empty ] +04)----FilterExec: Optional(DynamicFilter [ empty ]) 05)------DataSourceExec: partitions=1, partition_sizes=[0] # Test full join with limit @@ -4567,7 +4567,7 @@ physical_plan 04)------FilterExec: b@1 > 3, projection=[a@0] 05)--------DataSourceExec: partitions=2, partition_sizes=[1, 1] 06)----SortExec: expr=[c@2 DESC], preserve_partitioning=[true] -07)------FilterExec: DynamicFilter [ empty ] +07)------FilterExec: Optional(DynamicFilter [ empty ]) 08)--------DataSourceExec: partitions=2, partition_sizes=[1, 1] query TT @@ -4588,7 +4588,7 @@ physical_plan 04)------FilterExec: b@1 > 3, projection=[a@0] 05)--------DataSourceExec: partitions=2, partition_sizes=[1, 1] 06)----SortExec: expr=[c@2 DESC NULLS LAST], preserve_partitioning=[true] -07)------FilterExec: DynamicFilter [ empty ] +07)------FilterExec: Optional(DynamicFilter [ empty ]) 08)--------DataSourceExec: partitions=2, partition_sizes=[1, 1] query III diff --git a/datafusion/sqllogictest/test_files/limit.slt b/datafusion/sqllogictest/test_files/limit.slt index 1ff6ca4fb0253..b706dbd952543 100644 --- a/datafusion/sqllogictest/test_files/limit.slt +++ b/datafusion/sqllogictest/test_files/limit.slt @@ -918,7 +918,7 @@ physical_plan 01)ProjectionExec: expr=[1 as foo] 02)--SortPreservingMergeExec: [part_key@0 ASC NULLS LAST], fetch=1 03)----SortExec: TopK(fetch=1), expr=[part_key@0 ASC NULLS LAST], preserve_partitioning=[true] -04)------DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/limit/test_limit_with_partitions/part-0.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/limit/test_limit_with_partitions/part-1.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/limit/test_limit_with_partitions/part-2.parquet]]}, projection=[part_key], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[part_key@0 ASC NULLS LAST], dynamic_rg_pruning=eligible +04)------DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/limit/test_limit_with_partitions/part-0.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/limit/test_limit_with_partitions/part-1.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/limit/test_limit_with_partitions/part-2.parquet]]}, projection=[part_key], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[part_key@0 ASC NULLS LAST], dynamic_rg_pruning=eligible query I with selection as ( diff --git a/datafusion/sqllogictest/test_files/limit_pruning.slt b/datafusion/sqllogictest/test_files/limit_pruning.slt index dd04506ed3e54..4d2e1b15be7d2 100644 --- a/datafusion/sqllogictest/test_files/limit_pruning.slt +++ b/datafusion/sqllogictest/test_files/limit_pruning.slt @@ -120,7 +120,7 @@ explain analyze select * from tracking_data where species > 'M' AND s >= 50 orde ---- Plan with Metrics 01)SortExec: TopK(fetch=3), expr=[species@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[species@0 < Nlpine Sheep], metrics=[output_rows=3, elapsed_compute=, output_bytes=] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/limit_pruning/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND DynamicFilter [ species@0 < Nlpine Sheep ], sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, elapsed_compute=, output_bytes=, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, bytes_processed=, bytes_scanned=, row_groups_pruned_dynamic_filter=0, metadata_load_time=, scan_efficiency_ratio= (/)] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/limit_pruning/data.parquet]]}, projection=[species, s], file_type=parquet, predicate=species@0 > M AND s@1 >= 50 AND Optional(DynamicFilter [ species@0 < Nlpine Sheep ]), sort_order_for_reorder=[species@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=species_null_count@1 != row_count@2 AND species_max@0 > M AND s_null_count@4 != row_count@2 AND s_max@3 >= 50 AND species_null_count@1 != row_count@2 AND species_min@5 < Nlpine Sheep, required_guarantees=[], metrics=[output_rows=3, elapsed_compute=, output_bytes=, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=4 total → 3 matched -> 1 fully matched, row_groups_pruned_bloom_filter=3 total → 3 matched, page_index_pages_pruned=2 total → 2 matched, page_index_pages_skipped_by_fully_matched=1, limit_pruned_row_groups=0 total → 0 matched, bytes_processed=, bytes_scanned=, row_groups_pruned_dynamic_filter=0, metadata_load_time=, scan_efficiency_ratio= (/)] statement ok drop table tracking_data; diff --git a/datafusion/sqllogictest/test_files/null_aware_mark_join.slt b/datafusion/sqllogictest/test_files/null_aware_mark_join.slt index c45d87b51efc2..ee5074916fed5 100644 --- a/datafusion/sqllogictest/test_files/null_aware_mark_join.slt +++ b/datafusion/sqllogictest/test_files/null_aware_mark_join.slt @@ -522,7 +522,7 @@ physical_plan 01)FilterExec: NOT mark@1 OR value@0 = zzz, projection=[value@0] 02)--HashJoinExec: mode=CollectLeft, join_type=LeftMark, on=[(id@0, id@0), (grp@1, grp@1)], projection=[value@2, mark@3], null_aware 03)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/null_aware_mark_join/outer.parquet]]}, projection=[id, grp, value], file_type=parquet -04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/null_aware_mark_join/inner.parquet]]}, projection=[id, grp], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/null_aware_mark_join/inner.parquet]]}, projection=[id, grp], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # `a` (grp 1 = {2, NULL}) is UNKNOWN and `b` matches, so only `c` (no match in # grp 2) and `d` (empty grp 3) pass. A dropped NULL probe row would wrongly diff --git a/datafusion/sqllogictest/test_files/preserve_file_partitioning.slt b/datafusion/sqllogictest/test_files/preserve_file_partitioning.slt index e2dd22cc82bba..94a05cc46455c 100644 --- a/datafusion/sqllogictest/test_files/preserve_file_partitioning.slt +++ b/datafusion/sqllogictest/test_files/preserve_file_partitioning.slt @@ -367,7 +367,7 @@ physical_plan 08)--------------FilterExec: service@2 = log 09)----------------RepartitionExec: partitioning=RoundRobinBatch(3), input_partitions=1 10)------------------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/dimension/data.parquet]]}, projection=[d_dkey, env, service], file_type=parquet, predicate=service@2 = log, pruning_predicate=service_null_count@2 != row_count@3 AND service_min@0 <= log AND log <= service_max@1, required_guarantees=[service in (log)] -11)------------DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/fact/f_dkey=A/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/fact/f_dkey=B/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/fact/f_dkey=C/data.parquet]]}, projection=[value, f_dkey], output_ordering=[f_dkey@1 ASC NULLS LAST], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +11)------------DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/fact/f_dkey=A/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/fact/f_dkey=B/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/fact/f_dkey=C/data.parquet]]}, projection=[value, f_dkey], output_ordering=[f_dkey@1 ASC NULLS LAST], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Verify results without optimization query TTTIR rowsort @@ -418,7 +418,7 @@ physical_plan 06)----------FilterExec: service@2 = log 07)------------RepartitionExec: partitioning=RoundRobinBatch(3), input_partitions=1 08)--------------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/dimension/data.parquet]]}, projection=[d_dkey, env, service], file_type=parquet, predicate=service@2 = log, pruning_predicate=service_null_count@2 != row_count@3 AND service_min@0 <= log AND log <= service_max@1, required_guarantees=[service in (log)] -09)--------DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/fact/f_dkey=A/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/fact/f_dkey=B/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/fact/f_dkey=C/data.parquet]]}, projection=[value, f_dkey], output_ordering=[f_dkey@1 ASC NULLS LAST], output_partitioning=Hash([f_dkey@1], 3), file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +09)--------DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/fact/f_dkey=A/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/fact/f_dkey=B/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/fact/f_dkey=C/data.parquet]]}, projection=[value, f_dkey], output_ordering=[f_dkey@1 ASC NULLS LAST], output_partitioning=Hash([f_dkey@1], 3), file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible query TTTIR rowsort SELECT f.f_dkey, MAX(d.env), MAX(d.service), count(*), sum(f.value) @@ -643,7 +643,7 @@ physical_plan 05)--------RepartitionExec: partitioning=Hash([d_dkey@1], 3), input_partitions=3 06)----------DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/dimension_partitioned/d_dkey=A/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/dimension_partitioned/d_dkey=B/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/dimension_partitioned/d_dkey=C/data.parquet]]}, projection=[env, d_dkey], file_type=parquet 07)--------RepartitionExec: partitioning=Hash([f_dkey@1], 3), input_partitions=3 -08)----------DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/fact/f_dkey=A/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/fact/f_dkey=B/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/fact/f_dkey=C/data.parquet]]}, projection=[value, f_dkey], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +08)----------DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/fact/f_dkey=A/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/fact/f_dkey=B/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/preserve_file_partitioning/fact/f_dkey=C/data.parquet]]}, projection=[value, f_dkey], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]) AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible query TTR rowsort SELECT f.f_dkey, d.env, sum(f.value) diff --git a/datafusion/sqllogictest/test_files/projection_pushdown.slt b/datafusion/sqllogictest/test_files/projection_pushdown.slt index c92c95fdfbc59..37fcf918a079b 100644 --- a/datafusion/sqllogictest/test_files/projection_pushdown.slt +++ b/datafusion/sqllogictest/test_files/projection_pushdown.slt @@ -445,7 +445,7 @@ logical_plan 03)----TableScan: simple_struct projection=[id, s] physical_plan 01)SortExec: TopK(fetch=3), expr=[id@0 ASC NULLS LAST], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, get_field(s@1, value) as simple_struct.s[value]], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, get_field(s@1, value) as simple_struct.s[value]], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible # Verify correctness query II @@ -468,7 +468,7 @@ logical_plan 03)----TableScan: simple_struct projection=[id, s] physical_plan 01)SortExec: TopK(fetch=3), expr=[id@0 ASC NULLS LAST], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, get_field(s@1, value) + 1 as simple_struct.s[value] + Int64(1)], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, get_field(s@1, value) + 1 as simple_struct.s[value] + Int64(1)], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible # Verify correctness query II @@ -491,7 +491,7 @@ logical_plan 03)----TableScan: simple_struct projection=[id, s] physical_plan 01)SortExec: TopK(fetch=3), expr=[id@0 ASC NULLS LAST], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, get_field(s@1, value) as simple_struct.s[value], get_field(s@1, label) as simple_struct.s[label]], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, get_field(s@1, value) as simple_struct.s[value], get_field(s@1, label) as simple_struct.s[label]], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible # Verify correctness query IIT @@ -514,7 +514,7 @@ logical_plan 03)----TableScan: nested_struct projection=[id, nested] physical_plan 01)SortExec: TopK(fetch=2), expr=[id@0 ASC NULLS LAST], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/nested.parquet]]}, projection=[id, get_field(nested@1, outer, inner) as nested_struct.nested[outer][inner]], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/nested.parquet]]}, projection=[id, get_field(nested@1, outer, inner) as nested_struct.nested[outer][inner]], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible # Verify correctness query II @@ -536,7 +536,7 @@ logical_plan 03)----TableScan: simple_struct projection=[id, s] physical_plan 01)SortExec: TopK(fetch=3), expr=[id@0 ASC NULLS LAST], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, get_field(s@1, label) || _suffix as simple_struct.s[label] || Utf8("_suffix")], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, get_field(s@1, label) || _suffix as simple_struct.s[label] || Utf8("_suffix")], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible # Verify correctness query IT @@ -596,7 +596,7 @@ physical_plan 01)ProjectionExec: expr=[id@1 as id, __datafusion_extracted_1@0 as simple_struct.s[value]] 02)--SortExec: TopK(fetch=2), expr=[__datafusion_extracted_1@0 ASC NULLS LAST], preserve_partitioning=[false] 03)----FilterExec: id@1 > 1 -04)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[get_field(s@1, value) as __datafusion_extracted_1, id], file_type=parquet, predicate=id@0 > 1 AND DynamicFilter [ empty ], dynamic_rg_pruning=eligible, pruning_predicate=id_null_count@1 != row_count@2 AND id_max@0 > 1, required_guarantees=[] +04)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[get_field(s@1, value) as __datafusion_extracted_1, id], file_type=parquet, predicate=id@0 > 1 AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible, pruning_predicate=id_null_count@1 != row_count@2 AND id_max@0 > 1, required_guarantees=[] # Verify correctness query II @@ -622,7 +622,7 @@ physical_plan 01)SortExec: TopK(fetch=2), expr=[id@0 ASC NULLS LAST], preserve_partitioning=[false] 02)--ProjectionExec: expr=[id@1 as id, __datafusion_extracted_1@0 + 1 as simple_struct.s[value] + Int64(1)] 03)----FilterExec: id@1 > 1 -04)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[get_field(s@1, value) as __datafusion_extracted_1, id], file_type=parquet, predicate=id@0 > 1 AND DynamicFilter [ empty ], dynamic_rg_pruning=eligible, pruning_predicate=id_null_count@1 != row_count@2 AND id_max@0 > 1, required_guarantees=[] +04)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[get_field(s@1, value) as __datafusion_extracted_1, id], file_type=parquet, predicate=id@0 > 1 AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible, pruning_predicate=id_null_count@1 != row_count@2 AND id_max@0 > 1, required_guarantees=[] # Verify correctness query II @@ -714,7 +714,7 @@ logical_plan physical_plan 01)SortPreservingMergeExec: [id@0 ASC NULLS LAST], fetch=3 02)--SortExec: TopK(fetch=3), expr=[id@0 ASC NULLS LAST], preserve_partitioning=[true] -03)----DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part1.parquet, WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part2.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part3.parquet, WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part4.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part5.parquet]]}, projection=[id, get_field(s@1, value) as multi_struct.s[value]], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible +03)----DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part1.parquet, WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part2.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part3.parquet, WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part4.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part5.parquet]]}, projection=[id, get_field(s@1, value) as multi_struct.s[value]], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible # Verify correctness query II @@ -738,7 +738,7 @@ logical_plan physical_plan 01)SortPreservingMergeExec: [id@0 ASC NULLS LAST], fetch=3 02)--SortExec: TopK(fetch=3), expr=[id@0 ASC NULLS LAST], preserve_partitioning=[true] -03)----DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part1.parquet, WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part2.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part3.parquet, WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part4.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part5.parquet]]}, projection=[id, get_field(s@1, value) + 1 as multi_struct.s[value] + Int64(1)], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible +03)----DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part1.parquet, WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part2.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part3.parquet, WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part4.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/multi/part5.parquet]]}, projection=[id, get_field(s@1, value) + 1 as multi_struct.s[value] + Int64(1)], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible # Verify correctness query II @@ -875,7 +875,7 @@ logical_plan 03)----TableScan: simple_struct projection=[id, s] physical_plan 01)SortExec: TopK(fetch=3), expr=[id@0 ASC NULLS LAST], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, get_field(s@1, value) as simple_struct.s[value], get_field(s@1, value) + 10 as simple_struct.s[value] + Int64(10), get_field(s@1, label) as simple_struct.s[label]], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, get_field(s@1, value) as simple_struct.s[value], get_field(s@1, value) + 10 as simple_struct.s[value] + Int64(10), get_field(s@1, label) as simple_struct.s[label]], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible # Verify correctness query IIIT @@ -898,7 +898,7 @@ logical_plan 03)----TableScan: simple_struct projection=[id] physical_plan 01)SortExec: TopK(fetch=3), expr=[id@0 ASC NULLS LAST], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, 42 as constant], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, 42 as constant], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible # Verify correctness query II @@ -920,7 +920,7 @@ logical_plan 02)--TableScan: simple_struct projection=[id] physical_plan 01)SortExec: TopK(fetch=3), expr=[id@0 ASC NULLS LAST], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible # Verify correctness query I @@ -948,7 +948,7 @@ logical_plan 03)----TableScan: simple_struct projection=[id] physical_plan 01)SortExec: TopK(fetch=3), expr=[id@0 ASC NULLS LAST], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, id@0 + 100 as computed], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, id@0 + 100 as computed], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible # Verify correctness query II @@ -1040,7 +1040,7 @@ logical_plan 03)----TableScan: simple_struct projection=[id, s] physical_plan 01)SortExec: TopK(fetch=3), expr=[id@0 ASC NULLS LAST], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, get_field(s@1, value) + id@0 as combined], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, get_field(s@1, value) + id@0 as combined], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible # Verify correctness query II @@ -1096,7 +1096,7 @@ logical_plan 03)----TableScan: simple_struct projection=[id, s] physical_plan 01)SortExec: TopK(fetch=2), expr=[id@0 ASC NULLS LAST], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, 42 as answer, get_field(s@1, label) as simple_struct.s[label]], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, 42 as answer, get_field(s@1, label) as simple_struct.s[label]], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible # Verify correctness query IIT @@ -1119,7 +1119,7 @@ logical_plan 03)----TableScan: simple_struct projection=[id, s] physical_plan 01)SortExec: TopK(fetch=2), expr=[id@0 ASC NULLS LAST], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, get_field(s@1, value) + 100 as simple_struct.s[value] + Int64(100), get_field(s@1, label) || _test as simple_struct.s[label] || Utf8("_test")], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, get_field(s@1, value) + 100 as simple_struct.s[value] + Int64(100), get_field(s@1, label) || _test as simple_struct.s[label] || Utf8("_test")], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible # Verify correctness query IIT @@ -1318,7 +1318,7 @@ logical_plan physical_plan 01)ProjectionExec: expr=[id@0 as id] 02)--SortExec: TopK(fetch=2), expr=[__datafusion_extracted_1@1 ASC NULLS LAST], preserve_partitioning=[false] -03)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, get_field(s@1, value) as __datafusion_extracted_1], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +03)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id, get_field(s@1, value) as __datafusion_extracted_1], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Verify correctness query I @@ -1425,7 +1425,7 @@ logical_plan physical_plan 01)HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(__datafusion_extracted_1@0, __datafusion_extracted_2 * Int64(10)@2)], projection=[id@1, id@3] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[get_field(s@1, value) as __datafusion_extracted_1, id], file_type=parquet -03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/join_right.parquet]]}, projection=[get_field(s@1, level) as __datafusion_extracted_2, id, get_field(s@1, level) * 10 as __datafusion_extracted_2 * Int64(10)], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/join_right.parquet]]}, projection=[get_field(s@1, level) as __datafusion_extracted_2, id, get_field(s@1, level) * 10 as __datafusion_extracted_2 * Int64(10)], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Verify correctness - value = level * 10 # simple_struct: (1,100), (2,200), (3,150), (4,300), (5,250) @@ -1461,7 +1461,7 @@ physical_plan 01)HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(id@0, id@0)] 02)--FilterExec: __datafusion_extracted_1@0 > 150, projection=[id@1] 03)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[get_field(s@1, value) as __datafusion_extracted_1, id], file_type=parquet, predicate=get_field(s@1, value) > 150 -04)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/join_right.parquet]]}, projection=[id], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +04)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/join_right.parquet]]}, projection=[id], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Verify correctness - id matches and value > 150 query II @@ -1501,7 +1501,7 @@ physical_plan 02)--FilterExec: __datafusion_extracted_1@0 > 100, projection=[id@1] 03)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[get_field(s@1, value) as __datafusion_extracted_1, id], file_type=parquet, predicate=get_field(s@1, value) > 100 04)--FilterExec: __datafusion_extracted_2@0 > 3, projection=[id@1] -05)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/join_right.parquet]]}, projection=[get_field(s@1, level) as __datafusion_extracted_2, id], file_type=parquet, predicate=get_field(s@1, level) > 3 AND DynamicFilter [ empty ], dynamic_rg_pruning=eligible +05)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/join_right.parquet]]}, projection=[get_field(s@1, level) as __datafusion_extracted_2, id], file_type=parquet, predicate=get_field(s@1, level) > 3 AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Verify correctness - id matches, value > 100, and level > 3 # Matching ids where value > 100: 2(200), 3(150), 4(300), 5(250) @@ -1537,7 +1537,7 @@ physical_plan 01)ProjectionExec: expr=[id@0 as id, __datafusion_extracted_1@1 as simple_struct.s[label], __datafusion_extracted_2@2 as join_right.s[role]] 02)--HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(id@1, id@1)], projection=[id@1, __datafusion_extracted_1@0, __datafusion_extracted_2@2] 03)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[get_field(s@1, label) as __datafusion_extracted_1, id], file_type=parquet -04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/join_right.parquet]]}, projection=[get_field(s@1, role) as __datafusion_extracted_2, id], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/join_right.parquet]]}, projection=[get_field(s@1, role) as __datafusion_extracted_2, id], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Verify correctness query ITT @@ -1569,7 +1569,7 @@ logical_plan physical_plan 01)HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(id@0, id@0)] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[id], file_type=parquet -03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/join_right.parquet]]}, projection=[id], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/join_right.parquet]]}, projection=[id], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Verify correctness query II @@ -1608,7 +1608,7 @@ physical_plan 02)--HashJoinExec: mode=CollectLeft, join_type=Left, on=[(id@1, id@0)], projection=[id@1, __datafusion_extracted_2@0, __datafusion_extracted_3@3] 03)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[get_field(s@1, value) as __datafusion_extracted_2, id], file_type=parquet 04)----FilterExec: __datafusion_extracted_1@0 > 5, projection=[id@1, __datafusion_extracted_3@2] -05)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/join_right.parquet]]}, projection=[get_field(s@1, level) as __datafusion_extracted_1, id, get_field(s@1, level) as __datafusion_extracted_3], file_type=parquet, predicate=get_field(s@1, level) > 5 AND DynamicFilter [ empty ], dynamic_rg_pruning=eligible +05)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/join_right.parquet]]}, projection=[get_field(s@1, level) as __datafusion_extracted_1, id, get_field(s@1, level) as __datafusion_extracted_3], file_type=parquet, predicate=get_field(s@1, level) > 5 AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Verify correctness - left join with level > 5 condition # Only join_right rows with level > 5 are matched: id=1 (level=10), id=4 (level=8) @@ -1902,7 +1902,7 @@ physical_plan 01)ProjectionExec: expr=[__datafusion_extracted_3@0 as s.s[value], __datafusion_extracted_4@1 as j.s[role]] 02)--HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(id@2, id@2)], filter=__datafusion_extracted_1@1 > __datafusion_extracted_2@0, projection=[__datafusion_extracted_3@4, __datafusion_extracted_4@1] 03)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/join_right.parquet]]}, projection=[get_field(s@1, level) as __datafusion_extracted_2, get_field(s@1, role) as __datafusion_extracted_4, id], file_type=parquet -04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[get_field(s@1, value) as __datafusion_extracted_1, get_field(s@1, value) as __datafusion_extracted_3, id], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[get_field(s@1, value) as __datafusion_extracted_1, get_field(s@1, value) as __datafusion_extracted_3, id], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Verify correctness - only admin roles match (ids 1 and 4) query II @@ -1938,7 +1938,7 @@ logical_plan physical_plan 01)HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(id@1, id@1)], filter=__datafusion_extracted_1@0 > __datafusion_extracted_2@1, projection=[id@1, id@3] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/simple.parquet]]}, projection=[get_field(s@1, value) as __datafusion_extracted_1, id], file_type=parquet -03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/join_right.parquet]]}, projection=[get_field(s@1, level) as __datafusion_extracted_2, id], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/projection_pushdown/join_right.parquet]]}, projection=[get_field(s@1, level) as __datafusion_extracted_2, id], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Verify correctness - all rows match since value >> level for all ids # simple_struct: (1,100), (2,200), (3,150), (4,300), (5,250) diff --git a/datafusion/sqllogictest/test_files/push_down_filter_parquet.slt b/datafusion/sqllogictest/test_files/push_down_filter_parquet.slt index 674f866fc6274..626b39dc039b3 100644 --- a/datafusion/sqllogictest/test_files/push_down_filter_parquet.slt +++ b/datafusion/sqllogictest/test_files/push_down_filter_parquet.slt @@ -158,7 +158,7 @@ physical_plan 01)HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(k@0, k@0)] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/small_table.parquet]]}, projection=[k], output_ordering=[k@0 ASC NULLS LAST], file_type=parquet 03)--RepartitionExec: partitioning=RoundRobinBatch(4), input_partitions=1, maintains_sort_order=true -04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/large_table.parquet]]}, projection=[k, v], output_ordering=[k@0 ASC NULLS LAST], file_type=parquet, predicate=v@1 >= 50 AND DynamicFilter [ empty ], dynamic_rg_pruning=eligible, pruning_predicate=v_null_count@1 != row_count@2 AND v_max@0 >= 50, required_guarantees=[] +04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/large_table.parquet]]}, projection=[k, v], output_ordering=[k@0 ASC NULLS LAST], file_type=parquet, predicate=v@1 >= 50 AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible, pruning_predicate=v_null_count@1 != row_count@2 AND v_max@0 >= 50, required_guarantees=[] statement ok drop table small_table; @@ -206,7 +206,7 @@ EXPLAIN ANALYZE SELECT t FROM topk_pushdown ORDER BY t * t LIMIT 10; ---- Plan with Metrics 01)SortExec: TopK(fetch=10), expr=[t@0 * t@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[t@0 * t@0 < 1884329474306198481], metrics=[output_rows=10, output_batches=1, row_replacements=10] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/topk_pushdown.parquet]]}, projection=[t], output_ordering=[t@0 ASC NULLS LAST], file_type=parquet, predicate=DynamicFilter [ t@0 * t@0 < 1884329474306198481 ], dynamic_rg_pruning=eligible, metrics=[output_rows=128, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=782 total → 782 matched, row_groups_pruned_bloom_filter=782 total → 782 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=128, pushdown_rows_pruned=99.87 K, predicate_cache_inner_records=128, predicate_cache_records=128, scan_efficiency_ratio=64.87% (258.7 K/398.8 K)] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/topk_pushdown.parquet]]}, projection=[t], output_ordering=[t@0 ASC NULLS LAST], file_type=parquet, predicate=Optional(DynamicFilter [ t@0 * t@0 < 1884329474306198481 ]), dynamic_rg_pruning=eligible, metrics=[output_rows=128, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=782 total → 782 matched, row_groups_pruned_bloom_filter=782 total → 782 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=128, pushdown_rows_pruned=99.87 K, predicate_cache_inner_records=128, predicate_cache_records=128, scan_efficiency_ratio=64.87% (258.7 K/398.8 K)] statement ok reset datafusion.explain.analyze_categories; @@ -257,7 +257,7 @@ EXPLAIN SELECT * FROM topk_single_col ORDER BY b DESC LIMIT 1; ---- physical_plan 01)SortExec: TopK(fetch=1), expr=[b@1 DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/topk_single_col.parquet]]}, projection=[a, b, c], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[b@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/topk_single_col.parquet]]}, projection=[a, b, c], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[b@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible statement ok set datafusion.explain.analyze_categories = 'rows'; @@ -268,7 +268,7 @@ EXPLAIN ANALYZE SELECT * FROM topk_single_col ORDER BY b DESC LIMIT 1; ---- Plan with Metrics 01)SortExec: TopK(fetch=1), expr=[b@1 DESC], preserve_partitioning=[false], filter=[b@1 IS NULL OR b@1 > bd], metrics=[output_rows=1, output_batches=1, row_replacements=1] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/topk_single_col.parquet]]}, projection=[a, b, c], file_type=parquet, predicate=DynamicFilter [ b@1 IS NULL OR b@1 > bd ], sort_order_for_reorder=[b@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible, pruning_predicate=b_null_count@0 > 0 OR b_null_count@0 != row_count@2 AND b_max@1 > bd, required_guarantees=[], metrics=[output_rows=4, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=4, pushdown_rows_pruned=0, predicate_cache_inner_records=4, predicate_cache_records=4, scan_efficiency_ratio=21.94% (222/1.01 K)] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/topk_single_col.parquet]]}, projection=[a, b, c], file_type=parquet, predicate=Optional(DynamicFilter [ b@1 IS NULL OR b@1 > bd ]), sort_order_for_reorder=[b@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible, pruning_predicate=b_null_count@0 > 0 OR b_null_count@0 != row_count@2 AND b_max@1 > bd, required_guarantees=[], metrics=[output_rows=4, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=4, pushdown_rows_pruned=0, predicate_cache_inner_records=4, predicate_cache_records=4, scan_efficiency_ratio=21.94% (222/1.01 K)] statement ok reset datafusion.explain.analyze_categories; @@ -319,7 +319,7 @@ EXPLAIN ANALYZE SELECT * FROM topk_multi_col ORDER BY b ASC NULLS LAST, a DESC L ---- Plan with Metrics 01)SortExec: TopK(fetch=2), expr=[b@1 ASC NULLS LAST, a@0 DESC], preserve_partitioning=[false], filter=[b@1 < bb OR b@1 = bb AND (a@0 IS NULL OR a@0 > ac)], metrics=[output_rows=2, output_batches=1, row_replacements=2] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/topk_multi_col.parquet]]}, projection=[a, b, c], file_type=parquet, predicate=DynamicFilter [ b@1 < bb OR b@1 = bb AND (a@0 IS NULL OR a@0 > ac) ], sort_order_for_reorder=[b@1 ASC NULLS LAST, a@0 DESC], dynamic_rg_pruning=eligible, pruning_predicate=b_null_count@1 != row_count@2 AND b_min@0 < bb OR b_null_count@1 != row_count@2 AND b_min@0 <= bb AND bb <= b_max@3 AND (a_null_count@4 > 0 OR a_null_count@4 != row_count@2 AND a_max@5 > ac), required_guarantees=[], metrics=[output_rows=4, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=4, pushdown_rows_pruned=0, predicate_cache_inner_records=8, predicate_cache_records=8, scan_efficiency_ratio=21.94% (222/1.01 K)] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/topk_multi_col.parquet]]}, projection=[a, b, c], file_type=parquet, predicate=Optional(DynamicFilter [ b@1 < bb OR b@1 = bb AND (a@0 IS NULL OR a@0 > ac) ]), sort_order_for_reorder=[b@1 ASC NULLS LAST, a@0 DESC], dynamic_rg_pruning=eligible, pruning_predicate=b_null_count@1 != row_count@2 AND b_min@0 < bb OR b_null_count@1 != row_count@2 AND b_min@0 <= bb AND bb <= b_max@3 AND (a_null_count@4 > 0 OR a_null_count@4 != row_count@2 AND a_max@5 > ac), required_guarantees=[], metrics=[output_rows=4, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=4, pushdown_rows_pruned=0, predicate_cache_inner_records=8, predicate_cache_records=8, scan_efficiency_ratio=21.94% (222/1.01 K)] statement ok reset datafusion.explain.analyze_categories; @@ -389,7 +389,7 @@ FROM join_probe p INNER JOIN join_build AS build Plan with Metrics 01)HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(a@0, a@0), (b@1, b@1)], projection=[a@3, b@4, c@2, e@5], metrics=[output_rows=2, output_batches=1, array_map_created_count=0, build_input_batches=1, build_input_rows=2, input_batches=1, input_rows=2, avg_fanout=100% (2/2), probe_hit_rate=100% (2/2)] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/join_build.parquet]]}, projection=[a, b, c], file_type=parquet, metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=0, pushdown_rows_pruned=0, predicate_cache_inner_records=0, predicate_cache_records=0, scan_efficiency_ratio=19.88% (196/986)] -03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/join_probe.parquet]]}, projection=[a, b, e], file_type=parquet, predicate=DynamicFilter [ a@0 >= aa AND a@0 <= ab AND b@1 >= ba AND b@1 <= bb AND struct(a@0, b@1) IN (SET) ([{c0:aa,c1:ba}, {c0:ab,c1:bb}]) ], dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_max@0 >= aa AND a_null_count@1 != row_count@2 AND a_min@3 <= ab AND b_null_count@5 != row_count@2 AND b_max@4 >= ba AND b_null_count@5 != row_count@2 AND b_min@6 <= bb, required_guarantees=[], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=2, predicate_cache_inner_records=8, predicate_cache_records=4, scan_efficiency_ratio=22.37% (228/1.02 K)] +03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/join_probe.parquet]]}, projection=[a, b, e], file_type=parquet, predicate=Optional(DynamicFilter [ a@0 >= aa AND a@0 <= ab AND b@1 >= ba AND b@1 <= bb AND struct(a@0, b@1) IN (SET) ([{c0:aa,c1:ba}, {c0:ab,c1:bb}]) ]), dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_max@0 >= aa AND a_null_count@1 != row_count@2 AND a_min@3 <= ab AND b_null_count@5 != row_count@2 AND b_max@4 >= ba AND b_null_count@5 != row_count@2 AND b_min@6 <= bb, required_guarantees=[], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=2, predicate_cache_inner_records=8, predicate_cache_records=4, scan_efficiency_ratio=22.37% (228/1.02 K)] statement ok reset datafusion.explain.analyze_categories; @@ -475,8 +475,8 @@ Plan with Metrics 01)HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(c@3, d@0)], metrics=[output_rows=2, output_batches=1, array_map_created_count=0, build_input_batches=1, build_input_rows=2, input_batches=1, input_rows=2, avg_fanout=100% (2/2), probe_hit_rate=100% (2/2)] 02)--HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(a@0, b@0)], metrics=[output_rows=2, output_batches=1, array_map_created_count=0, build_input_batches=1, build_input_rows=2, input_batches=1, input_rows=2, avg_fanout=100% (2/2), probe_hit_rate=100% (2/2)] 03)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/nested_t1.parquet]]}, projection=[a, x], file_type=parquet, metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=0, pushdown_rows_pruned=0, predicate_cache_inner_records=0, predicate_cache_records=0, scan_efficiency_ratio=17.72% (132/745)] -04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/nested_t2.parquet]]}, projection=[b, c, y], file_type=parquet, predicate=DynamicFilter [ b@0 >= aa AND b@0 <= ab AND b@0 IN (SET) ([aa, ab]) ], dynamic_rg_pruning=eligible, pruning_predicate=b_null_count@1 != row_count@2 AND b_max@0 >= aa AND b_null_count@1 != row_count@2 AND b_min@3 <= ab AND (b_null_count@1 != row_count@2 AND b_min@3 <= aa AND aa <= b_max@0 OR b_null_count@1 != row_count@2 AND b_min@3 <= ab AND ab <= b_max@0), required_guarantees=[b in (aa, ab)], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=1 total → 1 matched, page_index_rows_pruned=5 total → 5 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=3, predicate_cache_inner_records=5, predicate_cache_records=2, scan_efficiency_ratio=22.78% (234/1.03 K)] -05)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/nested_t3.parquet]]}, projection=[d, z], file_type=parquet, predicate=DynamicFilter [ d@0 >= ca AND d@0 <= cb AND hash_lookup ], dynamic_rg_pruning=eligible, pruning_predicate=d_null_count@1 != row_count@2 AND d_max@0 >= ca AND d_null_count@1 != row_count@2 AND d_min@3 <= cb, required_guarantees=[], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=1 total → 1 matched, page_index_rows_pruned=8 total → 8 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=6, predicate_cache_inner_records=8, predicate_cache_records=2, scan_efficiency_ratio=21.86% (172/787)] +04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/nested_t2.parquet]]}, projection=[b, c, y], file_type=parquet, predicate=Optional(DynamicFilter [ b@0 >= aa AND b@0 <= ab AND b@0 IN (SET) ([aa, ab]) ]), dynamic_rg_pruning=eligible, pruning_predicate=b_null_count@1 != row_count@2 AND b_max@0 >= aa AND b_null_count@1 != row_count@2 AND b_min@3 <= ab AND (b_null_count@1 != row_count@2 AND b_min@3 <= aa AND aa <= b_max@0 OR b_null_count@1 != row_count@2 AND b_min@3 <= ab AND ab <= b_max@0), required_guarantees=[b in (aa, ab)], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=1 total → 1 matched, page_index_rows_pruned=5 total → 5 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=3, predicate_cache_inner_records=5, predicate_cache_records=2, scan_efficiency_ratio=22.78% (234/1.03 K)] +05)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/nested_t3.parquet]]}, projection=[d, z], file_type=parquet, predicate=Optional(DynamicFilter [ d@0 >= ca AND d@0 <= cb AND hash_lookup ]), dynamic_rg_pruning=eligible, pruning_predicate=d_null_count@1 != row_count@2 AND d_max@0 >= ca AND d_null_count@1 != row_count@2 AND d_min@3 <= cb, required_guarantees=[], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=1 total → 1 matched, page_index_rows_pruned=8 total → 8 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=6, predicate_cache_inner_records=8, predicate_cache_records=2, scan_efficiency_ratio=21.86% (172/787)] statement ok reset datafusion.explain.analyze_categories; @@ -541,7 +541,7 @@ physical_plan 01)HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(a@0, d@0)] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/parent_build.parquet]]}, projection=[a, b, c], file_type=parquet, predicate=a@0 = aa, pruning_predicate=a_null_count@2 != row_count@3 AND a_min@0 <= aa AND aa <= a_max@1, required_guarantees=[a in (aa)] 03)--RepartitionExec: partitioning=RoundRobinBatch(4), input_partitions=1 -04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/parent_probe.parquet]]}, projection=[d, e, f], file_type=parquet, predicate=e@1 = ba AND d@0 = aa AND DynamicFilter [ empty ], dynamic_rg_pruning=eligible, pruning_predicate=e_null_count@2 != row_count@3 AND e_min@0 <= ba AND ba <= e_max@1 AND d_null_count@6 != row_count@3 AND d_min@4 <= aa AND aa <= d_max@5, required_guarantees=[d in (aa), e in (ba)] +04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/parent_probe.parquet]]}, projection=[d, e, f], file_type=parquet, predicate=e@1 = ba AND d@0 = aa AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible, pruning_predicate=e_null_count@2 != row_count@3 AND e_min@0 <= ba AND ba <= e_max@1 AND d_null_count@6 != row_count@3 AND d_min@4 <= aa AND aa <= d_max@5, required_guarantees=[d in (aa), e in (ba)] statement ok drop table parent_build; @@ -606,7 +606,7 @@ Plan with Metrics 01)SortExec: TopK(fetch=2), expr=[e@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[e@0 < bb], metrics=[output_rows=2, output_batches=1, row_replacements=2] 02)--HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(a@0, d@0)], projection=[e@2], metrics=[output_rows=2, output_batches=1, array_map_created_count=0, build_input_batches=1, build_input_rows=2, input_batches=1, input_rows=2, avg_fanout=100% (2/2), probe_hit_rate=100% (2/2)] 03)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/topk_join_build.parquet]]}, projection=[a], file_type=parquet, metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=0, pushdown_rows_pruned=0, predicate_cache_inner_records=0, predicate_cache_records=0, scan_efficiency_ratio=6.49% (64/986)] -04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/topk_join_probe.parquet]]}, projection=[d, e], file_type=parquet, predicate=DynamicFilter [ d@0 >= aa AND d@0 <= ab AND d@0 IN (SET) ([aa, ab]) ] AND DynamicFilter [ e@1 < bb ], dynamic_rg_pruning=eligible, pruning_predicate=d_null_count@1 != row_count@2 AND d_max@0 >= aa AND d_null_count@1 != row_count@2 AND d_min@3 <= ab AND (d_null_count@1 != row_count@2 AND d_min@3 <= aa AND aa <= d_max@0 OR d_null_count@1 != row_count@2 AND d_min@3 <= ab AND ab <= d_max@0) AND e_null_count@5 != row_count@2 AND e_min@4 < bb, required_guarantees=[d in (aa, ab)], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=1 total → 1 matched, page_index_rows_pruned=4 total → 4 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=2, predicate_cache_inner_records=8, predicate_cache_records=4, scan_efficiency_ratio=15.11% (154/1.02 K)] +04)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/topk_join_probe.parquet]]}, projection=[d, e], file_type=parquet, predicate=Optional(DynamicFilter [ d@0 >= aa AND d@0 <= ab AND d@0 IN (SET) ([aa, ab]) ]) AND Optional(DynamicFilter [ e@1 < bb ]), dynamic_rg_pruning=eligible, pruning_predicate=d_null_count@1 != row_count@2 AND d_max@0 >= aa AND d_null_count@1 != row_count@2 AND d_min@3 <= ab AND (d_null_count@1 != row_count@2 AND d_min@3 <= aa AND aa <= d_max@0 OR d_null_count@1 != row_count@2 AND d_min@3 <= ab AND ab <= d_max@0) AND e_null_count@5 != row_count@2 AND e_min@4 < bb, required_guarantees=[d in (aa, ab)], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=1 total → 1 matched, page_index_rows_pruned=4 total → 4 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=2, predicate_cache_inner_records=8, predicate_cache_records=4, scan_efficiency_ratio=15.11% (154/1.02 K)] statement ok reset datafusion.explain.analyze_categories; @@ -656,7 +656,7 @@ EXPLAIN ANALYZE SELECT b, a FROM topk_proj ORDER BY a LIMIT 2; Plan with Metrics 01)ProjectionExec: expr=[b@1 as b, a@0 as a], metrics=[output_rows=2, output_batches=1] 02)--SortExec: TopK(fetch=2), expr=[a@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[a@0 < 2], metrics=[output_rows=2, output_batches=1, row_replacements=2] -03)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/topk_proj.parquet]]}, projection=[a, b], file_type=parquet, predicate=DynamicFilter [ a@0 < 2 ], sort_order_for_reorder=[a@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_min@0 < 2, required_guarantees=[], metrics=[output_rows=3, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=3, pushdown_rows_pruned=0, predicate_cache_inner_records=3, predicate_cache_records=3, scan_efficiency_ratio=13.4% (141/1.05 K)] +03)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/topk_proj.parquet]]}, projection=[a, b], file_type=parquet, predicate=Optional(DynamicFilter [ a@0 < 2 ]), sort_order_for_reorder=[a@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_min@0 < 2, required_guarantees=[], metrics=[output_rows=3, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=3, pushdown_rows_pruned=0, predicate_cache_inner_records=3, predicate_cache_records=3, scan_efficiency_ratio=13.4% (141/1.05 K)] # Case 2: prune — `SELECT a` — filter stays as `a < 2` on the scan. query TT @@ -664,7 +664,7 @@ EXPLAIN ANALYZE SELECT a FROM topk_proj ORDER BY a LIMIT 2; ---- Plan with Metrics 01)SortExec: TopK(fetch=2), expr=[a@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[a@0 < 2], metrics=[output_rows=2, output_batches=1, row_replacements=2] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/topk_proj.parquet]]}, projection=[a], file_type=parquet, predicate=DynamicFilter [ a@0 < 2 ], sort_order_for_reorder=[a@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_min@0 < 2, required_guarantees=[], metrics=[output_rows=3, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=3, pushdown_rows_pruned=0, predicate_cache_inner_records=3, predicate_cache_records=3, scan_efficiency_ratio=6.94% (73/1.05 K)] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/topk_proj.parquet]]}, projection=[a], file_type=parquet, predicate=Optional(DynamicFilter [ a@0 < 2 ]), sort_order_for_reorder=[a@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_min@0 < 2, required_guarantees=[], metrics=[output_rows=3, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=3, pushdown_rows_pruned=0, predicate_cache_inner_records=3, predicate_cache_records=3, scan_efficiency_ratio=6.94% (73/1.05 K)] # Case 3: expression — `SELECT a+1 AS a_plus_1` — the TopK filter is on # `a_plus_1`, the scan predicate must read `a@0 + 1`. @@ -673,7 +673,7 @@ EXPLAIN ANALYZE SELECT a + 1 AS a_plus_1, b FROM topk_proj ORDER BY a_plus_1 LIM ---- Plan with Metrics 01)SortExec: TopK(fetch=2), expr=[a_plus_1@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[a_plus_1@0 < 3], metrics=[output_rows=2, output_batches=1, row_replacements=2] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/topk_proj.parquet]]}, projection=[CAST(a@0 AS Int64) + 1 as a_plus_1, b], file_type=parquet, predicate=DynamicFilter [ CAST(a@0 AS Int64) + 1 < 3 ], dynamic_rg_pruning=eligible, metrics=[output_rows=3, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=3, pushdown_rows_pruned=0, predicate_cache_inner_records=3, predicate_cache_records=3, scan_efficiency_ratio=13.4% (141/1.05 K)] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/topk_proj.parquet]]}, projection=[CAST(a@0 AS Int64) + 1 as a_plus_1, b], file_type=parquet, predicate=Optional(DynamicFilter [ CAST(a@0 AS Int64) + 1 < 3 ]), dynamic_rg_pruning=eligible, metrics=[output_rows=3, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=3, pushdown_rows_pruned=0, predicate_cache_inner_records=3, predicate_cache_records=3, scan_efficiency_ratio=13.4% (141/1.05 K)] # Case 4: alias shadowing — `SELECT a+1 AS a` — the projection renames # `a+1` to `a`, so the TopK's `a < 3` must still be rewritten to @@ -683,7 +683,7 @@ EXPLAIN ANALYZE SELECT a + 1 AS a, b FROM topk_proj ORDER BY a LIMIT 2; ---- Plan with Metrics 01)SortExec: TopK(fetch=2), expr=[a@0 ASC NULLS LAST], preserve_partitioning=[false], filter=[a@0 < 3], metrics=[output_rows=2, output_batches=1, row_replacements=2] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/topk_proj.parquet]]}, projection=[CAST(a@0 AS Int64) + 1 as a, b], file_type=parquet, predicate=DynamicFilter [ CAST(a@0 AS Int64) + 1 < 3 ], sort_order_for_reorder=[a@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, metrics=[output_rows=3, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=3, pushdown_rows_pruned=0, predicate_cache_inner_records=3, predicate_cache_records=3, scan_efficiency_ratio=13.4% (141/1.05 K)] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/topk_proj.parquet]]}, projection=[CAST(a@0 AS Int64) + 1 as a, b], file_type=parquet, predicate=Optional(DynamicFilter [ CAST(a@0 AS Int64) + 1 < 3 ]), sort_order_for_reorder=[a@0 ASC NULLS LAST], dynamic_rg_pruning=eligible, metrics=[output_rows=3, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=3, pushdown_rows_pruned=0, predicate_cache_inner_records=3, predicate_cache_records=3, scan_efficiency_ratio=13.4% (141/1.05 K)] statement ok reset datafusion.explain.analyze_categories; @@ -745,7 +745,7 @@ Plan with Metrics 04)----AggregateExec: mode=FinalPartitioned, gby=[a@0 as a], aggr=[min(join_agg_probe.value)], metrics=[output_rows=2, output_batches=2, spill_count=0, spilled_rows=0] 05)------RepartitionExec: partitioning=Hash([a@0], 4), input_partitions=1, metrics=[output_rows=2, output_batches=2, spill_count=0, spilled_rows=0] 06)--------AggregateExec: mode=Partial, gby=[a@0 as a], aggr=[min(join_agg_probe.value)], metrics=[output_rows=2, output_batches=1, spill_count=0, spilled_rows=0, skipped_aggregation_rows=0, reduction_factor=100% (2/2)] -07)----------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/join_agg_probe.parquet]]}, projection=[a, value], file_type=parquet, predicate=DynamicFilter [ a@0 >= h1 AND a@0 <= h2 AND a@0 IN (SET) ([h1, h2]) ], dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_max@0 >= h1 AND a_null_count@1 != row_count@2 AND a_min@3 <= h2 AND (a_null_count@1 != row_count@2 AND a_min@3 <= h1 AND h1 <= a_max@0 OR a_null_count@1 != row_count@2 AND a_min@3 <= h2 AND h2 <= a_max@0), required_guarantees=[a in (h1, h2)], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=1 total → 1 matched, page_index_rows_pruned=4 total → 4 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=2, predicate_cache_inner_records=4, predicate_cache_records=2, scan_efficiency_ratio=19.43% (151/777)] +07)----------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/join_agg_probe.parquet]]}, projection=[a, value], file_type=parquet, predicate=Optional(DynamicFilter [ a@0 >= h1 AND a@0 <= h2 AND a@0 IN (SET) ([h1, h2]) ]), dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_max@0 >= h1 AND a_null_count@1 != row_count@2 AND a_min@3 <= h2 AND (a_null_count@1 != row_count@2 AND a_min@3 <= h1 AND h1 <= a_max@0 OR a_null_count@1 != row_count@2 AND a_min@3 <= h2 AND h2 <= a_max@0), required_guarantees=[a in (h1, h2)], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=1 total → 1 matched, page_index_rows_pruned=4 total → 4 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=2, predicate_cache_inner_records=4, predicate_cache_records=2, scan_efficiency_ratio=19.43% (151/777)] statement ok reset datafusion.explain.analyze_categories; @@ -808,7 +808,7 @@ ON nulls_build.a = nulls_probe.a AND nulls_build.b = nulls_probe.b; Plan with Metrics 01)HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(a@0, a@0), (b@1, b@1)], metrics=[output_rows=1, output_batches=1, array_map_created_count=0, build_input_batches=1, build_input_rows=3, input_batches=1, input_rows=1, avg_fanout=100% (1/1), probe_hit_rate=100% (1/1)] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/nulls_build.parquet]]}, projection=[a, b], file_type=parquet, metrics=[output_rows=3, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=0, pushdown_rows_pruned=0, predicate_cache_inner_records=0, predicate_cache_records=0, scan_efficiency_ratio=18.6% (144/774)] -03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/nulls_probe.parquet]]}, projection=[a, b, c], file_type=parquet, predicate=DynamicFilter [ a@0 >= aa AND a@0 <= ab AND b@1 >= 1 AND b@1 <= 2 AND struct(a@0, b@1) IN (SET) ([{c0:aa,c1:1}, {c0:,c1:2}, {c0:ab,c1:}]) ], dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_max@0 >= aa AND a_null_count@1 != row_count@2 AND a_min@3 <= ab AND b_null_count@5 != row_count@2 AND b_max@4 >= 1 AND b_null_count@5 != row_count@2 AND b_min@6 <= 2, required_guarantees=[], metrics=[output_rows=1, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=1, pushdown_rows_pruned=3, predicate_cache_inner_records=8, predicate_cache_records=2, scan_efficiency_ratio=20.45% (225/1.10 K)] +03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/nulls_probe.parquet]]}, projection=[a, b, c], file_type=parquet, predicate=Optional(DynamicFilter [ a@0 >= aa AND a@0 <= ab AND b@1 >= 1 AND b@1 <= 2 AND struct(a@0, b@1) IN (SET) ([{c0:aa,c1:1}, {c0:,c1:2}, {c0:ab,c1:}]) ]), dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_max@0 >= aa AND a_null_count@1 != row_count@2 AND a_min@3 <= ab AND b_null_count@5 != row_count@2 AND b_max@4 >= 1 AND b_null_count@5 != row_count@2 AND b_min@6 <= 2, required_guarantees=[], metrics=[output_rows=1, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=1, pushdown_rows_pruned=3, predicate_cache_inner_records=8, predicate_cache_records=2, scan_efficiency_ratio=20.45% (225/1.10 K)] statement ok reset datafusion.explain.analyze_categories; @@ -874,7 +874,7 @@ ON lj_build.a = lj_probe.a AND lj_build.b = lj_probe.b; Plan with Metrics 01)HashJoinExec: mode=CollectLeft, join_type=Left, on=[(a@0, a@0), (b@1, b@1)], metrics=[output_rows=2, output_batches=1, array_map_created_count=0, build_input_batches=1, build_input_rows=2, input_batches=2, input_rows=2, avg_fanout=100% (2/2), probe_hit_rate=100% (2/2)] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/lj_build.parquet]]}, projection=[a, b, c], file_type=parquet, metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=0, pushdown_rows_pruned=0, predicate_cache_inner_records=0, predicate_cache_records=0, scan_efficiency_ratio=19.88% (196/986)] -03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/lj_probe.parquet]]}, projection=[a, b, e], file_type=parquet, predicate=DynamicFilter [ a@0 >= aa AND a@0 <= ab AND b@1 >= ba AND b@1 <= bb AND struct(a@0, b@1) IN (SET) ([{c0:aa,c1:ba}, {c0:ab,c1:bb}]) ], dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_max@0 >= aa AND a_null_count@1 != row_count@2 AND a_min@3 <= ab AND b_null_count@5 != row_count@2 AND b_max@4 >= ba AND b_null_count@5 != row_count@2 AND b_min@6 <= bb, required_guarantees=[], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=2, predicate_cache_inner_records=8, predicate_cache_records=4, scan_efficiency_ratio=22.37% (228/1.02 K)] +03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/lj_probe.parquet]]}, projection=[a, b, e], file_type=parquet, predicate=Optional(DynamicFilter [ a@0 >= aa AND a@0 <= ab AND b@1 >= ba AND b@1 <= bb AND struct(a@0, b@1) IN (SET) ([{c0:aa,c1:ba}, {c0:ab,c1:bb}]) ]), dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_max@0 >= aa AND a_null_count@1 != row_count@2 AND a_min@3 <= ab AND b_null_count@5 != row_count@2 AND b_max@4 >= ba AND b_null_count@5 != row_count@2 AND b_min@6 <= bb, required_guarantees=[], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=2, predicate_cache_inner_records=8, predicate_cache_records=4, scan_efficiency_ratio=22.37% (228/1.02 K)] # LEFT SEMI JOIN: only matching build rows are returned; probe scan still # receives the dynamic filter. @@ -890,7 +890,7 @@ WHERE EXISTS ( Plan with Metrics 01)HashJoinExec: mode=CollectLeft, join_type=LeftSemi, on=[(a@0, a@0), (b@1, b@1)], metrics=[output_rows=2, output_batches=1, array_map_created_count=0, build_input_batches=1, build_input_rows=2, input_batches=2, input_rows=4, avg_fanout=100% (2/2), probe_hit_rate=100% (2/2)] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/lj_build.parquet]]}, projection=[a, b, c], file_type=parquet, metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=0, pushdown_rows_pruned=0, predicate_cache_inner_records=0, predicate_cache_records=0, scan_efficiency_ratio=19.88% (196/986)] -03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/lj_probe.parquet]]}, projection=[a, b], file_type=parquet, predicate=DynamicFilter [ a@0 >= aa AND a@0 <= ab AND b@1 >= ba AND b@1 <= bb AND struct(a@0, b@1) IN (SET) ([{c0:aa,c1:ba}, {c0:ab,c1:bb}]) ], dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_max@0 >= aa AND a_null_count@1 != row_count@2 AND a_min@3 <= ab AND b_null_count@5 != row_count@2 AND b_max@4 >= ba AND b_null_count@5 != row_count@2 AND b_min@6 <= bb, required_guarantees=[], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=2, predicate_cache_inner_records=8, predicate_cache_records=4, scan_efficiency_ratio=15.11% (154/1.02 K)] +03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/lj_probe.parquet]]}, projection=[a, b], file_type=parquet, predicate=Optional(DynamicFilter [ a@0 >= aa AND a@0 <= ab AND b@1 >= ba AND b@1 <= bb AND struct(a@0, b@1) IN (SET) ([{c0:aa,c1:ba}, {c0:ab,c1:bb}]) ]), dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_max@0 >= aa AND a_null_count@1 != row_count@2 AND a_min@3 <= ab AND b_null_count@5 != row_count@2 AND b_max@4 >= ba AND b_null_count@5 != row_count@2 AND b_min@6 <= bb, required_guarantees=[], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=2, predicate_cache_inner_records=8, predicate_cache_records=4, scan_efficiency_ratio=15.11% (154/1.02 K)] statement ok reset datafusion.explain.analyze_categories; @@ -960,7 +960,7 @@ FROM hl_probe p INNER JOIN hl_build AS build Plan with Metrics 01)HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(a@0, a@0), (b@1, b@1)], projection=[a@3, b@4, c@2, e@5], metrics=[output_rows=2, output_batches=1, array_map_created_count=0, build_input_batches=1, build_input_rows=2, input_batches=1, input_rows=2, avg_fanout=100% (2/2), probe_hit_rate=100% (2/2)] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/hl_build.parquet]]}, projection=[a, b, c], file_type=parquet, metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=0, pushdown_rows_pruned=0, predicate_cache_inner_records=0, predicate_cache_records=0, scan_efficiency_ratio=19.88% (196/986)] -03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/hl_probe.parquet]]}, projection=[a, b, e], file_type=parquet, predicate=DynamicFilter [ a@0 >= aa AND a@0 <= ab AND b@1 >= ba AND b@1 <= bb AND hash_lookup ], dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_max@0 >= aa AND a_null_count@1 != row_count@2 AND a_min@3 <= ab AND b_null_count@5 != row_count@2 AND b_max@4 >= ba AND b_null_count@5 != row_count@2 AND b_min@6 <= bb, required_guarantees=[], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=2, predicate_cache_inner_records=8, predicate_cache_records=4, scan_efficiency_ratio=22.37% (228/1.02 K)] +03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/hl_probe.parquet]]}, projection=[a, b, e], file_type=parquet, predicate=Optional(DynamicFilter [ a@0 >= aa AND a@0 <= ab AND b@1 >= ba AND b@1 <= bb AND hash_lookup ]), dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_max@0 >= aa AND a_null_count@1 != row_count@2 AND a_min@3 <= ab AND b_null_count@5 != row_count@2 AND b_max@4 >= ba AND b_null_count@5 != row_count@2 AND b_min@6 <= bb, required_guarantees=[], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=2, predicate_cache_inner_records=8, predicate_cache_records=4, scan_efficiency_ratio=22.37% (228/1.02 K)] statement ok drop table hl_build; @@ -1009,7 +1009,7 @@ FROM int_build b INNER JOIN int_probe p Plan with Metrics 01)HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(id1@0, id1@0), (id2@1, id2@1)], projection=[id1@0, id2@1, value@2, data@5], metrics=[output_rows=2, output_batches=1, array_map_created_count=0, build_input_batches=1, build_input_rows=2, input_batches=1, input_rows=2, avg_fanout=100% (2/2), probe_hit_rate=100% (2/2)] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/int_build.parquet]]}, projection=[id1, id2, value], file_type=parquet, metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=0, pushdown_rows_pruned=0, predicate_cache_inner_records=0, predicate_cache_records=0, scan_efficiency_ratio=18.48% (204/1.10 K)] -03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/int_probe.parquet]]}, projection=[id1, id2, data], file_type=parquet, predicate=DynamicFilter [ id1@0 >= 1 AND id1@0 <= 2 AND id2@1 >= 10 AND id2@1 <= 20 AND hash_lookup ], dynamic_rg_pruning=eligible, pruning_predicate=id1_null_count@1 != row_count@2 AND id1_max@0 >= 1 AND id1_null_count@1 != row_count@2 AND id1_min@3 <= 2 AND id2_null_count@5 != row_count@2 AND id2_max@4 >= 10 AND id2_null_count@5 != row_count@2 AND id2_min@6 <= 20, required_guarantees=[], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=2, predicate_cache_inner_records=8, predicate_cache_records=4, scan_efficiency_ratio=20.67% (221/1.07 K)] +03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/int_probe.parquet]]}, projection=[id1, id2, data], file_type=parquet, predicate=Optional(DynamicFilter [ id1@0 >= 1 AND id1@0 <= 2 AND id2@1 >= 10 AND id2@1 <= 20 AND hash_lookup ]), dynamic_rg_pruning=eligible, pruning_predicate=id1_null_count@1 != row_count@2 AND id1_max@0 >= 1 AND id1_null_count@1 != row_count@2 AND id1_min@3 <= 2 AND id2_null_count@5 != row_count@2 AND id2_max@4 >= 10 AND id2_null_count@5 != row_count@2 AND id2_min@6 <= 20, required_guarantees=[], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=2, predicate_cache_inner_records=8, predicate_cache_records=4, scan_efficiency_ratio=20.67% (221/1.07 K)] statement ok reset datafusion.explain.analyze_categories; @@ -1061,7 +1061,7 @@ EXPLAIN ANALYZE SELECT nej_build.id, nej_probe.id FROM nej_build JOIN nej_probe Plan with Metrics 01)HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(id@0, id@0)], NullsEqual: true, metrics=[output_rows=2, output_batches=1, array_map_created_count=0, build_input_batches=1, build_input_rows=2, input_batches=1, input_rows=2, avg_fanout=100% (2/2), probe_hit_rate=100% (2/2)] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/nej_build.parquet]]}, projection=[id], file_type=parquet, metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=0, pushdown_rows_pruned=0, predicate_cache_inner_records=0, predicate_cache_records=0, scan_efficiency_ratio=12.92% (65/503)] -03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/nej_probe.parquet]]}, projection=[id], file_type=parquet, predicate=DynamicFilter [ id@0 IS NULL OR id@0 >= 11 AND id@0 <= 11 AND id@0 IN (SET) ([11, NULL]) ], dynamic_rg_pruning=eligible, pruning_predicate=id_null_count@0 > 0 OR id_null_count@0 != row_count@2 AND id_max@1 >= 11 AND id_null_count@0 != row_count@2 AND id_min@3 <= 11 AND (id_null_count@0 != row_count@2 AND id_min@3 <= 11 AND 11 <= id_max@1 OR id_null_count@0 != row_count@2 AND id_min@3 <= NULL AND NULL <= id_max@1), required_guarantees=[], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=1 total → 1 matched, page_index_rows_pruned=3 total → 3 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=1, predicate_cache_inner_records=3, predicate_cache_records=3, scan_efficiency_ratio=14.45% (74/512)] +03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/nej_probe.parquet]]}, projection=[id], file_type=parquet, predicate=Optional(DynamicFilter [ id@0 IS NULL OR id@0 >= 11 AND id@0 <= 11 AND id@0 IN (SET) ([11, NULL]) ]), dynamic_rg_pruning=eligible, pruning_predicate=id_null_count@0 > 0 OR id_null_count@0 != row_count@2 AND id_max@1 >= 11 AND id_null_count@0 != row_count@2 AND id_min@3 <= 11 AND (id_null_count@0 != row_count@2 AND id_min@3 <= 11 AND 11 <= id_max@1 OR id_null_count@0 != row_count@2 AND id_min@3 <= NULL AND NULL <= id_max@1), required_guarantees=[], metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=1 total → 1 matched, page_index_rows_pruned=3 total → 3 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=2, pushdown_rows_pruned=1, predicate_cache_inner_records=3, predicate_cache_records=3, scan_efficiency_ratio=14.45% (74/512)] statement ok reset datafusion.explain.analyze_categories; @@ -1104,7 +1104,7 @@ EXPLAIN ANALYZE SELECT mnej_build.a, mnej_build.b, mnej_probe.a, mnej_probe.b FR Plan with Metrics 01)HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(a@0, a@0), (b@1, b@1)], NullsEqual: true, metrics=[output_rows=2, output_batches=1, array_map_created_count=0, build_input_batches=1, build_input_rows=2, input_batches=1, input_rows=3, avg_fanout=100% (2/2), probe_hit_rate=66.67% (2/3)] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/mnej_build.parquet]]}, projection=[a, b], file_type=parquet, metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=0, pushdown_rows_pruned=0, predicate_cache_inner_records=0, predicate_cache_records=0, scan_efficiency_ratio=16.42% (133/810)] -03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/mnej_probe.parquet]]}, projection=[a, b], file_type=parquet, predicate=DynamicFilter [ a@0 IS NULL OR b@1 IS NULL OR a@0 >= 1 AND a@0 <= 2 AND b@1 >= 10 AND b@1 <= 10 AND struct(a@0, b@1) IN (SET) ([{c0:1,c1:10}, {c0:2,c1:}]) ], dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@0 > 0 OR b_null_count@1 > 0 OR a_null_count@0 != row_count@3 AND a_max@2 >= 1 AND a_null_count@0 != row_count@3 AND a_min@4 <= 2 AND b_null_count@1 != row_count@3 AND b_max@5 >= 10 AND b_null_count@1 != row_count@3 AND b_min@6 <= 10, required_guarantees=[], metrics=[output_rows=3, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=3, pushdown_rows_pruned=0, predicate_cache_inner_records=6, predicate_cache_records=6, scan_efficiency_ratio=18.16% (148/815)] +03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/mnej_probe.parquet]]}, projection=[a, b], file_type=parquet, predicate=Optional(DynamicFilter [ a@0 IS NULL OR b@1 IS NULL OR a@0 >= 1 AND a@0 <= 2 AND b@1 >= 10 AND b@1 <= 10 AND struct(a@0, b@1) IN (SET) ([{c0:1,c1:10}, {c0:2,c1:}]) ]), dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@0 > 0 OR b_null_count@1 > 0 OR a_null_count@0 != row_count@3 AND a_max@2 >= 1 AND a_null_count@0 != row_count@3 AND a_min@4 <= 2 AND b_null_count@1 != row_count@3 AND b_max@5 >= 10 AND b_null_count@1 != row_count@3 AND b_min@6 <= 10, required_guarantees=[], metrics=[output_rows=3, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=3, pushdown_rows_pruned=0, predicate_cache_inner_records=6, predicate_cache_records=6, scan_efficiency_ratio=18.16% (148/815)] statement ok reset datafusion.explain.analyze_categories; @@ -1145,7 +1145,7 @@ EXPLAIN ANALYZE SELECT nnb_build.id, nnb_probe.id FROM nnb_build JOIN nnb_probe Plan with Metrics 01)HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(id@0, id@0)], NullsEqual: true, metrics=[output_rows=1, output_batches=1, array_map_created_count=1, build_input_batches=1, build_input_rows=2, input_batches=1, input_rows=1, avg_fanout=100% (1/1), probe_hit_rate=100% (1/1)] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/nnb_build.parquet]]}, projection=[id], file_type=parquet, metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=0, pushdown_rows_pruned=0, predicate_cache_inner_records=0, predicate_cache_records=0, scan_efficiency_ratio=13.71% (68/496)] -03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/nnb_probe.parquet]]}, projection=[id], file_type=parquet, predicate=DynamicFilter [ id@0 >= 11 AND id@0 <= 22 AND id@0 IN (SET) ([11, 22]) ], dynamic_rg_pruning=eligible, pruning_predicate=id_null_count@1 != row_count@2 AND id_max@0 >= 11 AND id_null_count@1 != row_count@2 AND id_min@3 <= 22 AND (id_null_count@1 != row_count@2 AND id_min@3 <= 11 AND 11 <= id_max@0 OR id_null_count@1 != row_count@2 AND id_min@3 <= 22 AND 22 <= id_max@0), required_guarantees=[id in (11, 22)], metrics=[output_rows=1, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=1 total → 1 matched, page_index_rows_pruned=3 total → 3 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=1, pushdown_rows_pruned=2, predicate_cache_inner_records=3, predicate_cache_records=1, scan_efficiency_ratio=14.45% (74/512)] +03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/nnb_probe.parquet]]}, projection=[id], file_type=parquet, predicate=Optional(DynamicFilter [ id@0 >= 11 AND id@0 <= 22 AND id@0 IN (SET) ([11, 22]) ]), dynamic_rg_pruning=eligible, pruning_predicate=id_null_count@1 != row_count@2 AND id_max@0 >= 11 AND id_null_count@1 != row_count@2 AND id_min@3 <= 22 AND (id_null_count@1 != row_count@2 AND id_min@3 <= 11 AND 11 <= id_max@0 OR id_null_count@1 != row_count@2 AND id_min@3 <= 22 AND 22 <= id_max@0), required_guarantees=[id in (11, 22)], metrics=[output_rows=1, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=1 total → 1 matched, page_index_rows_pruned=3 total → 3 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=1, pushdown_rows_pruned=2, predicate_cache_inner_records=3, predicate_cache_records=1, scan_efficiency_ratio=14.45% (74/512)] statement ok reset datafusion.explain.analyze_categories; @@ -1186,7 +1186,7 @@ EXPLAIN ANALYZE SELECT nnp_build.id, nnp_probe.id FROM nnp_build JOIN nnp_probe Plan with Metrics 01)HashJoinExec: mode=CollectLeft, join_type=Inner, on=[(id@0, id@0)], NullsEqual: true, metrics=[output_rows=1, output_batches=1, array_map_created_count=0, build_input_batches=1, build_input_rows=2, input_batches=1, input_rows=1, avg_fanout=100% (1/1), probe_hit_rate=100% (1/1)] 02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/nnp_build.parquet]]}, projection=[id], file_type=parquet, metrics=[output_rows=2, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=0 total → 0 matched, page_index_rows_pruned=0 total → 0 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=0, pushdown_rows_pruned=0, predicate_cache_inner_records=0, predicate_cache_records=0, scan_efficiency_ratio=12.92% (65/503)] -03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/nnp_probe.parquet]]}, projection=[id], file_type=parquet, predicate=DynamicFilter [ id@0 >= 11 AND id@0 <= 11 AND id@0 IN (SET) ([11, NULL]) ], dynamic_rg_pruning=eligible, pruning_predicate=id_null_count@1 != row_count@2 AND id_max@0 >= 11 AND id_null_count@1 != row_count@2 AND id_min@3 <= 11 AND (id_null_count@1 != row_count@2 AND id_min@3 <= 11 AND 11 <= id_max@0 OR id_null_count@1 != row_count@2 AND id_min@3 <= NULL AND NULL <= id_max@0), required_guarantees=[id in (11, NULL)], metrics=[output_rows=1, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=1 total → 1 matched, page_index_rows_pruned=2 total → 2 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=1, pushdown_rows_pruned=1, predicate_cache_inner_records=2, predicate_cache_records=1, scan_efficiency_ratio=13.71% (68/496)] +03)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/nnp_probe.parquet]]}, projection=[id], file_type=parquet, predicate=Optional(DynamicFilter [ id@0 >= 11 AND id@0 <= 11 AND id@0 IN (SET) ([11, NULL]) ]), dynamic_rg_pruning=eligible, pruning_predicate=id_null_count@1 != row_count@2 AND id_max@0 >= 11 AND id_null_count@1 != row_count@2 AND id_min@3 <= 11 AND (id_null_count@1 != row_count@2 AND id_min@3 <= 11 AND 11 <= id_max@0 OR id_null_count@1 != row_count@2 AND id_min@3 <= NULL AND NULL <= id_max@0), required_guarantees=[id in (11, NULL)], metrics=[output_rows=1, output_batches=1, files_ranges_pruned_statistics=1 total → 1 matched, row_groups_pruned_statistics=1 total → 1 matched, row_groups_pruned_bloom_filter=1 total → 1 matched, page_index_pages_pruned=1 total → 1 matched, page_index_rows_pruned=2 total → 2 matched, limit_pruned_row_groups=0 total → 0 matched, batches_split=0, file_open_errors=0, file_scan_errors=0, files_opened=1, files_processed=1, num_predicate_creation_errors=0, predicate_evaluation_errors=0, pushdown_rows_matched=1, pushdown_rows_pruned=1, predicate_cache_inner_records=2, predicate_cache_records=1, scan_efficiency_ratio=13.71% (68/496)] statement ok reset datafusion.explain.analyze_categories; @@ -1237,7 +1237,7 @@ physical_plan 02)--RepartitionExec: partitioning=Hash([id@0], 4), input_partitions=2 03)----DataSourceExec: file_groups={2 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/pnej_build/1.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/pnej_build/2.parquet]]}, projection=[id], file_type=parquet 04)--RepartitionExec: partitioning=Hash([id@0], 4), input_partitions=2 -05)----DataSourceExec: file_groups={2 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/pnej_probe/1.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/pnej_probe/2.parquet]]}, projection=[id], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +05)----DataSourceExec: file_groups={2 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/pnej_probe/1.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_parquet/pnej_probe/2.parquet]]}, projection=[id], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]) AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible query II rowsort SELECT pnej_build.id, pnej_probe.id FROM pnej_build JOIN pnej_probe ON pnej_build.id IS NOT DISTINCT FROM pnej_probe.id diff --git a/datafusion/sqllogictest/test_files/push_down_filter_regression.slt b/datafusion/sqllogictest/test_files/push_down_filter_regression.slt index c260efae95334..ec3b7e4bde938 100644 --- a/datafusion/sqllogictest/test_files/push_down_filter_regression.slt +++ b/datafusion/sqllogictest/test_files/push_down_filter_regression.slt @@ -146,7 +146,7 @@ physical_plan 01)AggregateExec: mode=Final, gby=[], aggr=[max(agg_dyn_test.id)] 02)--CoalescePartitionsExec 03)----AggregateExec: mode=Partial, gby=[], aggr=[max(agg_dyn_test.id)] -04)------DataSourceExec: file_groups={2 groups: [[WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-01/j5fUeSDQo22oPyPU.parquet, WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-02/j5fUeSDQo22oPyPU.parquet], [WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-03/j5fUeSDQo22oPyPU.parquet, WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-04/j5fUeSDQo22oPyPU.parquet]]}, projection=[id], file_type=parquet, predicate=id@0 > 1 AND DynamicFilter [ empty ], dynamic_rg_pruning=eligible, pruning_predicate=id_null_count@1 != row_count@2 AND id_max@0 > 1, required_guarantees=[] +04)------DataSourceExec: file_groups={2 groups: [[WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-01/j5fUeSDQo22oPyPU.parquet, WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-02/j5fUeSDQo22oPyPU.parquet], [WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-03/j5fUeSDQo22oPyPU.parquet, WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-04/j5fUeSDQo22oPyPU.parquet]]}, projection=[id], file_type=parquet, predicate=id@0 > 1 AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible, pruning_predicate=id_null_count@1 != row_count@2 AND id_max@0 > 1, required_guarantees=[] query I select max(id) from agg_dyn_test where id > 1; @@ -161,7 +161,7 @@ physical_plan 01)AggregateExec: mode=Final, gby=[], aggr=[max(agg_dyn_test.id)] 02)--CoalescePartitionsExec 03)----AggregateExec: mode=Partial, gby=[], aggr=[max(agg_dyn_test.id)] -04)------DataSourceExec: file_groups={2 groups: [[WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-01/j5fUeSDQo22oPyPU.parquet, WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-02/j5fUeSDQo22oPyPU.parquet], [WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-03/j5fUeSDQo22oPyPU.parquet, WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-04/j5fUeSDQo22oPyPU.parquet]]}, projection=[id], file_type=parquet, predicate=CAST(id@0 AS Int64) + 1 > 1 AND DynamicFilter [ empty ], dynamic_rg_pruning=eligible +04)------DataSourceExec: file_groups={2 groups: [[WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-01/j5fUeSDQo22oPyPU.parquet, WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-02/j5fUeSDQo22oPyPU.parquet], [WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-03/j5fUeSDQo22oPyPU.parquet, WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-04/j5fUeSDQo22oPyPU.parquet]]}, projection=[id], file_type=parquet, predicate=CAST(id@0 AS Int64) + 1 > 1 AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Expect dynamic filter available inside data source query TT @@ -171,7 +171,7 @@ physical_plan 01)AggregateExec: mode=Final, gby=[], aggr=[max(agg_dyn_test.id), min(agg_dyn_test.id)] 02)--CoalescePartitionsExec 03)----AggregateExec: mode=Partial, gby=[], aggr=[max(agg_dyn_test.id), min(agg_dyn_test.id)] -04)------DataSourceExec: file_groups={2 groups: [[WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-01/j5fUeSDQo22oPyPU.parquet, WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-02/j5fUeSDQo22oPyPU.parquet], [WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-03/j5fUeSDQo22oPyPU.parquet, WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-04/j5fUeSDQo22oPyPU.parquet]]}, projection=[id], file_type=parquet, predicate=id@0 < 10 AND DynamicFilter [ empty ], dynamic_rg_pruning=eligible, pruning_predicate=id_null_count@1 != row_count@2 AND id_min@0 < 10, required_guarantees=[] +04)------DataSourceExec: file_groups={2 groups: [[WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-01/j5fUeSDQo22oPyPU.parquet, WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-02/j5fUeSDQo22oPyPU.parquet], [WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-03/j5fUeSDQo22oPyPU.parquet, WORKSPACE_ROOT/datafusion/core/tests/data/test_statistics_per_partition/date=2025-03-04/j5fUeSDQo22oPyPU.parquet]]}, projection=[id], file_type=parquet, predicate=id@0 < 10 AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible, pruning_predicate=id_null_count@1 != row_count@2 AND id_min@0 < 10, required_guarantees=[] # Dynamic filter should not be available for grouping sets query TT @@ -236,7 +236,7 @@ Plan with Metrics 01)AggregateExec: mode=Final, gby=[], aggr=[max(agg_dyn_e2e.column1)], metrics=[] 02)--CoalescePartitionsExec, metrics=[] 03)----AggregateExec: mode=Partial, gby=[], aggr=[max(agg_dyn_e2e.column1)], metrics=[] -04)------DataSourceExec: file_groups= projection=[column1], file_type=parquet, predicate=column1@0 > 1 AND DynamicFilter [ column1@0 > ], dynamic_rg_pruning=eligible, pruning_predicate=column1_null_count@1 != row_count@2 AND column1_max@0 > 1 AND column1_null_count@1 != row_count@2 AND column1_max@0 > , required_guarantees=[], metrics=[] +04)------DataSourceExec: file_groups= projection=[column1], file_type=parquet, predicate=column1@0 > 1 AND Optional(DynamicFilter [ column1@0 > ]), dynamic_rg_pruning=eligible, pruning_predicate=column1_null_count@1 != row_count@2 AND column1_max@0 > 1 AND column1_null_count@1 != row_count@2 AND column1_max@0 > , required_guarantees=[], metrics=[] statement ok reset datafusion.explain.analyze_categories; @@ -323,7 +323,7 @@ Plan with Metrics 01)AggregateExec: mode=Final, gby=[], aggr=[min(agg_dyn_single.a)], metrics=[] 02)--CoalescePartitionsExec, metrics=[] 03)----AggregateExec: mode=Partial, gby=[], aggr=[min(agg_dyn_single.a)], metrics=[] -04)------DataSourceExec: file_groups={2 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_single/file_0.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_single/file_1.parquet]]}, projection=[a], file_type=parquet, predicate=DynamicFilter [ a@0 < 1 ], dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_min@0 < 1, required_guarantees=[], metrics=[] +04)------DataSourceExec: file_groups={2 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_single/file_0.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_single/file_1.parquet]]}, projection=[a], file_type=parquet, predicate=Optional(DynamicFilter [ a@0 < 1 ]), dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_min@0 < 1, required_guarantees=[], metrics=[] # MAX(a) -> DynamicFilter [ a > 8 ] query TT @@ -333,7 +333,7 @@ Plan with Metrics 01)AggregateExec: mode=Final, gby=[], aggr=[max(agg_dyn_single.a)], metrics=[] 02)--CoalescePartitionsExec, metrics=[] 03)----AggregateExec: mode=Partial, gby=[], aggr=[max(agg_dyn_single.a)], metrics=[] -04)------DataSourceExec: file_groups={2 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_single/file_0.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_single/file_1.parquet]]}, projection=[a], file_type=parquet, predicate=DynamicFilter [ a@0 > 8 ], dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_max@0 > 8, required_guarantees=[], metrics=[] +04)------DataSourceExec: file_groups={2 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_single/file_0.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_single/file_1.parquet]]}, projection=[a], file_type=parquet, predicate=Optional(DynamicFilter [ a@0 > 8 ]), dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_max@0 > 8, required_guarantees=[], metrics=[] # MIN(a), MAX(a) -> DynamicFilter [ a < 1 OR a > 8 ] query TT @@ -343,7 +343,7 @@ Plan with Metrics 01)AggregateExec: mode=Final, gby=[], aggr=[min(agg_dyn_single.a), max(agg_dyn_single.a)], metrics=[] 02)--CoalescePartitionsExec, metrics=[] 03)----AggregateExec: mode=Partial, gby=[], aggr=[min(agg_dyn_single.a), max(agg_dyn_single.a)], metrics=[] -04)------DataSourceExec: file_groups={2 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_single/file_0.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_single/file_1.parquet]]}, projection=[a], file_type=parquet, predicate=DynamicFilter [ a@0 < 1 OR a@0 > 8 ], dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_min@0 < 1 OR a_null_count@1 != row_count@2 AND a_max@3 > 8, required_guarantees=[], metrics=[] +04)------DataSourceExec: file_groups={2 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_single/file_0.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_single/file_1.parquet]]}, projection=[a], file_type=parquet, predicate=Optional(DynamicFilter [ a@0 < 1 OR a@0 > 8 ]), dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_min@0 < 1 OR a_null_count@1 != row_count@2 AND a_max@3 > 8, required_guarantees=[], metrics=[] # MIN(a+1) -> no dynamic filter (expression input is not a plain column) query TT @@ -387,7 +387,7 @@ Plan with Metrics 01)AggregateExec: mode=Final, gby=[], aggr=[min(agg_dyn_two_col.a), max(agg_dyn_two_col.b)], metrics=[] 02)--CoalescePartitionsExec, metrics=[] 03)----AggregateExec: mode=Partial, gby=[], aggr=[min(agg_dyn_two_col.a), max(agg_dyn_two_col.b)], metrics=[] -04)------DataSourceExec: file_groups={2 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_two_col/file_0.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_two_col/file_1.parquet]]}, projection=[a, b], file_type=parquet, predicate=DynamicFilter [ a@0 < 1 OR b@1 > 9 ], dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_min@0 < 1 OR b_null_count@4 != row_count@2 AND b_max@3 > 9, required_guarantees=[], metrics=[] +04)------DataSourceExec: file_groups={2 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_two_col/file_0.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_two_col/file_1.parquet]]}, projection=[a, b], file_type=parquet, predicate=Optional(DynamicFilter [ a@0 < 1 OR b@1 > 9 ]), dynamic_rg_pruning=eligible, pruning_predicate=a_null_count@1 != row_count@2 AND a_min@0 < 1 OR b_null_count@4 != row_count@2 AND b_max@3 > 9, required_guarantees=[], metrics=[] statement ok drop table agg_dyn_two_col; @@ -470,7 +470,7 @@ Plan with Metrics 01)AggregateExec: mode=Final, gby=[], aggr=[min(agg_dyn_nulls.a)], metrics=[] 02)--CoalescePartitionsExec, metrics=[] 03)----AggregateExec: mode=Partial, gby=[], aggr=[min(agg_dyn_nulls.a)], metrics=[] -04)------DataSourceExec: file_groups={2 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_nulls/file_0.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_nulls/file_1.parquet]]}, projection=[a], file_type=parquet, predicate=DynamicFilter [ true ], dynamic_rg_pruning=eligible, metrics=[] +04)------DataSourceExec: file_groups={2 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_nulls/file_0.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/push_down_filter_regression/agg_dyn_nulls/file_1.parquet]]}, projection=[a], file_type=parquet, predicate=Optional(DynamicFilter [ true ]), dynamic_rg_pruning=eligible, metrics=[] statement ok reset datafusion.explain.analyze_categories; diff --git a/datafusion/sqllogictest/test_files/range_partitioning.slt b/datafusion/sqllogictest/test_files/range_partitioning.slt index 463cd1799e857..3776aedad6f26 100644 --- a/datafusion/sqllogictest/test_files/range_partitioning.slt +++ b/datafusion/sqllogictest/test_files/range_partitioning.slt @@ -1903,10 +1903,12 @@ SELECT range_key, SUM(value) FROM ( ########## # TEST 48: Hash Join Dynamic Filter Pushdown on Compatible Range Inputs -# The Parquet-backed probe accepts the partition-routed dynamic filter. Matching +# The Parquet-backed probe accepts the partitioned dynamic filters. Matching # Range split points keep build filter i aligned with probe partition i. The -# build has rows only in partitions 0 and 2, so the sparse runtime CASE includes -# only those branches and lets empty partitions fall through to false. +# build has rows only in partitions 0 and 2. The bounds filter keeps the +# disjoint ranges of the two partitions. Both partitions push an IN list, so +# the routed membership CASE collapses into one IN list over the union of +# their keys. ########## statement ok @@ -1933,7 +1935,7 @@ JOIN range_partitioned p ON b.range_key = p.range_key; physical_plan 01)HashJoinExec: mode=Partitioned, join_type=Inner, on=[(range_key@0, range_key@0)], projection=[range_key@0, value@1, value@3] 02)--DataSourceExec: file_groups=, projection=[range_key, value], output_partitioning=Range([range_key@0 ASC], [(10), (20), (30)], 4), file_type=parquet, predicate=range_key@0 = 5 OR range_key@0 = 20, pruning_predicate=range_key_null_count@2 != row_count@3 AND range_key_min@0 <= 5 AND 5 <= range_key_max@1 OR range_key_null_count@2 != row_count@3 AND range_key_min@0 <= 20 AND 20 <= range_key_max@1, required_guarantees=[range_key in (20, 5)] -03)--DataSourceExec: file_groups=, projection=[range_key, value], output_partitioning=Range([range_key@0 ASC], [(10), (20), (30)], 4), file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +03)--DataSourceExec: file_groups=, projection=[range_key, value], output_partitioning=Range([range_key@0 ASC], [(10), (20), (30)], 4), file_type=parquet, predicate=Optional(DynamicFilter [ empty ]) AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible query TT EXPLAIN ANALYZE SELECT b.range_key, b.value, p.value @@ -1947,7 +1949,7 @@ JOIN range_partitioned p ON b.range_key = p.range_key; Plan with Metrics 01)HashJoinExec: mode=Partitionedmetrics=[output_rows=2,] 02)--DataSourceExec: file_type=parquet, predicate=range_key@0 = 5 OR range_key@0 = 20metrics=[output_rows=2,] -03)--DataSourceExec: file_type=parquet, predicate=DynamicFilter [ CASE range_partition WHEN 0 THEN range_key@0 >= 5 AND range_key@0 <= 5 AND range_key@0 IN (SET) ([5]) WHEN 2 THEN range_key@0 >= 20 AND range_key@0 <= 20 AND range_key@0 IN (SET) ([20]) ELSE false END ]metrics=[output_rows=2,] +03)--DataSourceExec: file_type=parquet, predicate=Optional(DynamicFilter [ range_key@0 >= 5 AND range_key@0 <= 5 OR range_key@0 >= 20 AND range_key@0 <= 20 ]) AND Optional(DynamicFilter [ range_key@0 IN (SET) ([5, 20]) ])metrics=[output_rows=2,] query III SELECT b.range_key, b.value, p.value diff --git a/datafusion/sqllogictest/test_files/repartition_subset_satisfaction.slt b/datafusion/sqllogictest/test_files/repartition_subset_satisfaction.slt index 5371ca59beea1..842f5fd1de9e5 100644 --- a/datafusion/sqllogictest/test_files/repartition_subset_satisfaction.slt +++ b/datafusion/sqllogictest/test_files/repartition_subset_satisfaction.slt @@ -380,7 +380,7 @@ physical_plan 12)----------------------CoalescePartitionsExec 13)------------------------FilterExec: service@1 = log, projection=[env@0, d_dkey@2] 14)--------------------------DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/dimension/d_dkey=A/data.parquet, WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/dimension/d_dkey=D/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/dimension/d_dkey=B/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/dimension/d_dkey=C/data.parquet]]}, projection=[env, service, d_dkey], output_partitioning=Hash([d_dkey@2], 3), file_type=parquet, predicate=service@1 = log, pruning_predicate=service_null_count@2 != row_count@3 AND service_min@0 <= log AND log <= service_max@1, required_guarantees=[service in (log)] -15)----------------------DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/fact/f_dkey=A/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/fact/f_dkey=B/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/fact/f_dkey=C/data.parquet]]}, projection=[timestamp, value, f_dkey], output_ordering=[f_dkey@2 ASC NULLS LAST, timestamp@0 ASC NULLS LAST], output_partitioning=Hash([f_dkey@2], 3), file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +15)----------------------DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/fact/f_dkey=A/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/fact/f_dkey=B/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/fact/f_dkey=C/data.parquet]]}, projection=[timestamp, value, f_dkey], output_ordering=[f_dkey@2 ASC NULLS LAST, timestamp@0 ASC NULLS LAST], output_partitioning=Hash([f_dkey@2], 3), file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Verify results without subset satisfaction query TPR rowsort @@ -475,7 +475,7 @@ physical_plan 10)------------------CoalescePartitionsExec 11)--------------------FilterExec: service@1 = log, projection=[env@0, d_dkey@2] 12)----------------------DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/dimension/d_dkey=A/data.parquet, WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/dimension/d_dkey=D/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/dimension/d_dkey=B/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/dimension/d_dkey=C/data.parquet]]}, projection=[env, service, d_dkey], output_partitioning=Hash([d_dkey@2], 3), file_type=parquet, predicate=service@1 = log, pruning_predicate=service_null_count@2 != row_count@3 AND service_min@0 <= log AND log <= service_max@1, required_guarantees=[service in (log)] -13)------------------DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/fact/f_dkey=A/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/fact/f_dkey=B/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/fact/f_dkey=C/data.parquet]]}, projection=[timestamp, value, f_dkey], output_ordering=[f_dkey@2 ASC NULLS LAST, timestamp@0 ASC NULLS LAST], output_partitioning=Hash([f_dkey@2], 3), file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +13)------------------DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/fact/f_dkey=A/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/fact/f_dkey=B/data.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/repartition_subset_satisfaction/fact/f_dkey=C/data.parquet]]}, projection=[timestamp, value, f_dkey], output_ordering=[f_dkey@2 ASC NULLS LAST, timestamp@0 ASC NULLS LAST], output_partitioning=Hash([f_dkey@2], 3), file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Verify results match with subset satisfaction query TPR rowsort diff --git a/datafusion/sqllogictest/test_files/sort_pushdown.slt b/datafusion/sqllogictest/test_files/sort_pushdown.slt index a173e76d6c262..cebe2f5550817 100644 --- a/datafusion/sqllogictest/test_files/sort_pushdown.slt +++ b/datafusion/sqllogictest/test_files/sort_pushdown.slt @@ -60,7 +60,7 @@ logical_plan 02)--TableScan: sorted_parquet projection=[id, value, name] physical_plan 01)SortExec: TopK(fetch=3), expr=[id@0 DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/sorted_data.parquet]]}, projection=[id, value, name], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/sorted_data.parquet]]}, projection=[id, value, name], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible # Test 1.2: Verify results are correct query IIT @@ -91,7 +91,7 @@ logical_plan 02)--TableScan: sorted_parquet projection=[id, value, name] physical_plan 01)SortExec: TopK(fetch=3), expr=[id@0 DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/sorted_data.parquet]]}, projection=[id, value, name], output_ordering=[id@0 ASC NULLS LAST], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/sorted_data.parquet]]}, projection=[id, value, name], output_ordering=[id@0 ASC NULLS LAST], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Re-enable statement ok @@ -108,7 +108,7 @@ logical_plan physical_plan 01)GlobalLimitExec: skip=2, fetch=3 02)--SortExec: TopK(fetch=5), expr=[id@0 DESC], preserve_partitioning=[false] -03)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/sorted_data.parquet]]}, projection=[id, value, name], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible +03)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/sorted_data.parquet]]}, projection=[id, value, name], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible query IIT SELECT * FROM sorted_parquet ORDER BY id DESC LIMIT 3 OFFSET 2; @@ -172,7 +172,7 @@ logical_plan 03)----TableScan: multi_rg_sorted projection=[id, category, value], partial_filters=[multi_rg_sorted.category = Utf8View("alpha") OR multi_rg_sorted.category = Utf8View("gamma")] physical_plan 01)SortExec: TopK(fetch=5), expr=[id@0 DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/multi_rg_sorted.parquet]]}, projection=[id, category, value], file_type=parquet, predicate=(category@1 = alpha OR category@1 = gamma) AND DynamicFilter [ empty ], sort_order_for_reorder=[id@0 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible, pruning_predicate=category_null_count@2 != row_count@3 AND category_min@0 <= alpha AND alpha <= category_max@1 OR category_null_count@2 != row_count@3 AND category_min@0 <= gamma AND gamma <= category_max@1, required_guarantees=[category in (alpha, gamma)] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/multi_rg_sorted.parquet]]}, projection=[id, category, value], file_type=parquet, predicate=(category@1 = alpha OR category@1 = gamma) AND Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible, pruning_predicate=category_null_count@2 != row_count@3 AND category_min@0 <= alpha AND alpha <= category_max@1 OR category_null_count@2 != row_count@3 AND category_min@0 <= gamma AND gamma <= category_max@1, required_guarantees=[category in (alpha, gamma)] # Verify the results are correct despite reverse scanning with row selection # Expected: gamma values (6, 5) then alpha values (2, 1), in DESC order by id @@ -289,7 +289,7 @@ logical_plan physical_plan 01)SortPreservingMergeExec: [id@0 DESC], fetch=3 02)--SortExec: TopK(fetch=3), expr=[id@0 DESC], preserve_partitioning=[true] -03)----DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/sorted_multi/part1.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/sorted_multi/part2.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/sorted_multi/part3.parquet]]}, projection=[id, value, name], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible +03)----DataSourceExec: file_groups={3 groups: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/sorted_multi/part1.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/sorted_multi/part2.parquet], [WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/sorted_multi/part3.parquet]]}, projection=[id, value, name], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible # Verify correctness with repartitioning and multiple files query IIT @@ -398,7 +398,7 @@ logical_plan 03)----TableScan: timeseries_parquet projection=[timeframe, period_end, value], partial_filters=[timeseries_parquet.timeframe = Utf8View("quarterly")] physical_plan 01)SortExec: TopK(fetch=2), expr=[period_end@1 DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/timeseries_sorted.parquet]]}, projection=[timeframe, period_end, value], file_type=parquet, predicate=timeframe@0 = quarterly AND DynamicFilter [ empty ], sort_order_for_reorder=[period_end@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible, pruning_predicate=timeframe_null_count@2 != row_count@3 AND timeframe_min@0 <= quarterly AND quarterly <= timeframe_max@1, required_guarantees=[timeframe in (quarterly)] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/timeseries_sorted.parquet]]}, projection=[timeframe, period_end, value], file_type=parquet, predicate=timeframe@0 = quarterly AND Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[period_end@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible, pruning_predicate=timeframe_null_count@2 != row_count@3 AND timeframe_min@0 <= quarterly AND quarterly <= timeframe_max@1, required_guarantees=[timeframe in (quarterly)] # Test 2.2: Verify the results are correct query TIR @@ -457,7 +457,7 @@ logical_plan 02)--TableScan: timeseries_parquet projection=[timeframe, period_end, value] physical_plan 01)SortExec: TopK(fetch=3), expr=[timeframe@0 ASC NULLS LAST, period_end@1 DESC], preserve_partitioning=[false], sort_prefix=[timeframe@0 ASC NULLS LAST] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/timeseries_sorted.parquet]]}, projection=[timeframe, period_end, value], output_ordering=[timeframe@0 ASC NULLS LAST, period_end@1 ASC NULLS LAST], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/timeseries_sorted.parquet]]}, projection=[timeframe, period_end, value], output_ordering=[timeframe@0 ASC NULLS LAST, period_end@1 ASC NULLS LAST], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Test 2.7: Disable sort pushdown and verify filter still works statement ok @@ -475,7 +475,7 @@ logical_plan 03)----TableScan: timeseries_parquet projection=[timeframe, period_end, value], partial_filters=[timeseries_parquet.timeframe = Utf8View("quarterly")] physical_plan 01)SortExec: TopK(fetch=2), expr=[period_end@1 DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/timeseries_sorted.parquet]]}, projection=[timeframe, period_end, value], output_ordering=[timeframe@0 ASC NULLS LAST, period_end@1 ASC NULLS LAST], file_type=parquet, predicate=timeframe@0 = quarterly AND DynamicFilter [ empty ], dynamic_rg_pruning=eligible, pruning_predicate=timeframe_null_count@2 != row_count@3 AND timeframe_min@0 <= quarterly AND quarterly <= timeframe_max@1, required_guarantees=[timeframe in (quarterly)] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/timeseries_sorted.parquet]]}, projection=[timeframe, period_end, value], output_ordering=[timeframe@0 ASC NULLS LAST, period_end@1 ASC NULLS LAST], file_type=parquet, predicate=timeframe@0 = quarterly AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible, pruning_predicate=timeframe_null_count@2 != row_count@3 AND timeframe_min@0 <= quarterly AND quarterly <= timeframe_max@1, required_guarantees=[timeframe in (quarterly)] # Results should still be correct query TIR @@ -508,7 +508,7 @@ logical_plan 03)----TableScan: timeseries_parquet projection=[timeframe, period_end, value], partial_filters=[timeseries_parquet.timeframe = Utf8View("daily") OR timeseries_parquet.timeframe = Utf8View("weekly")] physical_plan 01)SortExec: TopK(fetch=3), expr=[period_end@1 DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/timeseries_sorted.parquet]]}, projection=[timeframe, period_end, value], file_type=parquet, predicate=(timeframe@0 = daily OR timeframe@0 = weekly) AND DynamicFilter [ empty ], sort_order_for_reorder=[period_end@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible, pruning_predicate=timeframe_null_count@2 != row_count@3 AND timeframe_min@0 <= daily AND daily <= timeframe_max@1 OR timeframe_null_count@2 != row_count@3 AND timeframe_min@0 <= weekly AND weekly <= timeframe_max@1, required_guarantees=[timeframe in (daily, weekly)] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/timeseries_sorted.parquet]]}, projection=[timeframe, period_end, value], file_type=parquet, predicate=(timeframe@0 = daily OR timeframe@0 = weekly) AND Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[period_end@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible, pruning_predicate=timeframe_null_count@2 != row_count@3 AND timeframe_min@0 <= daily AND daily <= timeframe_max@1 OR timeframe_null_count@2 != row_count@3 AND timeframe_min@0 <= weekly AND weekly <= timeframe_max@1, required_guarantees=[timeframe in (daily, weekly)] # Test 2.9: Complex case - literal constant in sort expression itself # The literal 'constant' is ignored in sort analysis @@ -528,7 +528,7 @@ logical_plan 03)----TableScan: timeseries_parquet projection=[timeframe, period_end, value], partial_filters=[timeseries_parquet.timeframe = Utf8View("monthly")] physical_plan 01)SortExec: TopK(fetch=2), expr=[period_end@1 DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/timeseries_sorted.parquet]]}, projection=[timeframe, period_end, value], file_type=parquet, predicate=timeframe@0 = monthly AND DynamicFilter [ empty ], sort_order_for_reorder=[period_end@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible, pruning_predicate=timeframe_null_count@2 != row_count@3 AND timeframe_min@0 <= monthly AND monthly <= timeframe_max@1, required_guarantees=[timeframe in (monthly)] +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/timeseries_sorted.parquet]]}, projection=[timeframe, period_end, value], file_type=parquet, predicate=timeframe@0 = monthly AND Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[period_end@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible, pruning_predicate=timeframe_null_count@2 != row_count@3 AND timeframe_min@0 <= monthly AND monthly <= timeframe_max@1, required_guarantees=[timeframe in (monthly)] # Verify results query TIR @@ -617,7 +617,7 @@ logical_plan 02)--TableScan: timestamp_parquet projection=[id, ts, volume, price] physical_plan 01)SortExec: TopK(fetch=3), expr=[ts@1 DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/timestamp_sorted.parquet]]}, projection=[id, ts, volume, price], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[ts@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/timestamp_sorted.parquet]]}, projection=[id, ts, volume, price], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[ts@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible # Verify results query IPIR @@ -643,7 +643,7 @@ logical_plan 02)--TableScan: timestamp_parquet projection=[id, ts, volume, price] physical_plan 01)SortExec: TopK(fetch=3), expr=[date_trunc(day, ts@1) DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/timestamp_sorted.parquet]]}, projection=[id, ts, volume, price], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[date_trunc(day, ts@1) DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/timestamp_sorted.parquet]]}, projection=[id, ts, volume, price], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[date_trunc(day, ts@1) DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible # Verify results (descending day) query IPIR @@ -703,7 +703,7 @@ logical_plan 02)--TableScan: multi_month_parquet projection=[id, ts, volume, price] physical_plan 01)SortExec: TopK(fetch=2), expr=[ts@1 DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/multi_month_sorted.parquet]]}, projection=[id, ts, volume, price], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[ts@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/multi_month_sorted.parquet]]}, projection=[id, ts, volume, price], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[ts@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible query IPIR SELECT * FROM multi_month_parquet @@ -729,7 +729,7 @@ logical_plan 02)--TableScan: multi_month_parquet projection=[id, ts, volume, price] physical_plan 01)SortExec: TopK(fetch=2), expr=[date_trunc(month, ts@1) DESC, ts@1 DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/multi_month_sorted.parquet]]}, projection=[id, ts, volume, price], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[date_trunc(month, ts@1) DESC, ts@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/multi_month_sorted.parquet]]}, projection=[id, ts, volume, price], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[date_trunc(month, ts@1) DESC, ts@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible query IPIR SELECT * FROM multi_month_parquet @@ -771,7 +771,7 @@ logical_plan 02)--TableScan: int_parquet projection=[id, small_val, big_val] physical_plan 01)SortExec: TopK(fetch=2), expr=[CAST(small_val@1 AS Int64) DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/int_sorted.parquet]]}, projection=[id, small_val, big_val], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[CAST(small_val@1 AS Int64) DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/int_sorted.parquet]]}, projection=[id, small_val, big_val], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[CAST(small_val@1 AS Int64) DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible query III SELECT * FROM int_parquet @@ -813,7 +813,7 @@ logical_plan 02)--TableScan: float_parquet projection=[id, value] physical_plan 01)SortExec: TopK(fetch=3), expr=[ceil(value@1) DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/float_sorted.parquet]]}, projection=[id, value], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[ceil(value@1) DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/float_sorted.parquet]]}, projection=[id, value], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[ceil(value@1) DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible query IR SELECT * FROM float_parquet @@ -856,7 +856,7 @@ logical_plan 02)--TableScan: signed_parquet projection=[id, value] physical_plan 01)SortExec: TopK(fetch=3), expr=[abs(value@1) DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/signed_sorted.parquet]]}, projection=[id, value], output_ordering=[value@1 ASC NULLS LAST], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/signed_sorted.parquet]]}, projection=[id, value], output_ordering=[value@1 ASC NULLS LAST], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Results should still be correct (no optimization applied) query IR @@ -2005,7 +2005,7 @@ logical_plan 02)--TableScan: tb_overlap projection=[id, value] physical_plan 01)SortExec: TopK(fetch=5), expr=[id@0 DESC, value@1 DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/tb_overlap/file_z.parquet, WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/tb_overlap/file_y.parquet, WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/tb_overlap/file_x.parquet]]}, projection=[id, value], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 DESC, value@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/tb_overlap/file_z.parquet, WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/tb_overlap/file_y.parquet, WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/tb_overlap/file_x.parquet]]}, projection=[id, value], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 DESC, value@1 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible query II SELECT * FROM tb_overlap ORDER BY id DESC, value DESC LIMIT 5; @@ -2090,7 +2090,7 @@ logical_plan 02)--TableScan: tc_limit projection=[id, value] physical_plan 01)SortExec: TopK(fetch=3), expr=[id@0 DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/tc_limit/file_c.parquet, WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/tc_limit/file_b.parquet, WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/tc_limit/file_a.parquet]]}, projection=[id, value], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/tc_limit/file_c.parquet, WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/tc_limit/file_b.parquet, WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/tc_limit/file_a.parquet]]}, projection=[id, value], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible query II SELECT * FROM tc_limit ORDER BY id DESC LIMIT 3; @@ -2561,7 +2561,7 @@ logical_plan 02)--TableScan: th_reorder projection=[id, value] physical_plan 01)SortExec: TopK(fetch=3), expr=[id@0 ASC NULLS LAST], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/th_reorder/data.parquet]]}, projection=[id, value], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/th_reorder/data.parquet]]}, projection=[id, value], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 ASC NULLS LAST], dynamic_rg_pruning=eligible # Results must be correct regardless of RG reorder. query II @@ -2580,7 +2580,7 @@ logical_plan 02)--TableScan: th_reorder projection=[id, value] physical_plan 01)SortExec: TopK(fetch=3), expr=[id@0 DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/th_reorder/data.parquet]]}, projection=[id, value], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/th_reorder/data.parquet]]}, projection=[id, value], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible query II SELECT * FROM th_reorder ORDER BY id DESC LIMIT 3; @@ -2666,7 +2666,7 @@ logical_plan 02)--TableScan: tj_scrambled projection=[id, value] physical_plan 01)SortExec: TopK(fetch=3), expr=[id@0 DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/tj_scrambled/data.parquet]]}, projection=[id, value], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/tj_scrambled/data.parquet]]}, projection=[id, value], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible # Test J.2: Results must be correct query II @@ -2856,7 +2856,7 @@ logical_plan 02)--TableScan: tl_sorted projection=[id, value] physical_plan 01)SortExec: TopK(fetch=3), expr=[id@0 DESC, value@1 ASC NULLS LAST], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/tl_multikey/data.parquet]]}, projection=[id, value], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[id@0 DESC, value@1 ASC NULLS LAST], reverse_row_groups=true, dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/sort_pushdown/tl_multikey/data.parquet]]}, projection=[id, value], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[id@0 DESC, value@1 ASC NULLS LAST], reverse_row_groups=true, dynamic_rg_pruning=eligible query II SELECT id, value FROM tl_sorted ORDER BY id DESC, value ASC LIMIT 3; diff --git a/datafusion/sqllogictest/test_files/statistics_registry.slt b/datafusion/sqllogictest/test_files/statistics_registry.slt index 22209e5eb112f..7d197b1bcea84 100644 --- a/datafusion/sqllogictest/test_files/statistics_registry.slt +++ b/datafusion/sqllogictest/test_files/statistics_registry.slt @@ -104,7 +104,7 @@ physical_plan 04)--RepartitionExec: partitioning=Hash([small_id@2], 4), input_partitions=1, maintains_sort_order=true 05)----HashJoinExec: mode=Partitioned, join_type=Inner, on=[(customer_id@0, customer_id@1)], projection=[region_id@1, order_id@2, small_id@4] 06)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/statistics_registry/customers.parquet]]}, projection=[customer_id, region_id], output_ordering=[region_id@1 ASC NULLS LAST], file_type=parquet -07)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/statistics_registry/orders.parquet]]}, projection=[order_id, customer_id, small_id], output_ordering=[order_id@0 ASC NULLS LAST], file_type=parquet, predicate=DynamicFilter [ empty ] AND DynamicFilter [ empty ], dynamic_rg_pruning=eligible +07)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/statistics_registry/orders.parquet]]}, projection=[order_id, customer_id, small_id], output_ordering=[order_id@0 ASC NULLS LAST], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]) AND Optional(DynamicFilter [ empty ]) AND Optional(DynamicFilter [ empty ]) AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # -- Registry reflected in EXPLAIN statistics -------------------------------- # With show_statistics on, the displayed statistics column reflects the @@ -127,7 +127,7 @@ physical_plan 04)--RepartitionExec: partitioning=Hash([small_id@2], 4), input_partitions=1, maintains_sort_order=true, statistics=[Rows=Inexact(100), Bytes=Absent, [(Col[0]: Min=Exact(Int32(1)) Max=Exact(Int32(10)) Null=Exact(0) ScanBytes=Exact(40)),(Col[1]: Min=Inexact(Int32(1)) Max=Inexact(Int32(10)) Null=Inexact(0) ScanBytes=Inexact(40)),(Col[2]: Min=Inexact(Int32(1)) Max=Inexact(Int32(10)) Null=Inexact(0) ScanBytes=Inexact(40))]] 05)----HashJoinExec: mode=Partitioned, join_type=Inner, on=[(customer_id@0, customer_id@1)], projection=[region_id@1, order_id@2, small_id@4], statistics=[Rows=Inexact(100), Bytes=Absent, [(Col[0]: Min=Exact(Int32(1)) Max=Exact(Int32(10)) Null=Exact(0) ScanBytes=Exact(40)),(Col[1]: Min=Inexact(Int32(1)) Max=Inexact(Int32(10)) Null=Inexact(0) ScanBytes=Inexact(40)),(Col[2]: Min=Inexact(Int32(1)) Max=Inexact(Int32(10)) Null=Inexact(0) ScanBytes=Inexact(40))]] 06)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/statistics_registry/customers.parquet]]}, projection=[customer_id, region_id], output_ordering=[region_id@1 ASC NULLS LAST], file_type=parquet, statistics=[Rows=Exact(10), Bytes=Exact(80), [(Col[0]: Min=Exact(Int32(1)) Max=Exact(Int32(3)) Null=Exact(0) ScanBytes=Exact(40)),(Col[1]: Min=Exact(Int32(1)) Max=Exact(Int32(10)) Null=Exact(0) ScanBytes=Exact(40))]] -07)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/statistics_registry/orders.parquet]]}, projection=[order_id, customer_id, small_id], output_ordering=[order_id@0 ASC NULLS LAST], file_type=parquet, predicate=DynamicFilter [ empty ] AND DynamicFilter [ empty ], dynamic_rg_pruning=eligible, statistics=[Rows=Inexact(10), Bytes=Inexact(120), [(Col[0]: Min=Inexact(Int32(1)) Max=Inexact(Int32(10)) Null=Inexact(0) ScanBytes=Inexact(40)),(Col[1]: Min=Inexact(Int32(1)) Max=Inexact(Int32(3)) Null=Inexact(0) ScanBytes=Inexact(40)),(Col[2]: Min=Inexact(Int32(1)) Max=Inexact(Int32(10)) Null=Inexact(0) ScanBytes=Inexact(40))]] +07)------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/statistics_registry/orders.parquet]]}, projection=[order_id, customer_id, small_id], output_ordering=[order_id@0 ASC NULLS LAST], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]) AND Optional(DynamicFilter [ empty ]) AND Optional(DynamicFilter [ empty ]) AND Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible, statistics=[Rows=Inexact(10), Bytes=Inexact(120), [(Col[0]: Min=Inexact(Int32(1)) Max=Inexact(Int32(10)) Null=Inexact(0) ScanBytes=Inexact(40)),(Col[1]: Min=Inexact(Int32(1)) Max=Inexact(Int32(3)) Null=Inexact(0) ScanBytes=Inexact(40)),(Col[2]: Min=Inexact(Int32(1)) Max=Inexact(Int32(10)) Null=Inexact(0) ScanBytes=Inexact(40))]] # -- Registry reflected in EXPLAIN ANALYZE statistics ------------------------ # EXPLAIN ANALYZE renders the same registry-refined statistics as plain EXPLAIN diff --git a/datafusion/sqllogictest/test_files/topk.slt b/datafusion/sqllogictest/test_files/topk.slt index 41a454db29699..125eb1fbbef5e 100644 --- a/datafusion/sqllogictest/test_files/topk.slt +++ b/datafusion/sqllogictest/test_files/topk.slt @@ -316,7 +316,7 @@ explain select number, letter, age from partial_sorted order by number desc, let ---- physical_plan 01)SortExec: TopK(fetch=3), expr=[number@0 DESC, letter@1 ASC NULLS LAST, age@2 DESC], preserve_partitioning=[false], sort_prefix=[number@0 DESC, letter@1 ASC NULLS LAST] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/topk/partial_sorted/1.parquet]]}, projection=[number, letter, age], output_ordering=[number@0 DESC, letter@1 ASC NULLS LAST], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/topk/partial_sorted/1.parquet]]}, projection=[number, letter, age], output_ordering=[number@0 DESC, letter@1 ASC NULLS LAST], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Explain variations of the above query with different orderings, and different sort prefixes. @@ -326,28 +326,28 @@ explain select number, letter, age from partial_sorted order by age desc limit 3 ---- physical_plan 01)SortExec: TopK(fetch=3), expr=[age@2 DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/topk/partial_sorted/1.parquet]]}, projection=[number, letter, age], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[age@2 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/topk/partial_sorted/1.parquet]]}, projection=[number, letter, age], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[age@2 DESC], reverse_row_groups=true, dynamic_rg_pruning=eligible query TT explain select number, letter, age from partial_sorted order by number desc, letter desc limit 3; ---- physical_plan 01)SortExec: TopK(fetch=3), expr=[number@0 DESC, letter@1 DESC], preserve_partitioning=[false], sort_prefix=[number@0 DESC] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/topk/partial_sorted/1.parquet]]}, projection=[number, letter, age], output_ordering=[number@0 DESC, letter@1 ASC NULLS LAST], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/topk/partial_sorted/1.parquet]]}, projection=[number, letter, age], output_ordering=[number@0 DESC, letter@1 ASC NULLS LAST], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible query TT explain select number, letter, age from partial_sorted order by number asc limit 3; ---- physical_plan 01)SortExec: TopK(fetch=3), expr=[number@0 ASC NULLS LAST], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/topk/partial_sorted/1.parquet]]}, projection=[number, letter, age], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[number@0 ASC NULLS LAST], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/topk/partial_sorted/1.parquet]]}, projection=[number, letter, age], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[number@0 ASC NULLS LAST], dynamic_rg_pruning=eligible query TT explain select number, letter, age from partial_sorted order by letter asc, number desc limit 3; ---- physical_plan 01)SortExec: TopK(fetch=3), expr=[letter@1 ASC NULLS LAST, number@0 DESC], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/topk/partial_sorted/1.parquet]]}, projection=[number, letter, age], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[letter@1 ASC NULLS LAST, number@0 DESC], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/topk/partial_sorted/1.parquet]]}, projection=[number, letter, age], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[letter@1 ASC NULLS LAST, number@0 DESC], dynamic_rg_pruning=eligible # Explicit NULLS ordering cases (reversing the order of the NULLS on the number and letter orderings) query TT @@ -355,14 +355,14 @@ explain select number, letter, age from partial_sorted order by number desc, let ---- physical_plan 01)SortExec: TopK(fetch=3), expr=[number@0 DESC, letter@1 ASC], preserve_partitioning=[false], sort_prefix=[number@0 DESC] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/topk/partial_sorted/1.parquet]]}, projection=[number, letter, age], output_ordering=[number@0 DESC, letter@1 ASC NULLS LAST], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/topk/partial_sorted/1.parquet]]}, projection=[number, letter, age], output_ordering=[number@0 DESC, letter@1 ASC NULLS LAST], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible query TT explain select number, letter, age from partial_sorted order by number desc NULLS LAST, letter asc limit 3; ---- physical_plan 01)SortExec: TopK(fetch=3), expr=[number@0 DESC NULLS LAST, letter@1 ASC NULLS LAST], preserve_partitioning=[false] -02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/topk/partial_sorted/1.parquet]]}, projection=[number, letter, age], file_type=parquet, predicate=DynamicFilter [ empty ], sort_order_for_reorder=[number@0 DESC NULLS LAST, letter@1 ASC NULLS LAST], reverse_row_groups=true, dynamic_rg_pruning=eligible +02)--DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/topk/partial_sorted/1.parquet]]}, projection=[number, letter, age], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), sort_order_for_reorder=[number@0 DESC NULLS LAST, letter@1 ASC NULLS LAST], reverse_row_groups=true, dynamic_rg_pruning=eligible # Verify that the sort prefix is correctly computed on the normalized ordering (removing redundant aliased columns) @@ -372,7 +372,7 @@ explain select number, letter, age, number as column4, letter as column5 from pa physical_plan 01)ProjectionExec: expr=[number@0 as number, letter@1 as letter, age@2 as age, number@0 as column4, letter@1 as column5] 02)--SortExec: TopK(fetch=3), expr=[number@0 DESC, letter@1 ASC NULLS LAST, age@2 DESC], preserve_partitioning=[false], sort_prefix=[number@0 DESC, letter@1 ASC NULLS LAST] -03)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/topk/partial_sorted/1.parquet]]}, projection=[number, letter, age], output_ordering=[number@0 DESC, letter@1 ASC NULLS LAST], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +03)----DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/topk/partial_sorted/1.parquet]]}, projection=[number, letter, age], output_ordering=[number@0 DESC, letter@1 ASC NULLS LAST], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # `number + 1` is not order-maintaining (addition can overflow and wrap), so # no sort prefix can be computed over the projected expression. @@ -385,7 +385,7 @@ physical_plan 03)----SortExec: TopK(fetch=3), expr=[__common_expr_1@0 DESC, number@1 DESC, age@2 ASC NULLS LAST], preserve_partitioning=[true] 04)------ProjectionExec: expr=[CAST(number@0 AS Int64) + 1 as __common_expr_1, number@0 as number, age@1 as age] 05)--------RepartitionExec: partitioning=RoundRobinBatch(4), input_partitions=1, maintains_sort_order=true -06)----------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/topk/partial_sorted/1.parquet]]}, projection=[number, age], output_ordering=[number@0 DESC], file_type=parquet, predicate=DynamicFilter [ empty ], dynamic_rg_pruning=eligible +06)----------DataSourceExec: file_groups={1 group: [[WORKSPACE_ROOT/datafusion/sqllogictest/test_files/scratch/topk/partial_sorted/1.parquet]]}, projection=[number, age], output_ordering=[number@0 DESC], file_type=parquet, predicate=Optional(DynamicFilter [ empty ]), dynamic_rg_pruning=eligible # Cleanup statement ok diff --git a/docs/source/user-guide/configs.md b/docs/source/user-guide/configs.md index 1af6680bf64f9..bfff746c4d445 100644 --- a/docs/source/user-guide/configs.md +++ b/docs/source/user-guide/configs.md @@ -145,6 +145,7 @@ The following configuration settings are available: | datafusion.execution.objectstore_writer_buffer_size | 10485760 | Size (bytes) of data buffer DataFusion uses when writing output files. This affects the size of the data chunks that are uploaded to remote object stores (e.g. AWS S3). If very large (>= 100 GiB) output files are being written, it may be necessary to increase this size to avoid errors from the remote end point. | | datafusion.execution.enable_ansi_mode | false | Whether to enable ANSI SQL mode. The flag is experimental and relevant only for DataFusion Spark built-in functions When `enable_ansi_mode` is set to `true`, the query engine follows ANSI SQL semantics for expressions, casting, and error handling. This means: - **Strict type coercion rules:** implicit casts between incompatible types are disallowed. - **Standard SQL arithmetic behavior:** operations such as division by zero, numeric overflow, or invalid casts raise runtime errors rather than returning `NULL` or adjusted values. - **Consistent ANSI behavior** for string concatenation, comparisons, and `NULL` handling. When `enable_ansi_mode` is `false` (the default), the engine uses a more permissive, non-ANSI mode designed for user convenience and backward compatibility. In this mode: - Implicit casts between types are allowed (e.g., string to integer when possible). - Arithmetic operations are more lenient — for example, `abs()` on the minimum representable integer value returns the input value instead of raising overflow. - Division by zero or invalid casts may return `NULL` instead of failing. # Default `false` — ANSI SQL mode is disabled by default. | | datafusion.execution.hash_join_buffering_capacity | 0 | How many bytes to buffer in the probe side of hash joins while the build side is concurrently being built. Without this, hash joins will wait until the full materialization of the build side before polling the probe side. This is useful in scenarios where the query is not completely CPU bounded, allowing to do some early work concurrently and reducing the latency of the query. Note that when hash join buffering is enabled, the probe side will start eagerly polling data, not giving time for the producer side of dynamic filters to produce any meaningful predicate. Queries with dynamic filters might see performance degradation. Disabled by default, set to a number greater than 0 for enabling it. | +| datafusion.execution.optional_filter_min_saving_ns_per_row | 20 | The assumed work, in nanoseconds, that each row removed by an optional filter saves downstream. Optional filters are filters that are not needed for correctness, such as the dynamic filters that hash joins and TopK push down into scans. When an operator evaluates optional filters adaptively, it pauses an optional filter whose evaluation costs more than the work that it saves. Consumers that can measure the saving (the Parquet scan) add their measured decode cost. The default is about the cost of a hash table probe for one row. The best value depends on the hardware. | | datafusion.optimizer.enable_distinct_aggregation_soft_limit | true | When set to true, the optimizer will push a limit operation into grouped aggregations which have no aggregate expressions, as a soft limit, emitting groups once the limit is reached, before all rows in the group are read. | | datafusion.optimizer.enable_round_robin_repartition | true | When set to true, the physical plan optimizer will try to add round robin repartitioning to increase parallelism to leverage more CPU cores | | datafusion.optimizer.enable_topk_aggregation | true | When set to true, the optimizer will attempt to perform limit operations during aggregations, if possible |