Skip to content

feat: keep optional filters pruning-only in the Parquet post-scan path - #25722

Draft
adriangb wants to merge 11 commits into
apache:mainfrom
pydantic:parquet-post-scan-optional
Draft

adriangb wants to merge 11 commits into
apache:mainfrom
pydantic:parquet-post-scan-optional

Conversation

@adriangb

@adriangb adriangb commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

feat: keep optional filters pruning-only in the Parquet post-scan path

Which issue does this PR close?

graph LR
  FS["filter_stats commit"]
  A1["#25673 Optional wrapper"]
  A2["#25681 producers mark filters"]
  A3["#25674 gate"]
  A4["#25682 Parquet consumer"]
  B1["#25683 FilterExec consumer"]
  B2["#25721 FilterExec reordering"]
  C2["#25713 split join filter"]
  P["#22384 post-scan filter"]
  F["#25722 post-scan skips optional filters"]
  FS --> A3
  FS --> B2
  A1 --> A2
  C2 --> A2
  A1 --> A4
  A3 --> A4
  A1 --> B1
  A3 --> B1
  A1 --> F
  P --> F
  classDef this fill:#f6e7d6,stroke:#b25e12,stroke-width:3px
  class F this
Loading

Rationale for this change

With pushdown_filters = false, #22384 makes the scan evaluate all accepted filters for each row after the decode. This includes the hash join dynamic filter. The join checks the same rows again, so this work gives nothing, and it caused a 1.16x TPC-H regression on #22384.

Optional filters (#25673) are not needed for correctness. The post-scan filter must not evaluate them.

What changes are included in this PR?

The scan splits the predicate with split_optional (#25673):

Case Required conjunct Optional conjunct
pushdown_filters = false post-scan filter statistics, page index, bloom filter and file pruning only (as before #22384)
pushdown_filters = true, row filter accepts it row filter row filter
pushdown_filters = true, row filter rejects it for a file post-scan filter not used for that file
row filter build error for the whole file post-scan filter not used for that file
scan equivalences (FileSource::exact_filter, #25780 on main) exact (not for a pruning-only predicate, see #22384) never exact: the scan derives no equivalence from it
  // pushdown_filters = false
- post_scan = split_conjunction(predicate)
+ post_scan = split_optional(predicate).0   // required conjuncts only

The change is about 10 lines in opener/mod.rs, push_decoder.rs and row_filter.rs. It does not need the gate or optional_filter_mode (#25674, #25682).

What is the testing strategy for this PR?

optional_conjunct_is_never_evaluated_post_scan (opener unit test) reads a file with each predicate and counts the rows that the post-scan filter sees (post_scan_rows_pruned + post_scan_rows_matched):

Predicate pushdown_filters Output rows Post-scan rows
s IS NOT NULL (row filter rejects it) false 2 3
s IS NOT NULL true 2 3
id > 1 false 2 3
Optional(s IS NOT NULL) false 3 0
Optional(s IS NOT NULL) true 3 0
Optional(id > 1) false 3 0
Optional(id > 1) (row filter accepts it) true 2 0

datasource-parquet lib tests, parquet_integration, core_integration and the full sqllogictest suite pass.

Are there any user-facing changes?

No. Without optional filters (before #25681) nothing changes. With them, pushdown_filters = false scans do the same row-level work as before #22384 for dynamic filters.

🤖 Generated with Claude Code

adriangb and others added 2 commits September 24, 2026 18:17
…d conjuncts post-scan

Rebased onto main. The original commit 1 of this PR ("extract
DecoderProjection from build_stream") landed independently on main as
current `decoder_projection` / single-decoder + `rg_plan` model.

Two changes, both applied inside the parquet scan so the parent
`FilterExec` can be removed unconditionally for pushable filters:

1. Never drop conjuncts the `RowFilter` cannot place.
   `build_row_filter` previously `.flatten()`-ed away conjuncts that
   `FilterCandidateBuilder::build` rejected (whole-struct references,
   per-file physical-schema mismatches) and swallowed whole-build
   errors. By the time it runs, `try_pushdown_filters` has already
   removed the `FilterExec`, so those conjuncts were applied nowhere —
   wrong results. `build_row_filter` now returns
   `(Option<RowFilter>, Vec<rejected>)`, `RowFilterGenerator` exposes
   `rejected_conjuncts()`, and a whole-file build error routes every
   conjunct to the rejected list rather than relaxing the predicate.

2. Always accept pushable filters and run the remainder post-scan.
   `try_pushdown_filters` reports each pushable filter as accepted so the
   `FilterExec` is always removed; the scan owns the predicate. The
   opener routes conjuncts to two places, applying every one:
     - pushdown_filters=true  -> row-filterable conjuncts via the parquet
       `RowFilter`; rejected conjuncts via the in-scan post-scan filter.
     - pushdown_filters=false -> the whole predicate runs as a post-scan
       filter on decoded batches (behaviorally identical to `FilterExec`).

Implementation:
  - `DecoderProjection` (main's `decoder_projection` module) grows a
    `post_scan_conjuncts` parameter: it widens the decoder mask over
    (user projection ∪ post-scan filter columns), rebases the conjuncts
    onto the stream schema, and returns a `PostScanFilter` applied to
    every decoded batch with SQL `WHERE` semantics. Virtual-column
    conjuncts are stripped from the read-plan mask (they aren't file
    columns) but kept in the post-scan predicate, which sees the
    reader-appended virtual columns.
  - `PushDecoderStreamState` applies the post-scan filter in the decoded-
    batch arm, skips empty batches, and re-introduces a stream-level
    `remaining_limit` (main enforces LIMIT decoder-locally, which is
    unsafe once a post-scan filter can reject rows). The opener routes the
    limit to `remaining_limit` iff a post-scan filter is present.
  - New `post_scan_rows_pruned` / `post_scan_rows_matched` counters and
    `post_scan_filter_eval_time` on `ParquetFileMetrics`.

Tests:
  - `build_row_filter_surfaces_rejected_struct_conjunct` (row_filter.rs)
    asserts the rejected struct conjunct is returned, not dropped.
  - `rejected_struct_conjunct_runs_post_scan_not_dropped` (opener) is
    end-to-end: `s IS NOT NULL` over a struct column with pushdown on
    returns 2 (was 3 before the fix).
  - Parquet `.slt` files regenerated: `FilterExec` above `DataSourceExec`
    gone, predicate on the scan, `post_scan_rows_*` metrics on EXPLAIN
    ANALYZE. Opener / core insta / page_pruning assertions updated for the
    now-applied predicate.

Squashed follow-up commits (see their original messages on the
adriangb/parquet-post-scan-filter branch):
  - perf(parquet-datasource): narrow to the projector's columns before
    filtering
  - perf(parquet-datasource): coalesce post-scan filter output back to
    batch_size
  - perf(parquet-datasource): compact between conjuncts in the post-scan
    filter

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Tests and sqllogictest plans that landed on main after apache#22384 was
opened still expect the old "scan only uses the predicate for
pruning" behaviour. With the scan now applying every accepted filter:

- Two opener tests (`test_prune_all_null_column_equality_from_file_statistics`,
  `test_no_prune_when_missing_column_collapses_mixed_predicate`) now
  expect only the matching rows. The missing-column test also checks
  `post_scan_rows_pruned` so it still proves the file was read, not pruned.
- `string_in_list_pruning.rs` measured unpruned rows with the scan's
  `output_rows`. It now uses the post-scan matched + pruned counters,
  which count the rows that the scan decoded.
- Regenerated plans in `dynamic_filter_pushdown_config.slt`,
  `filter_without_sort_exec.slt`, `push_down_filter_parquet.slt`,
  `range_partitioning.slt` and `range_sorted_time_bin_agg.slt`: the
  `FilterExec` above parquet scans is gone and the new
  `post_scan_rows_*` metrics appear.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added physical-expr Changes to the physical-expr crates core Core DataFusion crate sqllogictest SQL Logic Tests (.slt) proto Related to proto crate datasource Changes to the datasource crate labels Sep 24, 2026
@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Thank you for opening this pull request!

Reviewer note: cargo-semver-checks reported the current version number is not SemVer-compatible with the changes in this pull request (compared against the base branch).

Details
     Cloning apache/main
    Building datafusion v55.1.0 (current)
       Built [  62.558s] (current)
     Parsing datafusion v55.1.0 (current)
      Parsed [   0.036s] (current)
    Building datafusion v55.1.0 (baseline)
       Built [  60.770s] (baseline)
     Parsing datafusion v55.1.0 (baseline)
      Parsed [   0.036s] (baseline)
    Checking datafusion v55.1.0 -> v55.1.0 (no change; assume patch)
     Checked [   0.604s] 223 checks: 223 pass, 31 skip
     Summary no semver update required
    Finished [ 125.629s] datafusion
    Building datafusion-datasource v55.1.0 (current)
       Built [  44.692s] (current)
     Parsing datafusion-datasource v55.1.0 (current)
      Parsed [   0.033s] (current)
    Building datafusion-datasource v55.1.0 (baseline)
       Built [  44.565s] (baseline)
     Parsing datafusion-datasource v55.1.0 (baseline)
      Parsed [   0.033s] (baseline)
    Checking datafusion-datasource v55.1.0 -> v55.1.0 (no change; assume patch)
     Checked [   0.271s] 223 checks: 223 pass, 31 skip
     Summary no semver update required
    Finished [  90.664s] datafusion-datasource
    Building datafusion-datasource-parquet v55.1.0 (current)
       Built [  50.836s] (current)
     Parsing datafusion-datasource-parquet v55.1.0 (current)
      Parsed [   0.035s] (current)
    Building datafusion-datasource-parquet v55.1.0 (baseline)
       Built [  51.484s] (baseline)
     Parsing datafusion-datasource-parquet v55.1.0 (baseline)
      Parsed [   0.036s] (baseline)
    Checking datafusion-datasource-parquet v55.1.0 -> v55.1.0 (no change; assume patch)
     Checked [   0.155s] 223 checks: 222 pass, 1 fail, 0 warn, 31 skip

--- failure constructible_struct_adds_field: struct exhaustively constructible through public API adds field ---

Description:
A pub struct that could be exhaustively constructed with a literal using only public API has a new pub field, breaking existing exhaustive literals.
        ref: https://doc.rust-lang.org/reference/expressions/struct-expr.html
       impl: https://github.com/obi1kenobi/cargo-semver-checks/tree/v0.50.0/src/lints/constructible_struct_adds_field.ron

Failed in:
  field ParquetFileMetrics.post_scan_rows_pruned in /home/runner/work/datafusion/datafusion/datafusion/datasource-parquet/src/metrics.rs:100
  field ParquetFileMetrics.post_scan_rows_matched in /home/runner/work/datafusion/datafusion/datafusion/datasource-parquet/src/metrics.rs:102
  field ParquetFileMetrics.post_scan_filter_eval_time in /home/runner/work/datafusion/datafusion/datafusion/datasource-parquet/src/metrics.rs:104

     Summary semver requires new major version: 1 major and 0 minor checks failed
    Finished [ 103.644s] datafusion-datasource-parquet
    Building datafusion-physical-expr v55.1.0 (current)
       Built [  30.234s] (current)
     Parsing datafusion-physical-expr v55.1.0 (current)
      Parsed [   0.050s] (current)
    Building datafusion-physical-expr v55.1.0 (baseline)
       Built [  30.289s] (baseline)
     Parsing datafusion-physical-expr v55.1.0 (baseline)
      Parsed [   0.051s] (baseline)
    Checking datafusion-physical-expr v55.1.0 -> v55.1.0 (no change; assume patch)
     Checked [   0.345s] 223 checks: 223 pass, 31 skip
     Summary no semver update required
    Finished [  62.303s] datafusion-physical-expr
    Building datafusion-physical-optimizer v55.1.0 (current)
       Built [  42.676s] (current)
     Parsing datafusion-physical-optimizer v55.1.0 (current)
      Parsed [   0.021s] (current)
    Building datafusion-physical-optimizer v55.1.0 (baseline)
       Built [  41.893s] (baseline)
     Parsing datafusion-physical-optimizer v55.1.0 (baseline)
      Parsed [   0.023s] (baseline)
    Checking datafusion-physical-optimizer v55.1.0 -> v55.1.0 (no change; assume patch)
     Checked [   0.117s] 223 checks: 223 pass, 31 skip
     Summary no semver update required
    Finished [  85.741s] datafusion-physical-optimizer
    Building datafusion-physical-plan v55.1.0 (current)
       Built [  39.549s] (current)
     Parsing datafusion-physical-plan v55.1.0 (current)
      Parsed [   0.172s] (current)
    Building datafusion-physical-plan v55.1.0 (baseline)
       Built [  39.845s] (baseline)
     Parsing datafusion-physical-plan v55.1.0 (baseline)
      Parsed [   0.178s] (baseline)
    Checking datafusion-physical-plan v55.1.0 -> v55.1.0 (no change; assume patch)
     Checked [   0.658s] 223 checks: 223 pass, 31 skip
     Summary no semver update required
    Finished [  81.519s] datafusion-physical-plan
    Building datafusion-proto v55.1.0 (current)
       Built [  56.923s] (current)
     Parsing datafusion-proto v55.1.0 (current)
      Parsed [   0.018s] (current)
    Building datafusion-proto v55.1.0 (baseline)
       Built [  57.022s] (baseline)
     Parsing datafusion-proto v55.1.0 (baseline)
      Parsed [   0.019s] (baseline)
    Checking datafusion-proto v55.1.0 -> v55.1.0 (no change; assume patch)
     Checked [   0.130s] 223 checks: 223 pass, 31 skip
     Summary no semver update required
    Finished [ 115.301s] datafusion-proto
    Building datafusion-proto-models v55.1.0 (current)
       Built [  26.019s] (current)
     Parsing datafusion-proto-models v55.1.0 (current)
      Parsed [   0.134s] (current)
    Building datafusion-proto-models v55.1.0 (baseline)
       Built [  25.883s] (baseline)
     Parsing datafusion-proto-models v55.1.0 (baseline)
      Parsed [   0.138s] (baseline)
    Checking datafusion-proto-models v55.1.0 -> v55.1.0 (no change; assume patch)
     Checked [   1.815s] 223 checks: 221 pass, 2 fail, 0 warn, 31 skip

--- failure constructible_struct_adds_field: struct exhaustively constructible through public API adds field ---

Description:
A pub struct that could be exhaustively constructed with a literal using only public API has a new pub field, breaking existing exhaustive literals.
        ref: https://doc.rust-lang.org/reference/expressions/struct-expr.html
       impl: https://github.com/obi1kenobi/cargo-semver-checks/tree/v0.50.0/src/lints/constructible_struct_adds_field.ron

Failed in:
  field ParquetScanExecNode.pruning_only_predicate in /home/runner/work/datafusion/datafusion/datafusion/proto-models/src/generated/prost.rs:2049
  field ParquetScanExecNode.pruning_only_predicate in /home/runner/work/datafusion/datafusion/datafusion/proto-models/src/generated/prost.rs:2049

--- failure enum_variant_added: enum variant added on exhaustive enum ---

Description:
A publicly-visible enum without #[non_exhaustive] has a new variant.
        ref: https://doc.rust-lang.org/cargo/reference/semver.html#enum-variant-new
       impl: https://github.com/obi1kenobi/cargo-semver-checks/tree/v0.50.0/src/lints/enum_variant_added.ron

Failed in:
  variant ExprType:OptionalFilter in /home/runner/work/datafusion/datafusion/datafusion/proto-models/src/generated/prost.rs:1686
  variant ExprType:OptionalFilter in /home/runner/work/datafusion/datafusion/datafusion/proto-models/src/generated/prost.rs:1686

     Summary semver requires new major version: 2 major and 0 minor checks failed
    Finished [  55.023s] datafusion-proto-models
    Building datafusion-pruning v55.1.0 (current)
       Built [  42.298s] (current)
     Parsing datafusion-pruning v55.1.0 (current)
      Parsed [   0.015s] (current)
    Building datafusion-pruning v55.1.0 (baseline)
       Built [  41.934s] (baseline)
     Parsing datafusion-pruning v55.1.0 (baseline)
      Parsed [   0.016s] (baseline)
    Checking datafusion-pruning v55.1.0 -> v55.1.0 (no change; assume patch)
     Checked [   0.074s] 223 checks: 223 pass, 31 skip
     Summary no semver update required
    Finished [  85.383s] datafusion-pruning
    Building datafusion-sqllogictest v55.1.0 (current)
       Built [ 105.868s] (current)
     Parsing datafusion-sqllogictest v55.1.0 (current)
      Parsed [   0.022s] (current)
    Building datafusion-sqllogictest v55.1.0 (baseline)
       Built [ 105.517s] (baseline)
     Parsing datafusion-sqllogictest v55.1.0 (baseline)
      Parsed [   0.022s] (baseline)
    Checking datafusion-sqllogictest v55.1.0 -> v55.1.0 (no change; assume patch)
     Checked [   0.098s] 223 checks: 223 pass, 31 skip
     Summary no semver update required
    Finished [ 214.210s] datafusion-sqllogictest

@github-actions github-actions Bot added the auto detected api change Auto detected API change label Sep 24, 2026
@codecov-commenter

codecov-commenter commented Sep 25, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 86.16106% with 177 lines in your changes missing coverage. Please review.
✅ Project coverage is 82.50%. Comparing base (1be6b04) to head (2246afc).
⚠️ Report is 21 commits behind head on main.

Files with missing lines Patch % Lines
datafusion/proto-models/src/generated/pbjson.rs 5.71% 64 Missing and 2 partials ⚠️
datafusion/datasource/src/file_scan_config/mod.rs 83.09% 24 Missing and 11 partials ⚠️
...n/physical-expr/src/expressions/optional_filter.rs 79.53% 17 Missing and 18 partials ⚠️
datafusion/datasource-parquet/src/push_decoder.rs 80.51% 9 Missing and 6 partials ⚠️
...usion/datasource-parquet/src/decoder_projection.rs 93.36% 3 Missing and 11 partials ⚠️
datafusion/datasource-parquet/src/row_filter.rs 95.74% 2 Missing and 2 partials ⚠️
datafusion/physical-expr/src/utils/mod.rs 96.47% 0 Missing and 3 partials ⚠️
datafusion/physical-expr/src/simplifier/mod.rs 88.88% 0 Missing and 2 partials ⚠️
...er/src/ensure_requirements/enforce_distribution.rs 71.42% 0 Missing and 2 partials ⚠️
datafusion/proto/src/physical_plan/from_proto.rs 0.00% 0 Missing and 1 partial ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main   #25722      +/-   ##
==========================================
+ Coverage   82.49%   82.50%   +0.01%     
==========================================
  Files        1140     1141       +1     
  Lines      438781   439700     +919     
  Branches   438781   439700     +919     
==========================================
+ Hits       361965   362790     +825     
- Misses      54938    54978      +40     
- Partials    21878    21932      +54     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

When the TopK dynamic filter prunes every remaining row group at a row
group boundary, `rebuild_decoder_at_boundary` returns `Ok(true)` and the
stream calls `finish()`. With a batch coalescer (a post-scan filter is
present, for example with the default `pushdown_filters = false`),
`finish()` flushes the coalescer and returns a batch. The next poll then
went back to the decoder, which still pointed at a row group that the
plan had dropped, and `sync_rg_plan_to_decoder_frontier` failed with
"push decoder frontier RG N is not in rg_plan; decoder and plan have
diverged". ClickBench Q23, Q24 and Q26 fail with this error.

After the flush, the stream now only drains the coalescer.

The new sqllogictest in `dynamic_row_group_pruning.slt` fails without
the fix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
adriangb and others added 4 commits September 25, 2026 21:08
…stribution

Move the check "a round-robin repartition of this input is useful for its
row count" from `EnforceDistribution` into
`repartition::round_robin_beneficial_for_rows`. The behavior does not
change. The next commit uses the same check in the file scan, so that the
scan and the optimizer make the same decision.

PR: apache#22384

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…et partitions

The Parquet scan accepts all pushable filters, thus `FilterPushdown`
removes the `FilterExec`. For a scan of one small file (one partition,
too small to split into byte ranges), main puts a round-robin
`RepartitionExec` between the scan and the `FilterExec`, and a
`CoalescePartitionsExec` above them. Without the `FilterExec`, the
optimizer adds neither. The filter then runs in one partition, and the
scans of the build sides of the hash joins run one after the other in the
task of the probe side, not in parallel tasks. On TPC-DS SF1 this made
short queries 5% to 30% slower than main (for example Q37 1.26x).

`FileScanConfig::try_pushdown_filters` now makes the same decision as
`EnforceDistribution`: if the scan has fewer than `target_partitions`
partitions, `repartitioned` cannot give more, and a round-robin
repartition is useful for the rows that the scan reads, the filters stay
above the scan (`PushedDown::No`). The scan still gets them, through the
new `FileSource::try_pushdown_pruning_filters`, and uses them only to
prune files, row groups and pages. This is what main does with all
filters when `pushdown_filters` is false. The plan is then the plan of
main for these scans.

- Only the filters of a `FilterExec` stay above the scan. A dynamic filter
  (of a join, a TopK or an aggregate) has no `FilterExec` above the scan,
  thus the scan applies it as before.
- The default of `try_pushdown_pruning_filters` returns `None`: other file
  sources get their filters as before.
- The Parquet scan applies all conjuncts of its predicate or none of them.
  A scan with a pruning-only predicate uses later filters only to prune
  too.
- `ParquetScanExecNode` gets `pruning_only_predicate`, thus a decoded scan
  does not apply its predicate again.
- An exact row count of at most one batch keeps the filter in the scan: a
  round-robin repartition cannot split one batch.

Tests:
- unit tests for the decision in `file_scan_config` and for the
  pruning-only predicate of `ParquetSource`;
- a proto round trip of the pruning-only predicate;
- a sqllogictest plan pin in `parquet_filter_pushdown.slt`;
- `parquet_statistics.slt` (no statistics, thus unknown rows): the plan
  is the plan of main again;
- two Parquet integration tests that check the filter in the scan use one
  target partition.

PR: apache#22384

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…orrectness

Add `OptionalFilterPhysicalExpr`, a transparent wrapper that marks a filter
as optional: a consumer can skip it without changing the query result. A
consumer can skip it only when the wrapper is a direct conjunct of the root
AND chain of its predicate. In all other positions the wrapper is
transparent, because `evaluate()` always evaluates the inner expression.
`snapshot()` returns the inner expression, so pruning sees through it.

Also add:
- `split_optional` and `is_optional_filter` helpers in
  `physical_expr::utils` for consumers
- `PhysicalOptionalFilterNode` proto message (field 29 in
  `PhysicalExprNode`) with self-encoding `try_to_proto`/`try_from_proto`

No producer uses the wrapper yet, so there is no behavior change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With `pushdown_filters = false`, the scan evaluates all accepted filters for each row after the decode. This includes the optional filters (hash join, TopK and aggregate dynamic filters, which the producers wrap in `Optional(...)`). The join dynamic filter evaluated for each row in the scan caused a 1.16x TPC-H regression.

Optional filters are not needed for correctness. Thus the post-scan filter now never gets an optional conjunct (a root `AND` conjunct found with `split_optional`):

- `pushdown_filters = false`: only the required conjuncts run post-scan. Optional conjuncts are used only for statistics, page index, bloom filter and file pruning.
- `pushdown_filters = true`: an optional conjunct that the row filter rejects for a file is not used for that file. A required conjunct that is rejected still runs post-scan.
- A whole-file row filter build error sends only the required conjuncts to the post-scan filter.

Accepted optional conjuncts with `pushdown_filters = true` stay row filter predicates, as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@adriangb
adriangb force-pushed the parquet-post-scan-optional branch from 827ead9 to 2246afc Compare September 26, 2026 03:02
@github-actions github-actions Bot added optimizer Optimizer rules physical-plan Changes to the physical-plan crate labels Sep 26, 2026
adriangb and others added 4 commits September 27, 2026 23:32
Conflict: one expected plan in cte.slt. apache#25780 changes the plan of main
(the `FilterExec` above the scan keeps its order); this PR removes the
`FilterExec` (the scan accepts the filter). This PR's plan is kept.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
apache#25780 on main added `FileSource::exact_filter`: the part of the filter
that every output row satisfies, the only part that the scan derives
equivalences from. Its `ParquetSource` version returns the pushable
conjuncts when `pushdown_filters` is on, and nothing otherwise. This PR
changes both cases:

- A pruning-only predicate (a filter that stays in a `FilterExec` above
  a scan that cannot give the target partitions) is used only to prune,
  also with `pushdown_filters = true`. `exact_filter` returned it, thus
  the scan claimed that `a` is constant for `a = 5`, the
  order-preserving repartition merged on `b` only, and
  `ORDER BY b LIMIT 1` returned 2 instead of 1. It now returns `None`.
- With `pushdown_filters = false` the scan applies the accepted
  conjuncts in the post-scan filter. They are exact, thus
  `exact_filter` returns them. This keeps the plans of this PR (for
  example no `SortExec` for `ORDER BY b` with `b = 2`).

Tests: a new case in `push_down_filter_parquet.slt` (a plan pin and two
results that were wrong: `2` for `LIMIT 1`, and `5 2 / 5 1` for the
order) and `exact_filter` checks in the pruning-only unit test. The plan
of the apache#25780 case with `pushdown_filters = false` changes: the scan
applies `a = 5` and there is no `FilterExec`.

PR: apache#22384

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
apache#22384 now has apache/main (with apache#25780, `FileSource::exact_filter`) and
its fix for a pruning-only predicate. No conflict.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`ParquetSource::exact_filter` (apache#25780) returns the conjuncts that every
output row satisfies. An optional conjunct is not one of them: the scan
does not evaluate it after the decode (this PR), and it drops it when
the `RowFilter` cannot evaluate it. `exact_filter` now skips optional
conjuncts, so that the scan claims no equivalence from them.

Test: `exact_filter_excludes_optional_conjuncts` (failed before this
change: `a@0 = 1 AND Optional(b@1 = 2)` with `pushdown_filters = false`).

PR: apache#25722

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Sep 28, 2026
adriangb added a commit to pydantic/datafusion that referenced this pull request Sep 28, 2026
apache#25722 now has the current apache#22384, apache/main (apache#25780,
`FileSource::exact_filter`) and the fixes for pruning-only predicates and
optional conjuncts in `ParquetSource::exact_filter`. No conflict.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
adriangb added a commit to pydantic/datafusion that referenced this pull request Sep 28, 2026
`ParquetSource::exact_filter` (apache#25780) returns the conjuncts that every
output row satisfies. An optional conjunct is not one of them: the scan
does not evaluate it after the decode (this PR), and it drops it when
the `RowFilter` cannot evaluate it. `exact_filter` now skips optional
conjuncts, so that the scan claims no equivalence from them.

Test: `exact_filter_excludes_optional_conjuncts` (failed before this
change: `a@0 = 1 AND Optional(b@1 = 2)` with `pushdown_filters = false`).

PR: apache#25722

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

auto detected api change Auto detected API change core Core DataFusion crate datasource Changes to the datasource crate documentation Improvements or additions to documentation optimizer Optimizer rules physical-expr Changes to the physical-expr crates physical-plan Changes to the physical-plan crate proto Related to proto crate sqllogictest SQL Logic Tests (.slt)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants