Repository navigation
Conversation
|
Thank you for opening this pull request! Reviewer note: cargo-semver-checks reported the current version number is not SemVer-compatible with the changes in this pull request (compared against the base branch). Details |
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #25328 +/- ##
==========================================
+ Coverage 82.60% 82.63% +0.02%
==========================================
Files 1144 1145 +1
Lines 442904 443867 +963
Branches 442904 443867 +963
==========================================
+ Hits 365881 366793 +912
- Misses 54870 54887 +17
- Partials 22153 22187 +34 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
@tschwarzinger Please do review when you get a chance and add other reviewers if required, thank you |
|
@athlcode thank you for working on this! Sorry for the late reply, I've been on vacation. The approach looks sound by quickly scrolling over it. I do have two reservations before doing a full review:
Relevant issues are: |
…te slt for exact distinct change
Which issue does this PR close?
Rationale for this change
Parquet row group, page index and bloom filter pruning only happens when a scan executes, so the optimizer plans with statistics for entire files. When data is clustered on a filtered column, the estimates can be far off. On the IMDB
movie_infotable sorted byinfo,info IN ('Bulgaria')is estimated at 14.8M rows, while page index pruning narrows the scan to 15,360 rows. Better estimates at planning time can improve join selection and enable follow-ups such as eager fetching of small scans (#24922).What changes are included in this PR?
datafusion.execution.parquet.eager_pruning(disabled(default),row_groups,page_index,bloom_filters) andeager_pruning_file_limit(default 256). Both can also be set per table viaformat.options.FileFormat::create_physical_plantakes a newfilters: &[Arc<dyn PhysicalExpr>]argument.ListingTablepasses its non-partition filters as hints.ParquetFormatruns the same pruning stages during planning:ParquetAccessPlanto each file, with fully matched flags cleared since the scan may execute with a different predicate,EXPLAINshows the outcome, e.g.eager_pruning=[level=row_groups, files_pruned=0/1, row_groups_pruned=8/10, rows_pruned=800/1000].configs.mdand the 56.0.0 upgrade guide.What is the testing strategy for this PR?
datasource-parquet/src/eager_pruning.rs(statistics rules, file limit, fully matched flags, display).core/tests/parquet/eager_pruning.rs:collect_statistics = false.parquet_eager_pruning.sltcovering EXPLAIN output for each level.movie_info, the estimate goes from 14.8M to 15,360 rows atpage_index, with under 0.5 ms of added planning. On this data query 3b got about 10% slower because the better estimates changed the join plan. That is worth discussing.Are there any user-facing changes?
FileFormat::create_physical_plan(needs theapi changelabel). Migration notes are in the upgrade guide.