Skip to content

perf: share one StatisticsContext across all physical optimizer rules - #26094

Open
asolimando wants to merge 8 commits into
apache:mainfrom
asolimando:asolimando/query-lifetime-stats-cache
Open

asolimando wants to merge 8 commits into
apache:mainfrom
asolimando:asolimando/query-lifetime-stats-cache

Conversation

@asolimando

@asolimando asolimando commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Which issue does this PR close?

Rationale for this change

Each physical optimizer rule that reads statistics creates its own StatisticsContext, so the statistics of the same plan nodes are computed again in every rule. JoinSelection creates a new context for every get_stats call, so it has no caching at all, even within its own pass.

#25098 showed the gain of sharing one context inside EnsureDistribution. #25929 made cache entries keep the plan node they were computed for, so a context can now be shared safely across plan rewrites. This PR uses that to share one StatisticsContext across all the rules of one optimize_physical_plan call.

On TPC-DS (98 queries) and TPC-H (21 queries), sf1 Parquet, planned with a local harness (not part of this PR):

main this PR
Statistics cache misses (statistics computed from scratch) 46,022 25,428 (-45%)
Physical planning time, sum over all queries 294.8 / 293.5 ms 274.3 / 279.7 ms (about -6%)
Physical plans identical for all 119 queries

The largest gain is TPC-DS q64: cache misses go from 3,785 to 531 and planning time drops by about 21%.

What changes are included in this PR?

  • StatisticsContext stores its cache in a parking_lot::Mutex instead of Rc<RefCell>, so it is Send + Sync. This is needed because PhysicalOptimizerContext: Send + Sync. The lock is held only for single map lookups and inserts, never across the recursive walk. The compute_statistics benchmark shows no difference from main (all cases within ±4%, in both directions).
  • New PhysicalOptimizerContext::statistics_context(), which returns None by default. DefaultPhysicalPlanner creates one context per optimize_physical_plan call, built from the session's statistics registry, and returns it to every rule.
  • ConfigOnlyContext also owns a StatisticsContext, so a rule called through optimize() shares one cache for its whole pass.
  • JoinSelection, EnsureRequirements (including PlanSize::from_plan), AggregateStatistics and LimitPushdown use the shared context. When a context does not share one (for example, a context received through FFI), the rules create a new context from the statistics registry, as they did before.
  • New default method PhysicalOptimizerContext::compute_statistics, which custom rules can call on the context they already receive. Without it, a custom rule would have to repeat the fallback itself (use the shared context, otherwise build one from statistics_registry()), and the obvious shortcut, StatisticsContext::new(), does not consult the registered statistics providers. With it, custom rules behave exactly like the built-in ones.
  • The fallback itself is the public function with_statistics_context. Rules that walk the plan (EnsureRequirements, LimitPushdown) use it to get one StatisticsContext for the whole traversal, so a fallback context still caches within the pass.
  • pushdown_limit_helper is deprecated in favour of the new pushdown_limit_helper_with_stats, which takes a &StatisticsContext. This is the same pattern as ensure_distribution_with_stats.
  • The shared cache keeps every plan node whose statistics were computed alive until optimize_physical_plan returns, including nodes that a later rule replaces. Only the rules that read statistics add entries, and the memory is released at the end of physical planning.

What is the testing strategy for this PR?

  • New optimizer_rules_share_statistics_context test in physical_planner.rs: two rules compute the root statistics through the shared context, and the second one gets the Arc cached by the first.
  • Existing tests cover the rewired rules. All sqllogictests pass with no plan changes.

Are there any user-facing changes?

No breaking changes.

  • New default methods PhysicalOptimizerContext::statistics_context() and PhysicalOptimizerContext::compute_statistics().
  • New public functions datafusion_session::with_statistics_context and pushdown_limit_helper_with_stats; pushdown_limit_helper is deprecated. All are described in the 56.0.0 upgrade guide.
  • StatisticsContext is now Send + Sync.
  • AggregateStatistics, LimitPushdown and PlanSize::from_plan now consult the session's statistics providers, as JoinSelection and EnsureRequirements already do. Without registered providers (the default), their results are unchanged.

Disclaimer: I used AI to assist in the code generation, I have manually reviewed the output and it matches my intention and understanding.

Create one StatisticsContext per optimize_physical_plan call and expose it
through the new PhysicalOptimizerContext::statistics_context, so statistics
computed by one rule are reused by later rules. JoinSelection,
EnsureRequirements, AggregateStatistics and LimitPushdown use it.

StatisticsContext stores its cache in a parking_lot::Mutex so it is
Send + Sync, as PhysicalOptimizerContext requires.
@github-actions github-actions Bot added optimizer Optimizer rules core Core DataFusion crate physical-plan Changes to the physical-plan crate labels Oct 6, 2026
@asolimando

Copy link
Copy Markdown
Member Author

@kosiew @zhuqi-lucas, could you run run benchmark sql_planner as I don't rights myself?

I figure you'd be interested in this PR as it builds on #25929, and it's the query-lifetime follow-up of what #25098 did for EnsureDistribution.

@codecov-commenter

codecov-commenter commented Oct 6, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 82.48588% with 31 lines in your changes missing coverage. Please review.
✅ Project coverage is 82.72%. Comparing base (db83fcc) to head (5a4626e).
⚠️ Report is 7 commits behind head on main.

Files with missing lines Patch % Lines
...atafusion/physical-optimizer/src/limit_pushdown.rs 59.37% 8 Missing and 5 partials ⚠️
...er/src/ensure_requirements/enforce_distribution.rs 83.78% 5 Missing and 1 partial ⚠️
...atafusion/physical-optimizer/src/join_selection.rs 58.33% 0 Missing and 5 partials ⚠️
datafusion/core/src/physical_planner.rs 91.11% 0 Missing and 4 partials ⚠️
.../physical-optimizer/src/ensure_requirements/mod.rs 85.71% 0 Missing and 1 partial ⚠️
datafusion/physical-plan/src/statistics.rs 80.00% 1 Missing ⚠️
datafusion/session/src/physical_optimizer.rs 95.65% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff            @@
##             main   #26094     +/-   ##
=========================================
  Coverage   82.72%   82.72%             
=========================================
  Files        1147     1147             
  Lines      448179   449304   +1125     
  Branches   448179   449304   +1125     
=========================================
+ Hits       370754   371701    +947     
- Misses      54895    54947     +52     
- Partials    22530    22656    +126     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@zhuqi-lucas

Copy link
Copy Markdown
Contributor

run benchmark sql_planner

@zhuqi-lucas

Copy link
Copy Markdown
Contributor

@kosiew @zhuqi-lucas, could you run run benchmark sql_planner as I don't rights myself?

I figure you'd be interested in this PR as it builds on #25929, and it's the query-lifetime follow-up of what #25098 did for EnsureDistribution.

@asolimando triggered now

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c6029846133-3104-clqsv 6.12.94+ #1 SMP Fri Aug 21 08:00:16 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing asolimando/query-lifetime-stats-cache (3179bd9) to db83fcc (merge-base) diff

Run configuration
run benchmark sql_planner

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing asolimando/query-lifetime-stats-cache (3179bd9) to db83fcc (merge-base) diff

Run configuration
run benchmark sql_planner
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                                 HEAD                                   asolimando_query-lifetime-stats-cache
-----                                                 ----                                   -------------------------------------
logical_aggregate_with_join                           1.01    348.9±1.61µs        ? ?/sec    1.00    347.0±1.25µs        ? ?/sec
logical_correlated_subquery_exists                    1.00    204.9±1.27µs        ? ?/sec    1.01    206.0±1.32µs        ? ?/sec
logical_correlated_subquery_in                        1.00    206.5±0.71µs        ? ?/sec    1.00    207.2±0.47µs        ? ?/sec
logical_distinct_many_columns                         1.00    292.4±0.86µs        ? ?/sec    1.00    292.2±0.89µs        ? ?/sec
logical_join_4_with_agg_and_filter                    1.04    198.8±1.28µs        ? ?/sec    1.00    191.2±1.44µs        ? ?/sec
logical_join_8_with_agg_sort_limit                    1.01    347.2±1.14µs        ? ?/sec    1.00    344.2±2.02µs        ? ?/sec
logical_join_chain_16                                 1.00    613.0±2.00µs        ? ?/sec    1.01    621.5±2.45µs        ? ?/sec
logical_join_chain_4                                  1.02     75.8±0.33µs        ? ?/sec    1.00     74.1±0.27µs        ? ?/sec
logical_join_chain_8                                  1.00    199.8±0.51µs        ? ?/sec    1.00    199.5±0.91µs        ? ?/sec
logical_multiple_subqueries                           1.00    392.4±3.30µs        ? ?/sec    1.00    392.6±7.08µs        ? ?/sec
logical_nested_cte_4_levels                           1.03    199.4±0.98µs        ? ?/sec    1.00    193.7±1.14µs        ? ?/sec
logical_plan_struct_join_agg_sort                     1.08    127.2±0.86µs        ? ?/sec    1.00    117.4±0.66µs        ? ?/sec
logical_plan_tpcds_all                                1.02     81.1±0.15ms        ? ?/sec    1.00     79.6±0.13ms        ? ?/sec
logical_plan_tpch_all                                 1.04      5.6±0.01ms        ? ?/sec    1.00      5.4±0.02ms        ? ?/sec
logical_scalar_subquery                               1.01    215.3±1.18µs        ? ?/sec    1.00    213.6±1.26µs        ? ?/sec
logical_select_all_from_1000                          1.00      6.9±0.02ms        ? ?/sec    1.05      7.2±0.02ms        ? ?/sec
logical_select_one_from_700                           1.00    229.5±3.49µs        ? ?/sec    1.00    229.4±1.33µs        ? ?/sec
logical_trivial_join_high_numbered_columns            1.00    208.5±1.53µs        ? ?/sec    1.00    208.7±1.20µs        ? ?/sec
logical_trivial_join_low_numbered_columns             1.00    196.1±1.42µs        ? ?/sec    1.01    197.7±1.26µs        ? ?/sec
logical_union_4_branches                              1.00    306.5±1.31µs        ? ?/sec    1.01    310.6±1.24µs        ? ?/sec
logical_union_8_branches                              1.00    619.6±1.64µs        ? ?/sec    1.02    629.4±3.04µs        ? ?/sec
logical_wide_aggregate_1000_exprs                     1.00     41.1±0.07ms        ? ?/sec    1.02     41.8±0.07ms        ? ?/sec
logical_wide_aggregate_100_exprs                      1.00   1548.0±7.52µs        ? ?/sec    1.00   1541.5±2.64µs        ? ?/sec
logical_wide_case_50_exprs                            1.00   1695.8±3.49µs        ? ?/sec    1.01   1704.9±5.24µs        ? ?/sec
logical_wide_filter_200_predicates                    1.00   1307.1±8.82µs        ? ?/sec    1.00   1313.6±6.29µs        ? ?/sec
logical_wide_filter_50_predicates                     1.01    343.9±1.76µs        ? ?/sec    1.00    340.4±1.20µs        ? ?/sec
optimizer_correlated_exists                           1.00    218.7±0.68µs        ? ?/sec    1.02    224.0±0.82µs        ? ?/sec
optimizer_join_4_with_agg_filter                      1.00    403.4±2.51µs        ? ?/sec    1.00    403.8±1.24µs        ? ?/sec
optimizer_join_chain_4                                1.00    156.6±0.28µs        ? ?/sec    1.02    159.8±0.39µs        ? ?/sec
optimizer_join_chain_8                                1.00    533.4±1.06µs        ? ?/sec    1.02    542.7±2.41µs        ? ?/sec
optimizer_select_all_from_1000                        1.00      6.8±0.01ms        ? ?/sec    1.05      7.1±0.01ms        ? ?/sec
optimizer_select_one_from_700                         1.00    240.8±0.69µs        ? ?/sec    1.01    243.5±0.47µs        ? ?/sec
optimizer_tpcds_all                                   1.01    277.7±0.39ms        ? ?/sec    1.00    275.9±0.36ms        ? ?/sec
optimizer_tpch_all                                    1.02     16.0±0.12ms        ? ?/sec    1.00     15.7±0.03ms        ? ?/sec
optimizer_wide_aggregate_100                          1.00   1750.6±7.40µs        ? ?/sec    1.00   1749.1±2.84µs        ? ?/sec
optimizer_wide_filter_200                             1.00      3.9±0.01ms        ? ?/sec    1.00      3.9±0.01ms        ? ?/sec
physical_intersection                                 1.01    520.7±1.73µs        ? ?/sec    1.00    517.8±2.74µs        ? ?/sec
physical_join_consider_sort                           1.00    964.8±2.05µs        ? ?/sec    1.00    966.8±7.02µs        ? ?/sec
physical_join_distinct                                1.00    189.8±1.26µs        ? ?/sec    1.01    192.1±1.18µs        ? ?/sec
physical_many_self_joins                              1.00      7.6±0.01ms        ? ?/sec    1.02      7.7±0.02ms        ? ?/sec
physical_plan_clickbench_all                          1.00    131.0±0.60ms        ? ?/sec    1.01    132.6±1.37ms        ? ?/sec
physical_plan_clickbench_q1                           1.00  1281.7±11.53µs        ? ?/sec    1.03   1314.9±7.82µs        ? ?/sec
physical_plan_clickbench_q10                          1.00   1875.3±6.23µs        ? ?/sec    1.02   1912.9±6.76µs        ? ?/sec
physical_plan_clickbench_q11                          1.00      2.0±0.00ms        ? ?/sec    1.01      2.0±0.01ms        ? ?/sec
physical_plan_clickbench_q12                          1.00      2.1±0.00ms        ? ?/sec    1.01      2.1±0.01ms        ? ?/sec
physical_plan_clickbench_q13                          1.00  1898.5±10.89µs        ? ?/sec    1.01   1917.0±5.97µs        ? ?/sec
physical_plan_clickbench_q14                          1.02      2.1±0.01ms        ? ?/sec    1.00      2.0±0.01ms        ? ?/sec
physical_plan_clickbench_q15                          1.00   1932.5±5.13µs        ? ?/sec    1.02   1974.1±6.90µs        ? ?/sec
physical_plan_clickbench_q16                          1.02   1701.9±6.43µs        ? ?/sec    1.00   1673.7±5.21µs        ? ?/sec
physical_plan_clickbench_q17                          1.00   1714.8±4.69µs        ? ?/sec    1.00  1719.8±12.16µs        ? ?/sec
physical_plan_clickbench_q18                          1.01   1568.7±4.85µs        ? ?/sec    1.00   1552.1±5.34µs        ? ?/sec
physical_plan_clickbench_q19                          1.01   1896.2±5.88µs        ? ?/sec    1.00   1881.8±9.36µs        ? ?/sec
physical_plan_clickbench_q2                           1.00   1689.5±8.98µs        ? ?/sec    1.00   1688.4±5.76µs        ? ?/sec
physical_plan_clickbench_q20                          1.02   1504.5±6.04µs        ? ?/sec    1.00   1473.8±4.88µs        ? ?/sec
physical_plan_clickbench_q21                          1.02   1705.4±5.51µs        ? ?/sec    1.00   1678.1±5.29µs        ? ?/sec
physical_plan_clickbench_q22                          1.02      2.1±0.01ms        ? ?/sec    1.00      2.0±0.00ms        ? ?/sec
physical_plan_clickbench_q23                          1.01      2.1±0.01ms        ? ?/sec    1.00      2.1±0.01ms        ? ?/sec
physical_plan_clickbench_q24                          1.01      4.9±0.01ms        ? ?/sec    1.00      4.8±0.01ms        ? ?/sec
physical_plan_clickbench_q25                          1.00   1805.2±4.89µs        ? ?/sec    1.01   1817.2±6.34µs        ? ?/sec
physical_plan_clickbench_q26                          1.00   1655.2±5.01µs        ? ?/sec    1.01   1668.8±5.59µs        ? ?/sec
physical_plan_clickbench_q27                          1.00   1827.8±5.70µs        ? ?/sec    1.00   1819.2±5.25µs        ? ?/sec
physical_plan_clickbench_q28                          1.00      2.1±0.01ms        ? ?/sec    1.01      2.1±0.01ms        ? ?/sec
physical_plan_clickbench_q29                          1.00      2.2±0.00ms        ? ?/sec    1.00      2.2±0.01ms        ? ?/sec
physical_plan_clickbench_q3                           1.00  1589.8±18.41µs        ? ?/sec    1.00   1587.9±4.09µs        ? ?/sec
physical_plan_clickbench_q30                          1.00     12.4±0.02ms        ? ?/sec    1.00     12.4±0.02ms        ? ?/sec
physical_plan_clickbench_q31                          1.00      2.2±0.00ms        ? ?/sec    1.00      2.2±0.01ms        ? ?/sec
physical_plan_clickbench_q32                          1.00      2.2±0.00ms        ? ?/sec    1.02      2.2±0.00ms        ? ?/sec
physical_plan_clickbench_q33                          1.00   1844.1±4.53µs        ? ?/sec    1.01   1862.0±5.15µs        ? ?/sec
physical_plan_clickbench_q34                          1.01   1703.3±4.94µs        ? ?/sec    1.00   1691.8±5.55µs        ? ?/sec
physical_plan_clickbench_q35                          1.00   1737.0±6.03µs        ? ?/sec    1.02  1765.8±26.72µs        ? ?/sec
physical_plan_clickbench_q36                          1.01   1982.0±6.94µs        ? ?/sec    1.00   1968.6±5.25µs        ? ?/sec
physical_plan_clickbench_q37                          1.01      2.4±0.00ms        ? ?/sec    1.00      2.4±0.01ms        ? ?/sec
physical_plan_clickbench_q38                          1.01      2.4±0.01ms        ? ?/sec    1.00      2.4±0.00ms        ? ?/sec
physical_plan_clickbench_q39                          1.02      2.5±0.01ms        ? ?/sec    1.00      2.4±0.01ms        ? ?/sec
physical_plan_clickbench_q4                           1.00  1409.6±17.78µs        ? ?/sec    1.00   1404.2±4.90µs        ? ?/sec
physical_plan_clickbench_q40                          1.02      3.1±0.02ms        ? ?/sec    1.00      3.1±0.02ms        ? ?/sec
physical_plan_clickbench_q41                          1.02      2.7±0.01ms        ? ?/sec    1.00      2.6±0.01ms        ? ?/sec
physical_plan_clickbench_q42                          1.01      2.7±0.01ms        ? ?/sec    1.00      2.7±0.01ms        ? ?/sec
physical_plan_clickbench_q43                          1.01      2.9±0.01ms        ? ?/sec    1.00      2.9±0.01ms        ? ?/sec
physical_plan_clickbench_q44                          1.03   1497.9±5.68µs        ? ?/sec    1.00   1454.6±4.87µs        ? ?/sec
physical_plan_clickbench_q45                          1.01   1499.4±2.77µs        ? ?/sec    1.00   1479.6±6.38µs        ? ?/sec
physical_plan_clickbench_q46                          1.00   1729.9±4.77µs        ? ?/sec    1.01   1745.1±5.08µs        ? ?/sec
physical_plan_clickbench_q47                          1.00      2.3±0.00ms        ? ?/sec    1.01      2.3±0.01ms        ? ?/sec
physical_plan_clickbench_q48                          1.00      2.4±0.00ms        ? ?/sec    1.00      2.5±0.00ms        ? ?/sec
physical_plan_clickbench_q49                          1.00      2.5±0.01ms        ? ?/sec    1.00      2.5±0.01ms        ? ?/sec
physical_plan_clickbench_q5                           1.01  1517.9±19.89µs        ? ?/sec    1.00   1507.5±5.02µs        ? ?/sec
physical_plan_clickbench_q50                          1.00      2.5±0.01ms        ? ?/sec    1.00      2.5±0.01ms        ? ?/sec
physical_plan_clickbench_q51                          1.00   1848.1±4.71µs        ? ?/sec    1.00   1843.0±5.58µs        ? ?/sec
physical_plan_clickbench_q52                          1.00      2.3±0.00ms        ? ?/sec    1.02      2.3±0.01ms        ? ?/sec
physical_plan_clickbench_q53                          1.00   1669.9±5.43µs        ? ?/sec    1.00   1674.4±6.61µs        ? ?/sec
physical_plan_clickbench_q54                          1.00  1671.6±13.04µs        ? ?/sec    1.00  1665.0±10.12µs        ? ?/sec
physical_plan_clickbench_q55                          1.02   1650.5±5.33µs        ? ?/sec    1.00  1611.0±12.31µs        ? ?/sec
physical_plan_clickbench_q56                          1.00   1633.1±5.36µs        ? ?/sec    1.00   1626.5±6.63µs        ? ?/sec
physical_plan_clickbench_q57                          1.00   1732.6±5.68µs        ? ?/sec    1.02  1769.3±16.62µs        ? ?/sec
physical_plan_clickbench_q58                          1.00      2.2±0.01ms        ? ?/sec    1.01      2.2±0.03ms        ? ?/sec
physical_plan_clickbench_q59                          1.00   1633.4±6.77µs        ? ?/sec    1.01  1654.8±11.75µs        ? ?/sec
physical_plan_clickbench_q6                           1.00   1493.9±7.48µs        ? ?/sec    1.01   1501.5±4.74µs        ? ?/sec
physical_plan_clickbench_q60                          1.01   1646.6±4.30µs        ? ?/sec    1.00   1634.9±8.75µs        ? ?/sec
physical_plan_clickbench_q7                           1.00   1343.2±5.40µs        ? ?/sec    1.04   1391.1±6.67µs        ? ?/sec
physical_plan_clickbench_q8                           1.00   1814.1±5.73µs        ? ?/sec    1.04  1890.4±11.33µs        ? ?/sec
physical_plan_clickbench_q9                           1.00   1755.2±9.96µs        ? ?/sec    1.02   1790.3±5.34µs        ? ?/sec
physical_plan_struct_join_agg_sort                    1.01   1214.6±2.01µs        ? ?/sec    1.00   1208.2±3.06µs        ? ?/sec
physical_plan_tpcds_all                               1.04    626.1±1.36ms        ? ?/sec    1.00    600.9±4.20ms        ? ?/sec
physical_plan_tpch_all                                1.02     41.4±0.22ms        ? ?/sec    1.00     40.4±0.07ms        ? ?/sec
physical_plan_tpch_q1                                 1.03   1437.2±4.69µs        ? ?/sec    1.00   1394.2±2.60µs        ? ?/sec
physical_plan_tpch_q10                                1.04      2.3±0.00ms        ? ?/sec    1.00      2.2±0.00ms        ? ?/sec
physical_plan_tpch_q11                                1.03      2.2±0.00ms        ? ?/sec    1.00      2.1±0.02ms        ? ?/sec
physical_plan_tpch_q12                                1.02   1216.7±2.10µs        ? ?/sec    1.00   1198.2±2.29µs        ? ?/sec
physical_plan_tpch_q13                                1.01   1005.9±2.46µs        ? ?/sec    1.00    994.5±2.96µs        ? ?/sec
physical_plan_tpch_q14                                1.05  1328.9±11.66µs        ? ?/sec    1.00   1266.2±2.56µs        ? ?/sec
physical_plan_tpch_q16                                1.05  1617.5±14.11µs        ? ?/sec    1.00   1545.3±3.55µs        ? ?/sec
physical_plan_tpch_q17                                1.02   1603.0±2.81µs        ? ?/sec    1.00   1569.9±2.55µs        ? ?/sec
physical_plan_tpch_q18                                1.02   1859.1±3.95µs        ? ?/sec    1.00   1824.6±3.01µs        ? ?/sec
physical_plan_tpch_q19                                1.02   1827.6±3.54µs        ? ?/sec    1.00  1792.0±10.76µs        ? ?/sec
physical_plan_tpch_q2                                 1.04      3.5±0.01ms        ? ?/sec    1.00      3.4±0.00ms        ? ?/sec
physical_plan_tpch_q20                                1.04      2.1±0.00ms        ? ?/sec    1.00  1986.8±16.64µs        ? ?/sec
physical_plan_tpch_q21                                1.05      2.6±0.01ms        ? ?/sec    1.00      2.5±0.01ms        ? ?/sec
physical_plan_tpch_q22                                1.02   1512.8±3.18µs        ? ?/sec    1.00   1485.2±2.52µs        ? ?/sec
physical_plan_tpch_q3                                 1.03   1734.0±3.06µs        ? ?/sec    1.00   1685.3±2.73µs        ? ?/sec
physical_plan_tpch_q4                                 1.04   1102.1±2.54µs        ? ?/sec    1.00   1055.1±3.59µs        ? ?/sec
physical_plan_tpch_q5                                 1.05      2.4±0.01ms        ? ?/sec    1.00      2.3±0.00ms        ? ?/sec
physical_plan_tpch_q6                                 1.03    599.0±1.29µs        ? ?/sec    1.00    580.3±1.47µs        ? ?/sec
physical_plan_tpch_q7                                 1.07      2.6±0.01ms        ? ?/sec    1.00      2.5±0.00ms        ? ?/sec
physical_plan_tpch_q8                                 1.06      3.6±0.01ms        ? ?/sec    1.00      3.4±0.00ms        ? ?/sec
physical_plan_tpch_q9                                 1.04      2.6±0.00ms        ? ?/sec    1.00      2.5±0.01ms        ? ?/sec
physical_select_aggregates_from_200                   1.00      7.4±0.01ms        ? ?/sec    1.01      7.5±0.01ms        ? ?/sec
physical_select_all_from_1000                         1.00     16.9±0.06ms        ? ?/sec    1.05     17.7±0.06ms        ? ?/sec
physical_select_one_from_700                          1.00    639.9±2.63µs        ? ?/sec    1.01    643.2±1.58µs        ? ?/sec
physical_sorted_union_order_by_10_int64               1.00      3.6±0.00ms        ? ?/sec    1.00      3.6±0.00ms        ? ?/sec
physical_sorted_union_order_by_10_uint64              1.01      6.8±0.01ms        ? ?/sec    1.00      6.8±0.03ms        ? ?/sec
physical_sorted_union_order_by_50_int64               1.00     78.7±0.20ms        ? ?/sec    1.00     78.8±0.22ms        ? ?/sec
physical_sorted_union_order_by_50_uint64              1.01    254.2±0.66ms        ? ?/sec    1.00    252.8±0.87ms        ? ?/sec
physical_theta_join_consider_sort                     1.00    993.9±2.34µs        ? ?/sec    1.00    991.4±6.72µs        ? ?/sec
physical_unnest_to_join                               1.00    554.1±2.33µs        ? ?/sec    1.00    555.8±1.81µs        ? ?/sec
physical_window_function_partition_by_12_on_values    1.04    665.1±2.98µs        ? ?/sec    1.00    640.0±1.45µs        ? ?/sec
physical_window_function_partition_by_30_on_values    1.00   1264.6±2.02µs        ? ?/sec    1.00   1259.4±1.65µs        ? ?/sec
physical_window_function_partition_by_4_on_values     1.03    401.6±1.08µs        ? ?/sec    1.00    388.3±1.14µs        ? ?/sec
physical_window_function_partition_by_7_on_values     1.02    491.6±1.53µs        ? ?/sec    1.00    482.0±0.87µs        ? ?/sec
physical_window_function_partition_by_8_on_values     1.02    528.1±1.13µs        ? ?/sec    1.00    516.6±1.17µs        ? ?/sec
with_param_values_many_columns                        1.00    368.7±2.10µs        ? ?/sec    1.02    377.2±2.01µs        ? ?/sec

Resource Usage

sql_planner — base (merge-base)

Metric Value
Wall time 2165.5s
Peak memory 141.2 MiB
Avg memory 85.1 MiB
CPU user 1980.6s
CPU sys 1.5s
Peak spill 0 B

sql_planner — branch

Metric Value
Wall time 2130.5s
Peak memory 140.5 MiB
Avg memory 87.0 MiB
CPU user 1985.4s
CPU sys 1.6s
Peak spill 0 B

File an issue against this benchmark runner

@asolimando

Copy link
Copy Markdown
Member Author

Thanks @zhuqi-lucas! My understanding of the benchmark run:

  • physical_plan_tpcds_all: 626.1 ms -> 600.9 ms (-4.0%)
  • physical_plan_tpch_all: 41.4 ms -> 40.4 ms (-2.4%), and every TPC-H query is faster (-1% to -7%)
  • ClickBench (single table, no joins) is within noise

The +5% on physical_select_all_from_1000 also appears in logical_select_all_from_1000 and optimizer_select_all_from_1000, which only run logical planning, so it is not related to this change.

The gain is little smaller than the sf1 numbers I had locally (6%, as reported in the description) because the benchmark tables are empty, so each statistics computation is cheap.

If you think this is interesting I can do another self-review pass on the PR and remove it from draft, wdyt?

@zhuqi-lucas

Copy link
Copy Markdown
Contributor

Thanks for the detailed breakdown, your reading matches mine. This looks valuable, please go ahead with the self-review and take it out of draft.

@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Thank you for opening this pull request!

Reviewer note: cargo-semver-checks reported the current version number is not SemVer-compatible with the changes in this pull request (compared against the base branch).

Details
     Cloning apache/main
    Building datafusion v55.1.0 (current)
       Built [  56.435s] (current)
     Parsing datafusion v55.1.0 (current)
      Parsed [   0.033s] (current)
    Building datafusion v55.1.0 (baseline)
       Built [  55.804s] (baseline)
     Parsing datafusion v55.1.0 (baseline)
      Parsed [   0.032s] (baseline)
    Checking datafusion v55.1.0 -> v55.1.0 (no change; assume patch)
     Checked [   0.624s] 223 checks: 223 pass, 31 skip
     Summary no semver update required
    Finished [ 114.719s] datafusion
    Building datafusion-physical-optimizer v55.1.0 (current)
       Built [  39.562s] (current)
     Parsing datafusion-physical-optimizer v55.1.0 (current)
      Parsed [   0.018s] (current)
    Building datafusion-physical-optimizer v55.1.0 (baseline)
       Built [  40.613s] (baseline)
     Parsing datafusion-physical-optimizer v55.1.0 (baseline)
      Parsed [   0.020s] (baseline)
    Checking datafusion-physical-optimizer v55.1.0 -> v55.1.0 (no change; assume patch)
     Checked [   0.107s] 223 checks: 222 pass, 1 fail, 0 warn, 31 skip

--- failure function_marked_deprecated: function #[deprecated] added ---

Description:
A function is now #[deprecated]. Downstream crates will get a compiler warning when using this function.
        ref: https://doc.rust-lang.org/reference/attributes/diagnostics.html#the-deprecated-attribute
       impl: https://github.com/obi1kenobi/cargo-semver-checks/tree/v0.50.0/src/lints/function_marked_deprecated.ron

Failed in:
  function datafusion_physical_optimizer::limit_pushdown::pushdown_limit_helper in /home/runner/work/datafusion/datafusion/datafusion/physical-optimizer/src/limit_pushdown.rs:171

     Summary semver requires new minor version: 0 major and 1 minor checks failed
    Finished [  81.534s] datafusion-physical-optimizer
    Building datafusion-physical-plan v55.1.0 (current)
       Built [  36.844s] (current)
     Parsing datafusion-physical-plan v55.1.0 (current)
      Parsed [   0.179s] (current)
    Building datafusion-physical-plan v55.1.0 (baseline)
       Built [  37.754s] (baseline)
     Parsing datafusion-physical-plan v55.1.0 (baseline)
      Parsed [   0.183s] (baseline)
    Checking datafusion-physical-plan v55.1.0 -> v55.1.0 (no change; assume patch)
     Checked [   0.634s] 223 checks: 223 pass, 31 skip
     Summary no semver update required
    Finished [  76.836s] datafusion-physical-plan
    Building datafusion-session v55.1.0 (current)
       Built [  36.861s] (current)
     Parsing datafusion-session v55.1.0 (current)
      Parsed [   0.008s] (current)
    Building datafusion-session v55.1.0 (baseline)
       Built [  37.455s] (baseline)
     Parsing datafusion-session v55.1.0 (baseline)
      Parsed [   0.008s] (baseline)
    Checking datafusion-session v55.1.0 -> v55.1.0 (no change; assume patch)
     Checked [   0.183s] 223 checks: 223 pass, 31 skip
     Summary no semver update required
    Finished [  75.481s] datafusion-session

@github-actions github-actions Bot added the auto detected api change Auto detected API change label Oct 7, 2026
@asolimando

Copy link
Copy Markdown
Member Author

@zhuqi-lucas thanks a lot for your feedback, after self-review I have pushed some test refactoring, doc improvements and improved API ergonomics (added PhysicalOptimizerContext::compute_statistics for custom physical rules), nothing affecting the benchmark (in principle), the PR is now reviewable if you have bandwidth!

I noticed the coverage warning, I could fix that easily but I don't see anything that genuinely needs more coverage, happy to add more tests if you see fit.

OT: is there a process to get into the allow-list to run benchmarks? What are the criteria for eligibility?

@asolimando
asolimando marked this pull request as ready for review October 7, 2026 13:05
/// [`Statistics`] only.
pub struct StatisticsContext {
cache: Rc<RefCell<StatsCache>>,
cache: Mutex<StatsCache>,

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note for reviewers: the Mutex is needed because PhysicalOptimizerContext: Send + Sync, and the shared context is held inside it and returned as &StatisticsContext. RefCell is not Sync (and the Rc was never cloned, so it was dropped rather than replaced).

The optimizer uses the context from one thread, so the lock is never contended, and it is briefly held (single map lookup/insert), never across the recursive walk. The compute_statistics benchmark shows no measurable difference from RefCell (all cases within ±4% of main, in both directions).

Alternatives I considered: a RefCell restricted to the creating thread (needs unsafe impl Sync), a thread-local cache (hidden global state), or removing Sync from PhysicalOptimizerContext (a breaking change). None of them seemed worth it for a lock that costs nothing measurable.

// Otherwise, fall back to standard distribution rather
// than choosing an arbitrary reference.
candidates
let best_satisfied_child: Option<(usize, Partitioning)> =

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note to reviewers: re-indented by rustfmt, the real change is passing stats_ctx to PlanSize::from_plan, but the diff can be confusing.

@zhuqi-lucas zhuqi-lucas left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice work — this is the piece #25098 and #25929 were building toward, and the numbers are convincing.

I checked the two things that worried me most and both hold up:

  • No lock held across the recursive walk. Every borrow() → lock() site is a single lookup or insert whose guard dies at the end of the statement, so the non-reentrant parking_lot::Mutex can't self-deadlock.
  • Pointer keys stay safe over the much longer sharing window. store_cache_entry is the only insertion point and always clones the owner into CacheEntry::_plan, so an address can't be recycled while its entry lives. That invariant is what makes widening the scope from one rule to the whole run sound — worth keeping in mind for anyone who later adds a second insertion path.

Also worth calling out, since it's easy to miss in the diff: the compute(x.as_ref()) → compute_arc(x) switches are not cosmetic. Borrowed roots are deliberately not memoized, so those lines are what actually lets the root of each traversal land in the shared cache.

Three small notes inline, none blocking.

Comment thread datafusion/physical-plan/src/statistics.rs
Comment thread datafusion/physical-optimizer/src/limit_pushdown.rs
Comment thread datafusion/physical-optimizer/src/optimizer.rs
@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Oct 8, 2026
@zhuqi-lucas

Copy link
Copy Markdown
Contributor

run benchmark sql_planner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c6061103335-3167-9bqt8 6.12.94+ #1 SMP Fri Aug 21 08:00:16 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing asolimando/query-lifetime-stats-cache (63aa427) to 3ed377a (merge-base) diff

Run configuration
run benchmark sql_planner

Results will be posted here when complete


File an issue against this benchmark runner

let optimizer_context = SessionOptimizerContext {
session: session_state,
};
let optimizer_context = SessionOptimizerContext::new(session_state);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One more thing this widens, worth stating somewhere: the cache is now exposed to in-place mutation for the whole planning run, not just one rule.

A pointer-keyed hit is only valid because a plan node never changes behind its Arc. That held trivially when each rule had its own context — anything a rule did produced new nodes. Now an entry computed by the first rule is still served to the last one, so a node that mutated its own statistics through interior mutability without changing identity would be read stale.

Nothing does that today (planning-time rules all rebuild nodes, and dynamic filters are updated at execution time, after this context is dropped), so this isn't a bug — but it's an invariant the design now leans on much harder. A line on StatisticsContext saying cached nodes must be immutable for the context's lifetime would make it checkable in review.

@asolimando asolimando Oct 8, 2026 •

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, it is an invariant the design now relies on much more. Added a paragraph in f7587af to the StatisticsContext docs: a plan node must not change its statistics in place while a context holds it, otherwise the cache returns stale values; optimizer rules satisfy this because they replace nodes instead of changing them.

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing asolimando/query-lifetime-stats-cache (63aa427) to 3ed377a (merge-base) diff

Run configuration
run benchmark sql_planner
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                                 HEAD                                   asolimando_query-lifetime-stats-cache
-----                                                 ----                                   -------------------------------------
logical_aggregate_with_join                           1.03    357.8±1.58µs        ? ?/sec    1.00    347.2±1.49µs        ? ?/sec
logical_correlated_subquery_exists                    1.02    210.4±6.76µs        ? ?/sec    1.00    205.3±1.15µs        ? ?/sec
logical_correlated_subquery_in                        1.01    208.3±0.91µs        ? ?/sec    1.00    206.9±0.53µs        ? ?/sec
logical_distinct_many_columns                         1.00    291.9±0.83µs        ? ?/sec    1.01    295.1±0.85µs        ? ?/sec
logical_join_4_with_agg_and_filter                    1.02    200.0±0.81µs        ? ?/sec    1.00    196.0±1.01µs        ? ?/sec
logical_join_8_with_agg_sort_limit                    1.00    343.9±1.34µs        ? ?/sec    1.00    343.7±1.38µs        ? ?/sec
logical_join_chain_16                                 1.00    616.7±1.62µs        ? ?/sec    1.00    617.9±1.76µs        ? ?/sec
logical_join_chain_4                                  1.01     75.9±0.63µs        ? ?/sec    1.00     75.1±0.41µs        ? ?/sec
logical_join_chain_8                                  1.01    199.7±0.46µs        ? ?/sec    1.00    198.7±0.62µs        ? ?/sec
logical_multiple_subqueries                           1.02    401.3±6.65µs        ? ?/sec    1.00    394.5±2.60µs        ? ?/sec
logical_nested_cte_4_levels                           1.00    200.1±0.69µs        ? ?/sec    1.00    199.8±1.29µs        ? ?/sec
logical_plan_struct_join_agg_sort                     1.00    121.2±0.66µs        ? ?/sec    1.03    124.5±1.51µs        ? ?/sec
logical_plan_tpcds_all                                1.00     81.4±0.24ms        ? ?/sec    1.00     81.1±0.21ms        ? ?/sec
logical_plan_tpch_all                                 1.00      5.6±0.02ms        ? ?/sec    1.01      5.6±0.02ms        ? ?/sec
logical_scalar_subquery                               1.00    216.7±1.43µs        ? ?/sec    1.00    216.4±1.44µs        ? ?/sec
logical_select_all_from_1000                          1.00      7.2±0.02ms        ? ?/sec    1.00      7.2±0.04ms        ? ?/sec
logical_select_one_from_700                           1.01    229.1±1.40µs        ? ?/sec    1.00    226.0±4.40µs        ? ?/sec
logical_trivial_join_high_numbered_columns            1.01    210.6±1.35µs        ? ?/sec    1.00    208.2±1.38µs        ? ?/sec
logical_trivial_join_low_numbered_columns             1.01    198.6±1.31µs        ? ?/sec    1.00    196.2±1.34µs        ? ?/sec
logical_union_4_branches                              1.00    312.4±1.26µs        ? ?/sec    1.01    314.6±6.94µs        ? ?/sec
logical_union_8_branches                              1.00    630.1±1.47µs        ? ?/sec    1.00    629.8±7.45µs        ? ?/sec
logical_wide_aggregate_1000_exprs                     1.00     42.7±0.08ms        ? ?/sec    1.00     42.5±0.09ms        ? ?/sec
logical_wide_aggregate_100_exprs                      1.00   1544.1±2.92µs        ? ?/sec    1.01   1558.5±4.23µs        ? ?/sec
logical_wide_case_50_exprs                            1.00   1720.6±3.54µs        ? ?/sec    1.01  1740.6±14.53µs        ? ?/sec
logical_wide_filter_200_predicates                    1.00   1301.7±5.54µs        ? ?/sec    1.01   1309.3±7.32µs        ? ?/sec
logical_wide_filter_50_predicates                     1.00    341.2±1.17µs        ? ?/sec    1.01    343.0±1.24µs        ? ?/sec
optimizer_correlated_exists                           1.00    217.0±0.63µs        ? ?/sec    1.00    216.3±0.66µs        ? ?/sec
optimizer_join_4_with_agg_filter                      1.01    401.3±1.42µs        ? ?/sec    1.00    397.8±0.93µs        ? ?/sec
optimizer_join_chain_4                                1.00    155.0±0.28µs        ? ?/sec    1.00    155.4±0.38µs        ? ?/sec
optimizer_join_chain_8                                1.00    539.0±1.30µs        ? ?/sec    1.00    538.7±0.98µs        ? ?/sec
optimizer_select_all_from_1000                        1.00      7.1±0.01ms        ? ?/sec    1.05      7.4±0.03ms        ? ?/sec
optimizer_select_one_from_700                         1.04    242.5±0.57µs        ? ?/sec    1.00    234.3±0.58µs        ? ?/sec
optimizer_tpcds_all                                   1.00    278.7±1.73ms        ? ?/sec    1.00    278.1±0.61ms        ? ?/sec
optimizer_tpch_all                                    1.01     15.9±0.05ms        ? ?/sec    1.00     15.7±0.05ms        ? ?/sec
optimizer_wide_aggregate_100                          1.00   1743.4±2.38µs        ? ?/sec    1.00   1751.6±3.51µs        ? ?/sec
optimizer_wide_filter_200                             1.00      3.9±0.01ms        ? ?/sec    1.01      4.0±0.01ms        ? ?/sec
physical_intersection                                 1.02    514.3±2.63µs        ? ?/sec    1.00    502.6±2.75µs        ? ?/sec
physical_join_consider_sort                           1.01    963.0±1.32µs        ? ?/sec    1.00    956.2±1.82µs        ? ?/sec
physical_join_distinct                                1.01    191.3±1.19µs        ? ?/sec    1.00    190.1±1.20µs        ? ?/sec
physical_many_self_joins                              1.01      7.7±0.02ms        ? ?/sec    1.00      7.6±0.03ms        ? ?/sec
physical_plan_clickbench_all                          1.00    137.7±0.69ms        ? ?/sec    1.00    137.8±1.44ms        ? ?/sec
physical_plan_clickbench_q1                           1.00   1312.2±8.31µs        ? ?/sec    1.03  1357.9±28.88µs        ? ?/sec
physical_plan_clickbench_q10                          1.00  1924.8±17.81µs        ? ?/sec    1.00  1921.4±15.12µs        ? ?/sec
physical_plan_clickbench_q11                          1.00      2.1±0.02ms        ? ?/sec    1.02      2.1±0.02ms        ? ?/sec
physical_plan_clickbench_q12                          1.00      2.2±0.03ms        ? ?/sec    1.01      2.2±0.02ms        ? ?/sec
physical_plan_clickbench_q13                          1.02  1961.9±21.55µs        ? ?/sec    1.00  1924.5±15.41µs        ? ?/sec
physical_plan_clickbench_q14                          1.01      2.1±0.02ms        ? ?/sec    1.00      2.1±0.02ms        ? ?/sec
physical_plan_clickbench_q15                          1.00      2.0±0.02ms        ? ?/sec    1.00      2.0±0.02ms        ? ?/sec
physical_plan_clickbench_q16                          1.00  1692.3±13.38µs        ? ?/sec    1.01  1704.8±20.13µs        ? ?/sec
physical_plan_clickbench_q17                          1.00  1742.3±11.00µs        ? ?/sec    1.00  1743.3±10.71µs        ? ?/sec
physical_plan_clickbench_q18                          1.00   1578.3±8.14µs        ? ?/sec    1.00  1584.1±10.96µs        ? ?/sec
physical_plan_clickbench_q19                          1.00  1933.2±15.38µs        ? ?/sec    1.00  1940.5±16.87µs        ? ?/sec
physical_plan_clickbench_q2                           1.00  1697.4±11.19µs        ? ?/sec    1.04  1763.8±50.61µs        ? ?/sec
physical_plan_clickbench_q20                          1.00   1480.4±5.35µs        ? ?/sec    1.01   1501.4±9.96µs        ? ?/sec
physical_plan_clickbench_q21                          1.00   1663.4±8.17µs        ? ?/sec    1.03  1705.4±14.04µs        ? ?/sec
physical_plan_clickbench_q22                          1.00      2.0±0.01ms        ? ?/sec    1.02      2.0±0.02ms        ? ?/sec
physical_plan_clickbench_q23                          1.00      2.2±0.01ms        ? ?/sec    1.02      2.2±0.03ms        ? ?/sec
physical_plan_clickbench_q24                          1.00      5.0±0.05ms        ? ?/sec    1.02      5.1±0.10ms        ? ?/sec
physical_plan_clickbench_q25                          1.02   1827.3±9.33µs        ? ?/sec    1.00  1796.2±18.27µs        ? ?/sec
physical_plan_clickbench_q26                          1.00   1671.7±7.04µs        ? ?/sec    1.00  1673.7±11.77µs        ? ?/sec
physical_plan_clickbench_q27                          1.00  1819.7±17.18µs        ? ?/sec    1.02  1848.2±18.13µs        ? ?/sec
physical_plan_clickbench_q28                          1.00      2.2±0.01ms        ? ?/sec    1.01      2.2±0.03ms        ? ?/sec
physical_plan_clickbench_q29                          1.00      2.3±0.04ms        ? ?/sec    1.02      2.4±0.03ms        ? ?/sec
physical_plan_clickbench_q3                           1.00  1570.2±13.00µs        ? ?/sec    1.03  1612.4±33.73µs        ? ?/sec
physical_plan_clickbench_q30                          1.00     13.0±0.07ms        ? ?/sec    1.00     12.9±0.04ms        ? ?/sec
physical_plan_clickbench_q31                          1.01      2.3±0.02ms        ? ?/sec    1.00      2.3±0.02ms        ? ?/sec
physical_plan_clickbench_q32                          1.01      2.3±0.02ms        ? ?/sec    1.00      2.3±0.02ms        ? ?/sec
physical_plan_clickbench_q33                          1.00  1897.5±16.81µs        ? ?/sec    1.01  1919.0±13.46µs        ? ?/sec
physical_plan_clickbench_q34                          1.00   1687.8±5.52µs        ? ?/sec    1.01   1697.0±9.83µs        ? ?/sec
physical_plan_clickbench_q35                          1.00  1759.9±15.08µs        ? ?/sec    1.00  1753.9±13.70µs        ? ?/sec
physical_plan_clickbench_q36                          1.00  1965.3±10.10µs        ? ?/sec    1.01  1993.9±27.71µs        ? ?/sec
physical_plan_clickbench_q37                          1.00      2.4±0.02ms        ? ?/sec    1.00      2.4±0.03ms        ? ?/sec
physical_plan_clickbench_q38                          1.00      2.4±0.01ms        ? ?/sec    1.01      2.4±0.03ms        ? ?/sec
physical_plan_clickbench_q39                          1.00      2.5±0.02ms        ? ?/sec    1.00      2.5±0.03ms        ? ?/sec
physical_plan_clickbench_q4                           1.00   1421.2±6.48µs        ? ?/sec    1.02  1447.3±36.77µs        ? ?/sec
physical_plan_clickbench_q40                          1.00      3.1±0.03ms        ? ?/sec    1.01      3.2±0.05ms        ? ?/sec
physical_plan_clickbench_q41                          1.00      2.7±0.02ms        ? ?/sec    1.02      2.8±0.04ms        ? ?/sec
physical_plan_clickbench_q42                          1.00      2.8±0.02ms        ? ?/sec    1.01      2.8±0.04ms        ? ?/sec
physical_plan_clickbench_q43                          1.00      3.0±0.03ms        ? ?/sec    1.02      3.1±0.05ms        ? ?/sec
physical_plan_clickbench_q44                          1.00   1480.7±6.95µs        ? ?/sec    1.02  1510.7±16.14µs        ? ?/sec
physical_plan_clickbench_q45                          1.00   1486.9±5.40µs        ? ?/sec    1.00   1492.5±7.23µs        ? ?/sec
physical_plan_clickbench_q46                          1.01  1742.2±10.68µs        ? ?/sec    1.00  1731.5±10.65µs        ? ?/sec
physical_plan_clickbench_q47                          1.03      2.3±0.02ms        ? ?/sec    1.00      2.3±0.02ms        ? ?/sec
physical_plan_clickbench_q48                          1.02      2.6±0.03ms        ? ?/sec    1.00      2.5±0.02ms        ? ?/sec
physical_plan_clickbench_q49                          1.02      2.6±0.03ms        ? ?/sec    1.00      2.6±0.03ms        ? ?/sec
physical_plan_clickbench_q5                           1.00   1547.5±8.13µs        ? ?/sec    1.00  1553.6±33.80µs        ? ?/sec
physical_plan_clickbench_q50                          1.00      2.6±0.03ms        ? ?/sec    1.00      2.6±0.03ms        ? ?/sec
physical_plan_clickbench_q51                          1.01   1870.7±9.11µs        ? ?/sec    1.00  1856.2±12.47µs        ? ?/sec
physical_plan_clickbench_q52                          1.01      2.4±0.02ms        ? ?/sec    1.00      2.4±0.03ms        ? ?/sec
physical_plan_clickbench_q53                          1.02  1733.4±12.23µs        ? ?/sec    1.00  1695.3±20.52µs        ? ?/sec
physical_plan_clickbench_q54                          1.00  1707.1±10.50µs        ? ?/sec    1.01  1717.0±14.67µs        ? ?/sec
physical_plan_clickbench_q55                          1.00   1623.0±8.95µs        ? ?/sec    1.01  1647.2±18.43µs        ? ?/sec
physical_plan_clickbench_q56                          1.00  1635.1±10.40µs        ? ?/sec    1.01  1656.2±15.90µs        ? ?/sec
physical_plan_clickbench_q57                          1.00   1779.3±7.81µs        ? ?/sec    1.00  1776.5±15.06µs        ? ?/sec
physical_plan_clickbench_q58                          1.00      2.2±0.02ms        ? ?/sec    1.03      2.3±0.03ms        ? ?/sec
physical_plan_clickbench_q59                          1.01   1680.7±8.23µs        ? ?/sec    1.00  1668.6±12.54µs        ? ?/sec
physical_plan_clickbench_q6                           1.00   1537.0±7.88µs        ? ?/sec    1.03  1588.6±65.41µs        ? ?/sec
physical_plan_clickbench_q60                          1.00   1640.7±6.92µs        ? ?/sec    1.01  1663.2±11.25µs        ? ?/sec
physical_plan_clickbench_q7                           1.00   1372.2±6.95µs        ? ?/sec    1.03  1407.8±21.04µs        ? ?/sec
physical_plan_clickbench_q8                           1.02  1944.8±15.79µs        ? ?/sec    1.00  1898.4±12.96µs        ? ?/sec
physical_plan_clickbench_q9                           1.00   1792.9±9.39µs        ? ?/sec    1.02  1834.3±17.10µs        ? ?/sec
physical_plan_struct_join_agg_sort                    1.02   1197.4±2.73µs        ? ?/sec    1.00   1179.0±3.97µs        ? ?/sec
physical_plan_tpcds_all                               1.02    649.5±2.66ms        ? ?/sec    1.00    636.3±5.33ms        ? ?/sec
physical_plan_tpch_all                                1.02     43.5±0.22ms        ? ?/sec    1.00     42.6±0.83ms        ? ?/sec
physical_plan_tpch_q1                                 1.01   1432.3±4.71µs        ? ?/sec    1.00   1411.3±8.73µs        ? ?/sec
physical_plan_tpch_q10                                1.01      2.3±0.01ms        ? ?/sec    1.00      2.3±0.04ms        ? ?/sec
physical_plan_tpch_q11                                1.00      2.2±0.01ms        ? ?/sec    1.00      2.2±0.04ms        ? ?/sec
physical_plan_tpch_q12                                1.00   1236.5±3.69µs        ? ?/sec    1.00   1230.5±8.81µs        ? ?/sec
physical_plan_tpch_q13                                1.00   1003.9±3.30µs        ? ?/sec    1.00   1007.2±4.78µs        ? ?/sec
physical_plan_tpch_q14                                1.03   1333.2±4.17µs        ? ?/sec    1.00  1296.4±16.78µs        ? ?/sec
physical_plan_tpch_q16                                1.00   1598.3±5.90µs        ? ?/sec    1.02  1623.0±19.86µs        ? ?/sec
physical_plan_tpch_q17                                1.01   1621.1±6.78µs        ? ?/sec    1.00  1612.3±18.58µs        ? ?/sec
physical_plan_tpch_q18                                1.00   1888.5±5.74µs        ? ?/sec    1.01  1916.3±33.13µs        ? ?/sec
physical_plan_tpch_q19                                1.01   1831.9±6.23µs        ? ?/sec    1.00  1805.2±17.73µs        ? ?/sec
physical_plan_tpch_q2                                 1.00      3.6±0.03ms        ? ?/sec    1.01      3.6±0.10ms        ? ?/sec
physical_plan_tpch_q20                                1.04      2.1±0.01ms        ? ?/sec    1.00      2.0±0.04ms        ? ?/sec
physical_plan_tpch_q21                                1.05      2.7±0.02ms        ? ?/sec    1.00      2.6±0.06ms        ? ?/sec
physical_plan_tpch_q22                                1.04   1545.2±7.87µs        ? ?/sec    1.00  1486.9±16.51µs        ? ?/sec
physical_plan_tpch_q3                                 1.02   1742.8±5.04µs        ? ?/sec    1.00  1706.4±23.35µs        ? ?/sec
physical_plan_tpch_q4                                 1.01   1087.0±3.55µs        ? ?/sec    1.00   1072.6±9.66µs        ? ?/sec
physical_plan_tpch_q5                                 1.02      2.4±0.01ms        ? ?/sec    1.00      2.4±0.06ms        ? ?/sec
physical_plan_tpch_q6                                 1.01    601.1±1.57µs        ? ?/sec    1.00    595.7±2.35µs        ? ?/sec
physical_plan_tpch_q7                                 1.03      2.7±0.02ms        ? ?/sec    1.00      2.6±0.06ms        ? ?/sec
physical_plan_tpch_q8                                 1.05      3.7±0.03ms        ? ?/sec    1.00      3.5±0.08ms        ? ?/sec
physical_plan_tpch_q9                                 1.01      2.7±0.01ms        ? ?/sec    1.00      2.7±0.08ms        ? ?/sec
physical_select_aggregates_from_200                   1.00      7.5±0.02ms        ? ?/sec    1.01      7.6±0.03ms        ? ?/sec
physical_select_all_from_1000                         1.00     17.9±0.06ms        ? ?/sec    1.01     18.1±0.04ms        ? ?/sec
physical_select_one_from_700                          1.04    649.2±2.38µs        ? ?/sec    1.00    625.4±1.89µs        ? ?/sec
physical_sorted_union_order_by_10_int64               1.00      3.6±0.01ms        ? ?/sec    1.00      3.6±0.02ms        ? ?/sec
physical_sorted_union_order_by_10_uint64              1.00      6.9±0.05ms        ? ?/sec    1.01      7.0±0.07ms        ? ?/sec
physical_sorted_union_order_by_50_int64               1.00     80.3±0.22ms        ? ?/sec    1.01     81.2±0.22ms        ? ?/sec
physical_sorted_union_order_by_50_uint64              1.00    265.7±0.93ms        ? ?/sec    1.01    267.1±1.46ms        ? ?/sec
physical_theta_join_consider_sort                     1.00    995.0±1.93µs        ? ?/sec    1.00    997.9±2.35µs        ? ?/sec
physical_unnest_to_join                               1.02    553.1±4.47µs        ? ?/sec    1.00    543.7±2.24µs        ? ?/sec
physical_window_function_partition_by_12_on_values    1.01    641.0±4.22µs        ? ?/sec    1.00    632.8±1.70µs        ? ?/sec
physical_window_function_partition_by_30_on_values    1.00   1259.5±7.56µs        ? ?/sec    1.00   1254.1±2.50µs        ? ?/sec
physical_window_function_partition_by_4_on_values     1.00    396.3±0.93µs        ? ?/sec    1.00    396.0±1.17µs        ? ?/sec
physical_window_function_partition_by_7_on_values     1.01    483.2±3.41µs        ? ?/sec    1.00    477.9±1.46µs        ? ?/sec
physical_window_function_partition_by_8_on_values     1.01    516.7±2.66µs        ? ?/sec    1.00    509.8±1.25µs        ? ?/sec
with_param_values_many_columns                        1.00    374.8±2.02µs        ? ?/sec    1.01    379.8±2.37µs        ? ?/sec

Resource Usage

sql_planner — base (merge-base)

Metric Value
Wall time 2065.5s
Peak memory 141.7 MiB
Avg memory 90.8 MiB
CPU user 1992.3s
CPU sys 1.6s
Peak spill 0 B

sql_planner — branch

Metric Value
Wall time 2085.5s
Peak memory 140.4 MiB
Avg memory 89.7 MiB
CPU user 1992.0s
CPU sys 1.6s
Peak spill 0 B

File an issue against this benchmark runner

@asolimando

Copy link
Copy Markdown
Member Author

Thanks @zhuqi-lucas for re-running the benchmark! Here is how I read the second run:

  • physical_plan_tpcds_all: 649.5 ms -> 636.3 ms (-2.0%)
  • physical_plan_tpch_all: 43.5 ms -> 42.6 ms (-2.1%), and most TPC-H queries are faster (up to -5%)
  • ClickBench (one table, no joins) does not change

optimizer_select_all_from_1000 is 5% slower, but it only does logical planning, which this PR does not change.

On TPC-DS we save less time than in the first run (13 ms instead of 25 ms). I checked the commits merged to main in between and I don't see anything that would change how statistics are computed, and on TPC-H we save the same time as before (0.9 ms vs 1.0 ms). As optimizer_select_all_from_1000 also moves by 5% without being related, I think this is just normal variation between runs (the TPC-DS error bar for the branch is within 5.3 ms drift in this run, which seems like a lot).

WDYT?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

auto detected api change Auto detected API change core Core DataFusion crate documentation Improvements or additions to documentation optimizer Optimizer rules physical-plan Changes to the physical-plan crate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants