Repository navigation
Consolidate built-in statistics estimation in one place #25571
Description
Activity
cc @asolimando, got bitten by this while making some stats improvement reports. Digging in the PR history, I did not find a good reason for us to maintain diverging implementations for statistics in the built-in StatisticsProviders VS the stats provided in the ExecutionPlans themselves, but maybe I missed something?
cc @asolimando, got bitten by this while making some stats improvement reports. Digging in the PR history, I did not find a good reason for us to maintain diverging implementations for statistics in the built-in StatisticsProviders VS the stats provided in the ExecutionPlans themselves, but maybe I missed something?
In the original design for #21443 proposed the
StatisticsRegistryoutside the built-in path to avoid breaking changes, over time we agreed to have it in the default path (usingExecutionPlan's defaults in absence of providers) in #23651, so "built-in providers" referred to optional theStatisticsRegistryitself, not for statistics propagation, as there was no fallback mechanism in theStatisticsRegistryto the statistics computation fromExecutionPlannodes.After #23651, the name should have been changed as it's indeed confusing, apologies for having missed that.
I see two possibilities for the future of these providers (more or less along the lines of what you suggested above):
- trimming them down as much as possible (there are some overlaps with
ExecutionPlanbuilt-ins for historical reasons), and use them to provide "advanced"/alternative statistics for some of the nodes (e.g.,which is based on the https://en.wikipedia.org/wiki/Coupon_collector%27s_problem)// Adjust distinct_count for each column using the selectivity ratio - remove them and leave custom providers (user-defined) and
ExecutionPlanbuilt-ins
Now that stats-bench has landed, it's probably easier to integrate improvements into
ExecutionPlanbuilt-ins (and 1. is less needed), as we have a way to show that statistics estimation is getting closer to the runtime statistics, while before it was harder to get such improvements in (see #21120 (comment) for instance).- trimming them down as much as possible (there are some overlaps with
Yeah, that's my intuition as well, just integrating the advanced stats calculation into
ExecutionPlanbuilt-ins sounds like a good option 👍Reacted by Alessandro Solimandotake
- added a commit that references this issue
on Oct 6, 2026
Is your feature request related to a problem or challenge?
DataFusion has statistics estimation inside individual operators and separate implementations in the bundled
StatisticsProviders. Enabling those providers can replace an operator's estimate with a different calculation for the same operation.This surfaced in #25570: for TPC-H SF1 Q14, the bundled providers produced a join estimate of 14.73 billion rows instead of 73,650, against 75,983 actual rows. Calling the optional set “built-in providers” also makes it hard to tell which implementation is used by default.
Describe the solution you'd like
Could we consolidate built-in statistics estimation so each operator has one authoritative implementation? The registry makes sense as an extension point for custom estimators, but DataFusion's own providers should share the operator's estimation logic rather than maintain a separate algorithm.
This would make it clearer where to fix estimation bugs and avoid improvements reaching only one implementation.
Describe alternatives you've considered
Renaming or documenting the optional providers would clarify their behavior, but would leave the duplicated estimation logic in place.
Additional context
#25570 aligns the CLI with the session's configuration. This issue is about consolidating the underlying estimators.