Skip to content

feat: optimize count star materialization - #25849

Open
wudidapaopao wants to merge 5 commits into
apache:mainfrom
wudidapaopao:optimize-count-star-materialization
Open

wudidapaopao wants to merge 5 commits into
apache:mainfrom
wudidapaopao:optimize-count-star-materialization

Conversation

@wudidapaopao

@wudidapaopao wudidapaopao commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

Rationale for this change

DataFusion represents COUNT(*) as COUNT(1) and expands the scalar 1 into an array for every input batch, although the accumulator only needs the number of rows.

What changes are included in this PR?

  • Rewrite non-DISTINCT COUNT(*) / COUNT(1) to a nullary COUNT() while preserving output names.
  • Pass input row counts to nullary aggregates in grouped and ungrouped execution.
  • Preserve SingleDistinctToGroupBy rewrites after name preservation adds an alias.
  • Unparse internal COUNT() as SQL COUNT(*).

The nullability-based COUNT(non_nullable_column) -> COUNT(1) rewrite was split into #25990.

What is the testing strategy for this PR?

Covered by unit, SQL logic, DataFrame, unparser, and Substrait tests, as well as formatting, Clippy, and extended workspace test suites.

The release-nonlto benchmark used a 5,000,000-row in-memory Int64 table, one thread, one partition, batch size 8192, and ran this query 100 times:

SELECT COUNT(*) FROM t WHERE v < 900

Median execution time: 1.263 ms → 0.947 ms (25.05% reduction).

Are there any user-facing changes?

Query results and output schemas are unchanged. Plans may display the internal aggregate as count(), while SQL unparsing emits COUNT(*). New accumulator methods have default implementations, so existing UDAFs remain source-compatible.

@github-actions github-actions Bot added sql SQL Planner logical-expr Logical plan and expressions physical-expr Changes to the physical-expr crates optimizer Optimizer rules core Core DataFusion crate sqllogictest SQL Logic Tests (.slt) substrait Changes to the substrait crate functions Changes to functions implementation physical-plan Changes to the physical-plan crate labels Sep 28, 2026
@codecov-commenter

codecov-commenter commented Sep 28, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 84.54545% with 17 lines in your changes missing coverage. Please review.
✅ Project coverage is 82.66%. Comparing base (c3ef346) to head (39ed527).

Files with missing lines Patch % Lines
datafusion/physical-expr-common/src/utils.rs 79.22% 0 Missing and 16 partials ⚠️
...plan/src/aggregates/aggregate_hash_table/common.rs 94.73% 0 Missing and 1 partial ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main   #25849      +/-   ##
==========================================
- Coverage   82.66%   82.66%   -0.01%     
==========================================
  Files        1147     1147              
  Lines      446357   446450      +93     
  Branches   446357   446450      +93     
==========================================
+ Hits       368971   369044      +73     
- Misses      54997    55000       +3     
- Partials    22389    22406      +17     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

Comment thread datafusion/core/tests/sql/unparser.rs Outdated
r#"FROM (SELECT sum("total_revenue") AS "alias2", "#,
r#"date_part('year', "signup_date") AS "group_alias_0", "#,
r#""customer_id" AS "alias1" "#
r#"FROM (SELECT date_part('year', "signup_date") AS "group_alias_0", "#,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm intrigued as to why this has changed, since the ISSUE_23317_QUERY is performing a COUNT(DISTINCT ...).

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SingleDistinctToGroupBy first rewrites COUNT(DISTINCT customer_id) into an inner GROUP BY customer_id and an outer COUNT(alias1). Since customer_id is non-nullable, this PR simplifies the outer COUNT(alias1) to COUNT(), allowing OptimizeProjections to remove alias1 from the inner output.

@mhilton mhilton left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This has buried within it a set of interface enhancements to Accumulator and GroupsAccumulator that add support for functions with no inputs but still have information about the number of rows. I think those changes are important enough to be assessed outside of the context here. Could you make a separate PR for the interface changes so they can be considered aside from the use case please?

}
}

fn retract_batch(&mut self, values: &[ArrayRef]) -> Result<()> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks like this will be broken if it was ever used with the zero-column version.

@wudidapaopao wudidapaopao Sep 29, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right that retract_batch would fail with nullary input. This PR intentionally does not support nullary aggregate windows: create_window_expr rejects empty arguments, and COUNT(*) used as a window aggregate still uses COUNT(1). So retract_batch cannot receive empty values on the normal path, and the PR does not introduce a regression.

@wudidapaopao

wudidapaopao commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor Author

This has buried within it a set of interface enhancements to Accumulator and GroupsAccumulator that add support for functions with no inputs but still have information about the number of rows. I think those changes are important enough to be assessed outside of the context here. Could you make a separate PR for the interface changes so they can be considered aside from the use case please?

That makes sense. I've split the interface changes into #25886.

@neilconway

neilconway commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

@wudidapaopao Thanks for this! I think this is a very useful optimization for some common query shapes.

If I understand correctly, there are two distinct optimizations here:

  1. Rewrite count(non-nullable-col) -> count(1), based on nullability. This seems very useful but I think it would make sense to split it off into a separate PR.
  2. Rewrite count(*) / count(1) -> count(). I agree that there is wasted work in the current evaluation of count(1), but instead of rewriting count(1) to a nullary aggregate (and adding support for nullary aggregates), what if we just optimized the evaluation of count(1) and other aggregates passed constant inputs? This would also benefit other situations, like the common case where string_agg is called with a constant separator, or nth_value(col, const-n). I suspect it would also be simpler than the current PR: no API changes, EXPLAIN churn, unparser handling, etc.

I had Claude Code prototype a quick implementation of caching for const agg args -- based on some quick benchmarks, it matches the performance of this PR for the count(1) case, and also is a small win for other cases like nth_value(col, const-n).

@wudidapaopao wudidapaopao changed the title feat: optimize count over non-null arguments feat: optimize count star materialization Oct 3, 2026
@neilconway

Copy link
Copy Markdown
Contributor

@wudidapaopao What do you think of the alternative approach to implementing this optimization that I suggested?

@github-actions github-actions Bot removed sql SQL Planner logical-expr Logical plan and expressions optimizer Optimizer rules core Core DataFusion crate sqllogictest SQL Logic Tests (.slt) substrait Changes to the substrait crate labels Oct 4, 2026
@github-actions github-actions Bot removed the functions Changes to functions implementation label Oct 4, 2026
…aterialization

# Conflicts:
#	datafusion/physical-plan/src/aggregates/aggregate_stream.rs
@wudidapaopao

Copy link
Copy Markdown
Contributor Author

@neilconway Thanks for the suggestion! I tested the aggregate argument caching approach and found that it achieves nearly identical performance. I’ve updated this PR to use the new approach and moved the COUNT(non-nullable column) → COUNT(1) rewrite into a separate PR.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

physical-expr Changes to the physical-expr crates physical-plan Changes to the physical-plan crate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Optimize COUNT(*) using a nullary aggregate

4 participants