Skip to content

DataFrame API: allow aggregate functions in select() 2 (#17874) - #25482

Open
cj-zhukov wants to merge 2 commits into
apache:mainfrom
cj-zhukov:cj-zhukov/DataFrame-API-allow-aggregate-functions-in-select2
Open

cj-zhukov wants to merge 2 commits into
apache:mainfrom
cj-zhukov:cj-zhukov/DataFrame-API-allow-aggregate-functions-in-select2

Conversation

@cj-zhukov

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

Rationale for this change

DataFrame::select currently does not support aggregate expressions directly. Users must explicitly call DataFrame::aggregate. This PR improves ergonomics by allowing aggregate expressions inside select, while preserving existing behavior and validation rules.

What changes are included in this PR?

  • DataFrame::select lifts a global Aggregate when every expression is an aggregate or a scalar over aggregates (and literals).
  • Existing tests were not rewritten.
  • Examples are left for a follow-up.

What is the testing strategy for this PR?

Are there any user-facing changes?

Yes — additive behavior only. The public select() signature is unchanged. There are no breaking API changes.

Users can write a global aggregation with select():

ctx.table("aggregate_test_100")
    .await?
    .select(vec![
        approx_distinct(col("c9")).alias("count_c9"),
        approx_distinct(cast(col("c9"), DataType::Utf8View))
            .alias("count_c9_str"),
    ])?
    .show()
    .await?;

Previously this required:

ctx.table("aggregate_test_100")
    .await?
    .aggregate(
        vec![],
        vec![
            approx_distinct(col("c9")).alias("count_c9"),
            approx_distinct(cast(col("c9"), DataType::Utf8View))
                .alias("count_c9_str"),
        ],
    )?
    .show()
    .await?;

Both remain valid. Grouped aggregation is still aggregate(group, aggs).

@github-actions github-actions Bot added the core Core DataFusion crate label Sep 18, 2026
@codecov-commenter

codecov-commenter commented Sep 18, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 87.12871% with 13 lines in your changes missing coverage. Please review.
✅ Project coverage is 82.73%. Comparing base (c149764) to head (197d541).
⚠️ Report is 348 commits behind head on main.

Files with missing lines Patch % Lines
datafusion/core/src/dataframe/mod.rs 87.12% 3 Missing and 10 partials ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main   #25482      +/-   ##
==========================================
+ Coverage   81.91%   82.73%   +0.81%     
==========================================
  Files        1134     1147      +13     
  Lines      425703   449552   +23849     
  Branches   425703   449552   +23849     
==========================================
+ Hits       348725   371937   +23212     
+ Misses      56300    54942    -1358     
- Partials    20678    22673    +1995     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@cj-zhukov

Copy link
Copy Markdown
Contributor Author

hey @alamb @Jefffrey @martin-g
could you please take a look at this PR? We were discussing the same issue in #21021 but it was closed because of unwanted breaking changes. I took another step without breaking changes now.

@kosiew kosiew left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@cj-zhukov,

Thanks for working on this. The aggregate support looks useful, but I think the new output-name deduplication changes an existing validation rule, so I'd like that addressed before approval.

Comment thread datafusion/core/src/dataframe/mod.rs Outdated
.aggregate(Vec::<Expr>::new(), aggr_with_alias)?
.build()?;
expr_list = rewrite_select_aggs(expr_list, &input, &rewrite_map)?;
expr_list = uniquify_select_expr_names(expr_list)?;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This now silently renames duplicate explicit aliases before the existing projection validator can reject them, so select behaves differently from ordinary projections and explicit aggregate. Please remove the public output-name uniquification and let the existing projection validation reject duplicate names, while keeping unique internal aggregate aliases where needed.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree with you

}

#[tokio::test]
async fn test_dataframe_api_select_semantics() -> Result<()> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you also add a small empty-input case using filter(lit(false)).select(...) with distinct aliases? It would be useful to verify that global aggregation still returns exactly one row, with COUNT = 0, SUM = NULL, and the literal value preserved.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed. Let me add new tests.

@cj-zhukov

Copy link
Copy Markdown
Contributor Author

@kosiew Thank you for the valuable review!

I've removed the public output-name uniquification and added empty-input test case. Could you please take another look when you have a chance?

@kosiew kosiew left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@cj-zhukov,

Thanks for addressing the review comments. The follow-up restores the existing duplicate projection name validation while keeping internal aggregate aliases unique. The new empty-input test also verifies the expected global aggregation behavior.

I've reviewed the changes and have no further concerns. LGTM!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core Core DataFusion crate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

DataFrame API: allow aggregate functions in select()

3 participants