Repository navigation
fix: prevent panic and incorrect results for COUNT with ORDER BY - #24997
Merged
Merged
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #24997 +/- ##
==========================================
+ Coverage 81.67% 81.74% +0.06%
==========================================
Files 1126 1128 +2
Lines 414842 416648 +1806
Branches 414842 416648 +1806
==========================================
+ Hits 338841 340578 +1737
+ Misses 56070 56001 -69
- Partials 19931 20069 +138 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
geoffreyclaude
force-pushed
the
fix/count-order-by
branch
from
September 6, 2026 20:13
49809cd to
0c5548e
Compare
This was referenced Sep 8, 2026
geoffreyclaude
force-pushed
the
fix/count-order-by
branch
from
September 8, 2026 09:20
befcf09 to
af53d29
Compare
geoffreyclaude
marked this pull request as ready for review
September 8, 2026 09:22
This was referenced Sep 26, 2026
lukekim
pushed a commit
to spiceai/datafusion
that referenced
this pull request
Sep 27, 2026
…che#24997) ## Which issue does this PR close? - Closes apache#25055. ## Rationale for this change `COUNT` does not depend on input order, but it inherited the default `AggregateOrderSensitivity::HardRequirement`. As detailed in apache#25055, this could pass ordering keys to count accumulators as additional arguments, causing incorrect results or a panic, and could introduce unnecessary sorting. ## What changes are included in this PR? Override `Count::order_sensitivity` to return `AggregateOrderSensitivity::Insensitive`. This makes aggregate ordering keys ineffective for physical planning: they are excluded from accumulator inputs and no `SortExec` is required solely for `COUNT`'s `ORDER BY`. ## What is the testing strategy for this PR? Regression tests in `aggregate.slt` cover grouped, multi-argument, and non-grouped counts. The data includes both a null ordering key, which must not affect the count, and a null counted argument, which must still be excluded. A bare global `COUNT(a ORDER BY b)` over the in-memory test table is folded by `AggregateStatistics` to `num_rows - null_count(a)`, so `CountAccumulator` never runs. The test uses `a + 0`, which preserves `a`'s nullness but prevents that fold, ensuring the regression exercises the accumulator. An `EXPLAIN` assertion verifies that the count does not require a sort. ## Are there any user-facing changes? `COUNT(... ORDER BY ...)` returns the correct count without panicking. Nulls in counted arguments retain their usual behavior; nulls in ordering keys no longer incorrectly exclude rows. (cherry picked from commit 224cc56) Conflict resolved when applying to spiceai-54 (DataFusion 54.1): - datafusion/sqllogictest/test_files/aggregate.slt: upstream appends this commit's COUNT ... ORDER BY block after other blocks that are not on 54.1, including the nested-aggregate block. Added only this commit's 43 lines, at the end of the 54.1 file.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Which issue does this PR close?
Rationale for this change
COUNTdoes not depend on input order, but it inherited the defaultAggregateOrderSensitivity::HardRequirement. As detailed in #25055, this could pass ordering keys to count accumulators as additional arguments, causing incorrect results or a panic, and could introduce unnecessary sorting.What changes are included in this PR?
Override
Count::order_sensitivityto returnAggregateOrderSensitivity::Insensitive. This makes aggregate ordering keys ineffective for physical planning: they are excluded from accumulator inputs and noSortExecis required solely forCOUNT'sORDER BY.What is the testing strategy for this PR?
Regression tests in
aggregate.sltcover grouped, multi-argument, and non-grouped counts. The data includes both a null ordering key, which must not affect the count, and a null counted argument, which must still be excluded.A bare global
COUNT(a ORDER BY b)over the in-memory test table is folded byAggregateStatisticstonum_rows - null_count(a), soCountAccumulatornever runs. The test usesa + 0, which preservesa's nullness but prevents that fold, ensuring the regression exercises the accumulator.An
EXPLAINassertion verifies that the count does not require a sort.Are there any user-facing changes?
COUNT(... ORDER BY ...)returns the correct count without panicking. Nulls in counted arguments retain their usual behavior; nulls in ordering keys no longer incorrectly exclude rows.