Repository navigation
Conversation
|
Thank you for opening this pull request! Reviewer note: cargo-semver-checks reported the current version number is not SemVer-compatible with the changes in this pull request (compared against the base branch). Details |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #25358 +/- ##
==========================================
+ Coverage 81.92% 82.61% +0.69%
==========================================
Files 1135 1147 +12
Lines 427772 445654 +17882
Branches 427772 445654 +17882
==========================================
+ Hits 350456 368190 +17734
+ Misses 56373 55106 -1267
- Partials 20943 22358 +1415 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
@yashrb24 |
| } | ||
| arrow::compute::take_record_batch(batch, &UInt32Array::from(selection))? | ||
| }; | ||
| let sort_columns = keys |
There was a problem hiding this comment.
Staged sorting can skip later fallible or volatile ORDER BY expressions when earlier keys already resolve the order, so queries can behave differently depending on tie distribution. Please either document this experimental/error/volatile-evaluation contract and add tests that pin it, including that disabled staging stays eager, or preserve eager evaluation.
| } | ||
| } | ||
|
|
||
| fn staged_sort_batch_chunked( |
There was a problem hiding this comment.
The new staged sorting path does not have focused behavioral tests. Please add coverage for staged versus eager equivalence, successive tie groups, null ordering, chunked output, the zero-tie/error case, and both single-batch and multiple in-memory-run paths.
Which issue does this PR close?
Related to #25033. This PR is mostly an ideation of the way we can bring in this change since it's kind of an invasive change to the sorting mechanics and thus wanted feedback.
Rationale for this change
Today, sorting evaluates all
ORDER BYexpressions before it compares rows. Some of this work is unnecessary when earlier keys already determine the order.This change evaluates later expressions only for rows tied on earlier keys. This could help when later expressions are expensive and earlier keys resolve most rows. The drawback is that the extra sorting steps can outweigh these savings when expressions are cheap or many rows remain tied.
What changes are included in this POC PR?
Are there any user-facing changes?
had added these flags for my local testing, can remove if needed
datafusion.execution.enable_staged_sort: defaultfalse.datafusion.execution.sort_key_group_size: default value is5Currently not extending to TopK or spill codepath