Skip to content

perf: PartitionedTopKRank on the shared store with decide-then-gather - #25968

Open
SubhamSinghal wants to merge 4 commits into
apache:mainfrom
SubhamSinghal:partitioned-topk-rank-shared-store
Open

SubhamSinghal wants to merge 4 commits into
apache:mainfrom
SubhamSinghal:partitioned-topk-rank-shared-store

Conversation

@SubhamSinghal

@SubhamSinghal SubhamSinghal commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

Part of #6899. Follows #25730 (merged), which applied this design to ROW_NUMBER. This PR brings RANK to the same design.

Rationale for this change

With enable_window_topn, RANK() ... WHERE rk <= K on high-cardinality PARTITION BY is slower than the sort plan it replaces: at 100K partitions the rewrite costs 2.76× the sort plan's CPU. PartitionedTopKRank::insert_batch:

  • gathers one take_record_batch sub-batch per (input batch, partition) pair, roughly one small RecordBatch per input row at 100K partitions;
  • does a 1-row take_record_batch for every row evicted into the tie list;
  • runs a full TopKHeap, with its own RecordBatchStore, per partition;
  • sums size() over every partition on every batch.

Profiling main at 1M partitions put about 58% of samples in creating and destroying those small batches. Payload width made almost no difference, so the cost is the number of calls, not the bytes copied. #25730 removed exactly this for ROW_NUMBER. This PR does the same for RANK, plus boundary ties.

What changes are included in this PR?

In datafusion/physical-plan/src/topk/mod.rs, plus a doc update in sorts/partitioned_topk.rs:

  • PartitionedTopKRank uses perf: PartitionedTopK hold retained rows in one shared RecordBatchStore #25730's design. It encodes the partition and ORDER BY columns once per batch, keeps or drops each row in its partition's PartitionHeap in one pass, then gathers the admitted rows once per input batch into an operator-wide RecordBatchStore. Output is interleaved out of that store in batch_size chunks by EmitState, so the BatchCoalescer goes away.
  • Ties are 8-byte StoreRefs (batch_id, row) in a per-partition ties list. Every tie shares the heap root's key, so none is stored.
    • Evicted but still tied: the row's reference and its store use move from the heap to the tie list. No 1-row take_record_batch.
    • Cutoff improves: the evicted row and all ties are released, one unuse each, so amortised O(1) per admitted row.
  • Memory stays proportional to rows retained (the property fix: reduce RANK window top-K memory from O(input) to O(K + ties) per partition #24591 established). The store holds only admitted rows, and the same ratio-2 compaction as perf: PartitionedTopK hold retained rows in one shared RecordBatchStore #25730 keeps store.total_rows ≤ 2 × live_slots after every batch, where live_slots counts heap rows and ties. Ties from every partition a batch touches share that batch's one store entry, so they are charged once, not once per partition (the over-count in PartitionedTopKRank over-accounts memory when a single batch has boundary ties across many partitions #23326).
  • size() is O(1), from running totals plus store.size().
  • Shared with ROW_NUMBER. Compaction (RecordBatchStore::compact, replacing PartitionedTopK::compact_store), the in-flight batch's bookkeeping (pending / release / insert_rows) and stream construction (EmitState::stream) are now used by both operators. No behaviour change for ROW_NUMBER.
  • EvictedRow is removed. RANK was its last reader. TopKHeap::add no longer looks up and clones the evicted row's batch.

Are these changes tested?

Yes. All existing PartitionedTopKRank tests pass, including the three memory-bound tests from #24591. New tests:

  • test_partitioned_topk_rank_bookkeeping_tracks_recompute: 256 randomized shapes, checking after every batch that each store entry's uses matches the rows pointing into it, total_rows ≤ 2 × live_slots, the running totals and reservation match a recompute, and the output matches brute force in (pk, val) order. It replaces test_partitioned_topk_rank_matches_bruteforce.
  • test_partitioned_topk_rank_store_bounded_when_ties_spread_thinly, test_partitioned_topk_rank_boundary_move_releases_ties_across_batches, and test_partitioned_topk_rank_ties_share_one_store_entry (the PartitionedTopKRank over-accounts memory when a single batch has boundary ties across many partitions #23326 shape).

Benchmarks

benchmarks/queries/h2o/window.sql RANK queries Q18–Q23 (h2o J1 large, 10M rows, K=2). User CPU, median of 7 alternating runs, 14 cores. base = #25730's tip, whose RANK code is main's; sort plan = flag off.

partitions main, ON vs OFF this branch, ON vs OFF
100 (Q18) 5.87× faster 6.07× faster
1,000 (Q19) 5.31× faster 6.00× faster
1,000 (Q20) 4.78× faster 5.24× faster
10,000 (Q21) 2.02× faster 5.44× faster
10,000 (Q22) 2.19× faster 4.74× faster
100,000 (Q23) 2.76× slower 1.83× faster

Are there any user-facing changes?

No. enable_window_topn stays default-false,

@github-actions github-actions Bot added the physical-plan Changes to the physical-plan crate label Oct 2, 2026
@codecov-commenter

codecov-commenter commented Oct 2, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 82.67%. Comparing base (4d167a1) to head (a499fc3).
⚠️ Report is 36 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main   #25968      +/-   ##
==========================================
+ Coverage   82.60%   82.67%   +0.07%     
==========================================
  Files        1145     1147       +2     
  Lines      444057   446767    +2710     
  Branches   444057   446767    +2710     
==========================================
+ Hits       366796   369348    +2552     
+ Misses      54996    54995       -1     
- Partials    22265    22424     +159     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@SubhamSinghal

Copy link
Copy Markdown
Contributor Author

@kosiew @jayzhan211 @kumarUjjawal can you help in reviewing this PR?

@kumarUjjawal kumarUjjawal left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you @SubhamSinghal

Left few comments please take a look.

Comment on lines +2322 to +2323
released_slots += state.ties.len();
state.ties.clear();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Release tie-list capacity when the boundary improves

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in f9d7919

Comment on lines +2393 to +2398
EmitState::stream(
schema,
futures::stream::iter(out),
)))
metrics,
reservation,
batch_size,
&store,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Limit interleave inputs to batches referenced by each output chunk

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in f9d7919

match state.heap.classify(k, key) {
// Strictly worse than the boundary: drop the row.
Some(Ordering::Greater) => continue,
Some(Ordering::Equal) => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In PartitionedTopKRank::insert_batch, rows admitted as ties (the Equal arm) do not increment replacements, but every other admission does, and every admission does in PartitionedTopK, so row_replacements undercounts rows the operator retained.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in f9d7919

// no two slots share a store row, so the key identifies exactly one
// slot.
let mut coords: Vec<(usize, usize)> = Vec::with_capacity(live_slots);
let mut moved: HashMap<StoreRef, StoreRef> =

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RecordBatchStore::compact builds a HashMap<StoreRef, StoreRef> with live_slots entries just to repoint slots, although partitions.values()/values_mut() and store_rows()/repoint() walk the slots in the same order, so a running counter would give each slot its new place.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in f9d7919

// compacted ones are built, drop before the store is rewritten.
let first_id = self.next_batch_id();
let (moved, chunks) = {
let (old, array_pos) = self.positional();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

compact calls positional(), which clones every stored RecordBatch (an Arc bump per column) and builds a fresh HashMap, but compact only needs borrowed &RecordBatch refs that it drops before the store is rewritten.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in f9d7919

Some(Ordering::Less) => {
// Replacing the root overwrites its key in place, so
// keep a copy to tell whether the boundary moved.
evicted_key.clear();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Every strictly-better eviction in RANK memcpys the old root key into evicted_key only to compare it with the new root. PartitionHeap::add could hand the old key back by swapping Vecs instead.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in f9d7919


// 1. Evaluate + encode partition columns into the reusable
// scratch (cleared then appended).
// 1. Evaluate the partition and ORDER BY columns and encode each once

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PartitionedTopKRank still duplicates PartitionedTopK almost line for line: the same ~14 fields, try_new, the phase-1 evaluate/encode block, the phase-3 insert_rows/compact/try_resize tail, the emit destructuring and size(). Only the per-row admission rule differs.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in 2901db2

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

// Collect the batches into a vec and store the "batch_id -> array_pos" mapping, to then

PR adds RecordBatchStore::positional to build the batches/batch_id -> position pair, but TopKHeap::emit_with_state still builds the same mapping inline.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in f9d7919

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

/// RANK-specific: heap fills with K rows tied at value V, equal_indices

Comments left stale by the rewrite: the doc and step comments of test_partitioned_topk_rank_boundary_shifts_clears_ties still describe the removed in-flight equal_indices list ('must clear both state.ties and the in-flight equal_indices'), and assert_store_matches_slots (line 4948) still names compact_store, which this PR replaced with RecordBatchStore::compact.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in f9d7919

@jayzhan211 jayzhan211 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @SubhamSinghal! Nice work — good to go from my side.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

physical-plan Changes to the physical-plan crate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants