Repository navigation
fix: leave memory for aggregate spill replay - #25383
Conversation
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #25383 +/- ##
==========================================
+ Coverage 81.92% 82.02% +0.09%
==========================================
Files 1135 1136 +1
Lines 427772 439539 +11767
Branches 427772 439539 +11767
==========================================
+ Hits 350456 360522 +10066
- Misses 56373 58242 +1869
+ Partials 20943 20775 -168 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
There was a problem hiding this comment.
Thanks @sunchao! One suggestion below, non-blocking:
check_headroom is enforced on every selection, but only the pass that drains sorted_spill_files feeds replay; intermediate passes spill back to disk and don't need headroom. With many runs this halves fan-in on every pass, roughly doubling the number of intermediate re-spills and peak temp-disk usage. Fine to handle in a follow-up: "Skip replay headroom for intermediate multi-level merge passes".
Sketch: let the caller control headroom per call and cap the number of admitted files. Probe with headroom; if that doesn't admit every remaining run, re-select without headroom while holding one run back so the following pass is still the final one (otherwise a full-pool selection could become the final pass with no headroom).
fn get_sorted_spill_files_to_merge(
&mut self,
buffer_len: usize,
minimum_number_of_required_streams: usize,
reservation: &mut MemoryReservation,
allow_minimum_without_headroom: bool,
+ headroom: bool,
+ max_files: Option<usize>,
) -> Result<SpillFilesToMerge> {
- let max_spill_files = effective_spill_merge_fan_in(configured_fan_in);
+ let max_spill_files = effective_spill_merge_fan_in(configured_fan_in)
+ .min(max_files.unwrap_or(usize::MAX));
...
- let check_headroom = self.reserve_replay_headroom && !skip_headroom;
+ let check_headroom = headroom && !skip_headroom;
...
- if self.reserve_replay_headroom {
+ if headroom {
reservation.shrink(reservation.size() - accepted_memory);
}The recursive buffer_len - 1 call passes headroom and max_files through unchanged. Then in merge_sorted_runs_within_mem_limit:
let total = self.sorted_spill_files.len();
let selection = self.get_sorted_spill_files_to_merge(
2,
minimum_number_of_required_streams,
&mut memory_reservation,
allow_minimum_without_headroom,
self.reserve_replay_headroom,
None,
)?;
let selection = match selection {
SpillFilesToMerge::Ready(spills, _)
if self.reserve_replay_headroom && spills.len() < total =>
{
// Intermediate pass: no replay follows, so use the full pool but
// keep one run back so the next pass remains the final one.
self.sorted_spill_files.splice(0..0, spills);
memory_reservation.free();
self.get_sorted_spill_files_to_merge(
2,
minimum_number_of_required_streams,
&mut memory_reservation,
allow_minimum_without_headroom,
false,
Some(total - 1),
)?
}
other => other,
};|
Thanks @jayzhan211 for the stamp! Agreed. Intermediate passes don’t feed aggregate replay, so reserving headroom there is unnecessarily conservative. I’ll handle this in a separate follow-up, ensuring any pass that feeds replay still reserves headroom. I’ll also cover the split/retry path and measure merge passes and spill I/O to quantify the benefit. |
## Which issue does this PR close? Prerequisite for apache#25172, which fixes `FairSpillPool` accounting across reservations belonging to the same memory consumer. This PR prepares aggregate spill replay to work within that shared allowance and should merge first. ## Rationale for this change **Spilling an aggregation to disk only helps if there is enough memory to read it back and finish the aggregation.** During spill replay, DataFusion merges sorted spill files and feeds their rows into an aggregate that combines the intermediate states into final results. The merge needs buffers for its inputs and output, while the aggregate needs memory for group keys and accumulators. Those allocations coexist: producing the next batch and processing it are both part of the same operator's memory budget. For example, consider: ```sql SELECT customer_id, ARRAY_AGG(event_id) FROM events GROUP BY customer_id; ``` With many customers and a small memory limit, the aggregate can spill several times. A customer's events may then be spread across several spill files. Merging those files brings the customer's intermediate states together, but the aggregate still has to build the final array. Even if the merge delivers one row at a time, that array grows until the customer's group is complete. Small input batches therefore do not eliminate the need for replay memory. The current merge tries to merge as many spill files as its reservation permits, without leaving room for this downstream work. As a simplified example, suppose the operator has a 100 MiB allowance and no other reservations. An 80 MiB merge fits by itself, but if the aggregate then needs 24 MiB to process its output, their combined 104 MiB does not fit. A pool enforcing the shared allowance must reject that growth, even though using a smaller merge could let the same aggregation make progress. This is particularly relevant to apache#25172. The merge and aggregate use separate reservations under the same memory consumer. Today, `FairSpillPool` checks those reservations individually, which can hide their combined usage exceeding the consumer's allowance. Once apache#25172 enforces that allowance across both reservations, replay must already leave space for the aggregate. This PR addresses that requirement before the accounting fix lands. ## What changes are included in this PR? **Aggregate spill merges now leave room for the aggregate consuming their output.** When choosing how many spill files to merge and how much to buffer, the merge asks the configured memory pool to admit its buffers plus an equal amount of spare capacity for replay. Once admitted, it releases the spare reservation before replay starts, retaining only the reservation for the merge buffers. This uses the pool's existing allocation checks, so the decision follows its admission policy, including any fair-share restrictions, without adding a public memory-pool API. In the example above, the 80 MiB merge would first need approval for 160 MiB and would be rejected. A smaller 40 MiB merge would need approval for 80 MiB; after returning the 40 MiB of temporary headroom, it would leave 60 MiB available, enough for the illustrative 24 MiB of aggregate state. The equal-sized headroom is a practical budgeting policy, not an estimate of the exact accumulator size or a guarantee that every aggregation will fit. The existing multi-level merge adapts by merging fewer files at once, reducing read-ahead, or splitting oversized spill batches into smaller batches. This PR makes those adjustments account for replay as well. It also handles the point where a batch cannot shrink further: a single wide row may still fit the real pool even when equal headroom does not. In that case, only the minimum merge may retry without the extra headroom, with read-ahead disabled and all merge memory still subject to the pool's checks. Inspecting decoded batches before rewriting a spill file avoids unnecessary writes when the rows already fit or cannot be split, including when the original spill files fill the disk quota. All supported aggregate replay implementations use this policy, including hash and ordered aggregation and the legacy implementation selected by `enable_migration_aggregate=false`. The legacy path also releases unused initial grouping capacity before replay. The headroom policy is private to aggregate replay; ordinary sort callers retain their existing admission behavior. ## Are these changes tested? Regression tests exercise replay across the supported aggregate implementations; the migrated hash and ordered paths also run with another registered spilling consumer. In particular, the legacy `ARRAY_AGG` test uses 64 groups with 64 values per group, one-row batches, and an 8 KiB pool to check that accumulator state can grow across replay batches. The tests verify exact aggregate results, reservation bounds, and release of memory and spill files. Merge-level coverage checks oversized batches, short and odd-sized batches, indivisible rows, decoded string-view sizes, and spill files that already fill the disk quota. It also checks that temporary headroom is released after a candidate merge is rejected and that the fallback still fails when the actual merge cannot fit. <details> <summary>Validation results and environment</summary> Validated replay alone, with the existing `FairSpillPool` implementation: - 2,263 physical-plan tests and 2,281 core/CLI tests passed. - The existing permanent-pressure regression passed with its original success expectation. - All 520 SQL logic test files passed. - `cargo fmt --all` and strict all-target/all-feature Clippy passed. - Full `./dev/rust_lint.sh` passed, including private Rust documentation and local Markdown links. Local validation uses Rust 1.98.1 and upstream revision `22651d24` with its unchanged dependency lockfile; newer main's dependency versions are unavailable in the local registry. Merge compatibility with current main is checked separately, and GitHub CI validates its merged revision. </details> ## Are there any user-facing changes? Aggregations that spill leave memory available for processing the merged rows, reducing avoidable replay failures under constrained memory. Achieving this can require smaller batches or additional merge passes. SQL semantics and public APIs are unchanged. The temporary headroom reservations can increase recorded reservation peaks without allocating additional data buffers. That headroom is released before replay, so other concurrent allocations or aggregate state that outgrows the available memory can still cause `ResourcesExhausted`.
## Which issue does this PR close? Follow-up to [the review of apache#25383](apache#25383 (review)), which suggested reducing intermediate aggregate spill work. ## Rationale for this change An intermediate aggregate merge can rewrite rows that could remain on disk until final replay. With eight uniform full-batch runs and an admitted fan-in of seven, merging six inputs leaves three for the final pass. The ordinary selection merges seven and leaves two, unnecessarily rewriting one run. The final pass must still fit if another partition consumes available memory during the intermediate write. The retained seven-input buffer reservation can cover three final inputs plus their replay headroom; it cannot guarantee seven final inputs with that headroom. ## What changes are included in this PR? Eligible spill-only intermediate merges select only enough inputs to leave a final pass whose expected buffers and replay headroom fit the existing buffer reservation. An actually trimmed pass retains that reservation in the merge builder through intermediate EOF and spill completion, transferring it to the next ordinary admission. Sizing acquires no additional pool memory and preserves read-ahead and output batching. Passes that are not trimmed keep ordinary reservation ownership. Only the newer aggregate implementation (`AggregateSpill`, enabled by default) opts in. Legacy aggregation keeps its original selection. Sizing requires equal recorded maximum batch memory and batch-size limits across admitted and pending runs. Every original run must contain a full target-size batch; short original runs disable sizing for their replay. Intermediate and split spill writes also check actual largest-batch row counts against their output limit. These checks use private metadata; public APIs are unchanged. The sizing budget assumes the produced run is no wider than the guarded inputs. Variable-width output can still require further intermediate work. Final admission and its normal headroom release remain unchanged. ## What is the testing strategy for this PR? Head: `66569072850d91476b042010906ac6461ebbbd08`. Comparison base: `991fd23dd0be0046af5945b8c3905612859b657a`. The regression tests cover competing memory consumers at replay-headroom release and intermediate EOF, both beneficial and unchanged merge selections, short batches, and heterogeneous input budgets. They check exact ordered results, intermediate rewrite counts, actual competing reservations, memory and spill-file cleanup, and reservation ownership when the intermediate stream and builder are dropped. CI runs for this head: [push Rust run](https://github.com/apache/datafusion/actions/runs/36461591867) and [pull-request Rust run](https://github.com/apache/datafusion/actions/runs/36461597627). ### Performance validation The updated head has not been rebenchmarked against `991fd23dd0be0046af5945b8c3905612859b657a`. ## Are there any user-facing changes? Eligible spilling aggregations can perform fewer intermediate row writes. Results and configuration remain unchanged. Short runs and selections without sufficient retained replay capacity use the ordinary selection.
… memory share (apache#6544) * fix: let a spilled final aggregate read its spill files back past its memory share DataFusion 55's FinalHashAggregateStream replays its merged spill files through a stream that can't spill, so a refused memory request there fails the task. The merge's read buffers are a sibling reservation of the same consumer and take as many spill files as fit, so the replay often finds the consumer's share already taken (apache#6254). Both Comet pools now record that request as overcommit instead of refusing it. They recognize it as a request from a FinalHashAggregateStream consumer while another of its reservations holds memory, which in DataFusion 55.1 happens only during the replay. The replay emits its finished groups after every batch, so the overcommit stays around one batch of groups, and releases repay it first. Refusals while the aggregate reads its input, and while the merge picks its files, are unchanged. Remove this once Comet's DataFusion includes apache/datafusion#25383. * fix: let an ordered final aggregate read its spill files back too DataFusion runs a final aggregate whose input is sorted on some of its grouping keys as an OrderedFinalAggregateStream. It merges and replays its spill files the same way FinalHashAggregateStream does, so its replay failed the task the same way. Treat its consumer as a final aggregate too. Explain in spill_replay.rs why recording the replay's request is safe: the replay asks for memory only after it has aggregated a batch, so the memory already exists, as it does for a grow. Describe the exception in the memory management guide, which said the fair pool always refuses a request that fails its local checks. * docs: tighten the spill replay comments and guide The fair pool is the only one with a share and a pool total to skip, so say so. The greedy pool takes its tracking lock for other consumers when Spark refuses them, not never. Point whoever removes the workaround at the tests that show whether the replay still needs it. * test: share the spill replay tests between the pools fair_pool.rs and unified_pool.rs each had a copy of the same two spill replay scenarios. Run them once, in spill_replay.rs, against both pool types as createPlan builds them, which also checks that the wrappers pass register through to the greedy pool's tracking. Keep a greedy pool test for its per-consumer map, and drop a test import that the pool's own imports now cover. Share the final aggregate spill checks between the two apache#6254 tests in CometAggregateSuite, and use checkCometAnswer, which collects once and labels the Comet answer as Comet's. * refactor: move the spill replay workaround into a pool wrapper The workaround for apache#6254 lived in both Comet pools. The fair pool sent its three refusals through a check, and the greedy pool tracked each final aggregate across its reservations. apache#5613 reworks the fair pool's try_grow, and apache#6583 has to take the workaround out again, so both would have had to rework the pools. SpillReplayPool, in spill_replay.rs, now wraps the tracked Comet pool instead. It keeps a total for each final aggregate's consumer. When the pool refuses a request from one while another of its reservations holds memory, it calls the pool's grow, which skips the fair pool's limits and carries what Spark doesn't grant as overcommit. fair_pool.rs and unified_pool.rs are back to main's versions, and overcommit() looks through the wrapper. The Rust tests run against both pool types as createPlan builds them. They now reach the fair pool's pool-limit refusal too, and check that a failed Spark call during the replay is returned rather than recorded. The guide adds the wrapper to the pool stack and keeps a short section. The DataFusion 55.1 facts that an upgrade has to re-check stay in spill_replay.rs, which names apache#6583 next to the removal condition.
… memory share (#6544) (#6713) * fix: let a spilled final aggregate read its spill files back past its memory share DataFusion 55's FinalHashAggregateStream replays its merged spill files through a stream that can't spill, so a refused memory request there fails the task. The merge's read buffers are a sibling reservation of the same consumer and take as many spill files as fit, so the replay often finds the consumer's share already taken (#6254). Both Comet pools now record that request as overcommit instead of refusing it. They recognize it as a request from a FinalHashAggregateStream consumer while another of its reservations holds memory, which in DataFusion 55.1 happens only during the replay. The replay emits its finished groups after every batch, so the overcommit stays around one batch of groups, and releases repay it first. Refusals while the aggregate reads its input, and while the merge picks its files, are unchanged. Remove this once Comet's DataFusion includes apache/datafusion#25383. * fix: let an ordered final aggregate read its spill files back too DataFusion runs a final aggregate whose input is sorted on some of its grouping keys as an OrderedFinalAggregateStream. It merges and replays its spill files the same way FinalHashAggregateStream does, so its replay failed the task the same way. Treat its consumer as a final aggregate too. Explain in spill_replay.rs why recording the replay's request is safe: the replay asks for memory only after it has aggregated a batch, so the memory already exists, as it does for a grow. Describe the exception in the memory management guide, which said the fair pool always refuses a request that fails its local checks. * docs: tighten the spill replay comments and guide The fair pool is the only one with a share and a pool total to skip, so say so. The greedy pool takes its tracking lock for other consumers when Spark refuses them, not never. Point whoever removes the workaround at the tests that show whether the replay still needs it. * test: share the spill replay tests between the pools fair_pool.rs and unified_pool.rs each had a copy of the same two spill replay scenarios. Run them once, in spill_replay.rs, against both pool types as createPlan builds them, which also checks that the wrappers pass register through to the greedy pool's tracking. Keep a greedy pool test for its per-consumer map, and drop a test import that the pool's own imports now cover. Share the final aggregate spill checks between the two #6254 tests in CometAggregateSuite, and use checkCometAnswer, which collects once and labels the Comet answer as Comet's. * refactor: move the spill replay workaround into a pool wrapper The workaround for #6254 lived in both Comet pools. The fair pool sent its three refusals through a check, and the greedy pool tracked each final aggregate across its reservations. #5613 reworks the fair pool's try_grow, and #6583 has to take the workaround out again, so both would have had to rework the pools. SpillReplayPool, in spill_replay.rs, now wraps the tracked Comet pool instead. It keeps a total for each final aggregate's consumer. When the pool refuses a request from one while another of its reservations holds memory, it calls the pool's grow, which skips the fair pool's limits and carries what Spark doesn't grant as overcommit. fair_pool.rs and unified_pool.rs are back to main's versions, and overcommit() looks through the wrapper. The Rust tests run against both pool types as createPlan builds them. They now reach the fair pool's pool-limit refusal too, and check that a failed Spark call during the replay is returned rather than recorded. The guide adds the wrapper to the pool stack and keeps a short section. The DataFusion 55.1 facts that an upgrade has to re-check stay in spill_replay.rs, which names #6583 next to the removal condition. (cherry picked from commit ba7c892) Adapted for branch-1.1: - The tests build the pools over a fake Spark through create_pool, a seam that #6271 added to memory_pools. Only that seam is ported, not #6271's memory usage log change, which is not on branch-1.1: create_memory_pool builds the pools through create_pool, with_spark is pub(super), and the pools' new constructors, which nothing calls any more, are removed. - overcommit() and unwrap_task_shared(), which only that log reads, are not ported, so nothing looks through the wrapper. spill_replay.rs drops unwrap_spill_replay, and its tests drop their overcommit(&pool) assertions, which restate pool.reserved() minus what the fake Spark holds. The test that shrinks the replay checks pool.reserved() instead. - memory_management.md: the paragraph about the task's shared CometTaskMemoryManager, from #6261, is not on branch-1.1, so the new section follows the task-shared pools section directly. - CometAggregateSuite: only its imports differ. This branch does not import Column, CometBaseAggregateExec or CometSortAggregateExec.
Which issue does this PR close?
Prerequisite for #25172, which fixes
FairSpillPoolaccounting across reservations belonging to the same memory consumer. This PR prepares aggregate spill replay to work within that shared allowance and should merge first.Rationale for this change
Spilling an aggregation to disk only helps if there is enough memory to read it back and finish the aggregation. During spill replay, DataFusion merges sorted spill files and feeds their rows into an aggregate that combines the intermediate states into final results. The merge needs buffers for its inputs and output, while the aggregate needs memory for group keys and accumulators. Those allocations coexist: producing the next batch and processing it are both part of the same operator's memory budget.
For example, consider:
With many customers and a small memory limit, the aggregate can spill several times. A customer's events may then be spread across several spill files. Merging those files brings the customer's intermediate states together, but the aggregate still has to build the final array. Even if the merge delivers one row at a time, that array grows until the customer's group is complete. Small input batches therefore do not eliminate the need for replay memory.
The current merge tries to merge as many spill files as its reservation permits, without leaving room for this downstream work. As a simplified example, suppose the operator has a 100 MiB allowance and no other reservations. An 80 MiB merge fits by itself, but if the aggregate then needs 24 MiB to process its output, their combined 104 MiB does not fit. A pool enforcing the shared allowance must reject that growth, even though using a smaller merge could let the same aggregation make progress.
This is particularly relevant to #25172. The merge and aggregate use separate reservations under the same memory consumer. Today,
FairSpillPoolchecks those reservations individually, which can hide their combined usage exceeding the consumer's allowance. Once #25172 enforces that allowance across both reservations, replay must already leave space for the aggregate. This PR addresses that requirement before the accounting fix lands.What changes are included in this PR?
Aggregate spill merges now leave room for the aggregate consuming their output. When choosing how many spill files to merge and how much to buffer, the merge asks the configured memory pool to admit its buffers plus an equal amount of spare capacity for replay. Once admitted, it releases the spare reservation before replay starts, retaining only the reservation for the merge buffers. This uses the pool's existing allocation checks, so the decision follows its admission policy, including any fair-share restrictions, without adding a public memory-pool API.
In the example above, the 80 MiB merge would first need approval for 160 MiB and would be rejected. A smaller 40 MiB merge would need approval for 80 MiB; after returning the 40 MiB of temporary headroom, it would leave 60 MiB available, enough for the illustrative 24 MiB of aggregate state. The equal-sized headroom is a practical budgeting policy, not an estimate of the exact accumulator size or a guarantee that every aggregation will fit.
The existing multi-level merge adapts by merging fewer files at once, reducing read-ahead, or splitting oversized spill batches into smaller batches. This PR makes those adjustments account for replay as well. It also handles the point where a batch cannot shrink further: a single wide row may still fit the real pool even when equal headroom does not. In that case, only the minimum merge may retry without the extra headroom, with read-ahead disabled and all merge memory still subject to the pool's checks. Inspecting decoded batches before rewriting a spill file avoids unnecessary writes when the rows already fit or cannot be split, including when the original spill files fill the disk quota.
All supported aggregate replay implementations use this policy, including hash and ordered aggregation and the legacy implementation selected by
enable_migration_aggregate=false. The legacy path also releases unused initial grouping capacity before replay. The headroom policy is private to aggregate replay; ordinary sort callers retain their existing admission behavior.Are these changes tested?
Regression tests exercise replay across the supported aggregate implementations; the migrated hash and ordered paths also run with another registered spilling consumer. In particular, the legacy
ARRAY_AGGtest uses 64 groups with 64 values per group, one-row batches, and an 8 KiB pool to check that accumulator state can grow across replay batches. The tests verify exact aggregate results, reservation bounds, and release of memory and spill files.Merge-level coverage checks oversized batches, short and odd-sized batches, indivisible rows, decoded string-view sizes, and spill files that already fill the disk quota. It also checks that temporary headroom is released after a candidate merge is rejected and that the fallback still fails when the actual merge cannot fit.
Validation results and environment
Validated replay alone, with the existing
FairSpillPoolimplementation:cargo fmt --alland strict all-target/all-feature Clippy passed../dev/rust_lint.shpassed, including private Rust documentation and local Markdown links.Local validation uses Rust 1.98.1 and upstream revision
22651d24with its unchanged dependency lockfile; newer main's dependency versions are unavailable in the local registry. Merge compatibility with current main is checked separately, and GitHub CI validates its merged revision.Are there any user-facing changes?
Aggregations that spill leave memory available for processing the merged rows, reducing avoidable replay failures under constrained memory. Achieving this can require smaller batches or additional merge passes. SQL semantics and public APIs are unchanged.
The temporary headroom reservations can increase recorded reservation peaks without allocating additional data buffers. That headroom is released before replay, so other concurrent allocations or aggregate state that outgrows the available memory can still cause
ResourcesExhausted.