perf(rm): replace per-iteration sort in distributedAlloc with a min-heap - #1826
Conversation
82b1b23 to
b9da143
Compare
b9da143 to
21cd040
Compare
8333b12 to
5efcd18
Compare
|
Hi @rajatchopra — following up on your comment in #1788 ("Lets attack the loop optimization in #1826 next"). This PR is out of draft and ready for review: a single commit with the heap-based |
5efcd18 to
2319fe5
Compare
|
Heap logic looks good. Though this PR must wait for #1621 first. Will need a rebase after that. |
2319fe5 to
6848b06
Compare
|
Rebased on top of #1621 (now merged). The heap now sits on top of the new Ready for another look when you have a moment, @rajatchopra. |
06c33eb to
36e7ad1
Compare
954879b to
2af157d
Compare
2af157d to
3a62327
Compare
3a62327 to
8b0a351
Compare
|
@jonathan-meiri Do you have the benchmark data that shows the improvement? |
ca923f7 to
3be20fb
Compare
|
@henry118 The improvement is real, added
plus ~10× fewer allocations. To be clear on scope: this isn't a hot path (runs once per pod admission), so it's not a latency fix. The change came out of @rajatchopra's suggestion on #1788 to do one sort instead of N |
|
Thanks @jonathan-meiri, numbers look legit. Allocs dropping is good. Raw us savings won't be noticeable in practice, but the alloc/GC reduction is worth it on its own. Please rebase and sign off the last commit. |
Follow-up on top of NVIDIA#1621, which introduced the shared greedyAlloc loop with a pluggable replicaComparator (distributed vs packed). The loop still sorts the full candidate slice inside the allocation loop, paying O(n log n) per iteration for n iterations and giving O(n² log n) overall. Since all annotated replicas from the same underlying physical device share the same sort key, sorting at the replica granularity is wasted work — only m (the number of distinct physical devices contributing candidates) needs to be reordered. Refactor greedyAlloc to bucket candidates by their underlying physical device into a small gpuAllocState per device, holding a shared *replicaCount, the pickedFrom counter, and the remaining candidate IDs. A gpuPriorityQueue defers to the caller-supplied replicaComparator on allocated() for primary ordering and to pickedFrom for the tie-break (unchanged semantics). Each iteration pops the best device, takes one of its remaining replicas, updates counters, and pushes it back if any remain. Total cost drops to O(n log m). Both allocation policies (distributed and packed) benefit; no behavior change — the existing test suite (TestDistributedAlloc, TestPackedAlloc, TestPackedVsDistributedContrast, TestDistributedAlloc_PartiallyAllocated_DistributesAcrossDistinctGPUs, etc.) passes unchanged. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-Authored-By: runatom-ai <258621014+runatom-ai@users.noreply.github.com> Signed-off-by: Jonathan Meiri <33288957+Meiri28@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-Authored-By: runatom-ai <258621014+runatom-ai@users.noreply.github.com> Signed-off-by: Jonathan Meiri <33288957+Meiri28@users.noreply.github.com>
3be20fb to
498aba4
Compare
|
@henry118 Done, rebased on latest main and signed off the benchmark commit (DCO green). |
|
/ok to test 498aba4 |
Summary
Follow-up optimization stacked on #1788. Until #1788 lands, this PR's cumulative diff includes that commit too — the new work here is in the second commit (the heap refactor). Not for review until #1788 lands.
Replaces the per-iteration
sort.SliceindistributedAllocwith a min-heap keyed by(used, pickedFrom). Brings the loop fromO(n² log n)toO(n log m)wherenis replicas requested andmis the number of physical GPUs touched in this allocation. Same correctness as #1788; same tests pass.Practically,
nandmare small in real configurations and the wall-clock impact is invisible — this is structural cleanliness, not a hot-path speedup.Opened as draft so it doesn't enter the review queue alongside #1788. Happy to mark it ready as a separate follow-up after #1788 lands, or fold the change into #1788 if that's preferred.
Contributed by @Meiri28 on behalf of @runatom-ai.