Repository navigation
fix(scheduler): wait for capacity to restore external checkpoints, fairly and bounded (#922 + review fixes) - #929
Merged
voipmonitor merged 7 commits intoSep 28, 2026
Conversation
…mission Signed-off-by: Costa Tsaousis <costa@netdata.cloud>
Signed-off-by: Costa Tsaousis <costa@netdata.cloud>
Follow-up to #922 review findings on the admission reservations: - Import credits gated every allocation, so a running decode at a block boundary could be preempted (often the admitted consumer itself) to protect credits of a restore that had not started. Credits now apply to waiting and preempted admissions only; running requests keep growing and resolve pressure through ordinary preemption. - A capacity waiter blocked every other new admission, including requests the connector does not handle and higher-priority ones. A refused restore now only holds back admissions considered after it in the same scheduler step, and those still run when they fit beside the blocks it waits for. Requests ordered ahead of it are unaffected. - A capacity wait is bounded by VLLM_CHECKPOINT_RESTORE_MAX_WAIT_S (default 60, 0 recomputes immediately). Once it expires, one WARNING records the wait, the reservation keeps returning None, and the scheduler admits the request to recompute its prompt even though the connector still defers it for capacity. - Cancelling an import mid-copy kept its credits and running slot until the copy drained. Only the staged pages must wait for the copy now; credits and the slot are released at once. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Follow-up to #922 review findings on the waiting pass: - A failed allocation no longer stopped the waiting pass while any import was pending, so every later request could overtake a blocked queue head (FCFS and priority alike). The first failure stops ordinary admission again; only restored imports, which already own their capacity and slot, are still admitted past it. - With an import pending, every step walked the whole waiting queue (connector poll, checkpoint lookup and allocation per request), the saved-logits check scanned both queues (copying the priority heap), and the admission context summed in-flight prefills per candidate. Ready imports are now found through the allocator's admissions, the context is refreshed only after admissions, and priority queues test membership without the ordered heap copy. With 500 blocked 2K-token requests a step costs the same as without imports (0.06 ms instead of 15.7 ms FCFS / 28.3 ms priority). - A ready full-prefix import anywhere in the queue suppressed the running pass; if an older request was admitted instead, decodes lost the step. The import now gets its isolated saved-logits step first, and the running pass runs, without preempting new admissions, whenever that step is not taken. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
reset_prefix_cache(reset_running_requests=True) preempts every running request so the pool can be cleared, but a restore that finished copying and was not yet admitted still pinned its pages, and the reset raised RuntimeError. pause_scheduler(mode="keep", clear_cache=True) and sleep(level>=1, mode="keep") reach this path. Such restores are now released with the running requests; their requests fall back to an ordinary lookup. A reset without preemption still reports failure while a restored import holds its pages. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Update the #922 fragment to the reviewed behavior: reservations hold back only new admissions, queue order and step cost are kept while imports are pending, saved-logits steps no longer cost decodes, cancellation frees credits at once, forced resets release restored imports, and VLLM_CHECKPOINT_RESTORE_MAX_WAIT_S bounds capacity waits. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. 🗂️ Base branches to auto review (1)
Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
voipmonitor
marked this pull request as ready for review
September 28, 2026 16:05
This was referenced Sep 28, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This is #922 by @ktsaou (two commits, rebased onto the current beta) plus fixes for our review findings. It pairs with LMCache #103.
Problem (ktsaou, Discord). When several recurrent sessions no longer fit in the GPU together, the log fills with "Recurrent checkpoint restore … deferred … not enough free GPU blocks; it is admitted without the restore". Every such request recomputes its whole prompt, and total decode on GLM TP2 collapses to about 20 tok/s.
What #922 does. A restore reserves its GPU blocks and a running slot before the copy starts, and keeps the restored pages owned until execution takes over. Restores wait for capacity instead of being admitted to recompute.
Review fixes on top:
VLLM_CHECKPOINT_RESTORE_MAX_WAIT_S(default 60, 0 recomputes at once) bounds the wait. After it,reserve_external_boundary_checkpointreturns None, one WARNING is logged, and the request recomputes. No exception is raised.reset_prefix_cache(reset_running_requests=True)releases restores that finished but were not admitted, instead of raising.The fragment
vllm-restore-admissionwas updated to describe this.Tests (CPU).
test_boundary_admission_continuation.py: 89 passed.tests/v1/core: the same environment-only failures as base, and 73 more passing tests.E2E. GLM-5.3-Flash Spark TP2/DCP2, beta
c10dcc1b, LMCache L1 64 GB / L2 256 GB, on-evict.🤖 Generated with Claude Code