Skip to content

fix(scheduler): wait for capacity to restore external checkpoints, fairly and bounded (#922 + review fixes) - #929

Merged
voipmonitor merged 7 commits into
integration/karmic-kraken-betafrom
fix/restore-admission-followup
Sep 28, 2026
Merged

voipmonitor merged 7 commits into
integration/karmic-kraken-betafrom
fix/restore-admission-followup

Conversation

@voipmonitor

@voipmonitor voipmonitor commented Sep 28, 2026 •

Copy link
Copy Markdown

This is #922 by @ktsaou (two commits, rebased onto the current beta) plus fixes for our review findings. It pairs with LMCache #103.

Problem (ktsaou, Discord). When several recurrent sessions no longer fit in the GPU together, the log fills with "Recurrent checkpoint restore … deferred … not enough free GPU blocks; it is admitted without the restore". Every such request recomputes its whole prompt, and total decode on GLM TP2 collapses to about 20 tok/s.

What #922 does. A restore reserves its GPU blocks and a running slot before the copy starts, and keeps the restored pages owned until execution takes over. Restores wait for capacity instead of being admitted to recompute.

Review fixes on top:

  1. FCFS is kept. The first failed allocation stops ordinary admission again. Only restores that already own their capacity may pass a blocked head. Before this, any request could overtake a blocked head while an import was pending.
  2. Step cost. Ready imports come from the allocator instead of a queue scan, and the admission context is recomputed only after an admission. With 500 waiting requests and one copying import, a step now takes 0.056 ms; it took 15.7 ms (FCFS) or 28.4 ms (priority).
  3. Running pass. A ready import that is not at the queue head no longer suppresses the running (decode) pass.
  4. Preemption. Credits held for a restore that has not started no longer preempt running decodes; reservations gate only new admissions.
  5. Liveness.
    • A capacity waiter holds back only admissions that would take its space.
    • VLLM_CHECKPOINT_RESTORE_MAX_WAIT_S (default 60, 0 recomputes at once) bounds the wait. After it, reserve_external_boundary_checkpoint returns None, one WARNING is logged, and the request recomputes. No exception is raised.
  6. Cancelling frees the slot and credits at once; only the staged pages wait for the copy.
  7. reset_prefix_cache(reset_running_requests=True) releases restores that finished but were not admitted, instead of raising.

The fragment vllm-restore-admission was updated to describe this.

Tests (CPU).

  • PR test plan: 110 passed (the PR head passed 93).
  • test_boundary_admission_continuation.py: 89 passed.
  • The 21 new or changed tests fail on the PR head (17 of 21) and pass here.
  • Full tests/v1/core: the same environment-only failures as base, and 73 more passing tests.
  • LMCache perf(spec decode): return variable draft ids through async output #103's bridge suite against this tree: 46 passed.

E2E. GLM-5.3-Flash Spark TP2/DCP2, beta c10dcc1b, LMCache L1 64 GB / L2 256 GB, on-evict.

  • Load: 8 single-threaded agents. Each starts at 110K tokens and adds 4K tokens per turn, so their contexts together outgrow the ~1M-token KV pool. Each turn sends the whole conversation (token IDs, greedy, 256 output tokens). 20 minutes per arm.
Beta This PR + LMCache #103
Turns completed 90 658
Aggregate decode 17.2 tok/s 138.8 tok/s
Prompt tokens from cache 14.3% 98.1%
Later turns with a restore 13/82 650/650
Mean turn time 112.4 s 14.6 s
Tokens recomputed 10.2M 3.5M (the eight 110K cold starts and the new 4K per turn)
"admitted without the restore" 16 0
  • VRAM mode (no connector), same image with this PR: mixed load (45 requests, images and long prompts) finished with 0 errors and 0 tracebacks. Decode steps/s were C1 69.3 and C8 215.1, against 69.1 and 217.0 on the beta, which is within the run-to-run noise.

🤖 Generated with Claude Code

ktsaou and others added 6 commits September 28, 2026 14:09
…mission

Signed-off-by: Costa Tsaousis <costa@netdata.cloud>
Signed-off-by: Costa Tsaousis <costa@netdata.cloud>
Follow-up to #922 review findings on the admission reservations:

- Import credits gated every allocation, so a running decode at a block
  boundary could be preempted (often the admitted consumer itself) to
  protect credits of a restore that had not started. Credits now apply
  to waiting and preempted admissions only; running requests keep
  growing and resolve pressure through ordinary preemption.
- A capacity waiter blocked every other new admission, including
  requests the connector does not handle and higher-priority ones.
  A refused restore now only holds back admissions considered after it
  in the same scheduler step, and those still run when they fit beside
  the blocks it waits for. Requests ordered ahead of it are unaffected.
- A capacity wait is bounded by VLLM_CHECKPOINT_RESTORE_MAX_WAIT_S
  (default 60, 0 recomputes immediately). Once it expires, one WARNING
  records the wait, the reservation keeps returning None, and the
  scheduler admits the request to recompute its prompt even though the
  connector still defers it for capacity.
- Cancelling an import mid-copy kept its credits and running slot until
  the copy drained. Only the staged pages must wait for the copy now;
  credits and the slot are released at once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Follow-up to #922 review findings on the waiting pass:

- A failed allocation no longer stopped the waiting pass while any
  import was pending, so every later request could overtake a blocked
  queue head (FCFS and priority alike). The first failure stops ordinary
  admission again; only restored imports, which already own their
  capacity and slot, are still admitted past it.
- With an import pending, every step walked the whole waiting queue
  (connector poll, checkpoint lookup and allocation per request), the
  saved-logits check scanned both queues (copying the priority heap),
  and the admission context summed in-flight prefills per candidate.
  Ready imports are now found through the allocator's admissions, the
  context is refreshed only after admissions, and priority queues test
  membership without the ordered heap copy. With 500 blocked 2K-token
  requests a step costs the same as without imports (0.06 ms instead of
  15.7 ms FCFS / 28.3 ms priority).
- A ready full-prefix import anywhere in the queue suppressed the
  running pass; if an older request was admitted instead, decodes lost
  the step. The import now gets its isolated saved-logits step first,
  and the running pass runs, without preempting new admissions,
  whenever that step is not taken.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
reset_prefix_cache(reset_running_requests=True) preempts every running
request so the pool can be cleared, but a restore that finished copying
and was not yet admitted still pinned its pages, and the reset raised
RuntimeError. pause_scheduler(mode="keep", clear_cache=True) and
sleep(level>=1, mode="keep") reach this path. Such restores are now
released with the running requests; their requests fall back to an
ordinary lookup. A reset without preemption still reports failure while
a restored import holds its pages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Update the #922 fragment to the reviewed behavior: reservations hold back
only new admissions, queue order and step cost are kept while imports
are pending, saved-logits steps no longer cost decodes, cancellation
frees credits at once, forced resets release restored imports, and
VLLM_CHECKPOINT_RESTORE_MAX_WAIT_S bounds capacity waits.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (1)
  • dev/*

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: c1238355-e9ed-45c5-ac26-6c53fadb548a

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants