Skip to content

[vLLM] Wait for recurrent checkpoint restore capacity instead of recomputing - #100

Merged
voipmonitor merged 2 commits into
local-inference-lab:integration/local-inference-labfrom
ktsaou:fix/restore-admission
Sep 28, 2026
Merged

voipmonitor merged 2 commits into
local-inference-lab:integration/local-inference-labfrom
ktsaou:fix/restore-admission

Conversation

@ktsaou

@ktsaou ktsaou commented Sep 27, 2026 •

Copy link
Copy Markdown

What this PR does / why we need it:

Companion PR: local-inference-lab/vllm#922

When a recurrent checkpoint exists but GPU capacity is temporarily unavailable, keep the request waiting instead of permitting admission that can recompute the cached prefix. Copy-task saturation follows the same rule. Retain the answered directory manifest while waiting, so capacity pressure neither creates repeated lookups nor consumes the timeout for an unanswered lookup.

Use the paired vLLM admission-reservation API to protect both imported pages and continuation capacity. Revalidate a selected local checkpoint if ordinary admission has to wait and that checkpoint is later evicted; restart external lookup instead of silently losing reuse. Local reuse remains distinct from external cache-hit accounting.

Genuine misses, incompatible checkpoints, unanswered directory requests, and exhausted transfer retries still permit recomputation. Submitted worker copies must drain before their destinations become reusable. Failed-copy retries remain bounded; each unanswered lookup gets its own timeout.

Special notes for your reviewers:

  • Requires the paired vLLM restore-admission change. Merge/deploy vLLM first; this connector calls the new reservation and ownership APIs.
  • Extends #75, which deliberately permits ordinary admission after capacity refusal, and preserves bounded unanswered-lookup handling from #82.
  • Does not address the independent checkpoint supersession/storage-loss issue. Missing-page failures retain their existing retry/fallback behavior.
  • Capacity waiting has no elapsed-time fallback to full prefill. Runnable owners must release space; cancellation releases waiting reservations. This can reduce simultaneous running requests to preserve checkpoint reuse.

Validation against the paired installed runtimes:

python -m pytest --noconftest -o addopts= -p no:cacheprovider \
  tests/v1/test_vllm_semantic_checkpoint_transfer.py -q
  • 28 passed, 2 GPU-dependent skips; paired vLLM boundary/admission suites: 93 passed.
  • Tests cover capacity and copy-task waiting, collective visibility, cancellation and request-ID reuse, shorter-checkpoint recovery after copy failure, lookup failure/miss/timeout fallback, and eviction of both partial and full local selections before admission.
  • The two local-eviction regressions fail against the preceding local implementation and pass after the correction, including successful restore, cache-hit attribution, admission, and resource cleanup.
  • Python lint, formatting, and diff checks pass.
  • With GLM-5.3-Flash TP2/DCP2, C8, MTP3 and approximately 1.02M aggregate GPU KV tokens, 16 parallel research agents continued making progress through capacity waits. Representative intervals showed 94–98% external-cache hits and roughly 180–300 aggregate decode tokens/s. Preemptions still occur. The run uses always checkpoint writes and is not a controlled before/after comparison or an output-equivalence evaluation.

The paired vLLM PR also corrects continuation accounting for recycling attention. That final correction has CPU regression coverage and was not part of the production observation.

If applicable:

  • This PR contains unit tests.
  • User-facing behavior is described in the integration release fragment and API documentation.

Signed-off-by: Costa Tsaousis <costa@netdata.cloud>
@coderabbitai

coderabbitai Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (1)
  • dev/*

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 9155c189-260b-471f-a1f2-633f37a96dcb

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Signed-off-by: Costa Tsaousis <costa@netdata.cloud>
@ktsaou
ktsaou marked this pull request as ready for review September 27, 2026 22:37
voipmonitor added a commit that referenced this pull request Sep 28, 2026
…n-followup

fix(checkpoint): wait for restore admission under capacity pressure (#100 + review fixes)
@voipmonitor
voipmonitor merged commit 65254d2 into local-inference-lab:integration/local-inference-lab Sep 28, 2026
1 check passed
@voipmonitor

Copy link
Copy Markdown

Thank you, this was the right fix for the "admitted without the restore" collapse. We merged it together with fixes for our review findings: vLLM LMCache#929 (your two commits rebased onto the current beta, authorship kept) and LMCache #103 (your commits unchanged).

What we changed on top, in short:

  • vLLM, queue order: FCFS is kept while imports are pending; only restores that already own their capacity may pass a blocked head.
  • vLLM, step cost: a step no longer rescans the waiting queue: with 500 waiting requests it takes 0.056 ms instead of 15.7 ms.
  • vLLM, decode: a ready import no longer suppresses the decode pass, and reservations no longer preempt running decodes.
  • vLLM, bounded wait: a waiting restore gives up after VLLM_CHECKPOINT_RESTORE_MAX_WAIT_S (default 60 s) and the prompt is recomputed.
  • vLLM, cancel and reset: cancelling frees the slot and credits at once, and a forced prefix-cache reset releases ready imports.
  • LMCache: the bridge checks for the vLLM API at startup and keeps the old behaviour on an older vLLM instead of crashing on the first finished request. It also revalidates only when the selected checkpoint was actually lost, logs every 30 s while a restore waits, and refreshes the lookup during long waits.

End-to-end on GLM-5.3-Flash Spark TP2 with 8 agents whose contexts outgrow the KV pool (your reproduction):

Beta With these fixes
Aggregate decode 17.2 tok/s 138.8 tok/s
Prompt tokens from cache 14.3% 98.1%
Later turns with a restore 13/82 650/650
Mean turn time 112 s 14.6 s

The next beta image will include both. Details and tests are in vLLM LMCache#929 and LMCache #103.

🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants