[vLLM] Wait for recurrent checkpoint restore capacity instead of recomputing - #100
Conversation
Signed-off-by: Costa Tsaousis <costa@netdata.cloud>
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. 🗂️ Base branches to auto review (1)
Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Signed-off-by: Costa Tsaousis <costa@netdata.cloud>
…n-followup fix(checkpoint): wait for restore admission under capacity pressure (#100 + review fixes)
65254d2
into
local-inference-lab:integration/local-inference-lab
|
Thank you, this was the right fix for the "admitted without the restore" collapse. We merged it together with fixes for our review findings: vLLM LMCache#929 (your two commits rebased onto the current beta, authorship kept) and LMCache #103 (your commits unchanged). What we changed on top, in short:
End-to-end on GLM-5.3-Flash Spark TP2 with 8 agents whose contexts outgrow the KV pool (your reproduction):
The next beta image will include both. Details and tests are in vLLM LMCache#929 and LMCache #103. 🤖 Generated with Claude Code |
What this PR does / why we need it:
Companion PR: local-inference-lab/vllm#922
When a recurrent checkpoint exists but GPU capacity is temporarily unavailable, keep the request waiting instead of permitting admission that can recompute the cached prefix. Copy-task saturation follows the same rule. Retain the answered directory manifest while waiting, so capacity pressure neither creates repeated lookups nor consumes the timeout for an unanswered lookup.
Use the paired vLLM admission-reservation API to protect both imported pages and continuation capacity. Revalidate a selected local checkpoint if ordinary admission has to wait and that checkpoint is later evicted; restart external lookup instead of silently losing reuse. Local reuse remains distinct from external cache-hit accounting.
Genuine misses, incompatible checkpoints, unanswered directory requests, and exhausted transfer retries still permit recomputation. Submitted worker copies must drain before their destinations become reusable. Failed-copy retries remain bounded; each unanswered lookup gets its own timeout.
Special notes for your reviewers:
Validation against the paired installed runtimes:
alwayscheckpoint writes and is not a controlled before/after comparison or an output-equivalence evaluation.The paired vLLM PR also corrects continuation accounting for recycling attention. That final correction has CPU regression coverage and was not part of the production observation.
If applicable: