Conversation
…mission Signed-off-by: Costa Tsaousis <costa@netdata.cloud>
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. 🗂️ Base branches to auto review (1)
Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Signed-off-by: Costa Tsaousis <costa@netdata.cloud>
…n-followup fix(scheduler): wait for capacity to restore external checkpoints, fairly and bounded (#922 + review fixes)
|
Thank you, this was the right fix for the "admitted without the restore" collapse. We merged it together with fixes for our review findings: vLLM #929 (your two commits rebased onto the current beta, authorship kept) and LMCache #103 (your commits unchanged). What we changed on top, in short:
End-to-end on GLM-5.3-Flash Spark TP2 with 8 agents whose contexts outgrow the KV pool (your reproduction):
The next beta image will include both. Details and tests are in vLLM #929 and LMCache #103. 🤖 Generated with Claude Code |
Purpose
Companion PR: local-inference-lab/LMCache#100
Keep an external recurrent-checkpoint restore waiting until both its imported pages and its continuation can fit. Under GPU pressure, ordinary admission could otherwise start recomputing the cached prefix. Reserving only the imported pages is insufficient: competing admissions can consume the space needed to resume execution, and publication can leave the restored checkpoint evictable before its consumer acquires it.
This change adds opt-in admission reservations to the external checkpoint API:
The paired LMCache change enables this contract and waits on capacity refusal. Existing callers keep the previous API behavior unless they request admission reservations. Merge vLLM before enabling the paired LMCache change.
Related: LMCache #75 retains manifests across temporary pressure but explicitly permits ordinary admission without restore. This change reserves execution capacity and protects ownership through admission. It is separate from vLLM #735, which targets a Jovian scheduler fairness path, and from missing checkpoint pages caused by storage supersession.
Test Plan
Run against the paired installed runtimes:
python -m pytest --noconftest -o addopts= -p no:cacheprovider \ tests/v1/core/test_boundary_admission.py \ tests/v1/core/test_prefix_caching.py -k 'boundary or external or recycling' -qCoverage includes execution-capacity ownership across chunked prefill, cancellation before and after copy completion, slot reservations, FIFO capacity waiting, watermark handling, impossible geometry, eight-session rotation, saved-logits scheduling behind a blocked head, streaming resumption, and physical-window reservations for sliding-window/chunked-local attention (including retained tails and aligned/partial short suffixes).
Test Result
The production run uses
alwayscheckpoint writes. From 22:12:13 to 22:29:28 UTC on September 27, metric deltas recorded 482 completions and 19,428,947 external-cache hit tokens out of 20,183,733 queried (96.26%). Journal coverage from 22:10:05 showed no restore failures or missing-page errors; the service did not restart. The final recycling-attention correction is PR-only and was not deployed during this observation. This change does not repair missing storage pages, guarantee zero preemptions, or establish model-output equivalence. Dedicated GPU roundtrip tests remain skipped in the CPU suite. The sliding-window/chunked-local continuation correction is covered by real-scheduler CPU regressions; the production observation qualifies only the deployed GLM MTP configuration.