Skip to content

[KV cache] Reserve execution capacity and ownership for checkpoint restores - #922

Closed
ktsaou wants to merge 2 commits into
local-inference-lab:integration/karmic-kraken-betafrom
ktsaou:fix/restore-admission-pr
Closed

ktsaou wants to merge 2 commits into
local-inference-lab:integration/karmic-kraken-betafrom
ktsaou:fix/restore-admission-pr

Conversation

@ktsaou

@ktsaou ktsaou commented Sep 27, 2026 •

Copy link
Copy Markdown

Purpose

Companion PR: local-inference-lab/LMCache#100

Keep an external recurrent-checkpoint restore waiting until both its imported pages and its continuation can fit. Under GPU pressure, ordinary admission could otherwise start recomputing the cached prefix. Reserving only the imported pages is insufficient: competing admissions can consume the space needed to resume execution, and publication can leave the restored checkpoint evictable before its consumer acquires it.

This change adds opt-in admission reservations to the external checkpoint API:

  • Reserve a running slot and continuation, capture, and replay capacity before copying; respect existing prefill reservations and the admission watermark.
  • Retain checkpoint ownership after all-rank publication until ordinary allocation takes over. Release credits and ownership on completion or cancellation, after submitted copies drain.
  • Let existing reservation owners progress past a blocked queue head, including saved-logits restores. Keep streaming resumes outside the new-admission gate.
  • Distinguish temporary resource pressure from checkpoint geometry that cannot fit an empty pool.

The paired LMCache change enables this contract and waits on capacity refusal. Existing callers keep the previous API behavior unless they request admission reservations. Merge vLLM before enabling the paired LMCache change.

Related: LMCache #75 retains manifests across temporary pressure but explicitly permits ordinary admission without restore. This change reserves execution capacity and protects ownership through admission. It is separate from vLLM #735, which targets a Jovian scheduler fairness path, and from missing checkpoint pages caused by storage supersession.

Test Plan

Run against the paired installed runtimes:

python -m pytest --noconftest -o addopts= -p no:cacheprovider \
  tests/v1/core/test_boundary_admission.py \
  tests/v1/core/test_prefix_caching.py -k 'boundary or external or recycling' -q

Coverage includes execution-capacity ownership across chunked prefill, cancellation before and after copy completion, slot reservations, FIFO capacity waiting, watermark handling, impossible geometry, eight-session rotation, saved-logits scheduling behind a blocked head, streaming resumption, and physical-window reservations for sliding-window/chunked-local attention (including retained tails and aligned/partial short suffixes).

Test Result

  • CPU validation on the deployed local image with the final PR allocator overlaid: 93 passed in the vLLM suites above. Paired LMCache transfer suite: 28 passed, 2 GPU-dependent skips.
  • Capacity-wait and retained-ownership regressions fail on the unpatched image.
  • Python lint, formatting, and diff checks pass.
  • Production observation: GLM-5.3-Flash on TP2/DCP2, C8, MTP3, approximately 1.02M aggregate GPU KV tokens, with 16 parallel research agents. Requests continued completing during capacity waits; representative intervals showed 94–98% external-cache hits and roughly 180–300 aggregate decode tokens/s. This is an observational workload, not a controlled performance comparison. Preemptions still occur.

The production run uses always checkpoint writes. From 22:12:13 to 22:29:28 UTC on September 27, metric deltas recorded 482 completions and 19,428,947 external-cache hit tokens out of 20,183,733 queried (96.26%). Journal coverage from 22:10:05 showed no restore failures or missing-page errors; the service did not restart. The final recycling-attention correction is PR-only and was not deployed during this observation. This change does not repair missing storage pages, guarantee zero preemptions, or establish model-output equivalence. Dedicated GPU roundtrip tests remain skipped in the CPU suite. The sliding-window/chunked-local continuation correction is covered by real-scheduler CPU regressions; the production observation qualifies only the deployed GLM MTP configuration.

…mission

Signed-off-by: Costa Tsaousis <costa@netdata.cloud>
@coderabbitai

coderabbitai Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (1)
  • dev/*

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: eb23b3a0-3cd9-4138-904c-fb5d0739edd4

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Signed-off-by: Costa Tsaousis <costa@netdata.cloud>
@voipmonitor

Copy link
Copy Markdown

Thank you, this was the right fix for the "admitted without the restore" collapse. We merged it together with fixes for our review findings: vLLM #929 (your two commits rebased onto the current beta, authorship kept) and LMCache #103 (your commits unchanged).

What we changed on top, in short:

  • vLLM, queue order: FCFS is kept while imports are pending; only restores that already own their capacity may pass a blocked head.
  • vLLM, step cost: a step no longer rescans the waiting queue: with 500 waiting requests it takes 0.056 ms instead of 15.7 ms.
  • vLLM, decode: a ready import no longer suppresses the decode pass, and reservations no longer preempt running decodes.
  • vLLM, bounded wait: a waiting restore gives up after VLLM_CHECKPOINT_RESTORE_MAX_WAIT_S (default 60 s) and the prompt is recomputed.
  • vLLM, cancel and reset: cancelling frees the slot and credits at once, and a forced prefix-cache reset releases ready imports.
  • LMCache: the bridge checks for the vLLM API at startup and keeps the old behaviour on an older vLLM instead of crashing on the first finished request. It also revalidates only when the selected checkpoint was actually lost, logs every 30 s while a restore waits, and refreshes the lookup during long waits.

End-to-end on GLM-5.3-Flash Spark TP2 with 8 agents whose contexts outgrow the KV pool (your reproduction):

Beta With these fixes
Aggregate decode 17.2 tok/s 138.8 tok/s
Prompt tokens from cache 14.3% 98.1%
Later turns with a restore 13/82 650/650
Mean turn time 112 s 14.6 s

The next beta image will include both. Details and tests are in vLLM #929 and LMCache #103.

🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants