Skip to content

fix(checkpoint): wait for restore admission under capacity pressure (#100 + review fixes) - #103

Merged
voipmonitor merged 5 commits into
integration/local-inference-labfrom
fix/restore-admission-followup
Sep 28, 2026
Merged

voipmonitor merged 5 commits into
integration/local-inference-labfrom
fix/restore-admission-followup

Conversation

@voipmonitor

@voipmonitor voipmonitor commented Sep 28, 2026 •

Copy link
Copy Markdown

This is #100 by @ktsaou (commits unchanged) plus fixes for our review findings. It pairs with vLLM LMCache#929.

What #100 does. When a restore cannot reserve its GPU pages, the bridge keeps the answered manifest and waits for capacity. Before, it admitted the request to recompute its whole prompt ("admitted without the restore").

Review fixes on top:

  1. Compatibility gate. At creation the bridge checks for vLLM's restore-admission API: can_admit_external_boundary_request, external_boundary_admission_ready, release_external_boundary_admission, and the reserve_admission parameter. Without them it logs one WARNING and keeps the previous behaviour. Before this, an older vLLM hit an AttributeError on every finished request, which killed the engine.
  2. Revalidation. It happens only if the GPU cache no longer holds a checkpoint at least as long as the selected one. A replaced or longer local checkpoint is used without another directory round trip.
  3. Visibility. A waiting restore is logged when the wait starts and every 30 s after that, with free blocks against needed blocks, or copies in flight. report_status() counts waits, and the waiter entry is cleared on every exit path.
  4. Fresh pages during long waits. A wait longer than half of lookup_timeout sends the lookup again, so the checkpoint's pages stay recent.
  5. Documentation and bypass. The per-attempt lookup timeout is documented. A complete local hit again skips a pending lookup, unless its own reserved copy is in flight.
  6. Smaller fixes. The manifest is parsed once per answer, log wording is corrected, and the fragment's models are GLM-5.3-Flash and Qwen3.8-Flash-Next.

Tests (CPU).

E2E (with vLLM LMCache#929; setup and table in vllm#929): 8 agents whose contexts outgrow the KV pool.

Beta With the fixes
Aggregate decode 17.2 tok/s 138.8 tok/s
Prompt tokens from cache 14.3% 98.1%
Later turns with a restore 13/82 650/650
Mean turn time 112 s 14.6 s

The bridge logged 648 "waiting for GPU capacity" lines and no "admitted without the restore".

🤖 Generated with Claude Code

ktsaou and others added 4 commits September 27, 2026 22:28
Signed-off-by: Costa Tsaousis <costa@netdata.cloud>
Signed-off-by: Costa Tsaousis <costa@netdata.cloud>
Follow-up review fixes for #100:
- Detect vLLM restore admission reservations once, when the bridge is
  created (the admission methods and the reserve_admission parameter).
  Without them, log one warning and keep the previous fallback: a restore
  without free GPU blocks is admitted and retried while unadmitted, one
  without a copy slot recomputes, and no missing method is ever called
  (finish_request no longer raises on every finished request).
- Look a selected checkpoint up again only when no local checkpoint at
  least as long is cached, not whenever a different one replaced it.
- Log a capacity wait when it starts and every 30 s with its age, the
  free GPU blocks and the blocks the checkpoint needs (or the copies in
  flight); report_status() counts current waits, the longest wait and the
  waits started. Every way out of a wait releases vLLM's waiter entry.
- A restore that retains its answer for half the lookup timeout sends the
  lookup again without waiting for it, keeping its pages recent; the reply
  replaces the retained answer, and an empty one recomputes the prompt.
- A complete local hit again bypasses a pending directory lookup, except
  behind the request's own reserved copy, and settles the lookup so a later
  eviction looks the prefix up again.
- Validate an answered manifest once per answer, and log an invalidated
  publication without the "missed; looking up a shorter checkpoint" line.
- Tests run on vLLM with and without the admission API: reservation-only
  cases skip with a reason, the fallback is tested through an allocator
  without the API, and exits assert that no restore stays queued.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scope the fragment to the request-boundary recurrent models, document the
fallback with an older vLLM, the per-reply lookup timeout (up to four
replies for a request whose restores fail), wait logging, lookup refresh
and the selection and local-hit rules. It still requires the paired vLLM
fragment vllm-restore-admission.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (1)
  • dev/*

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 78fa6e9f-165c-47fc-b684-adf1bad068c0

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@voipmonitor
voipmonitor marked this pull request as ready for review September 28, 2026 16:05
@voipmonitor
voipmonitor merged commit ab11b84 into integration/local-inference-lab Sep 28, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants