Skip to content

fix(checkpoint): keep intact disk checkpoints listed when a restore cannot get RAM (ktsaou's #105 + follow-ups) - #108

Merged
voipmonitor merged 4 commits into
integration/local-inference-labfrom
fix/checkpoint-ram-failure-listing
Sep 29, 2026
Merged

voipmonitor merged 4 commits into
integration/local-inference-labfrom
fix/checkpoint-ram-failure-listing

Conversation

@voipmonitor

Copy link
Copy Markdown

Contains ktsaou's #105 unchanged (both commits, authorship kept, merged onto the current head 75f2b59), plus follow-ups from review.

The bug #105 fixes is real on the current head.

  • checkpoint_storage.py only receives a bitmap from the prefetch. The L1 reservation failure (prefetch_controller.py:1025-1028) is dropped.
  • On the second lookup, the store therefore judges "room in RAM" from free RAM at poll time and retires the checkpoint.
  • Result: an intact disk checkpoint gets delisted when RAM was held briefly and freed before the poll, or when free RAM is fragmented into pieces smaller than a page. Both cases are reproduced by the new tests, which fail on 75f2b59.

Follow-ups on top of #105 (commit 0337b0e):

  1. Lost pages are retired again in every case. fix(checkpoint): preserve disk checkpoints after L1 allocation failure #105 moved the admission-timeout check ahead of the "repeat once when there is room" check. A checkpoint with a deleted page file then stayed listed when the first lookup answered after the timeout, or the timeout was 0, and the log wrongly said "found no room in RAM". Now the timeout bounds only repeats after a RAM failure.
  2. No hot resubmit loop. After a RAM failure fix(checkpoint): preserve disk checkpoints after L1 allocation failure #105 resubmitted the lookup on every poll: 700–900 disk lookups per second for up to 8 s per restore. It now pauses 20 ms between repeats.
  3. fix(checkpoint): a restore stuck in storage misses instead of killing the engine #106's test helper _stall_lookups patches query_prefetch_status_detailed, which the store now calls. Without this, 3 tests in test_checkpoint_cancel_drain.py fail after the merge.
  4. New test_checkpoint_restore_ram_failures.py:
    • a writer frees RAM before the poll;
    • fragmented RAM keeps the checkpoint listed, with repeats bounded (≤ 80 in 1 s);
    • a lost page is retired with a fast first lookup, a first lookup slower than the timeout, and timeout 0.
  5. Release fragment lmcache-105.

Tests (CPU, beta image, 25 checkpoint/storage suites plus #105's and the new file): 276 passed, 21 skipped (native/CUDA-only). On the current head the new file fails in the four "stays listed" cases. With #105 alone, the lost-page cases with a slow first lookup and with timeout 0 fail, and the bounded-resubmit assertion fails.

A GPU end-to-end run (GLM-5.3 TP2, lmcache on-evict, 6 GiB L1) is prepared and runs before merge.

🤖 Generated with Claude Code

ktsaou and others added 4 commits September 29, 2026 09:50
…allocation failure

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ost pages

On top of ktsaou's #105:
- the admission timeout bounds only repeats after a RAM reservation failure;
  other misses repeat once when there is room, as before, so a checkpoint
  with a deleted page file is still retired when the first lookup answers
  after the timeout or the timeout is zero;
- a repeat after a RAM failure waits 20 ms instead of resubmitting on every
  poll (700-900 disk lookups per second before);
- #106's stalled-lookup test helper patches the detailed prefetch query the
  store now uses;
- regression tests for a writer that frees RAM before the poll, fragmented
  RAM (listing kept, resubmits bounded) and lost pages under slow or zero
  timeouts; release fragment lmcache-105.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 29, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (1)
  • dev/*

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 94b55481-5d1c-436d-bb6a-6e8536839c77

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@voipmonitor

Copy link
Copy Markdown
Author

GPU E2E on RTX PRO 6000 (.4): GLM-5.3-Flash Spark TP2, lmcache on-evict, 6 GiB L1, 8 agents at 110K tokens plus 4K per turn, so every turn restores. Base is the integration head 75f2b59; the other run is this PR.

base this PR
turns served 288 288
prompt tokens cached 96.2 % 96.2 %
turns with a restore hit 280/280 280/280
mean turn 17.4 s 16.9 s
delisted after miss 0 0
tracebacks, unsafe, engine dead 0 0

Both runs also logged about 900 HTTP 400s. That is the load generator growing contexts past the model length, the same in both. The RAM-failure path did not trigger in this run, so the E2E shows no regression rather than the fix. The fix is covered by test_checkpoint_restore_ram_failures.py, which fails on the head and passes here. Merging; it supersedes #105.

@voipmonitor
voipmonitor merged commit 820af25 into integration/local-inference-lab Sep 29, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants