fix(checkpoint): keep intact disk checkpoints listed when a restore cannot get RAM (ktsaou's #105 + follow-ups) - #108
Conversation
…allocation failure Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ost pages On top of ktsaou's #105: - the admission timeout bounds only repeats after a RAM reservation failure; other misses repeat once when there is room, as before, so a checkpoint with a deleted page file is still retired when the first lookup answers after the timeout or the timeout is zero; - a repeat after a RAM failure waits 20 ms instead of resubmitting on every poll (700-900 disk lookups per second before); - #106's stalled-lookup test helper patches the detailed prefetch query the store now uses; - regression tests for a writer that frees RAM before the poll, fragmented RAM (listing kept, resubmits bounded) and lost pages under slow or zero timeouts; release fragment lmcache-105. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. 🗂️ Base branches to auto review (1)
Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
GPU E2E on RTX PRO 6000 (.4): GLM-5.3-Flash Spark TP2, lmcache on-evict, 6 GiB L1, 8 agents at 110K tokens plus 4K per turn, so every turn restores. Base is the integration head
Both runs also logged about 900 HTTP 400s. That is the load generator growing contexts past the model length, the same in both. The RAM-failure path did not trigger in this run, so the E2E shows no regression rather than the fix. The fix is covered by |
Contains ktsaou's #105 unchanged (both commits, authorship kept, merged onto the current head
75f2b59), plus follow-ups from review.The bug #105 fixes is real on the current head.
checkpoint_storage.pyonly receives a bitmap from the prefetch. The L1 reservation failure (prefetch_controller.py:1025-1028) is dropped.75f2b59.Follow-ups on top of #105 (commit
0337b0e):_stall_lookupspatchesquery_prefetch_status_detailed, which the store now calls. Without this, 3 tests intest_checkpoint_cancel_drain.pyfail after the merge.test_checkpoint_restore_ram_failures.py:lmcache-105.Tests (CPU, beta image, 25 checkpoint/storage suites plus #105's and the new file): 276 passed, 21 skipped (native/CUDA-only). On the current head the new file fails in the four "stays listed" cases. With #105 alone, the lost-page cases with a slow first lookup and with timeout 0 fail, and the bounded-resubmit assertion fails.
A GPU end-to-end run (GLM-5.3 TP2, lmcache on-evict, 6 GiB L1) is prepared and runs before merge.
🤖 Generated with Claude Code