fix(checkpoint): preserve disk checkpoints after L1 allocation failure - #105
Conversation
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. 🗂️ Base branches to auto review (1)
Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Thanks, the bug is real on the current head. The store only got a bitmap and dropped the L1 reservation failure, so a restore whose RAM was held briefly, or was fragmented below one page, delisted an intact disk checkpoint. Our tests reproduce both cases on Merged onto the current head, this PR needed three follow-ups. They are in #108, which carries your two commits unchanged:
#108 also adds regression tests and the |
…ailure-listing fix(checkpoint): keep intact disk checkpoints listed when a restore cannot get RAM (ktsaou's #105 + follow-ups)
c337fc8
into
local-inference-lab:integration/local-inference-lab
|
Landed through #108 ( |
A failed checkpoint restore can remove an intact disk checkpoint from the directory when L1 cannot allocate its load buffers. The retry code checks free RAM after the prefetch has completed; allocation padding or a concurrent writer releasing memory can make that later sample incorrectly look sufficient.
Carry the actual reservation-failure flag with the completed prefetch result, count aligned page allocations, and check the existing admission deadline before resubmitting. Capacity failures preserve the checkpoint for a later restore. Existing bitmap-only callers consume the same result through compatibility wrappers.
Regression tests cover both filesystem adapters: seven small pages whose aligned allocations exceed available RAM, and a writer that releases RAM between failed allocation and result consumption. In each case the checkpoint remains listed and restores byte-for-byte after pressure is removed. Both scenarios fail on the integration base and pass with this patch.
Validation: 55 existing checkpoint-storage tests and 8 new parameterized cases passed using real L1 allocation and both filesystem adapters. Nine independent edge-case checks also passed, including actual missing pages, partial contention, deadline expiry and both result APIs. The final image passed a 203-test CPU suite, including 160 forced disk restores within its progress probe. CI code-quality checks passed at
8a2d97b4. GPU-gated suites were not run against production GPUs.This change addresses false checkpoint invalidation after allocation failure. It does not claim to fix an independently observed storage prefetch that stays pending through cancellation.