Skip to content

[MP] Preserve checkpoint integrity during deferred-write eviction - #101

Closed
ktsaou wants to merge 3 commits into
local-inference-lab:integration/local-inference-labfrom
ktsaou:fix/checkpoint-retention
Closed

ktsaou wants to merge 3 commits into
local-inference-lab:integration/local-inference-labfrom
ktsaou:fix/checkpoint-retention

Conversation

@ktsaou

@ktsaou ktsaou commented Sep 28, 2026

Copy link
Copy Markdown

A newer recurrent checkpoint can supersede an older branch while the older manifest remains discoverable. With deferred writes, RAM eviction can then discard the older checkpoint's unique state pages before they reach disk. A later restore finds the attention pages but not the state, falls back through incomplete checkpoints and recomputes the prompt.

Keep checkpoint publication, payload ownership and eviction consistent across the RAM/disk lifecycle. Supersession becomes an eviction preference rather than permission to discard an unwritten retained branch. Cache loss remains allowed under bounded resource pressure: retire affected generations before deliberately reclaiming their last known copy, so they become ordinary misses rather than advertised incomplete restores.

The implementation covers the consequences of that ownership rule:

  • Protect pending publication, active transfers and queued writes until the corresponding work drains; bound deferred persistence by allocated bytes, deduplicate shared pages and retry only missing copies.
  • Coordinate RAM/disk deletion with generation ownership, pending I/O and replica availability. Preserve ownership across asynchronous completion, failure and cancellation.
  • Drain checkpoint writes within the shutdown budget, retire incomplete remaining entries, and reconcile recovered manifests against available payload sizes.
  • Preserve valid disk checkpoints after temporary restore-capacity failures. Retry strictly shorter per-request lookup candidates after failed copies without globally invalidating intact manifests; bounded publication-race retries cannot move beyond the selected candidate.
  • Document the availability-first policy and cover pressure, branching, publication races, slow/failed I/O, deletion ordering, replica selection, truncated recovery and bounded shutdown with regression tests.

Dependencies and review scope:

Validation:

  • The exact combined implementation passed 364 installed-image CPU tests; two GPU-dependent cases were skipped. No source overlays were used for that run. Publication sources are byte-identical to the independently reviewed/tested files.
  • Additional isolated two-rank tests reproduced 53/61 and 112/120 page survival on the prior image under supersession plus live eviction. Both Python FS and native FS adapters restored 61/61 and 120/120 from L2 on the patched image, with byte-pattern verification. These synthetic tests inject supersession marking; they are not replays of private user prompts. Related public reproducer: https://github.com/mrweiner/lmcache-supersession-repro
  • Live GLM TP2 on-evict traffic crossed RAM and disk eviction thresholds. By the 10:25 UTC observation, 34,076 on-evict page writes had persisted with zero write timeouts. Complete mixed-tier retrievals included 44/44 pages: 12 RAM + 32 disk, taking 117–118 ms. The 10:01–10:25 window had no missing-page/failed-restore errors or warnings, and no new service restart.
  • Both controlled original/sibling probes initially reused the exact 40,239-token response endpoint out of 40,241 input tokens. The superseded generation was subsequently retired under pressure; a later probe correctly observed a miss rather than an incomplete advertised restore. This is not a lossless-retention or full cold-disk recovery claim.
  • Changed Python files pass Ruff lint/format; whitespace checks pass. Targeted MyPy retains two reproduced baseline errors. Sphinx HTML rendered the changed page, but the strict full build still reports the same 31 diagnostics as the unchanged baseline, including an existing directive error. The repository-wide clean-doc gate is not claimed.

Draft qualification limits: a full production restore with every page read from disk, orderly restart with active GPU transfers, and sustained physical SSD-write reduction remain unqualified. A roughly 40-second zero-output interval cleared without intervention during the workload; its cause is unresolved. A nearby warned deletion completed before that interval, so no causal attribution is made. These live observations support the fixes but are not a completed production acceptance gate.

@coderabbitai

coderabbitai Bot commented Sep 28, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@voipmonitor

Copy link
Copy Markdown

Thanks for chasing this down. The idea at the core of this PR is right: a listed manifest must never lose the last copy of a page it needs. We reproduced the bug on the current beta with mrweiner's reproducer (GLM-5.3-Flash Spark TP2, L1 24 GB / L2 128 GB, on-evict). R1 (supersede + clean restart) fails with the exact fingerprint, 40239: 13 of 22 pages on both ranks, then the walk-down to 40000 and 36864, then a full recompute.

We are not merging this PR as it stands. Besides the fix, the commit changes a lot of shared storage-engine behaviour for ordinary KV and every adapter. Our review reproduced several regressions from that part:

  1. The store-queue byte cap makes write-through lossy (store_controller.py:444, rejection at :139).
    • Writes above 25% of L1 are dropped silently. This applies to ordinary KV, and to checkpoint pages under the default policy (LMCACHE_L2_CHECKPOINT_WRITES=always). Nothing re-offers them.
    • Probe: 5 checkpoints of 896 KiB into a 4 MiB L1 → 4 of 5 never reached disk.
    • test_restore_after_restart_waits_for_ram_held_by_other_restores[False] fails 3/3 on this branch and passes on base.
  2. A forced /cache/clear can kill the store thread.
    • release_internal_reads now raises on a missing entry (l1_manager.py:354-363). _retry_ready runs outside any try (store_controller.py:729), so _store_loop dies.
    • After that, L2 writes stop for good and on-evict pages can never be evicted.
  3. The dead-worker lease cleanup from fix(checkpoint): release leases and pending generations of dead workers #92/fix(checkpoint): configurable abandoned-lease age and remap retry (follow-up to #92, #93) #99 is reverted.
    • reclaim_abandoned (:587-622) only cancels lookups that were never handed out. A crashed worker's store leases and pinned L1 pages stay taken until LMCache restarts.
  4. Other problems:
    • failed writes retry forever while pinned (no attempt cap);
    • pressure retirement drops checkpoints that are already on disk;
    • the publication budget charges shared pages per rank;
    • dead manifests are never invalidated;
    • startup reconciliation holds the directory lock for about 0.6 ms per manifest (roughly 38 s at the 65,536 cap);
    • exact size matching breaks the raw_block and serde adapters;
    • every L1 deletion, including ordinary KV, now takes the directory lock.

The volume of disk writes is a separate question. test_checkpoint_write_volume shows that with retained branches, on-evict writes as much as write-through. That is exactly the flash wear #85 was added to avoid.

What we merged instead: #102, a focused fix for the loss you, hashspamjam and mrweiner reported.

  • Supersession now marks a sibling checkpoint of the earlier request only when the new prompt is longer than it, i.e. the conversation moved past it.
  • The shutdown flush writes superseded pages after the current ones, within the same budget. Their manifests stay listed, so after a restart they restore instead of failing.
  • On the same setup, the reproducer's R1 cell now restores (40239 tokens, external). All six cells restore, and the shutdown flush takes 0.2 s.

If you still see listed-but-unreadable checkpoints after that fix, the ownership guard (retire the dependent manifests before deleting a page's last copy) would be welcome as its own small PR. The store-queue, pinning and adapter changes should come in separate PRs, each with its own tests, so they can be judged on their own. We are happy to review them.

🤖 Generated with Claude Code

@voipmonitor

Copy link
Copy Markdown

Thanks. Two of this PR's core ideas are now merged in #104:

  • Retire before delete. A checkpoint is retired before the last copy of a page it needs is deleted. For any other loss, the lookup confirms the missing pages with the adapter and falls back to the next shorter checkpoint in the same call.
  • Drain before flush. The engine's last stores finish before the shutdown flush.

#104 also keeps the checkpoint that a side request or sub-agent fork extends. It does this without the storage-engine changes from our review:

  • the store-queue cap;
  • the raising release_internal_reads;
  • the lease-reclaim revert;
  • unbounded pinned retries;
  • startup reconciliation;
  • exact size matching;
  • the directory lock on every L1 delete.

It also does not write retained branches on eviction, which measured about as much as write-through.

End to end on GLM Spark TP2, the beta vs #104 (details in #104):

  • a side request followed by RAM pressure went from 13 of 22 pages to a full restore;
  • a sub-agent fork and docker stop -t 10 while storing went from a failed restore to a clean miss;
  • mrweiner's six cells all restore.

Closing in favour of #104. If you still see K of M pages were readable on a beta that includes #104, please post the log. The storage-engine items are welcome as separate PRs with tests.

🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants