Skip to content

fix(checkpoint): keep side-request branch points, retire lost checkpoints, drain in-flight stores at shutdown - #104

Merged
voipmonitor merged 3 commits into
integration/local-inference-labfrom
fix/checkpoint-side-requests-shutdown-drain
Sep 29, 2026
Merged

voipmonitor merged 3 commits into
integration/local-inference-labfrom
fix/checkpoint-side-requests-shutdown-drain

Conversation

@voipmonitor

Copy link
Copy Markdown

Problem

Three gaps remain in checkpoint_on_evict (the GLM-5.3-Flash / Qwen3.8 profile default) after #102:

  1. Side requests under RAM pressure. A request whose prompt extends a conversation's latest checkpoint without being its next turn (a title, summary or follow-up request that appends a task to the whole conversation, or a sub-agent forked from it) made that checkpoint an ancestor, so it was superseded. The next L1 eviction pass dropped its recurrent-state pages without a write while the manifest stayed listed. When the user then continued the conversation, the restore failed with K of M pages were readable (exactly the state pages missing) and the prompt was recomputed. fix(checkpoint): keep the checkpoints a branch was made before #102 covered only the branch-before-response and clean-restart cases.
  2. Listed but unreadable. When the last copy of a page went (a superseded drop, L2 eviction, an admin delete, pages not written before a restart), the manifests that need it stayed listed until a restore failed on them.
  3. Shutdown race. On SIGTERM the HTTP server closed the message queue first and the modules afterwards. A checkpoint store still in flight (the engine gets SIGTERM at the same time and may still be copying its last checkpoints) could never finish. CheckpointModule.close() then raised, and MPCacheServer.close() never reached the storage manager: nothing was flushed at all and the SHM arena was left behind.

Design

Side requests: keep the checkpoint a new prompt continues from.

  • When a prompt checkpoint is published, its longest published ancestor from an earlier request (the checkpoint it directly extends) stays current. The other ancestors, and the passed-over siblings of their requests, are superseded as before (fix(checkpoint): keep the checkpoints a branch was made before #102's rule for longer siblings is unchanged).
  • The kept checkpoint is superseded once a later prompt moves past it too, which in a normal conversation is one turn later. A title request no longer strips the conversation's tip: under RAM pressure the tip is written to L2 like any current page, and the next turn restores it. Superseded pages are still never written to L2 while serving.
  • Fork points: a superseded checkpoint that a prompt had continued from keeps its LRU position for --checkpoint-supersede-grace-seconds (default 300) instead of being dropped first. It is still never written.
  • A lookup that finds a superseded checkpoint makes it current again.
  • Rejected after measuring: (a) writing superseded pages evicted during a grace window (what [MP] Preserve checkpoint integrity during deferred-write eviction #101 does for retained branches), which writes about as much as write-through under pressure; (b) pinning in-grace fork points in RAM up to 10% of L1, which cost up to +57 MB per 180 turns even with a roomy L1, because the pinned share is always full of dead main-line checkpoints.

Retire before the last copy goes; lookups skip what is gone.

  • Eager: before L1 or an L2 adapter deletes a superseded page that no other tier holds, the superseded checkpoints that reference it are retired in one SQLite transaction. The key-to-owner map exists only for superseded pages, is bounded, and forgets pages that no tier holds. Ordinary KV and current checkpoint pages return after a model-name prefix check or a dict lookup. The directory lock is taken only when something is retired.
  • At lookup: find() checks the pages of the longest candidate. A page counts as lost only if L1 does not hold it, no L2 residency record lists it, and every L2 adapter confirms the absence through the new L2AdapterInterface.absent_keys(). The default confirms nothing. fs and fs_native stat the object file, so a page that another server sharing the directory wrote still counts; the mock checks its dict. A lost candidate is retired and the next shorter one is tried in the same call. This covers L2 LRU eviction, admin deletes and restarts with unflushed pages. It needs no startup scan and does no O(index) work: only the candidates a lookup examines are checked.

Shutdown: serve in-flight stores, then flush, within one budget.

  • MPCacheServer.drain_for_shutdown() runs in the HTTP lifespan before the ZMQ server closes. It keeps serving checkpoint RPCs, and admits new stores, until one of these happens:
    • no store lease or unpublished generation remains and no store RPC arrived for 0.5 s;
    • no store RPC arrived for 3 s (the producer is gone);
    • min(5 s, budget / 3) has passed.
  • The flush then uses the rest of --checkpoint-shutdown-flush-seconds in this order: current pages, then in-grace superseded pages, then stale superseded ones. It logs progress every 2 s and ends with how many pages it left (current / superseded). Writes that were requested at eviction and never completed are resubmitted.
  • MPCacheServer.close() logs a module that fails to close and still closes the storage manager.
  • This works with Docker's default 10 s. The supervisor sends SIGTERM to the model and to LMCache at the same time, the drain usually ends about 0.5 s after the last store, and current pages go first. blackwell-llm-docker fix(checkpoint): preserve disk checkpoints after L1 allocation failure #105's model-first order is therefore not needed: under the default timeout it leaves LMCache almost no time. The companion docker PR sets stop_grace_period: 60s in the generated Compose files.

Measured write-volume impact

The workload drives the real StorageManager and CheckpointModule (checkpoint_on_evict, fs_native L2) through the same RPC methods the vLLM bridge uses, including its retry after a failed restore:

  • 12 conversations × 16 turns;
  • a 240-token context plus 10 tokens per turn;
  • 64 KiB pages, 9 unique pages per checkpoint (7 recurrent groups, auxiliary page, partial page), as in GLM-5.3-Flash request boundaries.

The table shows L2 MB written while serving (plus MB written by the shutdown flush), failed restores (K of M pages readable), and continuations restored from the checkpoint they extend, out of 180. The working set is about 44 MB, so L1 = 48–64 MB is the edge where it barely fits.

workload L1 base (#102 + #103) this PR
agent loops 32 MB 240 (+8), 0 failed, 180 242 (+7), 0 failed, 180
48 MB 173 (+12), 0 failed, 180 228 (+11), 0 failed, 180
64 MB 0 (+58), 0 failed, 180 97 (+19), 0 failed, 180
96 MB 0 (+94), 0 failed, 180 0.5 (+69), 0 failed, 180
+ side request after half the turns 48 MB 227 (+10), 190 failed, 92 291 (+13), 0 failed, 180
64 MB 169 (+12), 247 failed, 96 224 (+18), 0 failed, 180
96 MB 144 (+25), 225 failed, 118 101 (+69), 0 failed, 180
+ 3-turn sub-agent fork after a quarter of the turns 64 MB 128 (+20), 139 failed, 136 218 (+22), 0 failed, 148
96 MB 78 (+56), 118 failed, 148 88 (+73), 0 failed, 178
chat (template rewrites the answer) + side requests 64 MB 221 (+14), 132 failed, 96 267 (+19), 0 failed, 180
96 MB 160 (+52), 136 failed, 124 149 (+41), 0 failed, 180

For reference, write-through (always) writes 248 MB for the agent loops and 356 MB with side requests, at any L1 size.

  • Cost. Keeping the parent one more turn adds one checkpoint's unique pages per active conversation to RAM. It adds writes only where L1 barely holds the active conversations: at most about one checkpoint's unique pages per turn (about 215 MB for GLM-5.3-Flash). Otherwise, with a roomy L1 or with L1 already overcommitted, it adds nothing, and superseded pages are still never written while serving.
  • Benefit. Side requests and forks no longer cause failed restores. Base failed 118 to 247 restores per 180 turns in those workloads and recomputed or fell back for up to half of the continuations.
  • Forks. A fork point is protected only while it stays in RAM. If RAM pressure evicts it, it is retired and the lookup misses cleanly instead of failing a restore.

Tests

CPU-only, in ghcr.io/local-inference-lab/vllm:karmic-kraken-beta-20260918-c2f43154efeef1e3.

New

  • test_checkpoint_side_requests.py:
    • conversation → side request extending its response → L1 pressure → the conversation continues and restores. This fails on base with 3 of 6 pages were readable.
    • retirement before a superseded page's last copy goes;
    • a fork point is not dropped first during the grace period, and with no grace it is retired cleanly;
    • a lookup retires checkpoints with lost pages and falls back in one call (admin delete; restart without flush);
    • pages another server wrote to a shared directory are trusted;
    • ordinary deletions never reach the directory.
  • test_checkpoint_shutdown_drain.py:
    • a store finishing 0.4 s into shutdown is published and flushed, and restores after the restart;
    • a worker that never finishes: the drain is bounded, close still flushes, and the storage manager is still closed;
    • the drain returns at once when idle;
    • the drain is bounded while stores keep arriving;
    • the HTTP lifespan drains before the message queue closes.
  • Retention unit tests (grace, un-supersede, forgetting, last-copy retirement for L1/L2, absence confirmation); absent_keys for the mock, fs and fs_native adapters.

Updated

  • The two supersession tests now expect the continued-from checkpoint to stay current for one turn.
  • The bridge fallback test is parametrized: with lookup validation the first attempt is already the shorter checkpoint ([(8, True)]); without it the old failed-then-retry path still holds.

Results

GPU E2E

  • Prepared, not yet run: mrweiner's reproducer extended with side-request, fork and stop-while-storing cells on GLM-5.3-Flash Spark TP2.

Relation to #101 / blackwell-llm-docker #105

This PR takes two ideas from ktsaou's PRs:

  • never offering a checkpoint after its last page copy is gone (retired first);
  • finishing the engine's last stores before the flush.

It does not take the storage-engine changes that our review found regressions in:

  • the store-queue byte cap;
  • raising release_internal_reads;
  • the revert of dead-worker lease reclaim;
  • unbounded retries with pinned pages;
  • retiring checkpoints that are on disk;
  • publication budgets;
  • the startup reconciliation scan;
  • exact size matching;
  • the directory lock on every L1 delete;
  • writing retained branches on eviction.

🤖 Generated with Claude Code

…ints, drain stores at shutdown

Side requests. A request whose prompt extends a conversation's latest
checkpoint without being its next turn (a title or summary request that
appends a task to the whole conversation, a sub-agent forked from it)
superseded that checkpoint. RAM pressure then dropped its recurrent-state
pages without a write while it stayed listed, and when the user continued
the conversation its restore failed ('K of M pages were readable') and the
prompt was recomputed. The checkpoint a new prompt continues from, its
longest published ancestor, now stays current until a later prompt moves
past it too. Older ancestors are superseded as before, and superseded pages
are still never written to L2 while serving. A superseded checkpoint that a
prompt had continued from (a fork point) keeps its LRU position for
--checkpoint-supersede-grace-seconds (default 300) instead of being dropped
first, and a lookup that finds a superseded checkpoint makes it current
again.

Lost pages. Before the last copy of a superseded page leaves L1 or L2, the
superseded checkpoints that need it are retired from the directory, so a
lookup misses them cleanly. A lookup also checks the pages of the longest
candidate: when L1 does not hold a page and every L2 adapter confirms it
does not either (new absent_keys(); fs and fs_native stat the object file,
so pages another server sharing the directory wrote still count), the
lookup retires the candidate and returns the next shorter complete one in
the same call. This covers L2 eviction, administrative deletes and pages
not written before a restart, with no startup scan. Ordinary KV deletions
return before any of this, and the directory is touched only when a
checkpoint is actually retired.

Shutdown. At SIGTERM the server now keeps serving checkpoint stores that
are still in flight before its message queue closes, for at most a third
of --checkpoint-shutdown-flush-seconds and 5 s, ending once no store is in
flight, so an engine stopping at the same time can publish the checkpoints
it is copying. The flush then writes current pages first, then superseded
ones, within the rest of the same budget, and logs progress and what it
left. A module that refuses to close (copy leases of a killed worker) no
longer skips the storage manager's close, which used to skip the whole
flush and leave the SHM arena behind.

Retiring checkpoints before their last page copy goes and finishing
in-flight stores before the flush follow ideas from ktsaou's LMCache #101
and blackwell-llm-docker #105.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (1)
  • dev/*

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 2438fc54-5b1f-445c-b9bc-aaef3aadb2ec

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

voipmonitor and others added 2 commits September 28, 2026 22:04
… fixes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The HTTP server stopped the event bus, the coordinator registration and the
runtime plugins before serving the engine's last checkpoint stores. Serve
them first, so the drain starts as early as possible and its events still
reach the bus.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@voipmonitor

Copy link
Copy Markdown
Author

GPU E2E on GLM-5.3-Flash Spark TP2: published beta 6bd83bf0 (LMCache ab11b84) vs the same image with this PR's LMCache files. LMCache ran with L1 24 GB, L2 128 GB and on-evict. The cells are mrweiner's reproducer plus side-request, fork and stop-while-storing cells.

cell beta this PR
S1: side request, then L1 pressure failed (40239: 13 of 22 pages, walk-down 40000/36864 also 13/22 and 12/20) restored (40239)
S2: side request, then clean restart restored restored
F1: sub-agent fork, then pressure failed (13 of 22) clean miss; the manifest was retired, so there was no failed restore
K1: docker stop -t 10 while 8 clients store failed (0 of 22) clean miss (lookup retired it; store drain 0.2 s)
K2: docker stop -t 60 while 8 clients store restored restored (drain, then flush 0.6 s)
mrweiner's P0–P3, R0–R1 all restored all restored

🤖 Generated with Claude Code

@voipmonitor
voipmonitor merged commit 6f7a45a into integration/local-inference-lab Sep 29, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant