fix(checkpoint): keep side-request branch points, retire lost checkpoints, drain in-flight stores at shutdown - #104
Conversation
…ints, drain stores at shutdown
Side requests. A request whose prompt extends a conversation's latest
checkpoint without being its next turn (a title or summary request that
appends a task to the whole conversation, a sub-agent forked from it)
superseded that checkpoint. RAM pressure then dropped its recurrent-state
pages without a write while it stayed listed, and when the user continued
the conversation its restore failed ('K of M pages were readable') and the
prompt was recomputed. The checkpoint a new prompt continues from, its
longest published ancestor, now stays current until a later prompt moves
past it too. Older ancestors are superseded as before, and superseded pages
are still never written to L2 while serving. A superseded checkpoint that a
prompt had continued from (a fork point) keeps its LRU position for
--checkpoint-supersede-grace-seconds (default 300) instead of being dropped
first, and a lookup that finds a superseded checkpoint makes it current
again.
Lost pages. Before the last copy of a superseded page leaves L1 or L2, the
superseded checkpoints that need it are retired from the directory, so a
lookup misses them cleanly. A lookup also checks the pages of the longest
candidate: when L1 does not hold a page and every L2 adapter confirms it
does not either (new absent_keys(); fs and fs_native stat the object file,
so pages another server sharing the directory wrote still count), the
lookup retires the candidate and returns the next shorter complete one in
the same call. This covers L2 eviction, administrative deletes and pages
not written before a restart, with no startup scan. Ordinary KV deletions
return before any of this, and the directory is touched only when a
checkpoint is actually retired.
Shutdown. At SIGTERM the server now keeps serving checkpoint stores that
are still in flight before its message queue closes, for at most a third
of --checkpoint-shutdown-flush-seconds and 5 s, ending once no store is in
flight, so an engine stopping at the same time can publish the checkpoints
it is copying. The flush then writes current pages first, then superseded
ones, within the rest of the same budget, and logs progress and what it
left. A module that refuses to close (copy leases of a killed worker) no
longer skips the storage manager's close, which used to skip the whole
flush and leave the SHM arena behind.
Retiring checkpoints before their last page copy goes and finishing
in-flight stores before the flush follow ideas from ktsaou's LMCache #101
and blackwell-llm-docker #105.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. 🗂️ Base branches to auto review (1)
Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
… fixes Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The HTTP server stopped the event bus, the coordinator registration and the runtime plugins before serving the engine's last checkpoint stores. Serve them first, so the drain starts as early as possible and its events still reach the bus. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
GPU E2E on GLM-5.3-Flash Spark TP2: published beta
🤖 Generated with Claude Code |
Problem
Three gaps remain in
checkpoint_on_evict(the GLM-5.3-Flash / Qwen3.8 profile default) after #102:K of M pages were readable(exactly the state pages missing) and the prompt was recomputed. fix(checkpoint): keep the checkpoints a branch was made before #102 covered only the branch-before-response and clean-restart cases.CheckpointModule.close()then raised, andMPCacheServer.close()never reached the storage manager: nothing was flushed at all and the SHM arena was left behind.Design
Side requests: keep the checkpoint a new prompt continues from.
--checkpoint-supersede-grace-seconds(default 300) instead of being dropped first. It is still never written.Retire before the last copy goes; lookups skip what is gone.
find()checks the pages of the longest candidate. A page counts as lost only if L1 does not hold it, no L2 residency record lists it, and every L2 adapter confirms the absence through the newL2AdapterInterface.absent_keys(). The default confirms nothing.fsandfs_nativestat the object file, so a page that another server sharing the directory wrote still counts; the mock checks its dict. A lost candidate is retired and the next shorter one is tried in the same call. This covers L2 LRU eviction, admin deletes and restarts with unflushed pages. It needs no startup scan and does no O(index) work: only the candidates a lookup examines are checked.Shutdown: serve in-flight stores, then flush, within one budget.
MPCacheServer.drain_for_shutdown()runs in the HTTP lifespan before the ZMQ server closes. It keeps serving checkpoint RPCs, and admits new stores, until one of these happens:min(5 s, budget / 3)has passed.--checkpoint-shutdown-flush-secondsin this order: current pages, then in-grace superseded pages, then stale superseded ones. It logs progress every 2 s and ends with how many pages it left (current / superseded). Writes that were requested at eviction and never completed are resubmitted.MPCacheServer.close()logs a module that fails to close and still closes the storage manager.stop_grace_period: 60sin the generated Compose files.Measured write-volume impact
The workload drives the real
StorageManagerandCheckpointModule(checkpoint_on_evict,fs_nativeL2) through the same RPC methods the vLLM bridge uses, including its retry after a failed restore:The table shows L2 MB written while serving (plus MB written by the shutdown flush), failed restores (
K of M pages readable), and continuations restored from the checkpoint they extend, out of 180. The working set is about 44 MB, so L1 = 48–64 MB is the edge where it barely fits.For reference, write-through (
always) writes 248 MB for the agent loops and 356 MB with side requests, at any L1 size.Tests
CPU-only, in
ghcr.io/local-inference-lab/vllm:karmic-kraken-beta-20260918-c2f43154efeef1e3.New
test_checkpoint_side_requests.py:3 of 6 pages were readable.test_checkpoint_shutdown_drain.py:absent_keysfor the mock,fsandfs_nativeadapters.Updated
[(8, True)]); without it the old failed-then-retry path still holds.Results
test_clear_delegates_to_management_module.GPU E2E
Relation to #101 / blackwell-llm-docker #105
This PR takes two ideas from ktsaou's PRs:
It does not take the storage-engine changes that our review found regressions in:
release_internal_reads;🤖 Generated with Claude Code