fix(checkpoint): a restore stuck in storage misses instead of killing the engine - #106
Conversation
… the engine
A retrieve lookup that storage has not answered 30 s after it began is
cancelled. When storage still had not answered 30 s after the cancel
(for example a lookup queued behind a large burst of on-evict disk
writes), the worker raised UnsafeCheckpointCopyError ("Checkpoint lookup
cancellation did not drain" -> "Cancelled checkpoint lease ownership
could not drain"), stopped admitting transfers and the engine died. The
server then refused to close with "Checkpoint worker copy leases must
drain before close".
A pending lookup has exposed no SHM slots to its worker, so no GPU copy
can touch its pages; only the server's locks are at stake. Now:
- The worker waits at most a second for a cancelled lookup to drain,
then reports a miss (the request looks the checkpoint up again or
recomputes) and keeps admitting transfers. A failed cancel or poll of
a pending lookup is not fatal either. Copies whose slots were exposed
keep the fail-safe behaviour.
- The payload store owns a cancelled lookup: it releases its locks on
its next call after the storage lookup ends, and a late poll from its
worker still gets the miss once. Cancelling an unknown or released
lease does nothing.
- Closing the checkpoint module waits only for store leases and
retrieve leases with exposed slots; lookups still waiting for storage
are logged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. 🗂️ Base branches to auto review (1)
Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Reviewed: the worker abandons only a cancelled lookup that exposed no slots, and returns a miss. The payload store keeps the cancelled lease and releases its read locks on its next call once storage answers. Unsafe handling stays for leases whose slots reached the GPU copy, and close() waits only for those. Unit: 7 new cancel/drain tests fail on 6f7a45a and pass here; checkpoint suites 184 passed. GPU (.4, GLM-5.3 TP2, stalled 150 s lookup): the base engine dies with the reported signature; with this PR the request recomputes, the conversation later restores 100,065/100,066 tokens and shutdown is clean. Merging. Follow-ups: the disk-lane change (#107) is held, and so is capping the on-evict write burst. |
Problem
ktsaou (Discord, Sep 29): GLM-5.3-Flash TP2 with LMCache (beta
karmic-kraken-beta-20260928-6bd83bf0760b8115, LMCache ab11b84, 140 GiB RAM tier, disk tier withon-evictwrites) died twice with the same signature. Both ranks waited 30 s for a 224,528-token checkpoint restore, cancelled it, and the cancellation did not drain:The current integration head (6f7a45a, #104) has the same code path.
Root cause
Why the lookup stays pending and ignores the cancel. A restore's server-side lease waits for a storage prefetch (
CheckpointPayloadStore.poll_retrievereturns pending whilequery_prefetch_statusisNone).cancel_retrieveonly setslease.cancelled; the flag is honoured only after the prefetch ends, and neither the prefetch controller nor the native connector can cancel a queued request. Withfs_nativeall operations (lookupEXISTS, loadGET, storeSET, delete) share one FIFO served bynum_workersthreads (FSConnectorpasses no lane config toConnectorBase). Withcheckpoint_on_evict, one L1 eviction pass asks to write every current checkpoint page it scans (select_chunk_coherent_victimskeeps scanning past pages that need a write until it finds enough evictable ones), which after the watermark can be most of L1 (tens to over 100 GiB for a 140 GiB RAM tier), submitted at once. A restore that needs disk pages then queues itsEXISTSandGETtiles behind that whole burst, for as long as the disk needs to write it. Measured with the realStorageManager+fs_native+checkpoint_on_evicton the test server (16 GiB RAM tier, O_DIRECT): one pass requested all 3,520 page writes (13.8 GiB) at once; a restore of a 256 MiB disk-only checkpoint took 0.07 s idle and 83-89 s during the burst, ending only when the last write finished, and cancelling its lease after 30 s changed nothing. #107 gives lookups and loads their own workers.Why that is fatal. A pending lookup has exposed no SHM slots to the worker, so no GPU copy can touch its pages. The worker still treated a cancellation that did not drain within 30 s (plus another 30 s in
_discard_uncopied_lease) as unknown copy ownership: it raisedUnsafeCheckpointCopyError, set_unsafe(every latersubmitreturnsNone,close()raises), and the connector re-raises that error inbuild_connector_worker_meta, killing the engine. The server kept the two cancelled leases (nobody polled them any more) and refused to close.Fix
checkpoint_transfer.py). A lookup still pending afterrpc_timeoutis cancelled; the worker waits at most 1 s for it to drain, then reports a miss. The scheduler looks the checkpoint up again or recomputes the prompt, and the worker keeps admitting transfers. A failed cancel or poll of a pending lookup is not fatal either (the poll error is reported as a failed transfer). Leases whose slots were exposed keep the fail-safe behaviour: a copy that cannot drain, a ready lease that cannot be finished, a store lease, and lost lease replies are stillUnsafeCheckpointCopyError.checkpoint_storage.py). The payload store owns a cancelled lookup: on each later call (begin/poll/cancel/prepare/reclaim) it releases the locks of cancelled lookups whose storage lookup has ended, so they no longer wait for the 600 s abandoned-lease reclaim. A late poll from the worker still gets the miss once (bounded memory of 4096 released ids). Cancelling an unknown or released lease does nothing.report_statusaddsretrieve_lookupsandcancelled_lookups.modules/checkpoint.py).close()waits only for store leases and retrieve leases with exposed slots; lookups still waiting for storage are logged.docs/source/mp/l2_storage/index.rst.No vLLM change: the connector that re-raises is
lmcache/integration/vllm/recurrent_checkpoint_connector.py, and it keeps re-raising genuinely unsafe errors.Tests
New
tests/v1/multiprocess/test_checkpoint_cancel_drain.py(7 tests): a client whose lookup never drains after cancel, a failing cancel (error and timeout), a failing poll, a realCheckpointModuleover MQ whose prefetch stays pending and answers late (the worker misses, the server releases the lease when storage answers, the next restore copies, pages are unlocked, the checkpoint stays listed), a late poll of a lookup another call released, andclose()with pending lookups vs an exposed lease.UnsafeCheckpointCopyError,must drain before close)tests/v1/multiprocess/test_checkpoint_*.py,test_checkpoint_retention.py,test_vllm_semantic_checkpoint_transfer.py)GPU reproduction (GLM-5.3-Flash-NVFP4-Spark TP2, test server, GPUs 5-8)
Image
karmic-kraken-beta-20260928-6bd83bf0760b8115with LMCache 6f7a45a (base) or this branch (fix) mounted over it; L1 24 GB, disk tier on-evict; a test-only fault injection makes a checkpoint prefetch report "in progress" for 150 s (like a lookup queued behind a write burst; cancelling does not end it). Conversation A (100k tokens) is stored, evicted from the GPU, then continued while the stall is armed.EngineCore encountered a fatal errorCheckpoint lookup cancellation did not drain→Cancelled checkpoint storage lookup did not drain→UnsafeCheckpointCopyError: Cancelled checkpoint lease ownership could not drainstill waiting for storage after 30 s; cancelling it→the LMCache server releases it once storage answersCheckpoint worker copy leases must drain before close (0 store, 2 retrieve)Closing with 2 checkpoint lookups still waiting for storage (2 cancelled); no worker received their pagesNotes for reviewers
resultsis a bool per rank); making it do so needs a worker metadata change and is left for a follow-up..lil/changes/lmcache-checkpoint-cancel-drain.json.If applicable:
🤖 Generated with Claude Code