Skip to content

fix(checkpoint): a restore stuck in storage misses instead of killing the engine - #106

Merged
voipmonitor merged 2 commits into
integration/local-inference-labfrom
fix/checkpoint-cancel-drain
Sep 29, 2026
Merged

voipmonitor merged 2 commits into
integration/local-inference-labfrom
fix/checkpoint-cancel-drain

Conversation

@voipmonitor

@voipmonitor voipmonitor commented Sep 29, 2026 •

Copy link
Copy Markdown

Problem

ktsaou (Discord, Sep 29): GLM-5.3-Flash TP2 with LMCache (beta karmic-kraken-beta-20260928-6bd83bf0760b8115, LMCache ab11b84, 140 GiB RAM tier, disk tier with on-evict writes) died twice with the same signature. Both ranks waited 30 s for a 224,528-token checkpoint restore, cancelled it, and the cancellation did not drain:

LMCacheTimeoutError: Checkpoint lookup cancellation did not drain          (checkpoint_transfer.py:269)
LMCacheTimeoutError: Cancelled checkpoint storage lookup did not drain     (checkpoint_transfer.py:217)
UnsafeCheckpointCopyError: Cancelled checkpoint lease ownership could not drain  (checkpoint_transfer.py:287)
  -> recurrent_checkpoint_connector.py:354 re-raises it in the model worker -> EngineCore fatal error
RuntimeError: Checkpoint worker copy leases must drain before close        (modules/checkpoint.py:375, at shutdown)

The current integration head (6f7a45a, #104) has the same code path.

Root cause

Why the lookup stays pending and ignores the cancel. A restore's server-side lease waits for a storage prefetch (CheckpointPayloadStore.poll_retrieve returns pending while query_prefetch_status is None). cancel_retrieve only sets lease.cancelled; the flag is honoured only after the prefetch ends, and neither the prefetch controller nor the native connector can cancel a queued request. With fs_native all operations (lookup EXISTS, load GET, store SET, delete) share one FIFO served by num_workers threads (FSConnector passes no lane config to ConnectorBase). With checkpoint_on_evict, one L1 eviction pass asks to write every current checkpoint page it scans (select_chunk_coherent_victims keeps scanning past pages that need a write until it finds enough evictable ones), which after the watermark can be most of L1 (tens to over 100 GiB for a 140 GiB RAM tier), submitted at once. A restore that needs disk pages then queues its EXISTS and GET tiles behind that whole burst, for as long as the disk needs to write it. Measured with the real StorageManager + fs_native + checkpoint_on_evict on the test server (16 GiB RAM tier, O_DIRECT): one pass requested all 3,520 page writes (13.8 GiB) at once; a restore of a 256 MiB disk-only checkpoint took 0.07 s idle and 83-89 s during the burst, ending only when the last write finished, and cancelling its lease after 30 s changed nothing. #107 gives lookups and loads their own workers.

Why that is fatal. A pending lookup has exposed no SHM slots to the worker, so no GPU copy can touch its pages. The worker still treated a cancellation that did not drain within 30 s (plus another 30 s in _discard_uncopied_lease) as unknown copy ownership: it raised UnsafeCheckpointCopyError, set _unsafe (every later submit returns None, close() raises), and the connector re-raises that error in build_connector_worker_meta, killing the engine. The server kept the two cancelled leases (nobody polled them any more) and refused to close.

Fix

  • Worker (checkpoint_transfer.py). A lookup still pending after rpc_timeout is cancelled; the worker waits at most 1 s for it to drain, then reports a miss. The scheduler looks the checkpoint up again or recomputes the prompt, and the worker keeps admitting transfers. A failed cancel or poll of a pending lookup is not fatal either (the poll error is reported as a failed transfer). Leases whose slots were exposed keep the fail-safe behaviour: a copy that cannot drain, a ready lease that cannot be finished, a store lease, and lost lease replies are still UnsafeCheckpointCopyError.
  • Server (checkpoint_storage.py). The payload store owns a cancelled lookup: on each later call (begin/poll/cancel/prepare/reclaim) it releases the locks of cancelled lookups whose storage lookup has ended, so they no longer wait for the 600 s abandoned-lease reclaim. A late poll from the worker still gets the miss once (bounded memory of 4096 released ids). Cancelling an unknown or released lease does nothing. report_status adds retrieve_lookups and cancelled_lookups.
  • Shutdown (modules/checkpoint.py). close() waits only for store leases and retrieve leases with exposed slots; lookups still waiting for storage are logged.
  • Docs: a short "Slow storage" paragraph in docs/source/mp/l2_storage/index.rst.

No vLLM change: the connector that re-raises is lmcache/integration/vllm/recurrent_checkpoint_connector.py, and it keeps re-raising genuinely unsafe errors.

Tests

New tests/v1/multiprocess/test_checkpoint_cancel_drain.py (7 tests): a client whose lookup never drains after cancel, a failing cancel (error and timeout), a failing poll, a real CheckpointModule over MQ whose prefetch stays pending and answers late (the worker misses, the server releases the lease when storage answers, the next restore copies, pages are unlocked, the checkpoint stays listed), a late poll of a lookup another call released, and close() with pending lookups vs an exposed lease.

base 6f7a45a this PR
new tests 7 failed (UnsafeCheckpointCopyError, must drain before close) 7 passed
checkpoint suites (tests/v1/multiprocess/test_checkpoint_*.py, test_checkpoint_retention.py, test_vllm_semantic_checkpoint_transfer.py) 177 passed, 2 skipped 177 + 7 passed, 2 skipped

GPU reproduction (GLM-5.3-Flash-NVFP4-Spark TP2, test server, GPUs 5-8)

Image karmic-kraken-beta-20260928-6bd83bf0760b8115 with LMCache 6f7a45a (base) or this branch (fix) mounted over it; L1 24 GB, disk tier on-evict; a test-only fault injection makes a checkpoint prefetch report "in progress" for 150 s (like a lookup queued behind a write burst; cancelling does not end it). Conversation A (100k tokens) is stored, evicted from the GPU, then continued while the stall is armed.

base fix
continuation A2 (restore stuck in storage) HTTP 500 after 90.4 s, EngineCore encountered a fatal error 200 after 136 s (4 lookups × 31 s, then recompute)
worker log Checkpoint lookup cancellation did not drain → Cancelled checkpoint storage lookup did not drain → UnsafeCheckpointCopyError: Cancelled checkpoint lease ownership could not drain still waiting for storage after 30 s; cancelling it → the LMCache server releases it once storage answers
unrelated request during the stall 0.4 s 0.4 s
later restore of A (A3) engine dead 100,065 of 100,066 tokens restored from LMCache in 0.3 s
cancelled lookups released by the server when storage answered – 6 of 8 (2 were still stalled at shutdown)
shutdown Checkpoint worker copy leases must drain before close (0 store, 2 retrieve) Closing with 2 checkpoint lookups still waiting for storage (2 cancelled); no worker received their pages

Notes for reviewers

  • After a miss the scheduler looks the prefix up again (up to 4 attempts) and finds the same checkpoint, so a restore that stays stuck costs about 4 × 31 s before the prompt is recomputed. It cannot tell a slow disk from a miss (results is a bool per rank); making it do so needs a worker metadata change and is left for a follow-up.
  • Release fragment: .lil/changes/lmcache-checkpoint-cancel-drain.json.

If applicable:

  • this PR contains user facing changes - docs added
  • this PR contains unit tests

🤖 Generated with Claude Code

… the engine

A retrieve lookup that storage has not answered 30 s after it began is
cancelled. When storage still had not answered 30 s after the cancel
(for example a lookup queued behind a large burst of on-evict disk
writes), the worker raised UnsafeCheckpointCopyError ("Checkpoint lookup
cancellation did not drain" -> "Cancelled checkpoint lease ownership
could not drain"), stopped admitting transfers and the engine died. The
server then refused to close with "Checkpoint worker copy leases must
drain before close".

A pending lookup has exposed no SHM slots to its worker, so no GPU copy
can touch its pages; only the server's locks are at stake. Now:

- The worker waits at most a second for a cancelled lookup to drain,
  then reports a miss (the request looks the checkpoint up again or
  recomputes) and keeps admitting transfers. A failed cancel or poll of
  a pending lookup is not fatal either. Copies whose slots were exposed
  keep the fail-safe behaviour.
- The payload store owns a cancelled lookup: it releases its locks on
  its next call after the storage lookup ends, and a late poll from its
  worker still gets the miss once. Cancelling an unknown or released
  lease does nothing.
- Closing the checkpoint module waits only for store leases and
  retrieve leases with exposed slots; lookups still waiting for storage
  are logged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (1)
  • dev/*

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 4f9d2950-b577-4ebf-b8d5-99ac84271053

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@voipmonitor

Copy link
Copy Markdown
Author

Reviewed: the worker abandons only a cancelled lookup that exposed no slots, and returns a miss. The payload store keeps the cancelled lease and releases its read locks on its next call once storage answers. Unsafe handling stays for leases whose slots reached the GPU copy, and close() waits only for those. Unit: 7 new cancel/drain tests fail on 6f7a45a and pass here; checkpoint suites 184 passed. GPU (.4, GLM-5.3 TP2, stalled 150 s lookup): the base engine dies with the reported signature; with this PR the request recomputes, the conversation later restores 100,065/100,066 tokens and shutdown is clean. Merging. Follow-ups: the disk-lane change (#107) is held, and so is capping the on-evict write burst.

@voipmonitor
voipmonitor merged commit 75f2b59 into integration/local-inference-lab Sep 29, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant