Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
25 changes: 25 additions & 0 deletions .lil/changes/lmcache-104.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
{
"schema": "local-inference-release-change/v1",
"id": "lmcache-104",
"category": "fix",
"summary": "Conversations restore after a title request or sub-agent fork, lost checkpoints are no longer offered, and a stopping container finishes and writes its last checkpoints",
"models": [
"Qwen3.8-Flash-Next",
"GLM-5.3-Flash"
],
"compatibility": "No user action required. New option --checkpoint-supersede-grace-seconds (default 300). --checkpoint-shutdown-flush-seconds now also covers a short wait (at most a third of it and 5 s) for checkpoint stores still in flight when the container stops. Give the container time to stop (docker run --stop-timeout 60 or Compose stop_grace_period: 60s); with Docker's default 10 s the current checkpoints are written first.",
"details": [
"A request that extends a conversation's latest checkpoint without being its next turn, such as a title, summary or follow-up request, or a sub-agent forked from the conversation, marked that checkpoint as superseded. Under RAM pressure its state pages were dropped without a disk write, and when the user continued, the restore failed with 'K of M pages were readable' and recomputed the prompt. The checkpoint a new prompt continues from now stays current until a later prompt moves past it too, so the next turn restores it. Superseded checkpoints are still never written to disk while serving.",
"A fork point that a sub-agent superseded keeps its place in RAM eviction order for 300 s instead of being dropped first, and a lookup that finds it makes it current again.",
"A checkpoint is retired from the directory before the last copy of a page it needs is dropped, and a lookup skips and retires checkpoints whose pages no tier holds (evicted from disk, deleted, or not written before a restart). The next shorter checkpoint is restored in the same lookup instead of after a failed restore. Pages that another server sharing the cache directory wrote still count.",
"When the container stops, the cache keeps serving the checkpoint stores the model is still copying for a few seconds, then writes checkpoints that are only in RAM, current ones first, and logs what it could not write. Before, a store in flight at that moment made the cache skip the whole shutdown write.",
"In a synthetic 12-conversation workload with side requests, failed restores fell from 118-247 per 180 turns to 0. Disk writes grow only where RAM barely holds the active conversations, by at most about one checkpoint per turn."
],
"pull_requests": [
104
],
"evidence": [
"https://github.com/mrweiner/lmcache-supersession-repro"
],
"requires": []
}
9 changes: 9 additions & 0 deletions docs/design/v1/distributed/l2_adapters/overall.md
Original file line number Diff line number Diff line change
Expand Up @@ -528,6 +528,15 @@ existing `close()` (to keep data on disk) path. See
`nixl_store_dynamic_l2_adapter.py` for a reference implementation and
[`nixl_store.md`](nixl_store.md) for design details.

`absent_keys(keys)` returns the keys an adapter can confirm it does not
hold, answered synchronously from cheap local state; the default confirms
nothing. The storage manager treats a recurrent checkpoint page as lost only
when L1 does not hold it and every adapter confirms it absent, and a
checkpoint lookup then retires the checkpoints that need it instead of
offering a restore that fails. The check must see objects that other
processes sharing the backend wrote: `fs` and `fs_native` stat the object
files, and the in-process mock checks its dictionary.

### Native (C++/Rust) Storage Backends

For high-performance backends written in C++ or Rust, use the shared native
Expand Down
2 changes: 1 addition & 1 deletion docs/design/v1/mp_observability/METRICS.md
Original file line number Diff line number Diff line change
Expand Up @@ -471,7 +471,7 @@ the same backend type — same shape as the existing
|---|---|---|---|---|
| `lmcache_mp.l1_memory_usage_bytes` | `lmcache_mp_l1_memory_usage_bytes` | ObservableGauge | `L1Manager.get_memory_usage()` | Bytes currently held in L1 at scrape time |
| `lmcache_mp.l2_usage_bytes` | `lmcache_mp_l2_usage_bytes` | ObservableGauge (attr: `l2_name`) | `StorageManager.get_l2_usages()` (calls `L2AdapterInterface.get_usage().total_bytes_used`) | Per-adapter bytes currently held in L2 at scrape time; one observation per configured adapter. Adapters whose `get_usage()` raises are skipped silently. |
| `lmcache_mp.checkpoint_retention` | `lmcache_mp_checkpoint_retention` | ObservableGauge (attr: `stat`) | `CheckpointRetention.observations()` | Recurrent checkpoint retention: `superseded_checkpoints`, `superseded_pages`, `superseded_pages_tracked`, `l1_superseded_drops`, `l2_superseded_evictions`, `l2_superseded_eviction_bytes`, `write_on_evict_requests`, `write_on_evict_persisted`, `write_on_evict_timeouts`, `write_on_evict_pending`, `l2_resident_pages_tracked`, `l2_checkpoint_bytes`. Counters are cumulative since start. |
| `lmcache_mp.checkpoint_retention` | `lmcache_mp_checkpoint_retention` | ObservableGauge (attr: `stat`) | `CheckpointRetention.observations()` | Recurrent checkpoint retention: `superseded_checkpoints`, `superseded_pages`, `superseded_pages_tracked`, `superseded_pages_in_grace`, `l1_superseded_drops`, `l2_superseded_evictions`, `l2_superseded_eviction_bytes`, `retired_checkpoints`, `restored_checkpoints`, `write_on_evict_requests`, `write_on_evict_persisted`, `write_on_evict_timeouts`, `write_on_evict_pending`, `l2_resident_pages_tracked`, `l2_checkpoint_bytes`. Counters are cumulative since start. |
| `lmcache_mp.num_inflight_l2_stores` | `lmcache_mp_num_inflight_l2_stores` | ObservableGauge (attrs: `l2_name`, `adapter_index`) | `StoreController.get_inflight_count_by_adapter()` | Snapshot of in-flight L2 store tasks grouped by adapter |
| `lmcache_mp.num_inflight_l2_loads` | `lmcache_mp_num_inflight_l2_loads` | ObservableGauge (attrs: `l2_name`, `adapter_index`) | `PrefetchController.get_inflight_load_state_by_adapter()` | Per-adapter count from the same snapshot |
| `lmcache_mp.inflight_load_memory_usage_bytes` | `lmcache_mp_inflight_load_memory_usage_bytes` | ObservableGauge (attrs: `l2_name`, `adapter_index`) | `PrefetchController.get_inflight_load_state_by_adapter()` | Per-adapter reserved bytes from the same snapshot |
Expand Down
19 changes: 15 additions & 4 deletions docs/source/mp/configuration.rst
Original file line number Diff line number Diff line change
Expand Up @@ -493,10 +493,21 @@ Source: ``lmcache/v1/distributed/config.py``
evicted without an L2 copy and a warning is logged.
* - ``--checkpoint-shutdown-flush-seconds``
- ``30``
- With ``checkpoint_on_evict``, how long shutdown waits while current
checkpoint pages still only in L1 are written to L2, so the next
start can restore them. ``0`` skips the flush. Give the process
at least this much time between SIGTERM and SIGKILL.
- Budget of a clean shutdown for recurrent checkpoints. The server
first keeps serving checkpoint stores still in flight (at most a
third of the budget, at most 5 s, and only while stores are in
flight), so an engine stopping at the same time can publish them.
With ``checkpoint_on_evict`` it then writes checkpoint pages still
only in L1 to L2, current pages first, so the next start can restore
them. ``0`` skips both. Give the process at least this much time
between SIGTERM and SIGKILL.
* - ``--checkpoint-supersede-grace-seconds``
- ``300``
- How long a superseded recurrent checkpoint that a later prompt had
continued from (a branch point, such as a turn a sub-agent forked
from) keeps its eviction order before its pages are dropped first.
Superseded pages are never written to L2 while serving either way.
``0`` drops them first at once.
* - ``--l2-prefetch-policy``
- ``default``
- L2 prefetch policy. Determines which adapter loads each key
Expand Down
75 changes: 56 additions & 19 deletions docs/source/mp/l2_storage/index.rst
Original file line number Diff line number Diff line change
Expand Up @@ -450,29 +450,64 @@ auxiliary state are unique to each.

The vLLM integration reports each published checkpoint with its token
sequence and request. When the prompt checkpoint of a new request is
published, the server marks as *superseded* the unique pages of every
shorter checkpoint of that sequence produced by an earlier request, and of
the other checkpoints those requests published. That includes a response
endpoint that the chat template rewrote and the next prompt therefore does
not extend. Checkpoints of the same request and ``instruction``
checkpoints, which other conversations share, are never marked. With every
store policy:

* L1 eviction drops superseded pages before any LRU victim.
* L2 eviction deletes superseded pages before any LRU victim.
published, the longest earlier checkpoint that it extends is where it
continues from. That checkpoint stays current: the new request may be a
side request (a title or summary request that appends a task to the whole
conversation, or a sub-agent forked from it) while the conversation itself
goes on from the same point later. The server marks as *superseded* the
unique pages of every other shorter checkpoint of that sequence produced by
an earlier request (the conversation has moved past them twice), and of the
other checkpoints those requests published that the new prompt is longer
than. That includes a response endpoint that the chat template rewrote and
the next prompt therefore does not extend. A checkpoint at least as long as
the new prompt is kept: the new request branched off before it.
Checkpoints of the same request and ``instruction`` checkpoints, which other
conversations share, are never marked. With every store policy:

* L1 eviction drops stale superseded pages before any LRU victim.
* L2 eviction deletes stale superseded pages before any LRU victim.

A superseded page is stale at once, unless a later prompt had continued from
its checkpoint (a possible branch point, for example the end of a turn that
a sub-agent forked from and ran several turns on). Such pages keep their
normal eviction order for ``--checkpoint-supersede-grace-seconds`` (default
300), so the conversation can still continue from them. A lookup that finds
a superseded checkpoint makes it current again.

With ``--l2-store-policy checkpoint_on_evict`` in addition:

* A current checkpoint page is written to L2 once, when L1 is about to
evict it, instead of on every request. L1 keeps the page until the
write completes (at most ``--checkpoint-write-timeout-seconds``).
* A superseded page is never written to L2.
* Shutdown writes current pages still only in L1, so they can be restored
after a restart.

Superseded manifests stay listed, so a request that branches from an older
turn can still restore it while its pages last; when a page is gone, the
restore falls back to the longest remaining checkpoint.
* A superseded page is never written to L2 while serving, also during its
grace period.
* Shutdown writes pages still only in L1, current pages first, so they can
be restored after a restart.

A checkpoint is never offered once a page it needs is gone. Before the last
copy of a superseded page is deleted from L1 or L2, the superseded
checkpoints that reference it are retired from the directory. A lookup also
checks the pages of the longest candidate against L1 and the L2 inventory:
when a page was lost another way (L2 eviction, an administrative delete, a
page that was not written before a restart), the lookup retires that
candidate and returns the next shorter complete one in the same call. A page
counts as lost only when every L2 adapter confirms it does not hold it;
``fs`` and ``fs_native`` check the object file, so pages another server
sharing the directory wrote still count. With adapters that cannot confirm
an absence, the restore still validates every page and falls back to a
shorter checkpoint.

**Shutdown.** ``--checkpoint-shutdown-flush-seconds`` (default 30) is the
budget of a clean shutdown. At SIGTERM the server first keeps serving
checkpoint stores that are still in flight, so an engine stopping at the
same time can publish the checkpoints it is copying. This takes at most a
third of the budget (and at most 5 s) and ends as soon as no store is in
flight. With ``checkpoint_on_evict`` it then writes pages still only in L1
to L2 within the rest of the budget, current pages first, and logs how many
it could not write. Give the process that much time between SIGTERM and
SIGKILL; with a shorter stop timeout the most valuable pages are written
first, and the checkpoints left incomplete are retired by lookups after the
restart.

**Sizing.** With write-through (``default``) the L2 retention time is about::

Expand All @@ -485,5 +520,7 @@ per conversation that leaves L1, so retention is about::

and superseded pages are reclaimed first. The
``lmcache_mp_checkpoint_retention`` gauge reports, per ``stat``, superseded
pages, L1 drops and L2 evictions of superseded pages, write-on-evict
requests, completions and timeouts, and the checkpoint bytes held in L2.
pages (and those in their grace period), L1 drops and L2 evictions of
superseded pages, retired checkpoints, superseded checkpoints found again,
write-on-evict requests, completions and timeouts, and the checkpoint bytes
held in L2.
Loading
Loading