Skip to content

fix(fs_native): give lookups and loads their own workers - #107

Open
voipmonitor wants to merge 3 commits into
integration/local-inference-labfrom
fix/fs-native-read-lanes
Open

voipmonitor wants to merge 3 commits into
integration/local-inference-labfrom
fix/fs-native-read-lanes

Conversation

@voipmonitor

Copy link
Copy Markdown

Problem

Root cause of the storage stall behind ktsaou's crash (GLM-5.3-Flash TP2, 140 GiB RAM tier, disk tier with on-evict writes): a checkpoint restore that needed disk pages stayed pending for over 90 s. #106 makes that survivable; this PR removes the stall.

fs_native ran every operation (lookup EXISTS, load GET, store SET, delete) in one FIFO served by num_workers threads: FSConnector passed no lane config to ConnectorBase. With checkpoint_on_evict, the first L1 eviction pass above the watermark asks to write every current checkpoint page it scans, and nothing is on disk yet, so that is all of them: select_chunk_coherent_victims keeps scanning past pages that need a write until it finds enough evictable ones. They go to the adapter as one store task split into num_workers tiles, which occupies every worker. A restore that needs disk pages queues its EXISTS and GET tiles behind the whole burst. Cancelling the restore cannot shorten the wait: neither the prefetch controller nor the native queue can cancel a queued request.

Fix

ConnectorBase already supports per-operation worker lanes (Mooncake uses them). FSConnector and LMCacheFSClient now accept per_op_workers, and fs_native defaults to:

  • lookup: 2 workers (a lookup only stats files);
  • retrieve: num_workers;
  • store: max(2, num_workers // 4): fewer concurrent writes leave a saturated disk some room for loads, without writing more slowly in the measurement below;
  • delete: the shared num_workers pool.

per_op_workers in the adapter config overrides the default; {} restores the single shared pool. Docs: docs/source/mp/l2_storage/fs_native.rst.

Measurement (real LMCache components, CPU only)

StorageManager with a 16 GiB RAM tier, fs_native (8 workers, O_DIRECT) on the test server's md RAID5 of 4 NVMe partitions, checkpoint_on_evict, and CheckpointPayloadStore (the server side of a restore). Checkpoint A (64 × 4 MiB) is on disk only. The server then publishes checkpoints to 85 % of RAM: one eviction pass requests all 3,520 page writes (13.8 GiB) at once, and A is restored while they drain (its lease is cancelled after 30 s, like the worker does). Each row is one run.

fs_native workers restore of A, idle restore of A during the burst burst written in
one shared FIFO (before) 0.07 s 88.8 s, miss (0 writes left: it waited for the whole burst; the cancel at 30 s changed nothing) 90.2 s
one shared FIFO (before) 0.07 s 83.5 s, miss (0 writes left) 84.9 s
lanes, 8 store workers 0.07 s 44.8 s miss / 15.1 s ready (all 3,520 writes still queued) 96.2 / 92.5 s
lanes, default (2 store workers) 0.42 s 25.2 s, ready (all writes still queued) 73.5 s
lanes, default (2 store workers) 0.19 s 18.1 s, ready (all writes still queued) 79.7 s

Before, a restore behind a write burst waits as long as the whole burst takes, which grows with the RAM tier (a 140 GiB tier can queue over 100 GiB). With lanes it no longer waits in the queue. On this RAID5, O_DIRECT reads still compete with parity writes on the device, which fewer concurrent stores reduce; other runs gave 2.1 s and 44.6 s with the default lanes and 1.25 s with 1 store worker (86 s burst). A normal NVMe disk should be far less affected by device contention.

Tests

  • tests/v1/storage_backend/test_fs_native_connector.py::test_reads_in_their_own_lanes_do_not_wait_behind_queued_writes: with one shared worker a lookup and a load submitted after a 512 MiB write complete after it; with read lanes they complete before it.
  • tests/v1/distributed/test_l2_adapter_factory.py::TestFSNativeLanes: default lanes, explicit and empty per_op_workers, invalid lanes rejected, and the factory forwards the lanes to the native client.
  • Built lmcache_fs from this branch in the beta image: tests/v1/distributed, tests/v1/storage_backend/test_fs_native_connector.py, tests/v1/multiprocess: 3 failures, the same 3 as on the base (test_engine_passthroughs, test_event_ipc_handle_path, test_qstore: CUDA-only paths in a CPU container). Seven test_async_engine_driven_transfer_context.py failures in one run under concurrent disk load passed on rerun (22/22 alone; the multiprocess suite again showed only the 3 base failures).

Notes for reviewers

  • Independent of fix(checkpoint): a restore stuck in storage misses instead of killing the engine #106, which makes a restore that storage cannot answer in time a miss instead of an engine crash.
  • Not changed here, worth a follow-up: the first on-evict pass above the watermark requests writes for every current page (here 3,520 of 3,520), not just the pages it wants to evict, and it submits them as one store task. No page of that task becomes evictable until the whole task is written (the write-on-evict "persisted" count stayed at 0 for the whole 70-90 s), so RAM cannot free space for new checkpoints during the burst. Requesting at most one pass's eviction target, in smaller store tasks, would bound the burst and avoid writing hot pages that the next turn supersedes.
  • Release fragment: .lil/changes/lmcache-107.json.

If applicable:

  • this PR contains user facing changes - docs added
  • this PR contains unit tests

🤖 Generated with Claude Code

voipmonitor and others added 3 commits September 29, 2026 11:16
The native filesystem connector ran every operation in one FIFO served by
num_workers threads. With checkpoint_on_evict, the first L1 eviction
pass above the watermark asks to write every current checkpoint page it
scans (most of RAM, tens of GiB on a large RAM tier) as one store task
split over all workers. A restore that needed disk pages queued its
lookup (EXISTS) and load (GET) behind that whole burst, for as long as
the disk needed to write it. Cancelling the restore could not shorten
the wait, and the engine gave up on it after 30 s.

ConnectorBase already supports per-operation worker lanes; FSConnector
now accepts them (LMCacheFSClient per_op_workers). fs_native defaults to
2 lookup workers, num_workers load workers and max(2, num_workers // 4)
store workers, so reads never queue behind writes, and fewer concurrent
writes leave a saturated disk some room for them. Deletes keep the
shared pool. per_op_workers in the adapter config overrides the default;
{} restores the single shared pool.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (1)
  • dev/*

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: abf9093b-d7a0-4f53-affc-173119178b27

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant