Skip to content

Delayed production CUDA OOM on ds4-vision-jovian-judgement-r9: long-context prefill exceeds profiled peak by ~3 GiB #103

Description

@Seth777777777

Short version: the known "delayed production OOM" reproduced a third time (11.09, 13.09, 15.09), this time with PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True already applied. It is not fragmentation: at the moment of death only ~100 MiB was reserved-but-unallocated, i.e. essentially all of the 93.46 GiB was live.

Two hard numbers:

  • The allocation that failed is a routine 128 MiB torch.empty for the output of the query projection wq_b in deepseek_v4/attention.py:544 (DeepGEMM block-scaled MM). Nothing pathological about the victim — memory was already gone when it asked.
  • At that moment torch held ~3.0 GiB more than the startup memory profiler predicted as peak (93.46 GiB allocated vs. 90.49 GiB predicted), on a card where the profiler's own accounting leaves only ~2.5 GiB of headroom.

So the profiler's peak-activation figure (1.24 GiB) does not bound what a real production step actually allocates, and the safety margin it leaves is smaller than the error. The step that died was a chunked prefill with 149 760 tokens of KV context behind it; we suspect something whose size scales with KV length rather than with chunk size, but the traceback does not prove that — see §6, where we also note that similar long-context prefills survive routinely on this box, so the trigger is evidently narrower than "long context alone".

Full log (crash + automatic restart + 30 min of subsequent normal traffic) attached.


1. Environment

Version string 0.26.1rc0+jovian.judgement.cu133.r9.vllmf66599d.b12x15b6813
Model deepseek-ai/DeepSeek-V4-Flash-Vision-Exp, rev 6821d6ad3681a4b137b066b76094fa82ebd0a380
GPUs 2× RTX PRO 6000 Blackwell, 94.97 GiB each, SM120, PCIe (no NVLink), TP=2
Host Ubuntu 24.04, kernel 7.0.0-28 HWE, iommu=pt
All-reduce B12X_PCIE (oneshot) + PYNCCL, nccl 2.31.2
Attention / MoE --attention-backend B12X, --moe-backend b12x, fp8 KV (fp8_ds_mla), FP8 indexer cache
Container ds4-vision-jovian-judgement-r9, restarted automatically after the crash

Minor thing worth confirming: the version string says jovian.judgement…r9, but every traceback path is under /opt/infernal-invocation/vllm/…. We assume this is just a leftover install prefix in the image and not an r15/r9 mix-up — please confirm.

2. Configuration

Launch line as logged:

DS4 launch: variant=vision mode=dspark depth=fixed backend=b12x-a8-dglin allreduce=b12x tp=2 dcp=1
max_seqs=4 graph=16 load_format=instanttensor instanttensor_backend=BUFFERED
lmcache_transfer=engine_driven direct_lmcache=0 lmcache_memory_profile=standard native_l2=0
allocator=expandable_segments:True model=deepseek-ai/DeepSeek-V4-Flash-Vision-Exp

Relevant server flags:

--tensor-parallel-size 2 --decode-context-parallel-size 1
--gpu-memory-utilization 0.975            # our compose override leaves this on the release default
--kv-cache-dtype fp8 --block-size 256
--max-model-len 1048576 --max-num-seqs 4 --max-num-batched-tokens 4096
--enable-chunked-prefill --enable-prefix-caching --prefix-cache-retention-interval 4096
--async-scheduling --no-scheduler-reserve-full-isl
--max-cudagraph-capture-size 16 --compilation-config {"cudagraph_mode":"FULL_AND_PIECEWISE","custom_ops":["all"]}
--speculative-config {"model":"…","method":"dspark","num_speculative_tokens":3,
                      "draft_sample_method":"probabilistic","rejection_sample_method":"standard"}

Everything else is stock r9. MAX_MODEL_LEN and GPU_MEMORY_UTILIZATION are deliberately left at the release defaults (we tried pinning GPU_MEMORY_UTILIZATION earlier and reverted).

3. History

Date Uptime at crash Config at the time Result
11.09.2026 ~2nd day stock r9 CUDA OOM during long prefill
13.09.2026 ~2nd day stock r9 CUDA OOM, same signature
15.09.2026 09:38:18 ~2nd day + expandable_segments:True CUDA OOM, same signature

After 13.09 we added only PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True (per the note in the r9 page) and left GPU memory utilisation on automatic KV admission. The workaround did not change the outcome or the interval.

4. The failing step

Single request in flight, no queue, chunked prefill, multimodal:

num_running_reqs=1, num_waiting_reqs=0
num_scheduled_tokens={chatcmpl-b111f33f24ae489f-b9da62cc: 4096}
num_computed_tokens=149760
PrefixCacheStats(requests=1, queries=156219, hits=149760)
kv_cache_usage=0.3824
num_common_prefix_blocks=[602, 0, 0, 0, 0]
num_spec_tokens_to_schedule=3, scheduled_spec_decode_tokens={}
…mm_hash='f4d34022e9880f192546ef1767c1657e745bb3709bf7af31837479f4b5a70fec'
sampling_params=SamplingParams(…, max_tokens=892357, …)

The dump also shows several MultiModalFeatureSpec(modality='image', …) entries with distinct mm_hashes, e.g. one with mm_position=PlaceholderRange(offset=135238, length=355). So the images are interleaved deep inside the context, not just at the head — relevant because this build forces --disable_chunked_mm_input for multimodal-bidirectional attention.

So: a ~156k-token prompt containing several images, 95.9 % prefix-cache hit, being prefilled in 4096-token chunks. KV context at the failing step = 149760 + 4096 = 153 856 tokens. KV cache itself was only 38 % full — this is not a KV-capacity problem.

Error:

RuntimeError: Worker failed with error 'CUDA out of memory. Tried to allocate 128.00 MiB.
GPU 0 has a total capacity of 94.97 GiB of which 100.19 MiB is free.
Including non-PyTorch memory, this process has 94.86 GiB memory in use.
Of the allocated memory 93.46 GiB is allocated by PyTorch, with 19.40 MiB allocated in
private pools (e.g., CUDA Graphs), and 100.28 MiB is reserved by PyTorch but unallocated.'

Both ranks then log:

[rank0] expandable_segments: memory mapping failed with OOM on device 0 while trying to map
        20971520 bytes (free: 19070976, total: 101973491712)
[rank1] … (free: 16973824, …)

Two things follow from this:

  • Only 100 MiB reserved-but-unallocated. There was nothing for the allocator to reclaim — this is exhaustion, not fragmentation, which is why expandable_segments could not save it. The segment mapper itself failed to map 20 MiB against 18 MiB of free device memory.
  • Both GPUs hit the wall within microseconds of each other, i.e. it is symmetric TP work, not a rank-local artefact.

Worker-side traceback — the actual allocation site

Both ranks die in the same place, in the model forward, at the very start of attention:

File "vllm/v1/worker/gpu/model_runner.py", line 1911, in execute_model
  model_output = self.model(**model_inputs)
File "vllm/models/deepseek_v4/nvidia/vl_model.py", line 311, in forward
  return self.language_model(
File "vllm/models/deepseek_v4/nvidia/model.py", line 1548, in forward
  hidden_states, residual, post_mix, res_mix = layer(
File "vllm/models/deepseek_v4/nvidia/model.py", line 1367, in forward
  x = self.attn(positions, x, None)
File "vllm/models/deepseek_v4/attention.py", line 585, in _prepare_and_attn
  q = project_query_and_cache_kv()
File "vllm/models/deepseek_v4/attention.py", line 544, in project_query_and_cache_kv
  q = self.wq_b(qr).view(-1, self.n_local_heads, self.head_dim)
File "vllm/model_executor/kernels/linear/scaled_mm/deep_gemm.py", line 117, in apply_block_scaled_mm
  output = torch.empty(
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 128.00 MiB. GPU 1 …

128 MiB is exactly the expected q-projection output for a 4096-token chunk, so this is an ordinary allocation that happened to arrive when the card was already full — it identifies the moment, not the culprit. Note that the whole per-step attention work (indexer, sparse MLA) happens after this point, so whatever the excess is, it was already resident before the failing layer's attention began.

The allocator was thrashing at the edge for a few milliseconds before giving up — the earliest mapping failures report only 8.2–10.2 MiB free:

[rank0] … expandable_segments: memory mapping failed with OOM on device 0 while trying to map
        20971520 bytes (free: 10682368, total: 101973491712).
[rank1] … (free: 8585216, …)

5. Memory accounting: profiled vs. actual

From the restart of the very same container (same flags, same weights):

Model loading took 81.11 GiB
Available KV cache memory: 7.94 GiB   → GPU KV cache size: 1,196,643 tokens
Graph capturing … took 0.20 GiB
"Actual usage is 83.1 GiB for consumed memory (weights + non-torch),
 1.24 GiB for peak activation, and 0.2 GiB for CUDAGraph memory."
GiB
Weights (torch) 81.11
KV cache 7.94
CUDA graph pool 0.20
Steady-state subtotal 89.25
Peak activation as profiled 1.24
Predicted peak (torch) 90.49
Actually allocated by torch at crash 93.46
Under-estimate ≈ 2.97

Headroom between the profiled peak and the physical card (incl. non-torch ~1.4–2 GiB) is ≈ 2.5 GiB. The step needed ~3 GiB more than predicted, so it died — with mathematical certainty, every time this input shape occurs.

6. Hypothesis: the profiler never sees a long-context prefill step

The memory profiling run uses max_num_batched_tokens tokens against an empty KV context, so it measures a 4096 × 4096 step. The step that crashed is 4096 × 153 856 — a 37.6× larger KV extent. Any workspace whose size is O(chunk × kv_len) is therefore under-profiled by that factor.

The obvious candidate in this build is the DSA / Lightning Indexer score+top-k buffer (indexer.py:581 DSA indexer decode path: use_flattening=True supports_varlen=False (next_n=4, use_fp4_cache=False); Using FP8 indexer cache for Lightning Indexer). Order-of-magnitude for a 4096 × 153 856 score tensor:

dtype size
fp32 2.35 GiB
bf16 1.17 GiB
fp8 0.59 GiB

An fp32 score buffer plus an index/top-k copy would land close to the observed ~3 GiB gap. We want to be explicit that this is arithmetic, not evidence — the traceback does not name the indexer, and we have not instrumented the build. Treat it as the first place we would look, not as a diagnosis.

Counter-evidence we should not hide. The attached log covers 8 minutes before the crash and 30 minutes after it. In that window, long-context chunked prefills of the same shape happen several times per hour and complete normally — e.g. Avg prompt throughput: 8574.1 tokens/s … GPU KV cache usage: 41.7% and 5086.5 tokens/s … 40.4%, both after the restart, plus 7892.5 tokens/s before it. So a ~150k-token context with a big prefix-cache hit is clearly not sufficient on its own to kill the engine — something narrower tips it over. Candidates we cannot separate from our side:

  • The failing request carried several images interleaved deep in the context (one at offset 135238). With --disable_chunked_mm_input forced, an image's 355 tokens must be scheduled inside one chunk; a chunk that carries image tokens and sits at 150k KV may be the specific bad combination. MM cache hit rate was 96.8 %, so on most of those surviving prefills the image work was served from cache.
  • --async-scheduling lets step n+1 be prepared while step n is still resident. With ~2.5 GiB of headroom, an overlap of two heavy steps would compound anything above.
  • Eagle3/DSpark auxiliary hidden states (layers 41, 42, 43) are retained across the forward pass for the drafter.

What we are confident about, and what we think is worth fixing regardless of which of these it is: the profiler's 1.24 GiB peak-activation figure is not an upper bound for production steps, and the margin it leaves (~2.5 GiB) is smaller than the observed error (~3.0 GiB). Any config that passes startup profiling can therefore still OOM later.

7. Suggested reproduction

We have not reproduced it on demand (see the counter-evidence in §6 — most long-context prefills survive). What we would try first, on a 2× 96 GB SM120 box with the r9 vision image and stock flags:

  1. Start with --max-num-batched-tokens 4096, --enable-chunked-prefill, --enable-prefix-caching, --gpu-memory-utilization 0.975.
  2. Send a ~150 000-token prompt (ideally with one image, to match ours) and let it finish so the prefix is cached.
  3. Re-send the same prompt plus a few thousand new tokens, so the scheduler resumes at num_computed_tokens ≈ 150k and prefills 4096 new tokens against that context.
  4. Include 2–3 images placed deep in the prompt (e.g. around token 135k) and make sure they are not in the MM cache, so the encoder runs in the same step as a 4096-token chunk at high KV occupancy.
  5. Watch torch.cuda.memory_allocated() at the step where num_computed_tokens is largest, rather than waiting for the OOM.

If it survives, bisect upward on prompt length and on the number of uncached images. Even a run that does not crash is useful to you: if memory_allocated() at 150k context exceeds the profiled peak by a gigabyte or more, the profiling gap is confirmed independently of what our production traffic did.

8. Questions

  1. Is the indexer / sparse-MLA per-step workspace accounted for in the startup memory profiler? If not, can the profiling run be made to profile a chunk against a max_model_len-sized KV extent (or the workspace be reserved up front) so that the KV admission figure is actually safe?
  2. Can the indexer score/top-k be tiled over the KV dimension so peak memory is O(chunk × tile) instead of O(chunk × kv_len)?
  3. Is the indexer score buffer fp32 in this build? If so, is bf16/fp8 viable there?
  4. Does --async-scheduling allow two long-context prefill steps' activations to be resident simultaneously?
  5. Is a chunk that carries uncached image tokens while sitting at a very large KV offset a known-heavy combination in this build, given that --disable_chunked_mm_input is forced for multimodal-bidirectional attention?
  6. Is there a recommended max_num_batched_tokens ceiling as a function of max_model_len for the vision checkpoint? The vision weights (81.11 GiB) leave far less headroom than the text-only checkpoint, so the safe operating point is presumably different.

9. What we are doing on our side meanwhile

  • Halving --max-num-batched-tokens to 2048 — halves anything that scales as chunk × kv_len, at the cost of slower long prefills.
  • Pinning --kv-cache-memory explicitly (rather than 0.975 auto-admission) to buy ~3 GiB of headroom; our KV usage peaked at 38 % of 1.19 M tokens, and the surplus is mostly evictable prefix-cache blocks.
  • Keeping expandable_segments:True (harmless, just not sufficient).
  • If it happens again with those two in place, dropping --async-scheduling as the next test.
  • Continuing the per-minute nvidia-smi memory sampling so that if the next crash is a gradual leak rather than a single oversized step, we will have the curve to prove it.

The attached log covers 09:30:09–10:08:37 — 8 minutes of normal traffic, the crash, the automatic restart with its full config and memory accounting, and 30 minutes of normal traffic afterwards. The only thing removed is 246 [dump_input.py:79] lines (multimodal is_embed masks and the 8.5k-entry block table, ~250 KB of tensor text); the removal point is marked inline and we can send that block raw if you want it. Happy to run any instrumented build or extra diagnostics on this host.


Appendix A — log extracts

The full container log is 970 lines (09:30:09–10:08:37). Below are the parts that matter.
Ask and we will send the whole file, including the 246 [dump_input.py:79] lines
(multimodal is_embed masks + the 8.5k-entry block table) omitted throughout this report.

A1. Run-up to the crash

A long prompt is admitted at 09:35:39 (KV jumps to 20.3 %), prefills at 7892 tok/s and
completes normally. Two minutes later the engine dies; note the 09:38:09 → 09:40:37 gap in
the stats series, which is the crash plus the automatic restart.

09-15 09:35:29  Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 93.4%, MM cache hit rate: 96.9%
09-15 09:35:39  Avg prompt throughput: 48.1 tokens/s, Avg generation throughput: 5.4 tokens/s, Running: 2 reqs, Waiting: 0 reqs, GPU KV cache usage: 20.3%, Prefix cache hit rate: 93.2%, MM cache hit rate: 96.8%
09-15 09:35:49  Avg prompt throughput: 7892.5 tokens/s, Avg generation throughput: 61.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.0%, Prefix cache hit rate: 93.2%, MM cache hit rate: 96.8%
09-15 09:35:59  Avg prompt throughput: 3286.5 tokens/s, Avg generation throughput: 126.0 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.6%, Prefix cache hit rate: 93.2%, MM cache hit rate: 96.8%
09-15 09:38:09  Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 296.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 7.5%, Prefix cache hit rate: 93.2%, MM cache hit rate: 96.8%
09-15 09:40:37  Avg prompt throughput: 515.7 tokens/s, Avg generation throughput: 2.6 tokens/s, Running: 2 reqs, Waiting: 0 reqs, GPU KV cache usage: 41.7%, Prefix cache hit rate: 0.0%

A2. Scheduler stats at the failing step

[dump_input.py:81] Dumping scheduler stats: SchedulerStats(num_running_reqs=1, num_waiting_reqs=0, num_skipped_waiting_reqs=0, step_counter=0, current_wave=0, kv_cache_usage=0.3824219666744896, iteration_details=None, prefix_cache_stats=PrefixCacheStats(reset=False, requests=1, queries=156219, hits=149760, preempted_requests=0, preempted_queries=0, preempted_hits=0), connector_prefix_cache_stats=None, kv_cache_eviction_events=[], spec_decoding_stats=None, kv_connector_stats=None, waiting_lora_adapters={}, running_lora_adapters={}, cudagraph_stats=None, perf_stats=None, decode_compute_seconds=0.0, prefill_compute_seconds=0.0, prefill_compute_share=0.0, scheduled_prefill_tokens=0, active_partial_prefills=0, decode_only_steps=0, fairness_bypasses=0)

A3. Worker traceback — the actual allocation site (both ranks identical)

WorkerProc hit an exception.
Traceback (most recent call last):
  File "/opt/infernal-invocation/vllm/vllm/v1/executor/multiproc_executor.py", line 1047, in _execute_worker_rpc
    output = func(*args, **kwargs)
             ^^^^^^^^^^^^^^^^^^^^^
  File "/opt/infernal-invocation/vllm/vllm/v1/worker/worker_base.py", line 369, in execute_model
    return self.worker.execute_model(scheduler_output)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/utils/_contextlib.py", line 124, in decorate_context
    return func(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^
  File "/opt/infernal-invocation/vllm/vllm/v1/worker/gpu_worker.py", line 1213, in execute_model
    output = self.model_runner.execute_model(
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/infernal-invocation/vllm/vllm/v1/worker/gpu/model_runner.py", line 1911, in execute_model
    model_output = self.model(**model_inputs)
                   ^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1778, in _wrapped_call_impl
    return self._call_impl(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/dist-packages/torch/nn/modules/module.py", line 1789, in _call_impl
    return forward_call(*args, **kwargs)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/infernal-invocation/vllm/vllm/models/deepseek_v4/nvidia/vl_model.py", line 311, in forward
    return self.language_model(
           ^^^^^^^^^^^^^^^^^^^^
  File "/opt/infernal-invocation/vllm/vllm/models/deepseek_v4/nvidia/model.py", line 1956, in forward
    hidden_states = self.model(
                    ^^^^^^^^^^^
  File "/opt/infernal-invocation/vllm/vllm/models/deepseek_v4/nvidia/model.py", line 1548, in forward
    hidden_states, residual, post_mix, res_mix = layer(
                                                 ^^^^^^
  File "/opt/infernal-invocation/vllm/vllm/models/deepseek_v4/nvidia/model.py", line 1367, in forward
    x = self.attn(positions, x, None)
        ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/infernal-invocation/vllm/vllm/models/deepseek_v4/attention.py", line 480, in forward
    self._prepare_and_attn_fn(
  File "/opt/infernal-invocation/vllm/vllm/models/deepseek_v4/attention.py", line 585, in _prepare_and_attn
    q = project_query_and_cache_kv()
        ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/infernal-invocation/vllm/vllm/models/deepseek_v4/attention.py", line 544, in project_query_and_cache_kv
    q = self.wq_b(qr).view(-1, self.n_local_heads, self.head_dim)
        ^^^^^^^^^^^^^
  File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/linear.py", line 587, in forward
    output_parallel = self.quant_method.apply(self, input_, bias)
                      ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/infernal-invocation/vllm/vllm/model_executor/layers/quantization/fp8.py", line 461, in apply
    return self.fp8_linear.apply_weights(layer, x, bias)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/infernal-invocation/vllm/vllm/model_executor/kernels/linear/scaled_mm/BlockScaledMMLinearKernel.py", line 132, in apply_weights
    output = self.apply_block_scaled_mm(
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/infernal-invocation/vllm/vllm/model_executor/kernels/linear/scaled_mm/deep_gemm.py", line 117, in apply_block_scaled_mm
    output = torch.empty(
             ^^^^^^^^^^^^
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 128.00 MiB. GPU 1 has a total capacity of 94.97 GiB of which 98.19 MiB is free. Including non-PyTorch memory, this process has 94.86 GiB memory in use. Of the allocated memory 93.46 GiB is allocated by PyTorch, with 19.40 MiB allocated in private pools (e.g., CUDA Graphs), and 102.28 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation.  See documentation for Memory Management  (https://docs.pytorch.org/docs/stable/notes/cuda.html#optimizing-memory-usage-with-pytorch-cuda-alloc-conf)
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 128.00 MiB. GPU 0 has a total capacity of 94.97 GiB of which 100.19 MiB is free. Including non-PyTorch memory, this process has 94.86 GiB memory in use. Of the allocated memory 93.46 GiB is allocated by PyTorch, with 19.40 MiB allocated in private pools (e.g., CUDA Graphs), and 100.28 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation.  See documentation for Memory Management  (https://docs.pytorch.org/docs/stable/notes/cuda.html#optimizing-memory-usage-with-pytorch-cuda-alloc-conf)

A4. EngineCore fatal error

EngineCore encountered a fatal error.
Traceback (most recent call last):
  File "/opt/infernal-invocation/vllm/vllm/v1/engine/core.py", line 1500, in run_engine_core
    engine_core.run_busy_loop()
  File "/opt/infernal-invocation/vllm/vllm/v1/fault_tolerance/engine_core_sentinel.py", line 179, in run_with_fault_tolerance
    busy_loop_func(self)
  File "/opt/infernal-invocation/vllm/vllm/v1/engine/core.py", line 1559, in run_busy_loop
    self._process_engine_step()
  File "/opt/infernal-invocation/vllm/vllm/v1/engine/core.py", line 1612, in _process_engine_step
    outputs, model_executed = self.step_fn()
                              ^^^^^^^^^^^^^^
  File "/opt/infernal-invocation/vllm/vllm/v1/engine/core.py", line 829, in step_with_batch_queue
    exec_model_fut.result()
  File "/opt/infernal-invocation/vllm/vllm/v1/executor/multiproc_executor.py", line 99, in result
    return super().result()
           ^^^^^^^^^^^^^^^^
  File "/usr/lib/python3.12/concurrent/futures/_base.py", line 449, in result
    return self.__get_result()
           ^^^^^^^^^^^^^^^^^^^
  File "/usr/lib/python3.12/concurrent/futures/_base.py", line 401, in __get_result
    raise self._exception
  File "/opt/infernal-invocation/vllm/vllm/v1/executor/multiproc_executor.py", line 103, in _wait_for_response
    response = self.aggregate(self.get_response())
                              ^^^^^^^^^^^^^^^^^^^
  File "/opt/infernal-invocation/vllm/vllm/v1/executor/multiproc_executor.py", line 437, in get_response
    raise RuntimeError(
RuntimeError: Worker failed with error 'CUDA out of memory. Tried to allocate 128.00 MiB. GPU 0 has a total capacity of 94.97 GiB of which 100.19 MiB is free. Including non-PyTorch memory, this process has 94.86 GiB memory in use. Of the allocated memory 93.46 GiB is allocated by PyTorch, with 19.40 MiB allocated in private pools (e.g., CUDA Graphs), and 100.28 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation.  See documentation for Memory Management  (https://docs.pytorch.org/docs/stable/notes/cuda.html#optimizing-memory-usage-with-pytorch-cuda-alloc-conf)', please check the stack trace above for the root cause

A5. Every expandable_segments mapping failure, in order

Note the first two: 10.2 MiB and 8.2 MiB free.

[rank0]:[W915 09:38:18.766921827 expandable_segments: memory mapping failed with OOM on device 0 while trying to map 20971520 bytes (free: 10682368, total: 101973491712).
[rank1]:[W915 09:38:18.767250333 expandable_segments: memory mapping failed with OOM on device 1 while trying to map 20971520 bytes (free: 8585216, total: 101973491712).
[rank1]:[W915 09:38:18.770274677 expandable_segments: memory mapping failed with OOM on device 1 while trying to map 20971520 bytes (free: 19070976, total: 101973491712).
[rank0]:[W915 09:38:18.770276981 expandable_segments: memory mapping failed with OOM on device 0 while trying to map 20971520 bytes (free: 21168128, total: 101973491712).
[rank0]:[W915 09:38:18.770693840 expandable_segments: memory mapping failed with OOM on device 0 while trying to map 20971520 bytes (free: 21168128, total: 101973491712).
[rank1]:[W915 09:38:18.770747411 expandable_segments: memory mapping failed with OOM on device 1 while trying to map 20971520 bytes (free: 19070976, total: 101973491712).
[rank0]:[W915 09:38:18.799656986 expandable_segments: memory mapping failed with OOM on device 0 while trying to map 20971520 bytes (free: 19070976, total: 101973491712).
[rank1]:[W915 09:38:18.800036468 expandable_segments: memory mapping failed with OOM on device 1 while trying to map 20971520 bytes (free: 16973824, total: 101973491712).
[rank0]:[W915 09:38:18.034262677 expandable_segments: memory mapping failed with OOM on device 0 while trying to map 20971520 bytes (free: 19070976, total: 101973491712).
[rank1]:[W915 09:38:18.035616631 expandable_segments: memory mapping failed with OOM on device 1 while trying to map 20971520 bytes (free: 16973824, total: 101973491712).

A6. Memory accounting from the automatic restart

[Worker_TP1] [model_runner.py:411] Model loading took 81.11 GiB memory and 53.455084 seconds
[Worker_TP0] [model_runner.py:411] Model loading took 81.11 GiB memory and 53.488019 seconds
[Worker_TP1] [model_runner.py:1048] Graph capturing finished in 4 secs, took 0.31 GiB
[Worker_TP0] [model_runner.py:1048] Graph capturing finished in 4 secs, took 0.31 GiB
[Worker_TP0] [gpu_worker.py:628] Available KV cache memory: 7.94 GiB
[EngineCore] [kv_cache_utils.py:2208] GPU KV cache size: 1,196,643 tokens, Maximum concurrency for 1,048,576 tokens per request: 1.14x
[Worker_TP1] [model_runner.py:1048] Graph capturing finished in 4 secs, took 0.20 GiB
[Worker_TP1] [gpu_worker.py:796] CUDA graph pool memory: 0.2 GiB (actual), 0.32 GiB (estimated), difference: 0.12 GiB (63.0%).
[Worker_TP1] [gpu_worker.py:859] Free memory on device (93.79/94.97 GiB) on startup. Desired GPU memory utilization is (0.975, 92.6 GiB). Actual usage is 83.1 GiB for consumed memory (weights + non-torch), 1.24 GiB for peak activation, and 0.2 GiB for CUDAGraph memory. Replace gpu_memory_utilization config with `--kv-cache-memory=8496266036` (7.91 GiB) to fit into requested memory, or `--kv-cache-memory=9781741568` (9.11 GiB) to fully utilize gpu memory. Current kv cache memory in use is 7.94 GiB.
[Worker_TP0] [model_runner.py:1048] Graph capturing finished in 4 secs, took 0.20 GiB
[Worker_TP0] [gpu_worker.py:796] CUDA graph pool memory: 0.2 GiB (actual), 0.32 GiB (estimated), difference: 0.12 GiB (63.0%).
[Worker_TP0] [gpu_worker.py:859] Free memory on device (93.79/94.97 GiB) on startup. Desired GPU memory utilization is (0.975, 92.6 GiB). Actual usage is 83.1 GiB for consumed memory (weights + non-torch), 1.24 GiB for peak activation, and 0.2 GiB for CUDAGraph memory. Replace gpu_memory_utilization config with `--kv-cache-memory=8496266036` (7.91 GiB) to fit into requested memory, or `--kv-cache-memory=9781741568` (9.11 GiB) to fully utilize gpu memory. Current kv cache memory in use is 7.94 GiB.
[EngineCore] [core.py:400] init engine (profile, create kv cache, warmup model) took 19.33 s

A7. Same-shape long-context prefills that survived, after the restart

In both cases a large prompt is admitted (KV ~40 %, two requests briefly resident), then
prefilled at 5–8.5k tok/s without incident. This is why we do not claim the crash is
deterministic for a given context length.

09-15 09:40:37  Avg prompt throughput: 515.7 tokens/s, Avg generation throughput: 2.6 tokens/s, Running: 2 reqs, Waiting: 0 reqs, GPU KV cache usage: 41.7%, Prefix cache hit rate: 0.0%
09-15 09:40:47  Avg prompt throughput: 8574.1 tokens/s, Avg generation throughput: 122.8 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.3%, Prefix cache hit rate: 45.5%
09-15 09:40:57  Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 224.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.4%, Prefix cache hit rate: 45.5%
...
09-15 10:08:07  Avg prompt throughput: 84.3 tokens/s, Avg generation throughput: 0.5 tokens/s, Running: 2 reqs, Waiting: 0 reqs, GPU KV cache usage: 40.4%, Prefix cache hit rate: 91.9%, MM cache hit rate: 83.6%
09-15 10:08:17  Avg prompt throughput: 5086.5 tokens/s, Avg generation throughput: 166.1 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.8%, Prefix cache hit rate: 91.9%, MM cache hit rate: 83.6%
09-15 10:08:27  Avg prompt throughput: 260.3 tokens/s, Avg generation throughput: 79.8 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 91.9%, MM cache hit rate: 83.6%

Activity

  1. Seth777777777 commented on Sep 15, 2026

    @Seth777777777
    Author

    Second occurrence the same day — and this one narrows it down considerably.

    Crash at 09-15 13:53:47, ~4 hours after the one described above. Same build, same flags — we had not changed anything yet (--max-num-batched-tokens 4096, --gpu-memory-utilization 0.975, expandable_segments:True). Same restart accounting: weights 81.11 GiB, KV 7.94 GiB, graphs 0.20 GiB, profiled peak activation 1.24 GiB.

    Same overshoot: 93.50 GiB allocated by torch vs. 90.49 GiB predicted peak → +3.01 GiB, essentially identical to the +2.97 GiB of the first crash. Only 120 MiB reserved-but-unallocated, so again live memory, not fragmentation.

    What is different, and why it matters

    1. No multimodal input at all. The engine-stats lines in this window carry no MM cache hit rate field whatsoever — this was a pure-text workload (an agent loop doing web search and tool calls). That removes the "uncached image tokens in the chunk" candidate we raised in §6 above. Context length alone is enough.

    2. Even longer context. Scheduler stats at the failing step:

    num_running_reqs=1, num_waiting_reqs=0
    kv_cache_usage=0.4963
    PrefixCacheStats(requests=1, queries=255549, hits=212992)
    

    A ~255k-token prompt with 212 992 tokens served from the prefix cache, being chunk-prefilled at 4096 tokens/step. The previous crash was 156k queries / 149 760 hits. Both crashes are long-context chunked prefills with a large prefix-cache hit; nothing else about the two workloads is alike.

    3. The traceback now lands inside the indexer. Last time it died in the q projection (wq_b). This time:

    File "vllm/models/deepseek_v4/attention.py", line 556, in _prepare_and_attn
      q, (indexer_inputs, _) = execute_in_parallel(
    File "vllm/utils/multi_stream_utils.py", line 146, in execute_in_parallel
      aux_results[i] = fn()
    File "vllm/models/deepseek_v4/attention.py", line 559, in <lambda>
      lambda: indexer(
    File "vllm/models/deepseek_v4/attention.py", line 1039, in forward
      (q_quant, weights), _ = maybe_execute_in_parallel(
    File "vllm/utils/multi_stream_utils.py", line 69, in maybe_execute_in_parallel
      result0 = fn0()
    File "vllm/models/deepseek_v4/attention.py", line 1027, in wq_b_and_q_quant
      return fused_indexer_q_rope_quant(
    File "vllm/models/deepseek_v4/common/ops/fused_indexer_q.py", line 427, in fused_indexer_q_rope_quant
      index_q_fp8 = torch.empty_like(index_q, dtype=fp8_dtype)
    torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 32.00 MiB.
    GPU 1 has a total capacity of 94.97 GiB of which 40.19 MiB is free …
    

    The victim is again tiny (32 MiB, and index_q_fp8 is per-token, not per-context), so it does not identify the hog by itself. But across the two crashes the failing allocation has now landed twice in a row inside the same per-step attention path, once just before the indexer and once inside it.

    4. Multi-stream execution. Both crashes go through execute_in_parallel / maybe_execute_in_parallel: the indexer runs on a separate CUDA stream concurrently with the query projection. Two streams' intermediates are live at the same time, and the caching allocator does not freely reuse blocks across streams. With ~2.5 GiB of headroom this is not a detail we can ignore, and we cannot evaluate it from outside.

    Order-of-magnitude, updated

    At this crash the chunk was 4096 tokens against ≥212 992 tokens of KV. A chunk × kv_len score tensor would be:

    dtype 4096 × 212 992
    fp32 3.25 GiB
    bf16 1.63 GiB

    The observed overshoot is 3.01 GiB. We are not claiming this identifies the buffer — we are pointing out that the only quantity in this step that grew between the two crashes (154k → 213k+ of KV) is the one whose scaling matches an overshoot that stayed at ~3 GiB in both. If anything in the indexer path is O(chunk × kv_len) and is not in the startup profile, that would explain both.

    Revised questions

    1. Is any buffer in the indexer path (fused_indexer_q_rope_quant, the score/top-k computation) sized as O(chunk × kv_len)? If so, can it be tiled over the KV dimension — same result, bounded peak?
    2. Does execute_in_parallel keep the q-projection and indexer intermediates simultaneously live, and is that accounted for in the startup memory profile?
    3. Does the profiling run exercise the indexer at a realistic KV extent at all? It profiles a 4096-token chunk against an empty cache, i.e. a 4096-token context — 52× shorter than this crash.

    Standing conclusion

    Two crashes, four hours apart, different workloads (one multimodal, one pure text), different failing allocations, identical overshoot of ~3.0 GiB over the profiled peak. Whatever the specific buffer turns out to be, the actionable defect is the same one stated above: the startup profiler's peak-activation figure does not bound a real long-context prefill step, and the margin it leaves is smaller than the error. We are happy to run an instrumented build or a torch.cuda.memory._record_memory_history() snapshot on this host if that would settle it faster — just tell us what you want captured.

  2. Seth777777777 commented on Sep 15, 2026

    @Seth777777777
    Author

    Third crash, same host, same day — different failure class, and this one comes with a Python traceback that names a concrete suspect. It also reproduces on demand.

    Crash at 09-15 17:56:12. Config at the time: everything as described above, except we had halved --max-num-batched-tokens to 2048 as a mitigation experiment (--gpu-memory-utilization still 0.975). Build, model and host unchanged.

    This one is not an OOM

    torch.AcceleratorError: CUDA error: an illegal memory access was encountered
    

    The worker then hung — No available shared memory broadcast block found in 60 seconds — and was SIGKILLed; EngineCore died with RuntimeError: cancelled.

    It is reproducible

    The workload is an agent loop doing web search and tool calls (pure text, no images). It crashed once, and re-opening the same conversation and continuing it from the point where it died crashed the engine again. The context at that point is ~209k tokens, largely served from the prefix cache. This is the first reproducer we have for any of these crashes, and it needs no soak time — just a conversation that has grown past ~200k tokens and a request that continues it.

    The traceback

    File "vllm/models/deepseek_v4/attention.py", line 480, in forward
      self._prepare_and_attn_fn(
    File "vllm/models/deepseek_v4/attention.py", line 587, in _prepare_and_attn
      self._sparse_indexer_and_attn(
    File "vllm/models/deepseek_v4/attention.py", line 691, in _sparse_indexer_and_attn
      self.forward_mqa(q, kv, positions, out)
    File "vllm/models/deepseek_v4/nvidia/b12x.py", line 744, in forward_mqa
      self._forward_prefill(
    File "vllm/models/deepseek_v4/nvidia/b12x.py", line 906, in _forward_prefill
      _run_compressed_sparse_mla(
    File "vllm/models/deepseek_v4/nvidia/b12x.py", line 485, in _run_compressed_sparse_mla
      scratch = current_workspace_manager().get_simultaneous(*plan.shapes_and_dtypes())
    File "vllm/v1/worker/workspace.py", line 172, in get_simultaneous
      current_workspace = self._ensure_workspace_size(total_bytes)
    File "vllm/v1/worker/workspace.py", line 291, in _ensure_workspace_size
      torch.accelerator.empty_cache()
    torch.AcceleratorError: CUDA error: an illegal memory access was encountered
    

    Why we think this matters for the OOM above as well

    _run_compressed_sparse_mla sizes a scratch buffer per step from plan.shapes_and_dtypes(), and _ensure_workspace_size grows that buffer at runtime, calling torch.accelerator.empty_cache() to free room for the larger allocation.

    A buffer that is allocated and grown during serving cannot appear in the startup memory profile, which is exactly the ~3 GiB discrepancy reported in this issue (93.46 / 93.50 GiB actually allocated vs. 90.49 GiB predicted peak). We are not in a position to prove it is the same buffer, but:

    • it lives in the sparse-MLA prefill path, which is where all three crashes happened;
    • the two earlier OOMs died two frames away from here, inside _prepare_and_attn — one in the q projection, one inside the indexer;
    • it is sized per-step by a plan rather than bounded at startup;
    • _ensure_workspace_size reacting by dropping the allocator cache is precisely the behaviour you would expect right before a nearby small allocation fails with OOM, which is what both earlier crashes looked like (128 MiB and 32 MiB victims).

    The failing step is also informative about what the size depends on:

    scheduled_cached_reqs=CachedRequestData(req_ids=['chatcmpl-a1b7cca990ca8d5e-aa2c5b35'], num_computed_tokens=[208896], …)
    total_num_scheduled_tokens=216
    kv_cache_usage=0.0969
    num_common_prefix_blocks=[817, 0, 0, 0, 0]
    

    216 scheduled tokens against 208 896 tokens of context, with the KV pool under 10 % full. So whatever is being sized here scales with the length of the context, not with the number of tokens in the step — which also explains why lowering --max-num-batched-tokens did not save us.

    Caveat on the illegal access itself

    It surfaced at empty_cache(), which is a synchronization point where CUDA reports errors from kernels that already finished. The faulting kernel therefore ran earlier and the trace does not name it. We can re-run with CUDA_LAUNCH_BLOCKING=1, or under compute-sanitizer, if you tell us which you prefer — the workload reproduces, so this is cheap for us.

    Questions

    1. Is the WorkspaceManager scratch used by _run_compressed_sparse_mla accounted for anywhere in the startup memory profile, or reserved up front? If not, is that the intended design?
    2. Is the size returned by plan.shapes_and_dtypes() bounded as a function of kv_len, or does it grow without limit as the context grows?
    3. Could the two earlier OOMs simply be this same workspace failing to grow — i.e. one defect with two presentations, depending on whether the allocator can satisfy the growth?
    4. Is _ensure_workspace_size safe against a concurrent stream still reading the previous workspace? Both earlier crashes went through execute_in_parallel / multi-stream execution in the same region, and dropping and reallocating a buffer that another stream is using would produce exactly this error.

    Tally so far

    Three crashes on one host on 2026-09-15: 09:38 (OOM, multimodal, 156k prompt), 13:53 (OOM, pure text, 255k prompt), 17:56 (illegal memory access, pure text, 209k context). All three in the sparse-MLA prefill path, all three at contexts above 150k. The last one reproduces on demand.

  3. Seth777777777 commented on Sep 16, 2026

    @Seth777777777
    Author

    We patched the build, ran it, and the reproducer that killed the engine now serves 872k-token contexts. Below: two corrections to what we said earlier, the measurements, the diff, and the one question we cannot answer from outside the compiled module.

    Same host and build as above (2× RTX PRO 6000 Blackwell, TP=2, jovian.judgement…r9, DeepSeek-V4-Flash-Vision-Exp, --max-num-batched-tokens 2048, --gpu-memory-utilization 0.968).

    Two corrections to our earlier comments

    1. We never actually disabled expandable_segments:True. We removed PYTORCH_CUDA_ALLOC_CONF from our compose file and reported that the engine then survived long contexts. That was wrong: the image's own entrypoint sets the allocator, and the launch banner shows allocator=expandable_segments:True regardless. The stability we saw came from the other two changes (2048 and 0.968), not from the allocator. We apologise for the noise — the point about empty_cache() unmapping pages still stands as a hazard, but we have not tested the engine without it.

    2. max_q_chunks is not the cause of the shortfall. We had flagged that _reserve_profile_workspace passes max_q_chunks while _run_compressed_sparse_mla does not. We tested this directly: our patch reserves for both Caps constructions — the profile-shaped one and the runtime-shaped one that omits max_q_chunks. The second reservation added zero bytes, i.e. the runtime-shaped plan is never larger. That hypothesis is dead; see the open question at the end.

    What we changed

    Three lines of substance, all in Python (full diff at the end):

    1. b12x.py:_reserve_profile_workspace — get_simultaneous → reserve_all.
    2. b12x_indexer.py:_reserve_profile_workspace — get_simultaneous → reserve_all.
    3. b12x.py:_reserve_profile_workspace — reserve an extra fixed headroom (we used 768 MB) on top of the planned scratch.

    We also restored lock_workspace() at the end of capture_model in the V2 runner, behind an env flag, currently off. (The r21 V2 runner calls it; the r9 one does not — noted in our previous comment.)

    Startup, with VLLM_DEBUG_WORKSPACE=1

    Before (lane 0 only; lane 1 never reached the heavy branch and stayed at 148.06 MB):

    b12x.py:_reserve_profile_workspace   12.50 →  128.75 MB
    b12x.py:_reserve_profile_workspace  284.14 →  474.10 MB
    b12x.py:_reserve_profile_workspace  474.10 → 1052.35 MB
    

    After, both ranks, identical:

    [WORKSPACE DEBUG] Reserved  896.75 MB in execution slots [0, 1]
    [WORKSPACE DEBUG] Reserved 1242.10 MB in execution slots [0, 1]
    [WORKSPACE DEBUG] Reserved 1820.35 MB in execution slots [0, 1]
    

    Each figure is exactly the old one plus our 768 MB headroom — which is how we know the runtime-shaped Caps contributed nothing.

    Under production traffic

    Previously, within five minutes of normal traffic:

    b12x.py:_run_compressed_sparse_mla  1052.35 → 1490.46 MB
    b12x.py:_run_compressed_sparse_mla  1490.46 → 1597.88 MB
    

    After the patch, over three hours of live multi-user traffic including agent loops with tool calls:

    $ docker logs … | grep -c "Resized workspace"
    2
    

    Both remaining lines are b12x.py:_binding (mHC) at startup, one per rank. Zero runtime growth.

    The request that previously reproduced the illegal memory access on demand — re-opening a conversation past ~200k tokens — now completes. Largest prompt served so far: 872,747 tokens (prompt_tokens from the API response). The crash we reported happened at 208,896.

    nvidia-smi under load holds steady at 95,410 / 97,887 MiB per card, ~2.4 GiB free, unchanged across hours.

    The cost

      stock r9 patched
    workspace, lane 0 / lane 1 1052.35 / 148.06 MB 1820.35 / 1820.35 MB
    KV cache 8.3 GiB, 1,612,371 tokens 5.92 GiB, 1,149,700 tokens
    max concurrency at 1,048,576 1.54x 1.10x

    Correct reservation costs about 2.4 GiB of KV per card, and the headroom is paid twice because it is paid per slot. The full 1M window still fits. We consider this a good trade against an engine that dies once or twice a day, but it is a real cost and a proper fix inside the module would presumably be cheaper than our blanket 768 MB.

    The open question

    Since max_q_chunks is excluded, the only remaining difference between the two Caps is max_width:

    • _reserve_profile_workspace (b12x.py:668–678) computes width = swa_width + indexed_width, where swa_width is self.window_size (widened via get_dspark_swa_index_width when DSpark is on) and indexed_width comes from topk_indices_buffer.shape[-1] / indexer.topk_tokens.
    • _run_compressed_sparse_mla (b12x.py:464–466) computes width = swa_indices.shape[-1] + indexed_indices.shape[-1], from the actual tensors.

    Can swa_metadata.prefill_swa_indices.shape[-1] exceed the configured window_size (after the DSpark widening) on a long chunked prefill? If yes, that is the whole bug, and the fix is a one-line change to how the profile derives width — far better than our headroom. We cannot check this from outside module.plan.

    A related question: rows in the profile is max_num_batched_tokens, while _max_q_chunks iterates range(1, rows+1) — is the reservation meant to bound every possible rows, or only the maximum?

    Suggested fixes, updated

    1. reserve_all instead of get_simultaneous in both _reserve_profile_workspace implementations. This one is unambiguous — workspace.py documents reserve_all as being for exactly this case, and lane 1 was sized at 14% of lane 0 without it.
    2. Reconcile width between the two Caps constructions (see above). max_q_chunks is a red herring.
    3. Call lock_workspace() after warmup in the V2 runner, as the r21 V2 runner does. A loud assertion naming the caller and the sizes beats an OOM in an unrelated 32 MiB allocation three hours later. Note this is only safe once 1 and 2 are in.
    4. Synchronize (or record_stream) before releasing a workspace that side streams may still be reading, given the empty_cache() in the resize path.

    Diff

    Applied over the files as shipped in ...-20260907-r9. We run it as a thin image layer over yours (FROM <r9 image> plus three COPY lines), so this is trivially revertible on our side and we are happy to test any variant you prefer.

    --- a/b12x.py
    +++ b/b12x.py
    @@ -2,6 +2,7 @@
     # SPDX-FileCopyrightText: Copyright contributors to the vLLM project
     """B12x compressed sparse MLA for DeepSeek V4."""
    

    +import os
    from collections.abc import Callable
    from functools import cache
    from typing import TYPE_CHECKING, Any, ClassVar, Literal, cast
    @@ -11,6 +12,7 @@
    from vllm.config import VllmConfig
    from vllm.distributed import tensor_model_parallel_all_reduce
    from vllm.forward_context import get_forward_context
    +from vllm.logger import init_logger
    from vllm.models.deepseek_v4.attention import DeepseekV4Attention
    from vllm.models.deepseek_v4.common.ops import (
    compute_global_topk_indices_and_lens,
    @@ -52,6 +54,19 @@
    _DSV4_CACHE_BYTES_PER_TOKEN = 584
    _C128A_TOPK_ALIGNMENT = 128

    +logger = init_logger(name)
    +
    +# --- MAGIKON PATCH (r9 workspace under-reservation) -------------------------
    +# Measured on 2x RTX PRO 6000 / TP=2 with VLLM_DEBUG_WORKSPACE=1:
    +# profile reserved 1052.35 MB, production grew to 1597.88 MB via
    +# b12x.py:_run_compressed_sparse_mla -> +545.53 MB taken AFTER the
    +# KV-cache budget was assigned. Extra bytes reserved per execution slot
    +# on top of the planned scratch. Set to 0 to disable.
    +_WORKSPACE_HEADROOM_BYTES = (

    • int(os.environ.get("MAGIKON_B12X_WORKSPACE_HEADROOM_MB", "768")) * 1024 * 1024
      +)
      +# ---------------------------------------------------------------------------

    def _require_b12x_compressed_sparse_mla() -> Any:
    module = get_b12x_compressed_sparse_mla()
    @@ -703,7 +718,43 @@
    decode_row_capacity=decode_row_capacity,
    )
    )

    •    current_workspace_manager().get_simultaneous(*plan.shapes_and_dtypes())
      
    •    # --- MAGIKON PATCH ---------------------------------------------------
      
    •    # (a) reserve_all instead of get_simultaneous: get_simultaneous only
      
    •    #     touches the current (ubatch, lane) slot, so lane 1 (DSpark draft)
      
    •    #     was left at 148 MB against lane 0's 1052 MB and grew at runtime.
      
    •    # (b) also reserve for the Caps that the RUNTIME path builds.
      
    •    #     _run_compressed_sparse_mla (see above) omits max_q_chunks; this
      
    •    #     bounds both constructions instead of guessing which is larger.
      
    •    # (c) add a fixed headroom for the residual discrepancy we cannot
      
    •    #     derive from outside the compiled module.
      
    •    manager = current_workspace_manager()
      
    •    headroom: tuple[tuple[tuple[int, ...], torch.dtype], ...] = ()
      
    •    if _WORKSPACE_HEADROOM_BYTES &gt; 0:
      
    •        headroom = (((_WORKSPACE_HEADROOM_BYTES,), torch.uint8),)
      
    •    manager.reserve_all(*plan.shapes_and_dtypes(), *headroom)
      
    •    try:
      
    •        runtime_plan = module.plan(
      
    •            module.Caps(
      
    •                device=q.device,
      
    •                num_q_heads=int(q.shape[1]),
      
    •                max_q_rows=rows,
      
    •                max_width=width,
      
    •                head_dim=_DSV4_HEAD_DIM,
      
    •                v_head_dim=_DSV4_HEAD_DIM,
      
    •                page_size=int(self.swa_cache_layer.block_size),
      
    •                max_chunks_per_row=max_chunks_per_row,
      
    •                decode_row_capacity=decode_row_capacity,
      
    •            )
      
    •        )
      
    •    except Exception:  # noqa: BLE001 - best effort, never block startup
      
    •        logger.warning(
      
    •            "[MAGIKON PATCH] could not plan the runtime-shaped Caps; "
      
    •            "reserving the profile shape only.",
      
    •            exc_info=True,
      
    •        )
      
    •    else:
      
    •        manager.reserve_all(*runtime_plan.shapes_and_dtypes(), *headroom)
      
    •    # --- END MAGIKON PATCH -----------------------------------------------
      

      def forward_mqa(
      self,
      --- a/b12x_indexer.py
      +++ b/b12x_indexer.py
      @@ -283,7 +283,11 @@
      )
      if shared_page_table:
      _assert_prefill_route(plan)

    •        current_workspace_manager().get_simultaneous(*plan.shapes_and_dtypes())
      
    •        # --- MAGIKON PATCH ---
      
    •        # reserve every (ubatch, lane) slot, not just the current one;
      
    •        # get_simultaneous left the DSpark draft lane undersized.
      
    •        current_workspace_manager().reserve_all(*plan.shapes_and_dtypes())
      
    •        # --- END MAGIKON PATCH ---
      

      def reserve_profile_workspace(self, q: torch.Tensor) -> None:
      self._reserve_profile_workspace(q)
      --- a/model_runner.py
      +++ b/model_runner.py
      @@ -19,6 +19,7 @@

    import functools
    import gc
    +import os
    import time
    from copy import deepcopy
    from typing import Any, NamedTuple
    @@ -165,7 +166,7 @@
    copy_kv_cache_blocks_inplace,
    get_uniform_decode_token_count,
    )
    -from vllm.v1.worker.workspace import use_workspace_lane
    +from vllm.v1.worker.workspace import lock_workspace, use_workspace_lane

    logger = init_logger(name)

    @@ -1050,6 +1051,17 @@
    elapsed_time,
    cuda_graph_size / (1 << 30),
    )

    •    # --- MAGIKON PATCH ---
      
    •    # Freeze the workspace after warmup, as the r21 V2 runner did.
      
    •    # Any later growth now raises an AssertionError naming the caller and
      
    •    # the sizes, instead of silently taking memory already promised to the
      
    •    # KV cache and surfacing hours later as an unrelated OOM.
      
    •    # OPT-IN: set MAGIKON_B12X_LOCK_WORKSPACE=1 to arm it. Off by default
      
    •    # so this file can be mounted together with the b12x patches and the
      
    •    # canary enabled only once the reservation fix is confirmed.
      
    •    if os.environ.get("MAGIKON_B12X_LOCK_WORKSPACE", "0") == "1":
      
    •        lock_workspace()
      
    •    # --- END MAGIKON PATCH ---
         return cuda_graph_size
      

      def _remove_request(self, req_id: str) -> bool:

    Happy to run CUDA_LAUNCH_BLOCKING=1, compute-sanitizer, a memory-history snapshot, or a build of your own on this host. The reproducer is still available to us.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions