Repository navigation
Delayed production CUDA OOM on ds4-vision-jovian-judgement-r9: long-context prefill exceeds profiled peak by ~3 GiB #103
Description
Activity
Second occurrence the same day — and this one narrows it down considerably.
Crash at 09-15 13:53:47, ~4 hours after the one described above. Same build, same flags — we had not changed anything yet (
--max-num-batched-tokens 4096,--gpu-memory-utilization 0.975,expandable_segments:True). Same restart accounting: weights 81.11 GiB, KV 7.94 GiB, graphs 0.20 GiB, profiled peak activation 1.24 GiB.Same overshoot: 93.50 GiB allocated by torch vs. 90.49 GiB predicted peak → +3.01 GiB, essentially identical to the +2.97 GiB of the first crash. Only 120 MiB reserved-but-unallocated, so again live memory, not fragmentation.
What is different, and why it matters
1. No multimodal input at all. The engine-stats lines in this window carry no
MM cache hit ratefield whatsoever — this was a pure-text workload (an agent loop doing web search and tool calls). That removes the "uncached image tokens in the chunk" candidate we raised in §6 above. Context length alone is enough.2. Even longer context. Scheduler stats at the failing step:
num_running_reqs=1, num_waiting_reqs=0 kv_cache_usage=0.4963 PrefixCacheStats(requests=1, queries=255549, hits=212992)A ~255k-token prompt with 212 992 tokens served from the prefix cache, being chunk-prefilled at 4096 tokens/step. The previous crash was 156k queries / 149 760 hits. Both crashes are long-context chunked prefills with a large prefix-cache hit; nothing else about the two workloads is alike.
3. The traceback now lands inside the indexer. Last time it died in the q projection (
wq_b). This time:File "vllm/models/deepseek_v4/attention.py", line 556, in _prepare_and_attn q, (indexer_inputs, _) = execute_in_parallel( File "vllm/utils/multi_stream_utils.py", line 146, in execute_in_parallel aux_results[i] = fn() File "vllm/models/deepseek_v4/attention.py", line 559, in <lambda> lambda: indexer( File "vllm/models/deepseek_v4/attention.py", line 1039, in forward (q_quant, weights), _ = maybe_execute_in_parallel( File "vllm/utils/multi_stream_utils.py", line 69, in maybe_execute_in_parallel result0 = fn0() File "vllm/models/deepseek_v4/attention.py", line 1027, in wq_b_and_q_quant return fused_indexer_q_rope_quant( File "vllm/models/deepseek_v4/common/ops/fused_indexer_q.py", line 427, in fused_indexer_q_rope_quant index_q_fp8 = torch.empty_like(index_q, dtype=fp8_dtype) torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 32.00 MiB. GPU 1 has a total capacity of 94.97 GiB of which 40.19 MiB is free …The victim is again tiny (32 MiB, and
index_q_fp8is per-token, not per-context), so it does not identify the hog by itself. But across the two crashes the failing allocation has now landed twice in a row inside the same per-step attention path, once just before the indexer and once inside it.4. Multi-stream execution. Both crashes go through
execute_in_parallel/maybe_execute_in_parallel: the indexer runs on a separate CUDA stream concurrently with the query projection. Two streams' intermediates are live at the same time, and the caching allocator does not freely reuse blocks across streams. With ~2.5 GiB of headroom this is not a detail we can ignore, and we cannot evaluate it from outside.Order-of-magnitude, updated
At this crash the chunk was 4096 tokens against ≥212 992 tokens of KV. A
chunk × kv_lenscore tensor would be:dtype 4096 × 212 992 fp32 3.25 GiB bf16 1.63 GiB The observed overshoot is 3.01 GiB. We are not claiming this identifies the buffer — we are pointing out that the only quantity in this step that grew between the two crashes (154k → 213k+ of KV) is the one whose scaling matches an overshoot that stayed at ~3 GiB in both. If anything in the indexer path is
O(chunk × kv_len)and is not in the startup profile, that would explain both.Revised questions
- Is any buffer in the indexer path (
fused_indexer_q_rope_quant, the score/top-k computation) sized asO(chunk × kv_len)? If so, can it be tiled over the KV dimension — same result, bounded peak? - Does
execute_in_parallelkeep the q-projection and indexer intermediates simultaneously live, and is that accounted for in the startup memory profile? - Does the profiling run exercise the indexer at a realistic KV extent at all? It profiles a 4096-token chunk against an empty cache, i.e. a 4096-token context — 52× shorter than this crash.
Standing conclusion
Two crashes, four hours apart, different workloads (one multimodal, one pure text), different failing allocations, identical overshoot of ~3.0 GiB over the profiled peak. Whatever the specific buffer turns out to be, the actionable defect is the same one stated above: the startup profiler's peak-activation figure does not bound a real long-context prefill step, and the margin it leaves is smaller than the error. We are happy to run an instrumented build or a
torch.cuda.memory._record_memory_history()snapshot on this host if that would settle it faster — just tell us what you want captured.- Is any buffer in the indexer path (
Third crash, same host, same day — different failure class, and this one comes with a Python traceback that names a concrete suspect. It also reproduces on demand.
Crash at 09-15 17:56:12. Config at the time: everything as described above, except we had halved
--max-num-batched-tokensto 2048 as a mitigation experiment (--gpu-memory-utilizationstill 0.975). Build, model and host unchanged.This one is not an OOM
torch.AcceleratorError: CUDA error: an illegal memory access was encounteredThe worker then hung —
No available shared memory broadcast block found in 60 seconds— and was SIGKILLed; EngineCore died withRuntimeError: cancelled.It is reproducible
The workload is an agent loop doing web search and tool calls (pure text, no images). It crashed once, and re-opening the same conversation and continuing it from the point where it died crashed the engine again. The context at that point is ~209k tokens, largely served from the prefix cache. This is the first reproducer we have for any of these crashes, and it needs no soak time — just a conversation that has grown past ~200k tokens and a request that continues it.
The traceback
File "vllm/models/deepseek_v4/attention.py", line 480, in forward self._prepare_and_attn_fn( File "vllm/models/deepseek_v4/attention.py", line 587, in _prepare_and_attn self._sparse_indexer_and_attn( File "vllm/models/deepseek_v4/attention.py", line 691, in _sparse_indexer_and_attn self.forward_mqa(q, kv, positions, out) File "vllm/models/deepseek_v4/nvidia/b12x.py", line 744, in forward_mqa self._forward_prefill( File "vllm/models/deepseek_v4/nvidia/b12x.py", line 906, in _forward_prefill _run_compressed_sparse_mla( File "vllm/models/deepseek_v4/nvidia/b12x.py", line 485, in _run_compressed_sparse_mla scratch = current_workspace_manager().get_simultaneous(*plan.shapes_and_dtypes()) File "vllm/v1/worker/workspace.py", line 172, in get_simultaneous current_workspace = self._ensure_workspace_size(total_bytes) File "vllm/v1/worker/workspace.py", line 291, in _ensure_workspace_size torch.accelerator.empty_cache() torch.AcceleratorError: CUDA error: an illegal memory access was encounteredWhy we think this matters for the OOM above as well
_run_compressed_sparse_mlasizes a scratch buffer per step fromplan.shapes_and_dtypes(), and_ensure_workspace_sizegrows that buffer at runtime, callingtorch.accelerator.empty_cache()to free room for the larger allocation.A buffer that is allocated and grown during serving cannot appear in the startup memory profile, which is exactly the ~3 GiB discrepancy reported in this issue (93.46 / 93.50 GiB actually allocated vs. 90.49 GiB predicted peak). We are not in a position to prove it is the same buffer, but:
- it lives in the sparse-MLA prefill path, which is where all three crashes happened;
- the two earlier OOMs died two frames away from here, inside
_prepare_and_attn— one in the q projection, one inside the indexer; - it is sized per-step by a plan rather than bounded at startup;
_ensure_workspace_sizereacting by dropping the allocator cache is precisely the behaviour you would expect right before a nearby small allocation fails with OOM, which is what both earlier crashes looked like (128 MiB and 32 MiB victims).
The failing step is also informative about what the size depends on:
scheduled_cached_reqs=CachedRequestData(req_ids=['chatcmpl-a1b7cca990ca8d5e-aa2c5b35'], num_computed_tokens=[208896], …) total_num_scheduled_tokens=216 kv_cache_usage=0.0969 num_common_prefix_blocks=[817, 0, 0, 0, 0]216 scheduled tokens against 208 896 tokens of context, with the KV pool under 10 % full. So whatever is being sized here scales with the length of the context, not with the number of tokens in the step — which also explains why lowering
--max-num-batched-tokensdid not save us.Caveat on the illegal access itself
It surfaced at
empty_cache(), which is a synchronization point where CUDA reports errors from kernels that already finished. The faulting kernel therefore ran earlier and the trace does not name it. We can re-run withCUDA_LAUNCH_BLOCKING=1, or undercompute-sanitizer, if you tell us which you prefer — the workload reproduces, so this is cheap for us.Questions
- Is the
WorkspaceManagerscratch used by_run_compressed_sparse_mlaaccounted for anywhere in the startup memory profile, or reserved up front? If not, is that the intended design? - Is the size returned by
plan.shapes_and_dtypes()bounded as a function ofkv_len, or does it grow without limit as the context grows? - Could the two earlier OOMs simply be this same workspace failing to grow — i.e. one defect with two presentations, depending on whether the allocator can satisfy the growth?
- Is
_ensure_workspace_sizesafe against a concurrent stream still reading the previous workspace? Both earlier crashes went throughexecute_in_parallel/ multi-stream execution in the same region, and dropping and reallocating a buffer that another stream is using would produce exactly this error.
Tally so far
Three crashes on one host on 2026-09-15: 09:38 (OOM, multimodal, 156k prompt), 13:53 (OOM, pure text, 255k prompt), 17:56 (illegal memory access, pure text, 209k context). All three in the sparse-MLA prefill path, all three at contexts above 150k. The last one reproduces on demand.
We patched the build, ran it, and the reproducer that killed the engine now serves 872k-token contexts. Below: two corrections to what we said earlier, the measurements, the diff, and the one question we cannot answer from outside the compiled module.
Same host and build as above (2× RTX PRO 6000 Blackwell, TP=2,
jovian.judgement…r9, DeepSeek-V4-Flash-Vision-Exp,--max-num-batched-tokens 2048,--gpu-memory-utilization 0.968).Two corrections to our earlier comments
1. We never actually disabled
expandable_segments:True. We removedPYTORCH_CUDA_ALLOC_CONFfrom our compose file and reported that the engine then survived long contexts. That was wrong: the image's own entrypoint sets the allocator, and the launch banner showsallocator=expandable_segments:Trueregardless. The stability we saw came from the other two changes (2048and0.968), not from the allocator. We apologise for the noise — the point aboutempty_cache()unmapping pages still stands as a hazard, but we have not tested the engine without it.2.
max_q_chunksis not the cause of the shortfall. We had flagged that_reserve_profile_workspacepassesmax_q_chunkswhile_run_compressed_sparse_mladoes not. We tested this directly: our patch reserves for bothCapsconstructions — the profile-shaped one and the runtime-shaped one that omitsmax_q_chunks. The second reservation added zero bytes, i.e. the runtime-shaped plan is never larger. That hypothesis is dead; see the open question at the end.What we changed
Three lines of substance, all in Python (full diff at the end):
b12x.py:_reserve_profile_workspace—get_simultaneous→reserve_all.b12x_indexer.py:_reserve_profile_workspace—get_simultaneous→reserve_all.b12x.py:_reserve_profile_workspace— reserve an extra fixed headroom (we used 768 MB) on top of the planned scratch.
We also restored
lock_workspace()at the end ofcapture_modelin the V2 runner, behind an env flag, currently off. (The r21 V2 runner calls it; the r9 one does not — noted in our previous comment.)Startup, with
VLLM_DEBUG_WORKSPACE=1Before (lane 0 only; lane 1 never reached the heavy branch and stayed at 148.06 MB):
b12x.py:_reserve_profile_workspace 12.50 → 128.75 MB b12x.py:_reserve_profile_workspace 284.14 → 474.10 MB b12x.py:_reserve_profile_workspace 474.10 → 1052.35 MBAfter, both ranks, identical:
[WORKSPACE DEBUG] Reserved 896.75 MB in execution slots [0, 1] [WORKSPACE DEBUG] Reserved 1242.10 MB in execution slots [0, 1] [WORKSPACE DEBUG] Reserved 1820.35 MB in execution slots [0, 1]Each figure is exactly the old one plus our 768 MB headroom — which is how we know the runtime-shaped
Capscontributed nothing.Under production traffic
Previously, within five minutes of normal traffic:
b12x.py:_run_compressed_sparse_mla 1052.35 → 1490.46 MB b12x.py:_run_compressed_sparse_mla 1490.46 → 1597.88 MBAfter the patch, over three hours of live multi-user traffic including agent loops with tool calls:
$ docker logs … | grep -c "Resized workspace" 2Both remaining lines are
b12x.py:_binding(mHC) at startup, one per rank. Zero runtime growth.The request that previously reproduced the
illegal memory accesson demand — re-opening a conversation past ~200k tokens — now completes. Largest prompt served so far: 872,747 tokens (prompt_tokensfrom the API response). The crash we reported happened at 208,896.nvidia-smiunder load holds steady at 95,410 / 97,887 MiB per card, ~2.4 GiB free, unchanged across hours.The cost
stock r9 patched workspace, lane 0 / lane 1 1052.35 / 148.06 MB 1820.35 / 1820.35 MB KV cache 8.3 GiB, 1,612,371 tokens 5.92 GiB, 1,149,700 tokens max concurrency at 1,048,576 1.54x 1.10x Correct reservation costs about 2.4 GiB of KV per card, and the headroom is paid twice because it is paid per slot. The full 1M window still fits. We consider this a good trade against an engine that dies once or twice a day, but it is a real cost and a proper fix inside the module would presumably be cheaper than our blanket 768 MB.
The open question
Since
max_q_chunksis excluded, the only remaining difference between the twoCapsismax_width:_reserve_profile_workspace(b12x.py:668–678) computeswidth = swa_width + indexed_width, whereswa_widthisself.window_size(widened viaget_dspark_swa_index_widthwhen DSpark is on) andindexed_widthcomes fromtopk_indices_buffer.shape[-1]/indexer.topk_tokens._run_compressed_sparse_mla(b12x.py:464–466) computeswidth = swa_indices.shape[-1] + indexed_indices.shape[-1], from the actual tensors.
Can
swa_metadata.prefill_swa_indices.shape[-1]exceed the configuredwindow_size(after the DSpark widening) on a long chunked prefill? If yes, that is the whole bug, and the fix is a one-line change to how the profile deriveswidth— far better than our headroom. We cannot check this from outsidemodule.plan.A related question:
rowsin the profile ismax_num_batched_tokens, while_max_q_chunksiteratesrange(1, rows+1)— is the reservation meant to bound every possiblerows, or only the maximum?Suggested fixes, updated
reserve_allinstead ofget_simultaneousin both_reserve_profile_workspaceimplementations. This one is unambiguous —workspace.pydocumentsreserve_allas being for exactly this case, and lane 1 was sized at 14% of lane 0 without it.- Reconcile
widthbetween the twoCapsconstructions (see above).max_q_chunksis a red herring. - Call
lock_workspace()after warmup in the V2 runner, as the r21 V2 runner does. A loud assertion naming the caller and the sizes beats an OOM in an unrelated 32 MiB allocation three hours later. Note this is only safe once 1 and 2 are in. - Synchronize (or
record_stream) before releasing a workspace that side streams may still be reading, given theempty_cache()in the resize path.
Diff
Applied over the files as shipped in
...-20260907-r9. We run it as a thin image layer over yours (FROM <r9 image>plus threeCOPYlines), so this is trivially revertible on our side and we are happy to test any variant you prefer.--- a/b12x.py +++ b/b12x.py @@ -2,6 +2,7 @@ # SPDX-FileCopyrightText: Copyright contributors to the vLLM project """B12x compressed sparse MLA for DeepSeek V4."""+import os
from collections.abc import Callable
from functools import cache
from typing import TYPE_CHECKING, Any, ClassVar, Literal, cast
@@ -11,6 +12,7 @@
from vllm.config import VllmConfig
from vllm.distributed import tensor_model_parallel_all_reduce
from vllm.forward_context import get_forward_context
+from vllm.logger import init_logger
from vllm.models.deepseek_v4.attention import DeepseekV4Attention
from vllm.models.deepseek_v4.common.ops import (
compute_global_topk_indices_and_lens,
@@ -52,6 +54,19 @@
_DSV4_CACHE_BYTES_PER_TOKEN = 584
_C128A_TOPK_ALIGNMENT = 128+logger = init_logger(name)
+
+# --- MAGIKON PATCH (r9 workspace under-reservation) -------------------------
+# Measured on 2x RTX PRO 6000 / TP=2 with VLLM_DEBUG_WORKSPACE=1:
+# profile reserved 1052.35 MB, production grew to 1597.88 MB via
+# b12x.py:_run_compressed_sparse_mla -> +545.53 MB taken AFTER the
+# KV-cache budget was assigned. Extra bytes reserved per execution slot
+# on top of the planned scratch. Set to 0 to disable.
+_WORKSPACE_HEADROOM_BYTES = (- int(os.environ.get("MAGIKON_B12X_WORKSPACE_HEADROOM_MB", "768")) * 1024 * 1024
+)
+# ---------------------------------------------------------------------------
def _require_b12x_compressed_sparse_mla() -> Any:
module = get_b12x_compressed_sparse_mla()
@@ -703,7 +718,43 @@
decode_row_capacity=decode_row_capacity,
)
)-
current_workspace_manager().get_simultaneous(*plan.shapes_and_dtypes())
-
# --- MAGIKON PATCH --------------------------------------------------- -
# (a) reserve_all instead of get_simultaneous: get_simultaneous only -
# touches the current (ubatch, lane) slot, so lane 1 (DSpark draft) -
# was left at 148 MB against lane 0's 1052 MB and grew at runtime. -
# (b) also reserve for the Caps that the RUNTIME path builds. -
# _run_compressed_sparse_mla (see above) omits max_q_chunks; this -
# bounds both constructions instead of guessing which is larger. -
# (c) add a fixed headroom for the residual discrepancy we cannot -
# derive from outside the compiled module. -
manager = current_workspace_manager() -
headroom: tuple[tuple[tuple[int, ...], torch.dtype], ...] = () -
if _WORKSPACE_HEADROOM_BYTES > 0: -
headroom = (((_WORKSPACE_HEADROOM_BYTES,), torch.uint8),) -
manager.reserve_all(*plan.shapes_and_dtypes(), *headroom) -
try: -
runtime_plan = module.plan( -
module.Caps( -
device=q.device, -
num_q_heads=int(q.shape[1]), -
max_q_rows=rows, -
max_width=width, -
head_dim=_DSV4_HEAD_DIM, -
v_head_dim=_DSV4_HEAD_DIM, -
page_size=int(self.swa_cache_layer.block_size), -
max_chunks_per_row=max_chunks_per_row, -
decode_row_capacity=decode_row_capacity, -
) -
) -
except Exception: # noqa: BLE001 - best effort, never block startup -
logger.warning( -
"[MAGIKON PATCH] could not plan the runtime-shaped Caps; " -
"reserving the profile shape only.", -
exc_info=True, -
) -
else: -
manager.reserve_all(*runtime_plan.shapes_and_dtypes(), *headroom) -
# --- END MAGIKON PATCH -----------------------------------------------def forward_mqa(
self,
--- a/b12x_indexer.py
+++ b/b12x_indexer.py
@@ -283,7 +283,11 @@
)
if shared_page_table:
_assert_prefill_route(plan)
-
current_workspace_manager().get_simultaneous(*plan.shapes_and_dtypes())
-
# --- MAGIKON PATCH --- -
# reserve every (ubatch, lane) slot, not just the current one; -
# get_simultaneous left the DSpark draft lane undersized. -
current_workspace_manager().reserve_all(*plan.shapes_and_dtypes()) -
# --- END MAGIKON PATCH ---def reserve_profile_workspace(self, q: torch.Tensor) -> None:
self._reserve_profile_workspace(q)
--- a/model_runner.py
+++ b/model_runner.py
@@ -19,6 +19,7 @@
import functools
import gc
+import os
import time
from copy import deepcopy
from typing import Any, NamedTuple
@@ -165,7 +166,7 @@
copy_kv_cache_blocks_inplace,
get_uniform_decode_token_count,
)
-from vllm.v1.worker.workspace import use_workspace_lane
+from vllm.v1.worker.workspace import lock_workspace, use_workspace_lanelogger = init_logger(name)
@@ -1050,6 +1051,17 @@
elapsed_time,
cuda_graph_size / (1 << 30),
)-
# --- MAGIKON PATCH --- -
# Freeze the workspace after warmup, as the r21 V2 runner did. -
# Any later growth now raises an AssertionError naming the caller and -
# the sizes, instead of silently taking memory already promised to the -
# KV cache and surfacing hours later as an unrelated OOM. -
# OPT-IN: set MAGIKON_B12X_LOCK_WORKSPACE=1 to arm it. Off by default -
# so this file can be mounted together with the b12x patches and the -
# canary enabled only once the reservation fix is confirmed. -
if os.environ.get("MAGIKON_B12X_LOCK_WORKSPACE", "0") == "1": -
lock_workspace() # --- END MAGIKON PATCH --- return cuda_graph_sizedef _remove_request(self, req_id: str) -> bool:
Happy to run
CUDA_LAUNCH_BLOCKING=1,compute-sanitizer, a memory-history snapshot, or a build of your own on this host. The reproducer is still available to us.
Short version: the known "delayed production OOM" reproduced a third time (11.09, 13.09, 15.09), this time with
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:Truealready applied. It is not fragmentation: at the moment of death only ~100 MiB was reserved-but-unallocated, i.e. essentially all of the 93.46 GiB was live.Two hard numbers:
torch.emptyfor the output of the query projectionwq_bindeepseek_v4/attention.py:544(DeepGEMM block-scaled MM). Nothing pathological about the victim — memory was already gone when it asked.So the profiler's peak-activation figure (1.24 GiB) does not bound what a real production step actually allocates, and the safety margin it leaves is smaller than the error. The step that died was a chunked prefill with 149 760 tokens of KV context behind it; we suspect something whose size scales with KV length rather than with chunk size, but the traceback does not prove that — see §6, where we also note that similar long-context prefills survive routinely on this box, so the trigger is evidently narrower than "long context alone".
Full log (crash + automatic restart + 30 min of subsequent normal traffic) attached.
1. Environment
0.26.1rc0+jovian.judgement.cu133.r9.vllmf66599d.b12x15b6813deepseek-ai/DeepSeek-V4-Flash-Vision-Exp, rev6821d6ad3681a4b137b066b76094fa82ebd0a380iommu=ptB12X_PCIE(oneshot) + PYNCCL, nccl 2.31.2--attention-backend B12X,--moe-backend b12x, fp8 KV (fp8_ds_mla), FP8 indexer cacheds4-vision-jovian-judgement-r9, restarted automatically after the crashMinor thing worth confirming: the version string says
jovian.judgement…r9, but every traceback path is under/opt/infernal-invocation/vllm/…. We assume this is just a leftover install prefix in the image and not an r15/r9 mix-up — please confirm.2. Configuration
Launch line as logged:
Relevant server flags:
Everything else is stock r9.
MAX_MODEL_LENandGPU_MEMORY_UTILIZATIONare deliberately left at the release defaults (we tried pinningGPU_MEMORY_UTILIZATIONearlier and reverted).3. History
expandable_segments:TrueAfter 13.09 we added only
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True(per the note in the r9 page) and left GPU memory utilisation on automatic KV admission. The workaround did not change the outcome or the interval.4. The failing step
Single request in flight, no queue, chunked prefill, multimodal:
The dump also shows several
MultiModalFeatureSpec(modality='image', …)entries with distinctmm_hashes, e.g. one withmm_position=PlaceholderRange(offset=135238, length=355). So the images are interleaved deep inside the context, not just at the head — relevant because this build forces--disable_chunked_mm_inputfor multimodal-bidirectional attention.So: a ~156k-token prompt containing several images, 95.9 % prefix-cache hit, being prefilled in 4096-token chunks. KV context at the failing step = 149760 + 4096 = 153 856 tokens. KV cache itself was only 38 % full — this is not a KV-capacity problem.
Error:
Both ranks then log:
Two things follow from this:
expandable_segmentscould not save it. The segment mapper itself failed to map 20 MiB against 18 MiB of free device memory.Worker-side traceback — the actual allocation site
Both ranks die in the same place, in the model forward, at the very start of attention:
128 MiB is exactly the expected q-projection output for a 4096-token chunk, so this is an ordinary allocation that happened to arrive when the card was already full — it identifies the moment, not the culprit. Note that the whole per-step attention work (indexer, sparse MLA) happens after this point, so whatever the excess is, it was already resident before the failing layer's attention began.
The allocator was thrashing at the edge for a few milliseconds before giving up — the earliest mapping failures report only 8.2–10.2 MiB free:
5. Memory accounting: profiled vs. actual
From the restart of the very same container (same flags, same weights):
Headroom between the profiled peak and the physical card (incl. non-torch ~1.4–2 GiB) is ≈ 2.5 GiB. The step needed ~3 GiB more than predicted, so it died — with mathematical certainty, every time this input shape occurs.
6. Hypothesis: the profiler never sees a long-context prefill step
The memory profiling run uses
max_num_batched_tokenstokens against an empty KV context, so it measures a 4096 × 4096 step. The step that crashed is 4096 × 153 856 — a 37.6× larger KV extent. Any workspace whose size isO(chunk × kv_len)is therefore under-profiled by that factor.The obvious candidate in this build is the DSA / Lightning Indexer score+top-k buffer (
indexer.py:581 DSA indexer decode path: use_flattening=True supports_varlen=False (next_n=4, use_fp4_cache=False);Using FP8 indexer cache for Lightning Indexer). Order-of-magnitude for a4096 × 153 856score tensor:An fp32 score buffer plus an index/top-k copy would land close to the observed ~3 GiB gap. We want to be explicit that this is arithmetic, not evidence — the traceback does not name the indexer, and we have not instrumented the build. Treat it as the first place we would look, not as a diagnosis.
Counter-evidence we should not hide. The attached log covers 8 minutes before the crash and 30 minutes after it. In that window, long-context chunked prefills of the same shape happen several times per hour and complete normally — e.g.
Avg prompt throughput: 8574.1 tokens/s … GPU KV cache usage: 41.7%and5086.5 tokens/s … 40.4%, both after the restart, plus7892.5 tokens/sbefore it. So a ~150k-token context with a big prefix-cache hit is clearly not sufficient on its own to kill the engine — something narrower tips it over. Candidates we cannot separate from our side:--disable_chunked_mm_inputforced, an image's 355 tokens must be scheduled inside one chunk; a chunk that carries image tokens and sits at 150k KV may be the specific bad combination. MM cache hit rate was 96.8 %, so on most of those surviving prefills the image work was served from cache.--async-schedulinglets step n+1 be prepared while step n is still resident. With ~2.5 GiB of headroom, an overlap of two heavy steps would compound anything above.What we are confident about, and what we think is worth fixing regardless of which of these it is: the profiler's 1.24 GiB peak-activation figure is not an upper bound for production steps, and the margin it leaves (~2.5 GiB) is smaller than the observed error (~3.0 GiB). Any config that passes startup profiling can therefore still OOM later.
7. Suggested reproduction
We have not reproduced it on demand (see the counter-evidence in §6 — most long-context prefills survive). What we would try first, on a 2× 96 GB SM120 box with the r9 vision image and stock flags:
--max-num-batched-tokens 4096,--enable-chunked-prefill,--enable-prefix-caching,--gpu-memory-utilization 0.975.num_computed_tokens ≈ 150kand prefills 4096 new tokens against that context.torch.cuda.memory_allocated()at the step wherenum_computed_tokensis largest, rather than waiting for the OOM.If it survives, bisect upward on prompt length and on the number of uncached images. Even a run that does not crash is useful to you: if
memory_allocated()at 150k context exceeds the profiled peak by a gigabyte or more, the profiling gap is confirmed independently of what our production traffic did.8. Questions
max_model_len-sized KV extent (or the workspace be reserved up front) so that the KV admission figure is actually safe?O(chunk × tile)instead ofO(chunk × kv_len)?--async-schedulingallow two long-context prefill steps' activations to be resident simultaneously?--disable_chunked_mm_inputis forced for multimodal-bidirectional attention?max_num_batched_tokensceiling as a function ofmax_model_lenfor the vision checkpoint? The vision weights (81.11 GiB) leave far less headroom than the text-only checkpoint, so the safe operating point is presumably different.9. What we are doing on our side meanwhile
--max-num-batched-tokensto 2048 — halves anything that scales aschunk × kv_len, at the cost of slower long prefills.--kv-cache-memoryexplicitly (rather than 0.975 auto-admission) to buy ~3 GiB of headroom; our KV usage peaked at 38 % of 1.19 M tokens, and the surplus is mostly evictable prefix-cache blocks.expandable_segments:True(harmless, just not sufficient).--async-schedulingas the next test.nvidia-smimemory sampling so that if the next crash is a gradual leak rather than a single oversized step, we will have the curve to prove it.The attached log covers 09:30:09–10:08:37 — 8 minutes of normal traffic, the crash, the automatic restart with its full config and memory accounting, and 30 minutes of normal traffic afterwards. The only thing removed is 246
[dump_input.py:79]lines (multimodalis_embedmasks and the 8.5k-entry block table, ~250 KB of tensor text); the removal point is marked inline and we can send that block raw if you want it. Happy to run any instrumented build or extra diagnostics on this host.Appendix A — log extracts
The full container log is 970 lines (09:30:09–10:08:37). Below are the parts that matter.
Ask and we will send the whole file, including the 246
[dump_input.py:79]lines(multimodal
is_embedmasks + the 8.5k-entry block table) omitted throughout this report.A1. Run-up to the crash
A long prompt is admitted at 09:35:39 (KV jumps to 20.3 %), prefills at 7892 tok/s and
completes normally. Two minutes later the engine dies; note the 09:38:09 → 09:40:37 gap in
the stats series, which is the crash plus the automatic restart.
A2. Scheduler stats at the failing step
A3. Worker traceback — the actual allocation site (both ranks identical)
A4. EngineCore fatal error
A5. Every
expandable_segmentsmapping failure, in orderNote the first two: 10.2 MiB and 8.2 MiB free.
A6. Memory accounting from the automatic restart
A7. Same-shape long-context prefills that survived, after the restart
In both cases a large prompt is admitted (KV ~40 %, two requests briefly resident), then
prefilled at 5–8.5k tok/s without incident. This is why we do not claim the crash is
deterministic for a given context length.