Skip to content

Karmic beta to canonical (Oct 5): rows of #868, through #982 - #894

Closed
voipmonitor wants to merge 496 commits into
dev/karmic-krakenfrom
canonical-candidate/karmic-20260924
Closed

voipmonitor wants to merge 496 commits into
dev/karmic-krakenfrom
canonical-candidate/karmic-20260924

Conversation

@voipmonitor

@voipmonitor voipmonitor commented Sep 24, 2026 •

Copy link
Copy Markdown

Snapshot of integration/karmic-kraken-beta at f1c2508f10 (Oct 5, 18:10 Prague), fast-forwarded from 980d84ef after #961, #963, #965, #970, #971, #974, #976, #977, #978, #980, #983 and #982. Canonical dev/karmic-kraken @ ab86b70734 is an ancestor of this branch, so the diff of this PR is exactly the beta-only changes. The full table with evidence, twins and audit is in #868; the b12x side is local-inference-lab/b12x#448.

New since the Oct 1 snapshot:

The online CSF scale encoder was taken out of beta on Oct 3 (history rewritten); nothing of it is in this branch. Squash-merge it to make dev/karmic-kraken equal to the beta.

🤖 Generated with Claude Code

MadeBy561 and others added 30 commits September 23, 2026 18:52
…at TP > 1

Two per-step costs of DFlash under tensor parallelism:

- The context projection fc (target aux hidden states -> draft hidden, e.g.
  20480 x 4096) was a ReplicatedLinear, so every rank streamed the whole
  weight each step. It is now a ColumnParallelLinear with gather_output=True
  at TP > 1: each output keeps its full dot product (bit-identical), each rank
  reads 1/tp of the weight. MiMo-V2.6-Flash TP4 + DFlash K7: C1 step time
  -60 us, +20K tokens of KV capacity.
- Probabilistic drafting all-gathered the drafter's [rows, vocab] logits
  (224 x 152576 BF16 at C32 K7, ~1.2 ms of a ~33 ms PCIe TP4 step). With
  VLLM_DFLASH_VOCAB_PARALLEL_DRAFT=1 each rank Gumbel-samples its vocab shard
  with the same absolute-token-id noise, and the per-rank winners are merged;
  verification uses the drawn token's logit and a global logsumexp combined
  from per-rank statistics, and rejection resampling runs per shard the same
  way. The all-gathers shrink to [rows, 5] and [reqs, 2]. It applies to plain
  vocab-parallel heads, standard rejection sampling and no watermarking;
  anything else keeps the full-vocab path. Default off.

Tests: tests/v1/spec_decode/test_vocab_parallel_drafting.py simulates TP4
shards against the full-vocab kernels: identical draft tokens and cache
shards, logsumexp within 1e-5 of FP64, identical rejection sampling (accept
lengths, resampled and bonus tokens) across greedy and 0.6-1.3 temperatures,
CUDA-graph pad rows and verification depths 0-7. MiMo-V2.6-Flash TP4 C32:
2182 -> 2263 tok/s.

Signed-off-by: MadeBy561 <madeby561@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Online `fp8_per_block` quantization (speculative_config.quantization)
quantizes the drafter's linears while they load, so the fused context K/V
projection concatenated raw FP8 rows without their block scales and
F.linear failed on a BF16 x FP8 matmul. Dequantize block-FP8 K/V rows with
their scales for the fused projection; each layer keeps its own FP8 GEMMs,
and serialized MXFP8 rows stay packed for their fused linear method.

Signed-off-by: MadeBy561 <madeby561@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…recoverable

The online acceptance estimator learns only from drafts that adaptive
verification admits. A draft predicted near-certain carries almost no IRLS
weight, so a data-poor refit that saw one confidently wrong draft divided
its residual by ~the ridge and moved an intercept by about -20. Predicted
acceptance then fell to ~1e-8, the budget admitted no drafts, and nothing
was ever graded again: speculation stopped for the rest of a long request
(reported as DFlash "falling off" during long thinking).

- Bound each refit's coefficient change (trust region) and keep the slope
  and intercepts within fixed limits, off logistic saturation.
- After EXPLORE_AFTER_IDLE_STEPS consecutive steps that schedule drafts but
  verify none, verify one draft per request for a step so the estimator
  keeps receiving evidence. Every rank makes the same choice.

Only the number of verified drafts changes; rejection sampling keeps the
output distribution exact.

Signed-off-by: MadeBy561 <madeby561@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
max_num_scheduled_tokens bounds a step so that decode streams keep moving
while long prompts prefill. When a step has no runnable decode (a single long
request, or every running request still prefilling), the cap only shrinks the
prefill chunk: more steps and more per-step overhead for the same work.

With VLLM_SCHEDULER_UNCAP_PREFILL_ONLY_STEPS=1, a step whose decode-aware path
finds no eligible decode uses max_num_batched_tokens as its token cap and
budget. Steps with runnable decode are unchanged, and the default (0) keeps
current scheduling for every deployment.

MiMo-V2.6-Flash TP4, batch 4096 / scheduled 2048: standalone prefill at 32K
goes from 11.8K to 12.7K tok/s (+8-12% across 8K-128K).

Tests: tests/v1/core/test_prefill_compute_share_scheduler.py (both settings).

Signed-off-by: MadeBy561 <madeby561@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
MiMo-V2 audio needs two torchaudio transforms: the media parser resamples
input audio to 24 kHz (default method "torchaudio"), and MiMoVLProcessor
computes a 128-bin mel spectrogram with torchaudio.transforms.MelSpectrogram.
Images built on a vendor PyTorch (for example NGC 2.14.0a0) have no matching
torchaudio wheel, so every audio request failed ("torchaudio is required for
audio").

vllm.multimodal.audio gains plain-PyTorch ports of the default
torchaudio.transforms.Resample (Hann-windowed sinc, width 6, rolloff 0.99,
kernel built in float64 and stored as float32) and
torchaudio.transforms.MelSpectrogram (periodic Hann window, reflect-padded
STFT, HTK mel scale, no filterbank normalization). They follow torchaudio's
arithmetic step for step, and are used only when torchaudio is not installed:

- _get_torchaudio_resampler returns the port (cached per rate pair), so
  AudioResampler(method="torchaudio") works everywhere with the same output.
- MiMoVLProcessor uses the ports for its mel spectrogram and its own
  resampler.

Deployments with torchaudio installed take exactly the same code paths as
before.

Tests (tests/multimodal/test_audio.py, tests/models/multimodal/
test_mimo_v2_omni.py):
- with torchaudio: torch.equal against torchaudio for resampling from 8-48
  kHz (mono and stereo) and for MiMo's and torchaudio's default mel settings;
- without torchaudio: pinned reference values, the "torchaudio" resampler
  path, and MiMo's audio preprocessing end to end (frames and token count).
On a torchaudio-less image: 75 passed, 13 skipped (the torchaudio
comparisons). Against torchaudio 2.11 / torch 2.14: 38/38 cases bitwise
identical.

Signed-off-by: MadeBy561 <madeby561@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…lags

serve-mimo26-flash.sh now starts the configuration the MiMo b12x work was
qualified with:

- DFlash drafting (7 tokens) from ${MODEL_PATH}/dflash, drafted across the
  target's TP ranks, with adaptive verification, probabilistic drafts and
  standard rejection sampling; the drafter keeps an FP8 KV cache
  (DFLASH_KV_CACHE_DTYPE).
- The MiMo b12x switches, each overridable: paged-decode KV layout
  (VLLM_KV_CACHE_LAYOUT=BLHNC), BF16 activations for the FP4 MoE, FP32
  router weights in the W4A16 top-k combine, exact dense split-K
  (B12X_DENSE_SPLITK_TURBO=0), BF16 GEMV plans, L2 weight prefetch,
  vocab-parallel drafting, uncapped prefill-only steps, and the PCIe
  all-reduce size thresholds.
- Scheduling: 32 sequences, 4096 batched / 2048 scheduled tokens, prefill
  compute share 0.8, async scheduling, FULL_AND_PIECEWISE graphs up to 256.
- Vision and audio through MiMoV2OmniForCausalLM.
- KV cache FP8 by default: twice the BF16 capacity, and what fits the
  model at TP2. KV_CACHE_DTYPE=bfloat16 selects the exact BF16 cache.
- MODEL_REVISION 5711b268: the first Hugging Face revision whose
  dflash/config.json is valid JSON (fixed in b2674c72); weights unchanged.

Signed-off-by: MadeBy561 <madeby561@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Signed-off-by: Martin Vit <martin@voipmonitor.org>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Signed-off-by: Martin Vit <martin@voipmonitor.org>
…d-store-tracking

[KV Connector] Track allocated blocks and keep CoW group store ranges aligned (from #867)
…ternal-prefix-hits

[KV Connector] Serve external prefix hits for align-mode Mamba hybrids (from #867)
#864 requires an atomic request-boundary checkpoint connector for every Qwen
KV transfer. That also rejected CACHE_MODE=native: vLLM maps native offload
to SimpleCPUOffloadConnector with --recurrent-checkpoint-policy aligned, so
every beta since #864 refused to start Qwen with the native CPU cache.

Aligned retention does not use request-boundary checkpoints. Accept it when
the connector declares supports_aligned_hybrid_transfer (SimpleCPU, and
OffloadingConnector with the #869/#870 fixes) and DCP is 1. Other connectors,
DCP and the auto/request_boundaries policies keep the atomic requirement.

Based on commit 4 of #867, narrowed to connectors that move aligned hybrid
state.

Co-authored-by: Carlos Augusto <mb@lab.how>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Signed-off-by: Martin Vit <martin@voipmonitor.org>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Signed-off-by: Martin Vit <martin@voipmonitor.org>
…load-gate

fix(config): let aligned Qwen offload through the atomic-checkpoint gate
The request timeline started at the renderer's arrival stamp, so a request
held after the server received it but before rendering, for example while
its body was still arriving, looked fast. An outermost ASGI middleware now
stamps the HTTP receipt and the end of the request body, the timeline
reports both spans, and a request that ends without any response after the
stall threshold is logged even though it never reached the engine.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Signed-off-by: Martin Vit <martin@voipmonitor.org>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Signed-off-by: Martin Vit <martin@voipmonitor.org>
…-timeline

feat(serve): time requests from HTTP receipt, not only from rendering
…hrough stale block-table rows (vllm-project#56734)

Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit d2983f2)

Signed-off-by: Martin Vit <martin@voipmonitor.org>
Qwen3.8-Flash-Next checkpoints are published under two model types that
load the same Qwen4Exp implementation: the official ones declare
qwen4_exp / qwen4_exp_text, renamed ones declare qwen3_8_flash_next /
qwen3_8_flash_next_text. Only Qwen3_8FlashNextTextConfig opted in to
supports_full_tp_dcp_with_kv_gather, so ModelConfig rejected TP4/DCP4
(two KV heads) for qwen4_exp checkpoints while accepting it for renamed
ones.

Declare the capability on Qwen4ExpTextConfig so both spellings inherit
it. Only the B12X QSA backend implements DCP; the non-B12X NVIDIA and
AMD QSA backends still refuse DCP > 1 when they are constructed.
voipmonitor and others added 7 commits October 4, 2026 05:54
…refill

b12x MoE: opt-in NVFP4-activation prefill for W4A16 experts, decode rows stay W4A16
Add the deepseek_v4_flash MXFP4-CSF family (43 target layers of 256
routed experts, hidden 4096, MoE intermediate 2048, DeepSeek-native
tensor names, TP1/2/4/8) and dispatch it to DeepseekV4Mxfp4CsfConfig.

The config derives from the class CUDA registers as deepseek_v4_fp8 and
parses the retained source block-FP8 fields, so attention, dense,
shared-expert, vision and MTP/DSpark draft weights keep their native
methods. Target routed experts read compressed scales with MXFP8
activations unless VLLM_B12X_MOE_FP4_FORCE_A16=1 requests BF16. Expert
parallelism is rejected: MegaMoE experts bypass quantization methods
and would never read the compressed experts.

The loader family table records each family's compressed layer range
and rejects reads outside it. The FP4-CSF guide documents the DeepSeek
MXFP8 activation default and the DeepSeek-V4-Flash serving contract.

Validation: tests/quantization/test_mxfp4_csf.py and test_nvfp4_csf.py
pass on CPU (74 tests, 13 new); ruff, mypy hook, typos pass. No GPU
serving run yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lash

Serve lossless MXFP4-CSF DeepSeek-V4-Flash and Vision-Exp checkpoints
An FP4-CSF serving directory holds only metadata; its config points at the
checkpoint root whose tensors/ directory holds the weights, PLE shards
included. The PLE dtype probe read the serving directory, found no shards,
and left the dtype at its default, so the Qwen3.8-Flash-Next QAD CSF
checkpoint (NVFP4 PLE table) failed to load with a PLE shape mismatch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
qwen4_exp: resolve the PLE storage dtype from an FP4-CSF checkpoint root
@voipmonitor voipmonitor changed the title Karmic beta to canonical (Oct 1): #861–#866, #869, #870, #883, #884, #886, #887, #889, #895, #897, #898, #903, #904, #908, MiMo-V2.6 (#905), #911, #918–#921, #928–#932, #934, #939, #946, #950, #951, #954, #958, #959 Karmic beta to canonical (Oct 4): rows of #868, through #974 Oct 4, 2026
voipmonitor and others added 12 commits October 4, 2026 13:07
W4A16_NVFP4 experts store no input scales (GLM-5.3-Flash keeps its MTP
layer's routed experts weight-only), so the registered parameters kept
whatever torch.empty returned. With B12X_W4A16_A4_PREFILL_MIN_TOKENS set,
the b12x MoE takes positive finite input scales for calibrated ones, and
the draft layer could run its prefill with NVFP4 activations under a
garbage global scale (5.2e-32 in a GLM Spark TP2 run). Zero reads as
absent: the existing guard keeps such layers on W4A16.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ed-input-scales

modelopt: zero-initialize NVFP4 MoE input scales so W4A16 layers never enable A4 prefill from garbage
With B12X_W4A16_A4_PREFILL_MIN_TOKENS set, the decode rows of a step run
W4A16 and its prefill rows run A4, whatever the call size (b12x
bind(a4_prefill=...)). Before, prefill rows took A4 only when the call had at
least the threshold's tokens, so a prompt's precision depended on how the
scheduler chunked and batched it. A captured CUDA graph keeps W4A16 for all
its rows, and layers without activation scales (the MTP draft layer) are not
split at all.

A GDN/KDA layer counts prefills on the host, so a decode-only step finds its
decode rows without a device read; steps with prefill rows read
query_start_loc once, as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-row-type

b12x MoE: choose the A4 prefill by row type, not by call size
Layers of an NVFP4-CSF model share one scale scratch, and every eager prefill
call expanded its experts' scales into it serially before its MoE (4% of a
GLM-5.3 glm53-tp2 A4 prefill, with nothing overlapping). After a layer's MoE
the method now expands the following layer's scales on a side stream, ordered
after that MoE, and the following layer waits for it and binds with
scales_expanded=True (b12x expand_scales). Only eager steps that expand -
A4 prefill rows or A16 calls above the stage-read limit - prefetch; CUDA graph
capture, decode-only steps and other forward passes are untouched, and any
pending expansion is waited for before the scratch is used.
VLLM_B12X_CSF_SCALE_PREFETCH=0 turns it off.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…prefetch

NVFP4-CSF: expand the next layer's scales during its attention
vllm-project#41268 sets max_split_size_mb:20 while loading weights so cached blocks do
not fragment. With expandable segments the allocator releases free pages
instead of whole blocks, and the limit instead strands the pages that
persistent weights share with freed loading temporaries: empty_cache()
cannot unmap them, and the memory profile then charges them against the KV
cache.

Measured with torch.cuda.memory_reserved() - memory_allocated() after
loading, per GPU:
- GLM-5.3-Flash glm53-tp2 (NVFP4-CSF, stage-read scales): 0.59 -> 0.08 GiB;
  with per-layer expansion 0.42 -> 0.12 GiB. Both start and serve; the
  preset's fixed KV cache leaves that much more headroom.
- DeepSeek-V4.1-Flash TP4 with b12x#479's inline scales before its own
  allocator fix: 1.01 -> 0.04 GiB, KV cache 11.96M -> 12.86M tokens, weight
  loading 143 -> 120 s.

The classic allocator keeps the limit.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Martin Vit <martin@voipmonitor.org>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Martin Vit <martin@voipmonitor.org>
…dable-segments

[Worker] Load weights without the split limit under expandable segments
@voipmonitor voipmonitor changed the title Karmic beta to canonical (Oct 4): rows of #868, through #974 Karmic beta to canonical (Oct 4): rows of #868, through #980 Oct 4, 2026
Defilan and others added 5 commits October 4, 2026 20:10
The deepseek_v41 reader admitted TP 1, 2, 4 and 8 only. At TP3 each rank
owns 768 of the 2304 routed-expert channels, which is 32-channel aligned
and falls on whole 16-row scale slabs, so the slicer and the b12x W4A8
CSF geometry already hold. Add 3 for deepseek_v41 only; the
deepseek_v4_flash (2048 channels) and kimi_k3 lists are unchanged.

Tests: the TP byte-slicing test also runs at TP3, and a new test pins
which families admit TP3.

Signed-off-by: Christopher Maher <chris@mahercode.io>
Signed-off-by: Christopher Maher <chris@mahercode.io>
FP4-CSF MoE layers logged 'native NVFP4 A16' even when their A4
prefill path was prepared, and the MTP draft layer, which has no
calibrated input scales, warned 'A4 prefill disabled for it'. Users
read that as hybrid A4/A16 having fallen back to full A16.

CSF layers now log 'W4A16 decode, A4 prefill'. The draft-layer message
names the layer, explains that it keeps A16, and is INFO.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
[b12x] Clarify A4 prefill logs for FP4-CSF and draft MoE layers
MXFP4-CSF: load DeepSeek-V4.1-Flash checkpoints at TP3
@voipmonitor voipmonitor changed the title Karmic beta to canonical (Oct 4): rows of #868, through #980 Karmic beta to canonical (Oct 5): rows of #868, through #982 Oct 5, 2026
@voipmonitor

Copy link
Copy Markdown
Author

Closing: canonical takes the beta PR by PR. The list of PRs, each with its canonical PR, is in #984. This snapshot of the whole beta (349 files, +29K lines) is too large to review or merge as one PR.

@voipmonitor voipmonitor closed this Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.