Karmic beta to canonical (Oct 5): rows of #868, through #982 - #894
Closed
voipmonitor wants to merge 496 commits into
Closed
voipmonitor wants to merge 496 commits into
voipmonitor wants to merge 496 commits into
Conversation
…at TP > 1 Two per-step costs of DFlash under tensor parallelism: - The context projection fc (target aux hidden states -> draft hidden, e.g. 20480 x 4096) was a ReplicatedLinear, so every rank streamed the whole weight each step. It is now a ColumnParallelLinear with gather_output=True at TP > 1: each output keeps its full dot product (bit-identical), each rank reads 1/tp of the weight. MiMo-V2.6-Flash TP4 + DFlash K7: C1 step time -60 us, +20K tokens of KV capacity. - Probabilistic drafting all-gathered the drafter's [rows, vocab] logits (224 x 152576 BF16 at C32 K7, ~1.2 ms of a ~33 ms PCIe TP4 step). With VLLM_DFLASH_VOCAB_PARALLEL_DRAFT=1 each rank Gumbel-samples its vocab shard with the same absolute-token-id noise, and the per-rank winners are merged; verification uses the drawn token's logit and a global logsumexp combined from per-rank statistics, and rejection resampling runs per shard the same way. The all-gathers shrink to [rows, 5] and [reqs, 2]. It applies to plain vocab-parallel heads, standard rejection sampling and no watermarking; anything else keeps the full-vocab path. Default off. Tests: tests/v1/spec_decode/test_vocab_parallel_drafting.py simulates TP4 shards against the full-vocab kernels: identical draft tokens and cache shards, logsumexp within 1e-5 of FP64, identical rejection sampling (accept lengths, resampled and bonus tokens) across greedy and 0.6-1.3 temperatures, CUDA-graph pad rows and verification depths 0-7. MiMo-V2.6-Flash TP4 C32: 2182 -> 2263 tok/s. Signed-off-by: MadeBy561 <madeby561@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Online `fp8_per_block` quantization (speculative_config.quantization) quantizes the drafter's linears while they load, so the fused context K/V projection concatenated raw FP8 rows without their block scales and F.linear failed on a BF16 x FP8 matmul. Dequantize block-FP8 K/V rows with their scales for the fused projection; each layer keeps its own FP8 GEMMs, and serialized MXFP8 rows stay packed for their fused linear method. Signed-off-by: MadeBy561 <madeby561@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…recoverable The online acceptance estimator learns only from drafts that adaptive verification admits. A draft predicted near-certain carries almost no IRLS weight, so a data-poor refit that saw one confidently wrong draft divided its residual by ~the ridge and moved an intercept by about -20. Predicted acceptance then fell to ~1e-8, the budget admitted no drafts, and nothing was ever graded again: speculation stopped for the rest of a long request (reported as DFlash "falling off" during long thinking). - Bound each refit's coefficient change (trust region) and keep the slope and intercepts within fixed limits, off logistic saturation. - After EXPLORE_AFTER_IDLE_STEPS consecutive steps that schedule drafts but verify none, verify one draft per request for a step so the estimator keeps receiving evidence. Every rank makes the same choice. Only the number of verified drafts changes; rejection sampling keeps the output distribution exact. Signed-off-by: MadeBy561 <madeby561@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
max_num_scheduled_tokens bounds a step so that decode streams keep moving while long prompts prefill. When a step has no runnable decode (a single long request, or every running request still prefilling), the cap only shrinks the prefill chunk: more steps and more per-step overhead for the same work. With VLLM_SCHEDULER_UNCAP_PREFILL_ONLY_STEPS=1, a step whose decode-aware path finds no eligible decode uses max_num_batched_tokens as its token cap and budget. Steps with runnable decode are unchanged, and the default (0) keeps current scheduling for every deployment. MiMo-V2.6-Flash TP4, batch 4096 / scheduled 2048: standalone prefill at 32K goes from 11.8K to 12.7K tok/s (+8-12% across 8K-128K). Tests: tests/v1/core/test_prefill_compute_share_scheduler.py (both settings). Signed-off-by: MadeBy561 <madeby561@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
MiMo-V2 audio needs two torchaudio transforms: the media parser resamples
input audio to 24 kHz (default method "torchaudio"), and MiMoVLProcessor
computes a 128-bin mel spectrogram with torchaudio.transforms.MelSpectrogram.
Images built on a vendor PyTorch (for example NGC 2.14.0a0) have no matching
torchaudio wheel, so every audio request failed ("torchaudio is required for
audio").
vllm.multimodal.audio gains plain-PyTorch ports of the default
torchaudio.transforms.Resample (Hann-windowed sinc, width 6, rolloff 0.99,
kernel built in float64 and stored as float32) and
torchaudio.transforms.MelSpectrogram (periodic Hann window, reflect-padded
STFT, HTK mel scale, no filterbank normalization). They follow torchaudio's
arithmetic step for step, and are used only when torchaudio is not installed:
- _get_torchaudio_resampler returns the port (cached per rate pair), so
AudioResampler(method="torchaudio") works everywhere with the same output.
- MiMoVLProcessor uses the ports for its mel spectrogram and its own
resampler.
Deployments with torchaudio installed take exactly the same code paths as
before.
Tests (tests/multimodal/test_audio.py, tests/models/multimodal/
test_mimo_v2_omni.py):
- with torchaudio: torch.equal against torchaudio for resampling from 8-48
kHz (mono and stereo) and for MiMo's and torchaudio's default mel settings;
- without torchaudio: pinned reference values, the "torchaudio" resampler
path, and MiMo's audio preprocessing end to end (frames and token count).
On a torchaudio-less image: 75 passed, 13 skipped (the torchaudio
comparisons). Against torchaudio 2.11 / torch 2.14: 38/38 cases bitwise
identical.
Signed-off-by: MadeBy561 <madeby561@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…lags
serve-mimo26-flash.sh now starts the configuration the MiMo b12x work was
qualified with:
- DFlash drafting (7 tokens) from ${MODEL_PATH}/dflash, drafted across the
target's TP ranks, with adaptive verification, probabilistic drafts and
standard rejection sampling; the drafter keeps an FP8 KV cache
(DFLASH_KV_CACHE_DTYPE).
- The MiMo b12x switches, each overridable: paged-decode KV layout
(VLLM_KV_CACHE_LAYOUT=BLHNC), BF16 activations for the FP4 MoE, FP32
router weights in the W4A16 top-k combine, exact dense split-K
(B12X_DENSE_SPLITK_TURBO=0), BF16 GEMV plans, L2 weight prefetch,
vocab-parallel drafting, uncapped prefill-only steps, and the PCIe
all-reduce size thresholds.
- Scheduling: 32 sequences, 4096 batched / 2048 scheduled tokens, prefill
compute share 0.8, async scheduling, FULL_AND_PIECEWISE graphs up to 256.
- Vision and audio through MiMoV2OmniForCausalLM.
- KV cache FP8 by default: twice the BF16 capacity, and what fits the
model at TP2. KV_CACHE_DTYPE=bfloat16 selects the exact BF16 cache.
- MODEL_REVISION 5711b268: the first Hugging Face revision whose
dflash/config.json is valid JSON (fixed in b2674c72); weights unchanged.
Signed-off-by: MadeBy561 <madeby561@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: MadeBy561 <madeby561@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Martin Vit <martin@voipmonitor.org>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Martin Vit <martin@voipmonitor.org>
…d-store-tracking [KV Connector] Track allocated blocks and keep CoW group store ranges aligned (from #867)
…ternal-prefix-hits [KV Connector] Serve external prefix hits for align-mode Mamba hybrids (from #867)
#864 requires an atomic request-boundary checkpoint connector for every Qwen KV transfer. That also rejected CACHE_MODE=native: vLLM maps native offload to SimpleCPUOffloadConnector with --recurrent-checkpoint-policy aligned, so every beta since #864 refused to start Qwen with the native CPU cache. Aligned retention does not use request-boundary checkpoints. Accept it when the connector declares supports_aligned_hybrid_transfer (SimpleCPU, and OffloadingConnector with the #869/#870 fixes) and DCP is 1. Other connectors, DCP and the auto/request_boundaries policies keep the atomic requirement. Based on commit 4 of #867, narrowed to connectors that move aligned hybrid state. Co-authored-by: Carlos Augusto <mb@lab.how> Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Martin Vit <martin@voipmonitor.org>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Martin Vit <martin@voipmonitor.org>
…load-gate fix(config): let aligned Qwen offload through the atomic-checkpoint gate
The request timeline started at the renderer's arrival stamp, so a request held after the server received it but before rendering, for example while its body was still arriving, looked fast. An outermost ASGI middleware now stamps the HTTP receipt and the end of the request body, the timeline reports both spans, and a request that ends without any response after the stall threshold is logged even though it never reached the engine. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Martin Vit <martin@voipmonitor.org>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: Martin Vit <martin@voipmonitor.org>
…-timeline feat(serve): time requests from HTTP receipt, not only from rendering
…hrough stale block-table rows (vllm-project#56734) Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit d2983f2) Signed-off-by: Martin Vit <martin@voipmonitor.org>
Qwen3.8-Flash-Next checkpoints are published under two model types that load the same Qwen4Exp implementation: the official ones declare qwen4_exp / qwen4_exp_text, renamed ones declare qwen3_8_flash_next / qwen3_8_flash_next_text. Only Qwen3_8FlashNextTextConfig opted in to supports_full_tp_dcp_with_kv_gather, so ModelConfig rejected TP4/DCP4 (two KV heads) for qwen4_exp checkpoints while accepting it for renamed ones. Declare the capability on Qwen4ExpTextConfig so both spellings inherit it. Only the B12X QSA backend implements DCP; the non-B12X NVIDIA and AMD QSA backends still refuse DCP > 1 when they are constructed.
…refill b12x MoE: opt-in NVFP4-activation prefill for W4A16 experts, decode rows stay W4A16
Add the deepseek_v4_flash MXFP4-CSF family (43 target layers of 256 routed experts, hidden 4096, MoE intermediate 2048, DeepSeek-native tensor names, TP1/2/4/8) and dispatch it to DeepseekV4Mxfp4CsfConfig. The config derives from the class CUDA registers as deepseek_v4_fp8 and parses the retained source block-FP8 fields, so attention, dense, shared-expert, vision and MTP/DSpark draft weights keep their native methods. Target routed experts read compressed scales with MXFP8 activations unless VLLM_B12X_MOE_FP4_FORCE_A16=1 requests BF16. Expert parallelism is rejected: MegaMoE experts bypass quantization methods and would never read the compressed experts. The loader family table records each family's compressed layer range and rejects reads outside it. The FP4-CSF guide documents the DeepSeek MXFP8 activation default and the DeepSeek-V4-Flash serving contract. Validation: tests/quantization/test_mxfp4_csf.py and test_nvfp4_csf.py pass on CPU (74 tests, 13 new); ruff, mypy hook, typos pass. No GPU serving run yet. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lash Serve lossless MXFP4-CSF DeepSeek-V4-Flash and Vision-Exp checkpoints
An FP4-CSF serving directory holds only metadata; its config points at the checkpoint root whose tensors/ directory holds the weights, PLE shards included. The PLE dtype probe read the serving directory, found no shards, and left the dtype at its default, so the Qwen3.8-Flash-Next QAD CSF checkpoint (NVFP4 PLE table) failed to load with a PLE shape mismatch. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
qwen4_exp: resolve the PLE storage dtype from an FP4-CSF checkpoint root
W4A16_NVFP4 experts store no input scales (GLM-5.3-Flash keeps its MTP layer's routed experts weight-only), so the registered parameters kept whatever torch.empty returned. With B12X_W4A16_A4_PREFILL_MIN_TOKENS set, the b12x MoE takes positive finite input scales for calibrated ones, and the draft layer could run its prefill with NVFP4 activations under a garbage global scale (5.2e-32 in a GLM Spark TP2 run). Zero reads as absent: the existing guard keeps such layers on W4A16. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ed-input-scales modelopt: zero-initialize NVFP4 MoE input scales so W4A16 layers never enable A4 prefill from garbage
With B12X_W4A16_A4_PREFILL_MIN_TOKENS set, the decode rows of a step run W4A16 and its prefill rows run A4, whatever the call size (b12x bind(a4_prefill=...)). Before, prefill rows took A4 only when the call had at least the threshold's tokens, so a prompt's precision depended on how the scheduler chunked and batched it. A captured CUDA graph keeps W4A16 for all its rows, and layers without activation scales (the MTP draft layer) are not split at all. A GDN/KDA layer counts prefills on the host, so a decode-only step finds its decode rows without a device read; steps with prefill rows read query_start_loc once, as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-row-type b12x MoE: choose the A4 prefill by row type, not by call size
Layers of an NVFP4-CSF model share one scale scratch, and every eager prefill call expanded its experts' scales into it serially before its MoE (4% of a GLM-5.3 glm53-tp2 A4 prefill, with nothing overlapping). After a layer's MoE the method now expands the following layer's scales on a side stream, ordered after that MoE, and the following layer waits for it and binds with scales_expanded=True (b12x expand_scales). Only eager steps that expand - A4 prefill rows or A16 calls above the stage-read limit - prefetch; CUDA graph capture, decode-only steps and other forward passes are untouched, and any pending expansion is waited for before the scratch is used. VLLM_B12X_CSF_SCALE_PREFETCH=0 turns it off. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…prefetch NVFP4-CSF: expand the next layer's scales during its attention
vllm-project#41268 sets max_split_size_mb:20 while loading weights so cached blocks do not fragment. With expandable segments the allocator releases free pages instead of whole blocks, and the limit instead strands the pages that persistent weights share with freed loading temporaries: empty_cache() cannot unmap them, and the memory profile then charges them against the KV cache. Measured with torch.cuda.memory_reserved() - memory_allocated() after loading, per GPU: - GLM-5.3-Flash glm53-tp2 (NVFP4-CSF, stage-read scales): 0.59 -> 0.08 GiB; with per-layer expansion 0.42 -> 0.12 GiB. Both start and serve; the preset's fixed KV cache leaves that much more headroom. - DeepSeek-V4.1-Flash TP4 with b12x#479's inline scales before its own allocator fix: 1.01 -> 0.04 GiB, KV cache 11.96M -> 12.86M tokens, weight loading 143 -> 120 s. The classic allocator keeps the limit. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Martin Vit <martin@voipmonitor.org>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Martin Vit <martin@voipmonitor.org>
…dable-segments [Worker] Load weights without the split limit under expandable segments
The deepseek_v41 reader admitted TP 1, 2, 4 and 8 only. At TP3 each rank owns 768 of the 2304 routed-expert channels, which is 32-channel aligned and falls on whole 16-row scale slabs, so the slicer and the b12x W4A8 CSF geometry already hold. Add 3 for deepseek_v41 only; the deepseek_v4_flash (2048 channels) and kimi_k3 lists are unchanged. Tests: the TP byte-slicing test also runs at TP3, and a new test pins which families admit TP3. Signed-off-by: Christopher Maher <chris@mahercode.io>
Signed-off-by: Christopher Maher <chris@mahercode.io>
FP4-CSF MoE layers logged 'native NVFP4 A16' even when their A4 prefill path was prepared, and the MTP draft layer, which has no calibrated input scales, warned 'A4 prefill disabled for it'. Users read that as hybrid A4/A16 having fallen back to full A16. CSF layers now log 'W4A16 decode, A4 prefill'. The draft-layer message names the layer, explains that it keeps A16, and is INFO. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
[b12x] Clarify A4 prefill logs for FP4-CSF and draft MoE layers
MXFP4-CSF: load DeepSeek-V4.1-Flash checkpoints at TP3
Author
|
Closing: canonical takes the beta PR by PR. The list of PRs, each with its canonical PR, is in #984. This snapshot of the whole beta (349 files, +29K lines) is too large to review or merge as one PR. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Snapshot of
integration/karmic-kraken-betaatf1c2508f10(Oct 5, 18:10 Prague), fast-forwarded from980d84efafter #961, #963, #965, #970, #971, #974, #976, #977, #978, #980, #983 and #982. Canonicaldev/karmic-kraken@ab86b70734is an ancestor of this branch, so the diff of this PR is exactly the beta-only changes. The full table with evidence, twins and audit is in #868; the b12x side is local-inference-lab/b12x#448.New since the Oct 1 snapshot:
max_split_size_mb:20when expandable segments are on. Twin [Worker] Load weights without the split limit under expandable segments (canonical twin of #980) #981.The online CSF scale encoder was taken out of beta on Oct 3 (history rewritten); nothing of it is in this branch. Squash-merge it to make
dev/karmic-krakenequal to the beta.🤖 Generated with Claude Code