| #861 (yatesdr) |
c6178167 |
fix(qsa): stable DCP top-k candidate order |
Same selected sets on 16/16 shapes |
| #862 (yatesdr) |
128f114e |
fix(scheduler): bound contended prefill compute steps |
Decode gap during a 229k-token prefill: max 286 ms |
| #863 (yatesdr) |
8223193e |
fix(b12x): skip missing multimodal stages during warmup |
Qwen vision on/off starts |
| #864 (yatesdr) |
ccfef884 |
feat(qwen): atomic external QSA checkpoints |
Take together with #883 |
| #866 |
2b86d3d8 |
feat(engine): where a request or the engine loop is held |
Diagnostics only |
| #869, #870 (llitz) |
a94eea09, b5d82593 |
OffloadingConnector fixes for hybrid models |
|
| #883 |
9e5c2f39 |
aligned-transfer capability gate |
Without it #864 stops CACHE_MODE=native Qwen from starting |
| #884 |
8a149a15 |
timeline from HTTP receipt |
Diagnostics only |
| #886 (yatesdr) |
4900f3eb |
full-TP DCP for qwen4_exp checkpoints |
Official Qwen TP4/DCP4 64/64 (with docker #76) |
| #887 |
2f04b107 |
DCP LSE correction for 3/6 ranks; NULL_BLOCK_ID padding rows |
5 failing tests before, 0 after |
| #889 |
5a689e3b |
GLM-5.3-Flash TP3; FP8 dense decode (TP3 only, flag); FlashInfer autotune rank-union save |
TP3 C1 82–85 steps/s. EP warm start hung before the autotune fix; DS4 cold/warm checked |
| #895 |
4f66cb92 |
startup plan saved under the pre-profiling fingerprint |
Opt-in VLLM_ENABLE_STARTUP_PLAN=1 now applies on DS4: warm init 202 s → 153–157 s |
| #898 |
a603af27 |
opt-in GPU memory savings: embedding table in host RAM, MXFP8 vision tower, shared PyNCCL communicator |
Spark TP2 frees 0.9 GiB/GPU: KV 998K → 1.15M tokens (docker #87 enables them). Decode unchanged, prefill −1–3 %, ChartQA 79.8 → 79.5 %. Until this lands, docker #89 keeps canonical Spark TP2 on the old 4-slot recipe |
| #903 |
d281c90f |
boundary-checkpoint copy wait bounded by VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS |
The only step RPC without a timeout; healthy path unchanged |
| #904 |
9575a0bc |
boundary checkpoints for image/video requests, placeholders keyed by content hash |
Needs LMCache #98 (already shared). Image chats restore after restart: 15.4K of 15.4K tokens, was 0; no cross-image restore |
| #905: MiMo-V2.6 series (#829, #831, #832 by logprobz; #871–#881 by MadeBy561) |
b31d20a1 |
MiMo-V2.6-Flash/Pro on b12x: paged_decode DiffKV decode/verify, fused all-reduce+RMSNorm sites, exact TP2 FP8 QKV load, L2 prefetch, audio without torchaudio, DFlash TP sharding and FP8 drafters, recoverable adaptive-verification estimate, uncapped prefill-only steps (opt-in), FP32 router, SM120 router graph guard, FP8 KV reaching MiMo, serve-mimo26-flash.sh |
Needs b12x #420–#424. Tracked in #882; the originals stay open against dev/karmic-kraken: merge them there. #829 is included as its four commits without its upstream merge commit. MiMo TP2: KV 1.31M, needles 32/32, vision 6/6. GLM MTP/DFlash2 and Qwen/DS4.1 unchanged (#905, #882) |
| #897 (yatesdr) |
4b103515 |
executor exits when EngineCore fails to start instead of hanging on its workers |
Test hangs without the fix; healthy path unchanged |
| #908 |
04c30fa9 |
image checkpoint keys carry the full image identity (fixes #904's 30-bit span seed) |
Reported with a real colliding PNG pair; GLM Spark TP2 image restore after restart unchanged (15.4K/15.4K) |
| #911 |
31c6c347 |
DeepSeek-V4.1: BF16 sparse-attention arithmetic by default (VLLM_DS41_ATTENTION_COMPUTE=bf16|reference|auto), FP32 logits (head_dtype=float32 on SM120; quantized heads widened) |
Pairs with b12x#435. Prefill KL vs reference .00841 → .00826; DSpark accepted/draft 1.748 → 1.755; C1/C8/C32 speed unchanged |
| #918 |
1d95593e |
DeepSeek-V4.1: decode page metadata shared across CUDA graphs instead of 117 MiB per graph |
ds41-flash TP4 at GMU 0.95 ran out of memory after start; with the fix KV 9.93–9.95 GiB, ~4.1 GiB free, loop tests 0 errors |
| #919 + #920 |
4cbfc247, 84f56052 |
V4.1 workspace tests refreshed for BF16 attention and L2 prefetch; #920 is #919's fragment |
Test-only: 50/50 pass (was 43/7) |
| #921 |
eba11af7 |
DSML tool parameters without string= are kept (import of vllm-project#56271) |
Python parser tests 128 pass; Rust parser 460 pass |
| #928 (includes tobymao's #917) + #950 |
8586695c, 5df66adc |
the engine stops at once when any rank fails mid-step; the API server then exits with status 1. #950 fixes #928's gather, which dropped replies that answered a later RPC (GLM/Qwen TP>1 died with TimeoutError in wait_for_boundary_checkpoint_copies) |
OOM injected on TP rank 1: before, it hung 332 s with /health 200 and exit 0; now it stops in the same second with exit 1. docker #107 (restart on failure) relies on it. Take #950 with #928: thecrimea's GLM TP4/DCP4 8×30K burst stopped the engine after 318 s without #950, 64/64 with it |
| #929 (replaces #922) |
6ea2fcf6 |
scheduler waits for capacity to restore external checkpoints, fairly and bounded (VLLM_CHECKPOINT_RESTORE_MAX_WAIT_S=60) |
Needs LMCache #103 (already shared). Agents 8×110K: restores 13/82 → 650/650, 17 → 139 tok/s, 0 "admitted without restore" |
| #931 |
635160cc |
structured output: MTP draft rows stay constrained after a skipped async step (#726) |
GLM Spark TP2 MTP JSON probe: open grammar rows 277 → 0, 0 rejections, acceptance 2.40 vs 2.41 |
| #932 |
590be52b |
GLM parser returns reasoning and content with the generated whitespace |
Pairs with docker #110 (already shared), the exact-history GLM template: response-checkpoint reuse 17/17 non-streamed, 20/20 streamed |
| #934 |
40a89cad |
tokenizer max_token_id is vocab_size - 1 |
Needed by docker #112 (already shared), which turns on fastokens in every profile. fastokens counts added tokens in vocab_size, so on DeepSeek-V4.x a prompt token id 129280 (one past the embedding table) passes validation without #934 |
| #930 |
02eae6f2 |
CUDA-graph memory estimate: samples are read after the warmup cache is flushed, count the memory each graph keeps, and include the smallest graph |
DS4.1 with #918 reverted: estimate 1.84 → 5.46 GiB (actual 5.52), so the KV cache shrinks instead of leaving 17 MiB free and running out of memory. Healthy DS4.1: KV 9.94 vs 9.93–9.95 GiB. GLM preset identical |
| #939 (canonical twin: #940, same diff) |
749b07e2 |
wheel build drops Python staged by the other branch from the shared native build cache |
Beta wheel 39ac37f carried 8 canonical-only modules; canonical wheel 502d6cb carried 7 beta-only ones. Fixes both channels |
| #946 (canonical twin: #948) |
82be583d |
decodes preempt in a prefill turn that scheduled nothing |
With prefill compute sharing (GLM default 0.4), a full pool plus requests waiting behind full slots froze generation until the waiters were aborted. mrweiner saw 9.5 min at 0 tok/s. GPU A/B with his setup: base stalled, fix ran 30 min with 0 stalls. Canonical is affected even without restores: #948's test reproduces the stall on ab86b70734 |
| #951 + #954 (canonical twin: #952, the same two commits) |
02ab5382, a2b44f52 |
wheels require CUTLASS DSL 4.7.1: runtime.lock and the wheel normalizer, which rejected anything but 4.6.2 (#951); the wheel build environment installs 4.7.1 (#954) |
Canonical publishing is blocked until #952 lands: b12x master 4bacd509 requires DSL 4.7.1, and the shared FlashInfer wheel and docker lock are already on it (flashinfer#4, docker #116). GPU gate GLM/Qwen/MiMo TP2 + DS4.1 TP4: all pass; decode C1 about −3% on DS4.1 and GLM (b12x launch heuristics); cold first start +2 to +15 min |
| #958 = #957 (randomvariable; replaces #955) |
85314135 |
ModelOpt MXFP8 MoE experts on native B12X (B12X_MXFP8): opt-in with "moe_backend": "b12x" in --speculative-config; auto keeps Marlin. An older b12x or a non-ModelOpt MXFP8 method is rejected at startup |
Needs b12x#452 (beta #453). Qwen QAD step 5500 TP2 with the drafter on B12X_MXFP8: smoke OK, engine step rate equal to the original heads, about 2.5% fewer steps/s than Marlin at C1; Qwen main revision unchanged. Tests: 26 MXFP8, 78 test_b12x.py, 3 spec-decode |
| #959 (canonical twin: #960, same commit) |
980d84ef |
prefill lanes go to queued requests only while the waiting pass can admit one |
With --max-parallel-prefills ≥ 2, queued requests took every lane while running was full (or PAUSED_NEW) and every step scheduled nothing (reported with root cause and fix by netwalker4370, DS4.1 TP4). Qwen TP2 repro: 0/8 in 600 s before, 8/8 in 20 s after. Canonical is affected too: #960's test fails on ab86b70734. Profiles keep lanes off (1) |
| #961 (port of #774 by René Honig; canonical twin: #962, same commit) |
4a379ed428 |
Qwen3.8-Flash-Next: one host-RAM PLE table shared by the processes of a host (VLLM_PLE_TABLE_MEMORY=shared, opt-in) |
Needs b12x#454 (master twin b12x#455). Two TP1 replicas held the 26.82 GiB table once: the first filled it in 40 s, the second attached read-only, a restart attached without reloading. Decode 2×TP1 vs TP2 on the same GPUs: C1 190 vs 243, C8 1024 vs 1022, C16 1523 vs 1491 tok/s. docker #117 (REPLICAS=N) turns it on when vLLM has it |
| #963: CSF loading (#956, stacked on #949) and the Kimi-K3 runtime units #841, #846, #858, #937, #953 (canonical twins: these PRs) |
1ae4c63afd |
vLLM loads MXFP4-CSF (DS4.1, Kimi-K3) and NVFP4-CSF (GLM, Qwen) checkpoints, validates their manifests and slices them per TP rank; DS4.1 CSF experts use native MXFP8 activations. The Kimi units bound InstantTensor staging, reuse native wheel payloads, write FlashInfer autotune caches atomically, keep DSpark/DFlash2 draft precision and DCP cache geometry, add B12X KDA prefill and FP32 MLA partials, and release MLA priming buffers |
Needs b12x#446 → #450 → #459 (in beta through b12x#456). #841, #846, #858, #937, #949 and #953 are open against dev/karmic-kraken. 61 CSF loader tests. Startup KV +34.6% DS4.1, +7.3% GLM, +5.1% Qwen; throughput C8–C32 −2 to −4% DS4.1 TP4, +0.5 to +2.8% GLM TP4, −3 to −5% Qwen TP2 (b12x#450, before #459 and #471). Kimi-K3 TP9 with the units reproduces the production control bit-for-bit at 32k and 524k. Since the Oct 3 rewrite: stored checkpoints only |
| #965 (canonical twin: #968, the same two commits, stacked on #956) |
1286ae9c9a |
GLM-5.3 Spark recipe from the QAD mtp-bf16 source: exactly the 531 Spark projections converted to MXFP8 on load, MTP routed experts to NVFP4 W4A16, strict target coverage. Opt-in; source files unchanged |
TP2 without MTP: loader allocation 92.07 → 88.55 GB per rank; decode BF16 → MXFP8 C1 121.7 → 145.1, C8 511.3 → 558.6 tok/s; 32K prefill 12119 → 12681 tok/s; KL to the source 0.099 nats/token. No profile uses it: docker #120 was withdrawn by #121. #968 also carries #841's commit |
| #970 + #976 + #977 (canonical twin: #972, same commits) |
3c6c8e44bf, be03c271fa, ad3743e358 |
opt-in NVFP4-activation prefill for W4A16 MoE (B12X_W4A16_A4_PREFILL_MIN_TOKENS): activation scales passed in W4A16 mode; the decode rows of a mixed step stay W4A16 |
Needs b12x#472 (master twin b12x#474). GLM-5.3-Flash QAD CSF TP2/DCP2 MTP3 on 2× RTX PRO 6000 Max-Q, threshold 1536: prefill 8K/32K 7490/7728 → 9528/9784 tok/s (+27%), decode unchanged; needle near misses 1.75% vs 0.53% with W4A16. 10 host tests. #976: test_modelopt_nvfp4_moe_input_scales_start_at_zero fails without it and passes with it. #977 (needs b12x#476): 1,154-token prefill 5,868 → 6,498 tok/s, 8K/32K and decode unchanged, needle 492/492/490 vs 490/492/497; 15 host tests |
| #971 (canonical twin: #973, same diff, stacked on #956) |
fba0d19476 |
MXFP4-CSF reader for DeepSeek-V4-Flash (0731) and DeepSeek-V4-Flash-Vision-Exp; routed experts keep MXFP8 activations; expert parallelism is rejected |
DS4-Flash TP2: model 77.29 vs 80.64 GiB, KV 1.85M vs 1.31M tokens (+41.7%). DS4.1 CSF TP4 with the same loader files: KV 12.95M vs 9.57M, prefill −1.3/−1.7% at 8K/32K, decode −1% to −3.5%. 65 tests |
| #974 (canonical twin: #975, same diff, stacked on #956) |
7f172870a0 |
Qwen3.8-Flash-Next takes the PLE storage dtype from checkpoint_root of an FP4-CSF serving directory |
Without it the Qwen QAD CSF checkpoint (NVFP4 PLE) stopped at load with shape mismatch for PLE shard 0 weight. Against the original checkpoint: model 72.39 vs 75.04 GiB, KV 750K vs 552K tokens (+36%), decode C1/C8/C16 190/768/1086 vs 199/781/879 tok/s. 43 tests |
#978 (canonical twin: #979, stacked on feat/kk-fp4-csf; A16 trigger only there, since that base has no A4 prefill) |
ff292a9c63 |
NVFP4-CSF layers expand the next layer's scales on a side stream during its attention (b12x expand_scales), instead of serially before the MoE; VLLM_B12X_CSF_SCALE_PREFETCH=0 turns it off |
Needs b12x#477 (twin b12x#478). glm53-tp2: A4 8K/32K prefill 10,032/10,312 → 10,254/10,620 tok/s, A16 7,672/7,925 → 7,763/8,058; decode unchanged; needle A4 495/490/492, A16 498/495/497. 10 host tests with CUDA streams |
| #980 (canonical twin: #981, same change) |
00a33e2314 |
weights load without upstream vllm-project#41268's max_split_size_mb:20 when PYTORCH_CUDA_ALLOC_CONF enables expandable segments, which otherwise strand pages shared by weights and freed loading temporaries |
Reserved minus allocated after loading, per GPU: glm53-tp2 0.59 → 0.08 GiB (per-layer expansion 0.42 → 0.12); DS4.1 TP4 with b12x#479 before its own fix 1.01 → 0.04 GiB, KV 11.96M → 12.86M tokens. 4 host tests; GLM servers start and serve |
| #983 |
8a5d364f02 |
the startup log names the prefill choice of every NVFP4-CSF layer (W4A16 decode, A4 prefill); a layer without calibrated activation scales (the GLM MTP draft layer) gets one INFO line with its name instead of a warning |
Log only; on beta b3523c2c GLM TP2 logs 84 layer lines and one draft INFO per rank. No canonical PR yet: the code it logs from comes with #970/#977 and #963 |
| #982 (Defilan) |
f1c2508f10 |
the MXFP4-CSF loader admits TP3 for deepseek_v41 checkpoints |
Unit tests in test_mxfp4_csf.py. Weights load at TP3, but on 96-GB GPUs about 93 GiB per GPU leaves no KV, so TP4 stays the minimum there; meant for larger GPUs (e.g. 3× DGX Spark). No canonical PR yet |
@lukealonso this lists everything in the Karmic Kraken beta that canonical does not have yet. Merging it makes
dev/karmic-krakenand b12xmasterequal to our beta branches.State on Oct 5, 18:10 Prague. Beta equals canonical plus exactly the tables below, in both repositories:
dev/karmic-kraken@ab86b70734is an ancestor of betaf1c2508f10, synced by Sync Karmic beta with canonical dev/karmic-kraken 502d6cb5ac #936, Sync Karmic beta with canonical dev/karmic-kraken 0296a817c8 #942 and Sync Karmic beta with canonical dev/karmic-kraken ab86b70734 #947;master@a321f9a6is an ancestor of b12x beta52640cb1, synced by b12x#439, b12x#443, b12x#447 and, inside b12x#456, byd3b8eae3(master's [II] Preserve Kimi parallel-draft KV group identity #441, feat(exl3): online K-quant embedding table (VLLM_EXL3_EMBED_ONLINE_BITS) #436, DSpark: qwen3 external drafts misrouted to deepseek_v4 by architectures-only check; draft FLASH_ATTN selection incompatible with fp8 KV targets #442 and the1.5.0 releasecommit; beta keeps version 1.3.0).Beta history was rewritten on Oct 3 to take out the online CSF scale encoder (compression of native FP4 expert scales while loading). b12x beta went from
3aaece6bto78ee52c3, vLLM beta from8851a645to1286ae9c9a:VLLM_B12X_MOE_FP4_CSF,b12x_scale_compression.py,csf_encode.pyoronline_scales.py. Stored CSF checkpoints still load. The GPU-built inline index from b12x#458 (nvfp4_csf_index.py) stays, because it serves stored checkpoints.New since Oct 4: #983 makes the startup log name the prefill choice of every NVFP4-CSF layer (log only), and #982 (Defilan) lets the MXFP4-CSF loader read DeepSeek-V4.1-Flash at TP3.
New since Oct 1:
dev/karmic-krakenare the twins.Canonical takes the beta PR by PR (list in #984). The snapshot PRs of the whole beta, #894 and b12x#448, were closed on Oct 5 as too large to review.
Taken one by one, the PRs need the conflict resolutions that beta holds: 8 files in vLLM, where #846, #858, #962 and #972 meet the rest of the beta, and two places in b12x (
_tuning.pybetween the CSF stack and master's #436;_impl.pyandplanning.pybetween b12x#474 and the CSF stack).Audit (Oct 5, 18:10 Prague): every file that differs between canonical and beta is traced to the rows below. That is 349 files in vLLM and 186 in b12x. The files not tied to a row are:
.lil/changes/) of the canonical syncs, now includingb12x-master-sync-20261001;vllm-813-withdrawal,b12x-413);pyproject.toml, which differs only by master's 1.5.0 version line.Findings to act on:
4bacd509(CUTLASS DSL 4.7.1). Every canonical build stops at the runtime dependency check:b12x requires nvidia-cutlass-dsl==4.7.1, but 4.6.2 is selected. The shared FlashInfer wheel and the docker runtime lock already moved to 4.7.1 (flashinfer#4, docker [GG] fix: align B12X NVFP4 sparse-MLA scratch API #116), and beta publishes with it (chore(release): wheels require CUTLASS DSL 4.7.1 with B12X #951). Canonical needs chore(release): wheels require CUTLASS DSL 4.7.1 with B12X (canonical twin of #951) #952, the same commits as chore(release): wheels require CUTLASS DSL 4.7.1 with B12X #951 and fix(release): install CUTLASS DSL 4.7.1 in the wheel build environment #954. Until it is merged, canonical keeps failing.TimeoutErrorinwait_for_boundary_checkpoint_copies(reported by two users; reproduced on GLM TP4 and fixed).0d6600e6reports residual heuristic cases 1.65–5.07% slower.output=toapply_mxfp8_marlin_linear; replicated-draft DCP needs backend support;test_mla_prefill_quant_outputpacked sequences;test_w4a16_e2ecases and 3tests/preparationtests.b12x-450-moe-routingand b12x#466b12x-collective-barrier-timeoutwith"models": [], which the publisher rejects. On Oct 4 they got the versions published in beta (fixed there by [II] Preserve CUDA rotary without vendored Python kernels #461 and [II] Preserve aliased LSE in native attention-state merge #468): b12x#450f838b9b6, [II] Bound Kimi vision memory to request inputs #459567b3b87, fix(cache): avoid target EAGLE drop for disaggregated DSpark #466138f4f90. Merged, they give canonical the same fragments as beta.dev/karmic-krakenare open.vLLM →
dev/karmic-krakenc6178167128f114e8223193eccfef8842b86d3d8a94eea09,b5d825939e5c2f39CACHE_MODE=nativeQwen from starting8a149a154900f3ebqwen4_expcheckpoints2f04b1075a689e3b4f66cb92VLLM_ENABLE_STARTUP_PLAN=1now applies on DS4: warm init 202 s → 153–157 sa603af27d281c90fVLLM_EXECUTE_MODEL_TIMEOUT_SECONDS9575a0bcb31d20a1serve-mimo26-flash.shdev/karmic-kraken: merge them there. #829 is included as its four commits without its upstream merge commit. MiMo TP2: KV 1.31M, needles 32/32, vision 6/6. GLM MTP/DFlash2 and Qwen/DS4.1 unchanged (#905, #882)4b10351504c30fa931c6c347VLLM_DS41_ATTENTION_COMPUTE=bf16|reference|auto), FP32 logits (head_dtype=float32on SM120; quantized heads widened)1d95593e4cbfc247,84f56052eba11af7string=are kept (import of vllm-project#56271)8586695c,5df66adcTimeoutErrorinwait_for_boundary_checkpoint_copies)/health200 and exit 0; now it stops in the same second with exit 1. docker #107 (restart on failure) relies on it. Take #950 with #928: thecrimea's GLM TP4/DCP4 8×30K burst stopped the engine after 318 s without #950, 64/64 with it6ea2fcf6VLLM_CHECKPOINT_RESTORE_MAX_WAIT_S=60)635160cc590be52b40a89cadmax_token_idisvocab_size - 1vocab_size, so on DeepSeek-V4.x a prompt token id 129280 (one past the embedding table) passes validation without #93402eae6f2749b07e282be583dab86b7073402ab5382,a2b44f52runtime.lockand the wheel normalizer, which rejected anything but 4.6.2 (#951); the wheel build environment installs 4.7.1 (#954)4bacd509requires DSL 4.7.1, and the shared FlashInfer wheel and docker lock are already on it (flashinfer#4, docker #116). GPU gate GLM/Qwen/MiMo TP2 + DS4.1 TP4: all pass; decode C1 about −3% on DS4.1 and GLM (b12x launch heuristics); cold first start +2 to +15 min85314135B12X_MXFP8): opt-in with"moe_backend": "b12x"in--speculative-config;autokeeps Marlin. An older b12x or a non-ModelOpt MXFP8 method is rejected at startuptest_b12x.py, 3 spec-decode980d84ef--max-parallel-prefills≥ 2, queued requests took every lane while running was full (orPAUSED_NEW) and every step scheduled nothing (reported with root cause and fix by netwalker4370, DS4.1 TP4). Qwen TP2 repro: 0/8 in 600 s before, 8/8 in 20 s after. Canonical is affected too: #960's test fails onab86b70734. Profiles keep lanes off (1)4a379ed428VLLM_PLE_TABLE_MEMORY=shared, opt-in)REPLICAS=N) turns it on when vLLM has it1ae4c63afddev/karmic-kraken. 61 CSF loader tests. Startup KV +34.6% DS4.1, +7.3% GLM, +5.1% Qwen; throughput C8–C32 −2 to −4% DS4.1 TP4, +0.5 to +2.8% GLM TP4, −3 to −5% Qwen TP2 (b12x#450, before #459 and #471). Kimi-K3 TP9 with the units reproduces the production control bit-for-bit at 32k and 524k. Since the Oct 3 rewrite: stored checkpoints only1286ae9c9amtp-bf16source: exactly the 531 Spark projections converted to MXFP8 on load, MTP routed experts to NVFP4 W4A16, strict target coverage. Opt-in; source files unchanged3c6c8e44bf,be03c271fa,ad3743e358B12X_W4A16_A4_PREFILL_MIN_TOKENS): activation scales passed in W4A16 mode; the decode rows of a mixed step stay W4A16test_modelopt_nvfp4_moe_input_scales_start_at_zerofails without it and passes with it. #977 (needs b12x#476): 1,154-token prefill 5,868 → 6,498 tok/s, 8K/32K and decode unchanged, needle 492/492/490 vs 490/492/497; 15 host testsfba0d194767f172870a0checkpoint_rootof an FP4-CSF serving directoryshape mismatch for PLE shard 0 weight. Against the original checkpoint: model 72.39 vs 75.04 GiB, KV 750K vs 552K tokens (+36%), decode C1/C8/C16 190/768/1086 vs 199/781/879 tok/s. 43 testsfeat/kk-fp4-csf; A16 trigger only there, since that base has no A4 prefill)ff292a9c63expand_scales), instead of serially before the MoE;VLLM_B12X_CSF_SCALE_PREFETCH=0turns it off00a33e2314max_split_size_mb:20whenPYTORCH_CUDA_ALLOC_CONFenables expandable segments, which otherwise strand pages shared by weights and freed loading temporaries8a5d364f02W4A16 decode, A4 prefill); a layer without calibrated activation scales (the GLM MTP draft layer) gets one INFO line with its name instead of a warningb3523c2cGLM TP2 logs 84 layer lines and one draft INFO per rank. No canonical PR yet: the code it logs from comes with #970/#977 and #963f1c2508f10deepseek_v41checkpointstest_mxfp4_csf.py. Weights load at TP3, but on 96-GB GPUs about 93 GiB per GPU leaves no KV, so TP4 stays the minimum there; meant for larger GPUs (e.g. 3× DGX Spark). No canonical PR yetNet-neutral, no action:
vllm-926fragment;b0b49023cc, so only beta's fragmentsvllm-943andvllm-943-long-contextremain;7a9de065a2in Sync Karmic beta with canonical dev/karmic-kraken 0296a817c8 #942; only thevllm-941fragment remains;b12x_pcie_all_reduce.py; Sync Karmic beta with canonical dev/karmic-kraken 0296a817c8 #942 keeps beta'sVLLM_DS41_ATTENTION_COMPUTE(DS4.1: BF16 sparse attention arithmetic and FP32 logits by default #911);93dabce32fand58d05bc762; on beta they are1ae4c63afdand1286ae9c9a.Queued for beta, not merged yet (not part of this list): #910 (yatesdr, needs b12x#434) and #913–#916 (tobymao); b12x#434 (yatesdr) and b12x#438 (tobymao). Review comments are on the PRs.
b12x →
masterBesides the rows below, b12x beta differs from master only by:
paged_decode/_preparation.py,_block_fp8_gemv.py), adapted to master's program-cache API in b12x#439;a4b0d1c4e39b437bf8069b2c914921daw8a8_mx,mxfp8_e8m0_k32), intermediate % 32 through split up/gate descriptors; launch views resolved once at preparation; FC1-extent regression testa99a28b2b557d878allocate_storage(..., host_allocator=...))79d3a2d2, #46173836dea; docs39eccb9f(= #459's last commit),78ee52c378ee52c3adds the beta qualification report79d3a2d2, #46736949728, #468566d70cf79d3a2d2nvfp4_csf_index.py, kept from closed b12x#458); Trellis lookup-table shared memory reserved only for Trellis weights; MoE tuning contract at query schema 17 / config schema 102cf055ecB12X_DYNAMIC_DETERMINISTIC_OUTPUT=1, GLM-5.3-Flash with uniform gate/up scales failed before startup. 128 tests; GLM TP2 starts on all three activation policies, NVFP4 repeat KLD exactly zero. Same diff9e2d9844getattrcheck reached beta in #473f8b5b5f1, #4732e44bb5f, #476dfe61efcB12X_W4A16_A4_PREFILL_MIN_TOKENStokens; #473 limits the shared input scale to W4A16 plans; #476 lets the caller choose it per call (bind(a4_prefill=...), used by vllm #977)d4b73cafexpand_scales(experts)expands every expert's NVFP4-CSF scales on the current stream;bind(..., scales_expanded=True)skips the per-call expansion (micro calls, whose barrier reset is fused, still expand)52640cb1GitHub shows b12x#456, #461, #464, #467 and #468 merged as
571c9e82,a77b3f85,66dba3ac,f3f23d69and3aaece6b; since the Oct 3 rewrite the commits on beta are the ones above.All of #420–#424, #427, #440, #446, #452, #455, #465, #466, #470 and #474 are open against
master, and GitHub reports each mergeable witha321f9a6; #450, #459 and #475 are stacked on #446. #417 and #431 are already in master.Shared, nothing to do
LMCache (
integration/local-inference-lab) and blackwell-llm-docker (main) serve both channels, and both branches already include these. Beta images carry all of them (LMCache #108 from betafd566a18on). The last canonical image,b174ad03(Sep 30), has LMCache #108 but not docker #114–#118; canonical gets those when it publishes again (finding 1):on-evictwrites, yatesdr's L1/prefetch fixes, restore waits for RAM, multimodal keys, branch-sibling supersession, restore admission, side-request branch points and shutdown drain, a stalled restore lookup missing instead of killing the engine (perf(dspark): add opt-in rowwise FP8 draft head #106), and intact disk checkpoints kept listed when a restore cannot get RAM (fix(indexer): remove duplicate RoPE quant helper #108, ktsaou's fix(spec decode): harden DSpark and DFlash edge paths #105);REPLICAS=Nsingle-GPU replicas behind one endpoint, [GG] fix: pin same-repository MTP to the target revision #118 MXFP8 MTP drafter experts back on b12x when vLLM and b12x support them. [gg-rebased] perf(dcp): split prefill queries and gather selected CKV #120 and its revert [gg-rebased] perf(nf3): integrate Grid188 hybrid decode #121 cancel out.Where a shared change relies on a beta-only vLLM PR, the row above says so (#928, #932, #934, #951, #961). docker #118 falls back to Marlin without #958 and b12x#453.
Kimi-K3 runtime units that entered beta through #963 are listed in its row; other Kimi-K3 TP9 work is not part of this and will be merged separately.
Previous versions of this issue (Sep 23, Sep 26, Oct 1) are superseded by the tables above.
🤖 Generated with Claude Code