Skip to content

Canonical review, detailed: per-PR table and file audit for dev/karmic-kraken and b12x master (updated Oct 5) #868

Description

@voipmonitor

The list of PRs to merge is #984. This issue keeps the commits, evidence and file audit behind it.

@lukealonso this lists everything in the Karmic Kraken beta that canonical does not have yet. Merging it makes dev/karmic-kraken and b12x master equal to our beta branches.

State on Oct 5, 18:10 Prague. Beta equals canonical plus exactly the tables below, in both repositories:

Beta history was rewritten on Oct 3 to take out the online CSF scale encoder (compression of native FP4 expert scales while loading). b12x beta went from 3aaece6b to 78ee52c3, vLLM beta from 8851a645 to 1286ae9c9a:

New since Oct 4: #983 makes the startup log name the prefill choice of every NVFP4-CSF layer (log only), and #982 (Defilan) lets the MXFP4-CSF loader read DeepSeek-V4.1-Flash at TP3.

New since Oct 1:

Canonical takes the beta PR by PR (list in #984). The snapshot PRs of the whole beta, #894 and b12x#448, were closed on Oct 5 as too large to review.

Taken one by one, the PRs need the conflict resolutions that beta holds: 8 files in vLLM, where #846, #858, #962 and #972 meet the rest of the beta, and two places in b12x (_tuning.py between the CSF stack and master's #436; _impl.py and planning.py between b12x#474 and the CSF stack).

Audit (Oct 5, 18:10 Prague): every file that differs between canonical and beta is traced to the rows below. That is 349 files in vLLM and 186 in b12x. The files not tied to a row are:

  • release fragments (.lil/changes/) of the canonical syncs, now including b12x-master-sync-20261001;
  • two older fragments (vllm-813-withdrawal, b12x-413);
  • b12x pyproject.toml, which differs only by master's 1.5.0 version line.

Findings to act on:

  1. Canonical images have not published since b12x 4bacd509 (CUTLASS DSL 4.7.1). Every canonical build stops at the runtime dependency check: b12x requires nvidia-cutlass-dsl==4.7.1, but 4.6.2 is selected. The shared FlashInfer wheel and the docker runtime lock already moved to 4.7.1 (flashinfer#4, docker [GG] fix: align B12X NVFP4 sparse-MLA scratch API #116), and beta publishes with it (chore(release): wheels require CUTLASS DSL 4.7.1 with B12X #951). Canonical needs chore(release): wheels require CUTLASS DSL 4.7.1 with B12X (canonical twin of #951) #952, the same commits as chore(release): wheels require CUTLASS DSL 4.7.1 with B12X #951 and fix(release): install CUTLASS DSL 4.7.1 in the wheel build environment #954. Until it is merged, canonical keeps failing.
  2. fix(executor): stop the engine at once when any rank fails mid-step #928 must go with fix(executor): keep replies that answer a later RPC (GLM/Qwen TP>1 RPC timeout) #950. fix(executor): stop the engine at once when any rank fails mid-step #928's gather dropped the replies of a collective RPC sent behind a single-rank one. On GLM and Qwen with TP > 1, the engine then died with TimeoutError in wait_for_boundary_checkpoint_copies (reported by two users; reproduced on GLM TP4 and fixed).
  3. GPU gate of the CUTLASS DSL 4.7.1 switch (GLM Spark TP2, Qwen TP2, MiMo TP2, DS4.1 TP4): all pass.
    • Decode C1 is about −3% on DS4.1 and GLM; b12x 0d6600e6 reports residual heuristic cases 1.65–5.07% slower.
    • The first start after the update compiles again: +2 to +15 min.
    • Details are in b12x#447.
  4. Failing tests on canonical. These fail on canonical vLLM + master too, not only on beta; details in Sync Karmic beta with canonical dev/karmic-kraken 502d6cb5ac #936:
    • vLLM: Kimi vision tests pass output= to apply_mxfp8_marlin_linear; replicated-draft DCP needs backend support; test_mla_prefill_quant_output packed sequences;
    • b12x: 70 test_w4a16_e2e cases and 3 tests/preparation tests.
  5. Release fragments of three b12x twins fixed. b12x#450 and [II] Bound Kimi vision memory to request inputs #459 carried b12x-450-moe-routing and b12x#466 b12x-collective-barrier-timeout with "models": [], which the publisher rejects. On Oct 4 they got the versions published in beta (fixed there by [II] Preserve CUDA rotary without vendored Python kernels #461 and [II] Preserve aliased LSE in native attention-state merge #468): b12x#450 f838b9b6, [II] Bound Kimi vision memory to request inputs #459 567b3b87, fix(cache): avoid target EAGLE drop for disaggregated DSpark #466 138f4f90. Merged, they give canonical the same fragments as beta.
  6. Three parts of b12x#456 have no master PR: the GPU-built inline NVFP4-CSF index, Trellis lookup-table shared memory reserved only for Trellis weights (master has the same unguarded condition from feat(exl3): online K-quant embedding table (VLLM_EXL3_EMBED_ONLINE_BITS) #436), and the MoE tuning contract at query schema 17 / config schema 10. They have no master PR yet.
  7. The beta includes the Kimi-K3 runtime units [KK] Bound InstantTensor staging and retain deferred checkpoint ownership #841, [KK] Reuse identical native payloads and recover incomplete autotune caches #846, [KK] Preserve parallel-draft precision and mixed DCP cache geometry #858, [KK] Prepare B12X Kimi KDA prefill and expose MLA partial precision #937, [KK] Load DS4.1 and Kimi-K3 lossless X4T checkpoints #949 and [KK] Release dense MLA priming buffers after preparation #953, which [beta] Integrate CSF loading and online scale compression with the Kimi runtime #963 brought in; their own PRs against dev/karmic-kraken are open.

vLLM → dev/karmic-kraken

PR Merge in beta Change Evidence / notes
#861 (yatesdr) c6178167 fix(qsa): stable DCP top-k candidate order Same selected sets on 16/16 shapes
#862 (yatesdr) 128f114e fix(scheduler): bound contended prefill compute steps Decode gap during a 229k-token prefill: max 286 ms
#863 (yatesdr) 8223193e fix(b12x): skip missing multimodal stages during warmup Qwen vision on/off starts
#864 (yatesdr) ccfef884 feat(qwen): atomic external QSA checkpoints Take together with #883
#866 2b86d3d8 feat(engine): where a request or the engine loop is held Diagnostics only
#869, #870 (llitz) a94eea09, b5d82593 OffloadingConnector fixes for hybrid models
#883 9e5c2f39 aligned-transfer capability gate Without it #864 stops CACHE_MODE=native Qwen from starting
#884 8a149a15 timeline from HTTP receipt Diagnostics only
#886 (yatesdr) 4900f3eb full-TP DCP for qwen4_exp checkpoints Official Qwen TP4/DCP4 64/64 (with docker #76)
#887 2f04b107 DCP LSE correction for 3/6 ranks; NULL_BLOCK_ID padding rows 5 failing tests before, 0 after
#889 5a689e3b GLM-5.3-Flash TP3; FP8 dense decode (TP3 only, flag); FlashInfer autotune rank-union save TP3 C1 82–85 steps/s. EP warm start hung before the autotune fix; DS4 cold/warm checked
#895 4f66cb92 startup plan saved under the pre-profiling fingerprint Opt-in VLLM_ENABLE_STARTUP_PLAN=1 now applies on DS4: warm init 202 s → 153–157 s
#898 a603af27 opt-in GPU memory savings: embedding table in host RAM, MXFP8 vision tower, shared PyNCCL communicator Spark TP2 frees 0.9 GiB/GPU: KV 998K → 1.15M tokens (docker #87 enables them). Decode unchanged, prefill −1–3 %, ChartQA 79.8 → 79.5 %. Until this lands, docker #89 keeps canonical Spark TP2 on the old 4-slot recipe
#903 d281c90f boundary-checkpoint copy wait bounded by VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS The only step RPC without a timeout; healthy path unchanged
#904 9575a0bc boundary checkpoints for image/video requests, placeholders keyed by content hash Needs LMCache #98 (already shared). Image chats restore after restart: 15.4K of 15.4K tokens, was 0; no cross-image restore
#905: MiMo-V2.6 series (#829, #831, #832 by logprobz; #871–#881 by MadeBy561) b31d20a1 MiMo-V2.6-Flash/Pro on b12x: paged_decode DiffKV decode/verify, fused all-reduce+RMSNorm sites, exact TP2 FP8 QKV load, L2 prefetch, audio without torchaudio, DFlash TP sharding and FP8 drafters, recoverable adaptive-verification estimate, uncapped prefill-only steps (opt-in), FP32 router, SM120 router graph guard, FP8 KV reaching MiMo, serve-mimo26-flash.sh Needs b12x #420–#424. Tracked in #882; the originals stay open against dev/karmic-kraken: merge them there. #829 is included as its four commits without its upstream merge commit. MiMo TP2: KV 1.31M, needles 32/32, vision 6/6. GLM MTP/DFlash2 and Qwen/DS4.1 unchanged (#905, #882)
#897 (yatesdr) 4b103515 executor exits when EngineCore fails to start instead of hanging on its workers Test hangs without the fix; healthy path unchanged
#908 04c30fa9 image checkpoint keys carry the full image identity (fixes #904's 30-bit span seed) Reported with a real colliding PNG pair; GLM Spark TP2 image restore after restart unchanged (15.4K/15.4K)
#911 31c6c347 DeepSeek-V4.1: BF16 sparse-attention arithmetic by default (VLLM_DS41_ATTENTION_COMPUTE=bf16|reference|auto), FP32 logits (head_dtype=float32 on SM120; quantized heads widened) Pairs with b12x#435. Prefill KL vs reference .00841 → .00826; DSpark accepted/draft 1.748 → 1.755; C1/C8/C32 speed unchanged
#918 1d95593e DeepSeek-V4.1: decode page metadata shared across CUDA graphs instead of 117 MiB per graph ds41-flash TP4 at GMU 0.95 ran out of memory after start; with the fix KV 9.93–9.95 GiB, ~4.1 GiB free, loop tests 0 errors
#919 + #920 4cbfc247, 84f56052 V4.1 workspace tests refreshed for BF16 attention and L2 prefetch; #920 is #919's fragment Test-only: 50/50 pass (was 43/7)
#921 eba11af7 DSML tool parameters without string= are kept (import of vllm-project#56271) Python parser tests 128 pass; Rust parser 460 pass
#928 (includes tobymao's #917) + #950 8586695c, 5df66adc the engine stops at once when any rank fails mid-step; the API server then exits with status 1. #950 fixes #928's gather, which dropped replies that answered a later RPC (GLM/Qwen TP>1 died with TimeoutError in wait_for_boundary_checkpoint_copies) OOM injected on TP rank 1: before, it hung 332 s with /health 200 and exit 0; now it stops in the same second with exit 1. docker #107 (restart on failure) relies on it. Take #950 with #928: thecrimea's GLM TP4/DCP4 8×30K burst stopped the engine after 318 s without #950, 64/64 with it
#929 (replaces #922) 6ea2fcf6 scheduler waits for capacity to restore external checkpoints, fairly and bounded (VLLM_CHECKPOINT_RESTORE_MAX_WAIT_S=60) Needs LMCache #103 (already shared). Agents 8×110K: restores 13/82 → 650/650, 17 → 139 tok/s, 0 "admitted without restore"
#931 635160cc structured output: MTP draft rows stay constrained after a skipped async step (#726) GLM Spark TP2 MTP JSON probe: open grammar rows 277 → 0, 0 rejections, acceptance 2.40 vs 2.41
#932 590be52b GLM parser returns reasoning and content with the generated whitespace Pairs with docker #110 (already shared), the exact-history GLM template: response-checkpoint reuse 17/17 non-streamed, 20/20 streamed
#934 40a89cad tokenizer max_token_id is vocab_size - 1 Needed by docker #112 (already shared), which turns on fastokens in every profile. fastokens counts added tokens in vocab_size, so on DeepSeek-V4.x a prompt token id 129280 (one past the embedding table) passes validation without #934
#930 02eae6f2 CUDA-graph memory estimate: samples are read after the warmup cache is flushed, count the memory each graph keeps, and include the smallest graph DS4.1 with #918 reverted: estimate 1.84 → 5.46 GiB (actual 5.52), so the KV cache shrinks instead of leaving 17 MiB free and running out of memory. Healthy DS4.1: KV 9.94 vs 9.93–9.95 GiB. GLM preset identical
#939 (canonical twin: #940, same diff) 749b07e2 wheel build drops Python staged by the other branch from the shared native build cache Beta wheel 39ac37f carried 8 canonical-only modules; canonical wheel 502d6cb carried 7 beta-only ones. Fixes both channels
#946 (canonical twin: #948) 82be583d decodes preempt in a prefill turn that scheduled nothing With prefill compute sharing (GLM default 0.4), a full pool plus requests waiting behind full slots froze generation until the waiters were aborted. mrweiner saw 9.5 min at 0 tok/s. GPU A/B with his setup: base stalled, fix ran 30 min with 0 stalls. Canonical is affected even without restores: #948's test reproduces the stall on ab86b70734
#951 + #954 (canonical twin: #952, the same two commits) 02ab5382, a2b44f52 wheels require CUTLASS DSL 4.7.1: runtime.lock and the wheel normalizer, which rejected anything but 4.6.2 (#951); the wheel build environment installs 4.7.1 (#954) Canonical publishing is blocked until #952 lands: b12x master 4bacd509 requires DSL 4.7.1, and the shared FlashInfer wheel and docker lock are already on it (flashinfer#4, docker #116). GPU gate GLM/Qwen/MiMo TP2 + DS4.1 TP4: all pass; decode C1 about −3% on DS4.1 and GLM (b12x launch heuristics); cold first start +2 to +15 min
#958 = #957 (randomvariable; replaces #955) 85314135 ModelOpt MXFP8 MoE experts on native B12X (B12X_MXFP8): opt-in with "moe_backend": "b12x" in --speculative-config; auto keeps Marlin. An older b12x or a non-ModelOpt MXFP8 method is rejected at startup Needs b12x#452 (beta #453). Qwen QAD step 5500 TP2 with the drafter on B12X_MXFP8: smoke OK, engine step rate equal to the original heads, about 2.5% fewer steps/s than Marlin at C1; Qwen main revision unchanged. Tests: 26 MXFP8, 78 test_b12x.py, 3 spec-decode
#959 (canonical twin: #960, same commit) 980d84ef prefill lanes go to queued requests only while the waiting pass can admit one With --max-parallel-prefills ≥ 2, queued requests took every lane while running was full (or PAUSED_NEW) and every step scheduled nothing (reported with root cause and fix by netwalker4370, DS4.1 TP4). Qwen TP2 repro: 0/8 in 600 s before, 8/8 in 20 s after. Canonical is affected too: #960's test fails on ab86b70734. Profiles keep lanes off (1)
#961 (port of #774 by René Honig; canonical twin: #962, same commit) 4a379ed428 Qwen3.8-Flash-Next: one host-RAM PLE table shared by the processes of a host (VLLM_PLE_TABLE_MEMORY=shared, opt-in) Needs b12x#454 (master twin b12x#455). Two TP1 replicas held the 26.82 GiB table once: the first filled it in 40 s, the second attached read-only, a restart attached without reloading. Decode 2×TP1 vs TP2 on the same GPUs: C1 190 vs 243, C8 1024 vs 1022, C16 1523 vs 1491 tok/s. docker #117 (REPLICAS=N) turns it on when vLLM has it
#963: CSF loading (#956, stacked on #949) and the Kimi-K3 runtime units #841, #846, #858, #937, #953 (canonical twins: these PRs) 1ae4c63afd vLLM loads MXFP4-CSF (DS4.1, Kimi-K3) and NVFP4-CSF (GLM, Qwen) checkpoints, validates their manifests and slices them per TP rank; DS4.1 CSF experts use native MXFP8 activations. The Kimi units bound InstantTensor staging, reuse native wheel payloads, write FlashInfer autotune caches atomically, keep DSpark/DFlash2 draft precision and DCP cache geometry, add B12X KDA prefill and FP32 MLA partials, and release MLA priming buffers Needs b12x#446 → #450 → #459 (in beta through b12x#456). #841, #846, #858, #937, #949 and #953 are open against dev/karmic-kraken. 61 CSF loader tests. Startup KV +34.6% DS4.1, +7.3% GLM, +5.1% Qwen; throughput C8–C32 −2 to −4% DS4.1 TP4, +0.5 to +2.8% GLM TP4, −3 to −5% Qwen TP2 (b12x#450, before #459 and #471). Kimi-K3 TP9 with the units reproduces the production control bit-for-bit at 32k and 524k. Since the Oct 3 rewrite: stored checkpoints only
#965 (canonical twin: #968, the same two commits, stacked on #956) 1286ae9c9a GLM-5.3 Spark recipe from the QAD mtp-bf16 source: exactly the 531 Spark projections converted to MXFP8 on load, MTP routed experts to NVFP4 W4A16, strict target coverage. Opt-in; source files unchanged TP2 without MTP: loader allocation 92.07 → 88.55 GB per rank; decode BF16 → MXFP8 C1 121.7 → 145.1, C8 511.3 → 558.6 tok/s; 32K prefill 12119 → 12681 tok/s; KL to the source 0.099 nats/token. No profile uses it: docker #120 was withdrawn by #121. #968 also carries #841's commit
#970 + #976 + #977 (canonical twin: #972, same commits) 3c6c8e44bf, be03c271fa, ad3743e358 opt-in NVFP4-activation prefill for W4A16 MoE (B12X_W4A16_A4_PREFILL_MIN_TOKENS): activation scales passed in W4A16 mode; the decode rows of a mixed step stay W4A16 Needs b12x#472 (master twin b12x#474). GLM-5.3-Flash QAD CSF TP2/DCP2 MTP3 on 2× RTX PRO 6000 Max-Q, threshold 1536: prefill 8K/32K 7490/7728 → 9528/9784 tok/s (+27%), decode unchanged; needle near misses 1.75% vs 0.53% with W4A16. 10 host tests. #976: test_modelopt_nvfp4_moe_input_scales_start_at_zero fails without it and passes with it. #977 (needs b12x#476): 1,154-token prefill 5,868 → 6,498 tok/s, 8K/32K and decode unchanged, needle 492/492/490 vs 490/492/497; 15 host tests
#971 (canonical twin: #973, same diff, stacked on #956) fba0d19476 MXFP4-CSF reader for DeepSeek-V4-Flash (0731) and DeepSeek-V4-Flash-Vision-Exp; routed experts keep MXFP8 activations; expert parallelism is rejected DS4-Flash TP2: model 77.29 vs 80.64 GiB, KV 1.85M vs 1.31M tokens (+41.7%). DS4.1 CSF TP4 with the same loader files: KV 12.95M vs 9.57M, prefill −1.3/−1.7% at 8K/32K, decode −1% to −3.5%. 65 tests
#974 (canonical twin: #975, same diff, stacked on #956) 7f172870a0 Qwen3.8-Flash-Next takes the PLE storage dtype from checkpoint_root of an FP4-CSF serving directory Without it the Qwen QAD CSF checkpoint (NVFP4 PLE) stopped at load with shape mismatch for PLE shard 0 weight. Against the original checkpoint: model 72.39 vs 75.04 GiB, KV 750K vs 552K tokens (+36%), decode C1/C8/C16 190/768/1086 vs 199/781/879 tok/s. 43 tests
#978 (canonical twin: #979, stacked on feat/kk-fp4-csf; A16 trigger only there, since that base has no A4 prefill) ff292a9c63 NVFP4-CSF layers expand the next layer's scales on a side stream during its attention (b12x expand_scales), instead of serially before the MoE; VLLM_B12X_CSF_SCALE_PREFETCH=0 turns it off Needs b12x#477 (twin b12x#478). glm53-tp2: A4 8K/32K prefill 10,032/10,312 → 10,254/10,620 tok/s, A16 7,672/7,925 → 7,763/8,058; decode unchanged; needle A4 495/490/492, A16 498/495/497. 10 host tests with CUDA streams
#980 (canonical twin: #981, same change) 00a33e2314 weights load without upstream vllm-project#41268's max_split_size_mb:20 when PYTORCH_CUDA_ALLOC_CONF enables expandable segments, which otherwise strand pages shared by weights and freed loading temporaries Reserved minus allocated after loading, per GPU: glm53-tp2 0.59 → 0.08 GiB (per-layer expansion 0.42 → 0.12); DS4.1 TP4 with b12x#479 before its own fix 1.01 → 0.04 GiB, KV 11.96M → 12.86M tokens. 4 host tests; GLM servers start and serve
#983 8a5d364f02 the startup log names the prefill choice of every NVFP4-CSF layer (W4A16 decode, A4 prefill); a layer without calibrated activation scales (the GLM MTP draft layer) gets one INFO line with its name instead of a warning Log only; on beta b3523c2c GLM TP2 logs 84 layer lines and one draft INFO per rank. No canonical PR yet: the code it logs from comes with #970/#977 and #963
#982 (Defilan) f1c2508f10 the MXFP4-CSF loader admits TP3 for deepseek_v41 checkpoints Unit tests in test_mxfp4_csf.py. Weights load at TP3, but on 96-GB GPUs about 93 GiB per GPU leaves no KV, so TP4 stays the minimum there; meant for larger GPUs (e.g. 3× DGX Spark). No canonical PR yet

Net-neutral, no action:

Queued for beta, not merged yet (not part of this list): #910 (yatesdr, needs b12x#434) and #913–#916 (tobymao); b12x#434 (yatesdr) and b12x#438 (tobymao). Review comments are on the PRs.

b12x → master

Besides the rows below, b12x beta differs from master only by:

  • the job-scoped compile cache of the two MiMo compile factories (paged_decode/_preparation.py, _block_fp8_gemv.py), adapted to master's program-cache API in b12x#439;
  • release fragments.
PR Merge in beta Change Evidence
b12x#427 #428 a4b0d1c4 PCIe one-shot and DMA all-reduce at world size 3; PDL flags take effect DMA 12/12 and one-shot RMS pass at world size 3. GLM TP4 A/B/A neutral, 32/32
b12x#420–#424 (MadeBy561) #432 e39b437b MiMo-V2.6 kernels: paged_decode split-KV decode/verify with KV append; tensor-core BF16 GEMV; small-row block-FP8 GEMV; optional FP32 W4A16 top-k weights; 16-byte gated activation Tests in #882. Needed by vllm #905. #420 and #422 also need b12x#439's compile-cache scoping
b12x#440 (= beta's #435, same diff) #435 f8069b2c V4.1 FP4 KV writer divides like the DeepSeek reference, so ties round the same way About 0.05% of written values differed from the reference (5.5% on tie data); now bit-exact. 70 tests pass. Pairs with vllm #911
b12x#452 (randomvariable; replaces b12x#449) #453 914921da Native MXFP8 W8A8 grouped MoE (w8a8_mx, mxfp8_e8m0_k32), intermediate % 32 through split up/gate descriptors; launch views resolved once at preparation; FC1-extent regression test 48 MXFP8/execution-model/FC1 tests; MoE regression 86/17, identical to beta. Needed by vllm #957. Beta has it cherry-picked without the 1.5.0 version commit
b12x#470 #469 a99a28b2 W4A16 large-M route blocks skip the MMAs of empty 16-row blocks (bit-identical) GLM-5.3 TP2 MoE layer 1–8% faster at 1K–8K tokens; prefill +2.1–2.4% at 8K/32K, decode unchanged. 33 tests on master; beta suites 249 passed / 70 failed, the same 70 as the beta tip
b12x#455 #454 b557d878 PLE tables in caller-provided mapped-host regions (allocate_storage(..., host_allocator=...)) 68 PLE tests. Needed by vllm #961 (two TP1 replicas map one 26.82 GiB table). Same diff
b12x#446 → b12x#450 → b12x#459 (stacked) #456 79d3a2d2, #461 73836dea; docs 39eccb9f (= #459's last commit), 78ee52c3 MXFP4-CSF (DS4.1, Kimi-K3) and NVFP4-CSF (GLM, Qwen) scales prepared through fused-MoE weight preparation; DS4.1 CSF experts on native MXFP8; NVFP4-CSF scale operands reconstructed cooperatively in dynamic MoE. #461 fixes a fragment; 78ee52c3 adds the beta qualification report Needed by vllm #963. After the rewrite: CSF suites 228 passed; W4A16 suites 330 passed / 68 failed, the same 68 as before. KV and speed are in the #963 row. #450 and #459 carry #461's fragment fix since Oct 4
b12x#466 (randomvariable's #384, plus #467) #456 79d3a2d2, #467 36949728, #468 566d70cf compiled-program ownership through the PCIe/RoCE launcher factories; collective priming through a deadline-enforced barrier, whose deadline is validated only for sessions that use it (#467). #468 fixes #467's fragment 116 preparation, launcher and RoCE tests; every other file of #466 equals beta's. #466 carries #468's fragment fix since Oct 4
none #456 79d3a2d2 beta-only parts of #456: inline NVFP4-CSF index built on the GPU (nvfp4_csf_index.py, kept from closed b12x#458); Trellis lookup-table shared memory reserved only for Trellis weights; MoE tuning contract at query schema 17 / config schema 10 The tuning contract joins the CSF stack (16/9) and master's #436 (15/8), which conflict. The Trellis guard keeps the IQ2_XS and Q8_0 candidates within their shared memory; master has the same unguarded condition from #436
b12x#465 #464 2cf055ec MoE output determinism resolved before candidate selection With B12X_DYNAMIC_DETERMINISTIC_OUTPUT=1, GLM-5.3-Flash with uniform gate/up scales failed before startup. 128 tests; GLM TP2 starts on all three activation policies, NVFP4 repeat KLD exactly zero. Same diff
b12x#475 (stacked on #459) #471 9e2d9844 W4A16 reads NVFP4-CSF expert scales per pipeline stage up to 1536 tokens instead of expanding them per layer GLM-5.3-Flash QAD CSF TP2/DCP2 MTP3 on Max-Q: decode C8 +10%, C1 +1.6% per step, prefill unchanged; needles 498/500/500 vs 498/492. Same code; its getattr check reached beta in #473
b12x#474 (starts with b12x#423's two commits) #472 f8b5b5f1, #473 2e44bb5f, #476 dfe61efc opt-in NVFP4-activation prefill over the packed W4A16 weights for calls of at least B12X_W4A16_A4_PREFILL_MIN_TOKENS tokens; #473 limits the shared input scale to W4A16 plans; #476 lets the caller choose it per call (bind(a4_prefill=...), used by vllm #977) Needed by vllm #970. MoE layer at 3072/8192 tokens: 1684/3201 vs 4114/8735 us; prefill +27%, decode unchanged, near misses 1.75% vs 0.53%; 73 A4 and CSF tests. #474 leaves out the CSF hooks, which master does not have
b12x#478 (stacked on b12x#475) #477 d4b73caf expand_scales(experts) expands every expert's NVFP4-CSF scales on the current stream; bind(..., scales_expanded=True) skips the per-call expansion (micro calls, whose barrier reset is fused, still expand) Needed by vllm #978. Bit-exact against native with poisoned scratch; 77 CSF tests on the twin
b12x#481 (stacked on b12x#478) #479 52640cb1 compact (N % 128 == 64) W4A8 kernels read MXFP4-CSF expert scales inline up to 1536-token plans; larger plans expand every expert's scales once from the same storage (Triton, compiled with the plan); storage built in a private memory pool DS4.1-Flash TP4 CSF vs per-call expansion: decode C8/C16/C32 +3.6/+3.6/+4.1% (= uncompressed checkpoint), prefill unchanged, KV 12.79M vs 13.11M tokens; bit-exact; 208 tests on the twin

GitHub shows b12x#456, #461, #464, #467 and #468 merged as 571c9e82, a77b3f85, 66dba3ac, f3f23d69 and 3aaece6b; since the Oct 3 rewrite the commits on beta are the ones above.

All of #420–#424, #427, #440, #446, #452, #455, #465, #466, #470 and #474 are open against master, and GitHub reports each mergeable with a321f9a6; #450, #459 and #475 are stacked on #446. #417 and #431 are already in master.

Shared, nothing to do

LMCache (integration/local-inference-lab) and blackwell-llm-docker (main) serve both channels, and both branches already include these. Beta images carry all of them (LMCache #108 from beta fd566a18 on). The last canonical image, b174ad03 (Sep 30), has LMCache #108 but not docker #114–#118; canonical gets those when it publishes again (finding 1):

Where a shared change relies on a beta-only vLLM PR, the row above says so (#928, #932, #934, #951, #961). docker #118 falls back to Marlin without #958 and b12x#453.

Kimi-K3 runtime units that entered beta through #963 are listed in its row; other Kimi-K3 TP9 work is not part of this and will be merged separately.


Previous versions of this issue (Sep 23, Sep 26, Oct 1) are superseded by the tables above.

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions