perf(GFX1100-TG200): T27 warp-cooperative QuantizeQ8KK for decode - #2796
perf(GFX1100-TG200): T27 warp-cooperative QuantizeQ8KK for decode#2796ghazni101 wants to merge 10 commits into
Conversation
d02bd01 to
61ac7fc
Compare
61ac7fc to
0b88e0c
Compare
|
Held back with the rest of the I reviewed this change on its own and have no objection to it. Once #2790's base Landing today from this set: #2782 (with the grouped-Q8_0 repair), #2777 and |
0b88e0c to
48f9390
Compare
|
Rebased onto the repaired stack (base b9f2ef4 = the external-contributor landing branch; chain tips: lever-C 0ac919d, T27 48f9390, T24 5d8a95d, T21 6b7deb7, T25 2ff6af4). Verification on gfx1100 (RX 7900 XTX, container rocm-dev:10.0.0, -DCMAKE_HIP_ARCHITECTURES=gfx1100): record gates green (check-env-doc, check-agent-record, check-rocm-dp4a-intrinsic); vllm-cli + test_rocm_quant_dot compile and link at -Werror; the DEFAULT path (all arms off) produces byte-identical completions to the staging baseline on Qwen3.5-4B Q4_K_M (canonical TG200 prompt, 256 tokens, greedy, seed 0) — leg-by-leg bisect: lever-C/T27/T24/T21/T25 all IDENTICAL. One default-path defect was found and fixed during the rebase: T21's VT_GDN_ROWPERM_KEEP_QUANT was default-ON and silently flipped trunk numerics; it is now opt-in by value like its COLPERM sibling, and the token-exact gate passes again. This branch's head is now 48f9390; content otherwise unchanged from the reviewed version (plus the rebase repairs each commit body discloses). Ping for re-review. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true |
48f9390 to
bce9ec4
Compare
… kROCM The GGUF loader routes a block-typed weight to MatmulBTQuant whenever the running device has the provider, so registering these two ops lights up keep-quant compute on every ROCm board with no model-path change: the dense and grouped MoE towers stage once through ResidentWeight and dispatch to the new device GEMM. Coverage mirrors the CUDA sibling exactly — the ten Q8_K-family encodings plus a native Q8_0 arm. The integer dots are the portable scalar forms of the CPU reference bodies in the CPU accumulation order, because gfx1100 exposes no signed byte dot (v_dot4_i32_iu8 is unsigned-only; sdot4 needs a feature this target does not offer), and the gate is bit-exactness against the CPU tier at NMSE 1e-6 with the f64 dequant band at 5e-4. Unsupported dtypes throw naming the dtype instead of silently falling back to a host kernel that cannot follow device pointers; VT_GGUF_KEEP_QUANT=0 restores load-time expansion. Gates on gfx1100 / ROCm 7.14.0: test_rocm_quant_dot 132,094 assertions green across all ten encodings (decode through prefill shapes, broadcast and per-row grouped arms over a poisoned output buffer), focused ctest 'rocm|cross_device|quant' 20/21 with only the pre-existing MoeSiluMul bf16 exactness failure (mudler#1588) remaining, and an end-to-end Qwen3.5-0.8B Q4_K_M decode that is deterministic on device. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:ox-alpha [omp]
The new kROCM provider takes over kMatmulBTQuantGrouped from the kernel in rocm_grouped_gemm.hip and delegates Q4_K/Q5_K/Q6_K back to it, but not Q8_0. Q8_0 has no arm in rocm_quant_dot.hip either -- it dots a Q8_0 activation rather than a Q8_K super-block, so IsRocmKeepQuantSupported answers no and a grouped Q8_0 expert GEMM throws on a path main serves today. Adds Q8_0 to the delegation list, and a q8_0 row to the test's kCases table so the grouped arm has a case that fails when the delegation is dropped. The table was the ten Q8_K-family encodings only, which is why nothing caught it. Closes mudler#2927. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:claude-opus-5-1m [claude-code]
…CM provider" This reverts commit 82d99ea. The fix is correct and mudler#2927 stays open for it, but this pull request is the base of a 22-branch stack and every later branch edits the same two files. Landing the repair here made 21 of them conflict; off this branch the stack merges clean. So the repair moves to its own branch on top of the landed stack, where it costs no conflict resolution at all. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:claude-opus-5-1m [claude-code]
The new kROCM provider takes over kMatmulBTQuantGrouped from the kernel in rocm_grouped_gemm.hip and delegates Q4_K/Q5_K/Q6_K back to it, but not Q8_0. Q8_0 has no arm in rocm_quant_dot.hip either -- it dots a Q8_0 activation rather than a Q8_K super-block, so IsRocmKeepQuantSupported answers no and a grouped Q8_0 expert GEMM throws on a path main serves today. Adds Q8_0 to the delegation list, and a q8_0 row to the test's kCases table so the grouped arm has a case that fails when the delegation is dropped. The table was the ten Q8_K-family encodings only, which is why nothing caught it. Closes mudler#2927. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:claude-opus-5-1m [claude-code]
Issue mudler#1588 still lacks active cache-state evidence and a three-mode ROCm correctness gate. This spec fixes the post-write probes, dtype audit, tolerance policy, tests, review mutations, and hardware evidence before implementation starts. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:gpt-5.6-sol [codex]
The rejected plan treated a red local gate as usable, invented a numerical envelope, and described oracle and provider paths that could not run. Bind the work to mudler#2773, keep both correctness prerequisites pending, and make the future evidence recipe executable without claiming unavailable results. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:gpt-6 [codex]
The mudler#2773 plan must describe the production caller and active oracle layout before instrumentation starts. Correct BF16 selector normalization, name the existing Qwen3.5 path, and record its shared-seam debt in mudler#2923. Separate SD storage from DS dump order and cite the active CPU attention test with its unchanged tolerances. Runtime acceptance remains pending. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:gpt-6-astra [codex]
Adds the VT_GEMV_MMVQ=1 opt-in K-quant decode GEMV arm for MatmulBTQuant, bit-exact vs the CPU oracle. The arm folds activation quant into the MMVQ GEMV prologue (deleting the standalone QuantizeQ8KK launch) and widens the gate to engine dtypes (bf16/f16 activations, bf16/f32 outputs). Sub-levers: - lever-B1: VT_GEMV_MMVQ_FOLD_MAX makes the fold crossover tunable at runtime - lever-B2: VT_SKINNY_BF16=1 f32-out decode-skinny arm for GDN BA projections - repair: m-gates the whole dispatch and makes the GEMV bit-equal to baseline - repair-2: host-side dispatch-route counters + F1/F2 routing-witness gates - lever-B2 test: red-first f32-out decode-skinny gate, true-unset routing window Architecture: F1 moved the live MatmulBTQuantKernelRocm to rocm_quant_dot.hip (anonymous namespace, internal linkage). T4a's MMVQ arm lives in rocm_grouped_gemm.hip's version (external linkage, renamed to *Gdn). This PR adds delegation: rocm_quant_dot.hip forwards Q4_K/Q5_K/Q6_K calls to the Gdn version, preserving F1's IQ-type providers while activating T4a's MMVQ arm. The default path (VT_GEMV_MMVQ unset) is byte-unchanged from F1. The arm is opt-in and validated by test_rocm_quant_dot (6/6 cases, 719 assertions) and test_rocm_skinny_f32 (2/2 cases, 51 assertions). Token-identical to upstream baseline on Qwen3.5-4B Q4_K, 32-token greedy decode, seed 0. Rebased onto the external-contributor landing branch (staging tip b9f2ef4), which already carries F1 (mudler#2782) and its grouped-Q8_0 repair (mudler#2927). This version carries NONE of the removals the previous base commit 51f5222 made: the ten-row kCases table, kMaxNmseErr/nmse_ref_max and the three F1 provider cases are restored beside this arm's kKQuantCases (mudler#2938); Dp4a keeps the __ockl_sdot4 hardware dot (mudler#2939); the documented VT_ROCM_Q8K_BLOCK selector (SelectQ8KQuantArm/LaunchQ8KQuantizer), the mudler#2472 cooperative gfx1100 default (QuantizeQ8KCooperativeK) and the VT_ROCM_Q6K_SMALL_PRIVATE A/B arm are restored with their witness helpers; the shared bench-evidence file keeps lever B1's section 14 record, whose truncation this branch had carried. Depends on mudler#2782 (F1 keep-quant GEMM infra). FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:GLM-5-2 [OMP] Assisted-by: AGENT:OMEN-ALPHA [OMP]
…Norm epilogue Lever-C adds an opt-in fused norm-quant epilogue (VT_NORM_QUANT_FUSED=1): RmsNormRowKernel emits the row's Q8_K blocks alongside its normal output, and MatmulBTQuant's K-quant branch skips the standalone QuantizeQ8KK when the consuming activation matches the producer token. Byte-identical to the standalone path by construction (shared QuantQ8KSBlock body). New files: - src/vt/rocm/rocm_act_quant.h: shared Q8_K quant-block body - src/vt/rocm/rocm_norm_quant_bridge.h: producer-consumer token contract Also fixes T4a routing counter placement (moved outside anonymous namespace for external linkage) and restores VT_GEMV_MMVQ_FOLD_MAX env var reading that was lost during cherry-pick conflict resolution. The default path (VT_NORM_QUANT_FUSED unset) is byte-unchanged. Validated by test_rocm_quant_dot (12/12 cases, 797 assertions). Token-identical to upstream baseline on Qwen3.5-4B Q4_K, 32-token greedy decode, seed 0. Depends on mudler#2782 (F1) and mudler#2790 (T4a). FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:GLM-5-2 [OMP]
The standalone QuantizeQ8KK kernel used 1 thread per 256-element superblock, each doing a serial scan of 256 elements (~800 instructions). For decode (m=1, nsb=10) only 10 of 128 threads were active, and on wave32 each thread is its own wave, so the kernel took ~13.4 us/call = 540 us/tok (6.0% of wall time). The new QuantizeQ8KKWarpCoop kernel uses 8 threads per superblock (32 elements each). The amax scan is done per-chunk (ascending, ax > amax first-occurrence), then reduced across 8 threads via __shfl_xor_sync with lower-chunk-index tie-break — equivalent to a sequential scan of all 256 elements. The quantization (iscale = -127/mx, DNearestInt, clamp 127) and bsums are order-independent. Output is BYTE-IDENTICAL to the original QuantQ8KSBlock, asserted by the gate test (16/16, 839 assertions) under VT_QUANT_Q8K_WARP=1. For m=1, nsb=10: 1 block, 80/128 threads active (vs 10/128), 3 waves of ~100 instructions (vs 10 waves of ~800) = ~8x fewer wave-cycles. A/B on acceptance workload (Qwen3.5-4B Q4_K_M, 256 tokens, temp 0, seed 0): OFF median: 91.532 tok/s ON median: 93.417 tok/s +2.06%, 5/5 pairs ON>OFF, all 5 byte-identical (1039 bytes) Gated by VT_QUANT_Q8K_WARP (default OFF, read per-call). Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:GLM-5-2 [OMP] FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:GLM [omp]
bce9ec4 to
fb8b60c
Compare
|
Rebased onto FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true |
Closes #2795.
Row:
GFX1100-TG200T27 adds the
VT_QUANT_Q8K_WARP=1opt-in warp-cooperativeQuantizeQ8KKWarpCoopkernel for the K-quant activation quantization. Thewarp-cooperative variant uses 16 super-blocks per block (vs 128 threads per
block in the standard kernel) and is byte-identical to the standard
QuantizeQ8KK. Measured +2.06% decode throughput improvement.The default path (
VT_QUANT_Q8K_WARPunset) is byte-unchanged. Validated bytest_rocm_quant_dot(12/12 cases, 797 assertions). Token-identical toupstream baseline on Qwen3.5-4B Q4_K, 32-token greedy decode, seed 0.
Depends on #2782 (F1), #2790 (T4a), and #2792 (lever-C).
FOLLOWING_AGENTS_PROTOCOL
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:GLM-5-2 [OMP]