Row: KERNEL-QUANT-CIQ-GEMM-ROCM
#1910 measured that KQuantGemmK gives one warp to each output element and
strides its 32 lanes over nsb = K/256 superblocks, so lanes nsb..31 execute
nothing. It measured K = 4096 on a 9B model. The same defect is on 61.7% of
Qwen3.8-27B Q4_K_M's per-token byte budget, and #2086's fix does not reach it.
The byte budget, from the artifact
Qwen3.8-27B-Q4_K_M.gguf, sha256 7e78da5d..., 866 tensors, qwen35,
block_count = 65, nextn_predict_layers = 1, embedding_length = 5120,
feed_forward_length = 17408, full_attention_interval = 4 (48 GDN layers, 16
full-attention layers). Header parsed directly; total tensor bytes computes to
17,095,778,304, which is byte-exact with the model_size llama-bench reports.
Both engines run 64 layers (#2497's fidelity term is discharged separately), so
one token sweeps blk.0..63 plus output.weight plus the final norm =
16,091,088,896 B.
Which arm each of those bytes takes at m = 1
Read from the tested revision, git show c796fea41:src/vt/rocm/rocm_grouped_gemm.hip.
nsb = K/256 where K is the row length, and KQuantDecodeCoopWarps (:777-788)
returns 8 only for m == 1 && nsb <= 32 && wdt == kQ6_K:
| arm |
tensors |
nsb |
share of swept bytes |
KQuantGemmK<OutT,0/1>, lanes idle |
ffn_gate, ffn_up, attn_gate, attn_q/k/v/output, attn_qkv (Q4_K half), ssm_out (Q5_K) |
20 or 24 |
61.73% |
KQuantGemmK<OutT,0/2>, lanes full |
ffn_down (Q4_K 32 layers, Q6_K 32 layers) |
68 |
24.51% |
KQuantGemmKCoopQ6K<OutT,8> (#2086's fix) |
attn_qkv Q6_K, attn_v Q6_K, lm_head |
20 |
13.11% |
| F32 (norms, ssm params) |
-- |
-- |
0.65% |
K = 5120 gives nsb = 20, so 12 of 32 lanes idle; K = 6144 gives
nsb = 24, so 8 of 32 idle. Every projection in this model except ffn_down
has K of 5120 or 6144.
#2086 kept Q4_K and Q5_K on the single-warp arm deliberately and by
measurement -- the kernel header says so at :506-512, and the sweep found the
cooperative split costs Q4_K/Q5_K 1.5x-1.8x at every width tried, because
splitting a superblock re-reads its header per warp and breaks one contiguous
128-byte read into eight 32-byte ones. That finding stands. It also means
Q4_K, which is 65.25% of this model's per-token bytes, has no decode arm at
all beyond the one #1910 diagnosed.
The second half of the mechanism: the access pattern
Lane L reads the whole of superblock L -- w_row + L * 144 for Q4_K -- and
consumes all 144 bytes itself (:457, DotQ4K at :225-251). Adjacent lanes
are 144 bytes apart, so one memory instruction across a wave touches 32
addresses spread over 4608 bytes instead of one contiguous burst, and each lane
then walks its own block with 4-byte loads. llama.cpp's vec_dot_q4_K_q8_1
splits a superblock 16 ways so consecutive lanes read consecutive dwords. The
kernel does use the hardware __ockl_sdot4 (:78-88, landed by 37255d1d5,
an ancestor of the tested build), so the integer core is not the issue.
Why this is worth the row
Effective read bandwidth per token, from the figures already published in
#3002 and #2940 and the byte budget above:
| engine |
tok/s |
effective GB/s |
ms/token |
llama.cpp b10451 |
12.219 (decode) |
196.6 |
81.8 |
vLLM 5559679229 |
11.056 (derived decode) |
177.9 |
90.4 |
vllm.cpp c796fea41 |
5.397 (derived decode) |
86.8 |
185.3 |
Both oracles land within 10% of each other near the part's published 256 GB/s
LPDDR5X ceiling. We are at 44% of llama.cpp's achieved bandwidth while holding
the board at 100% busy and a higher shader clock (#3015). The bytes are not in
dispute and the board is not idle, so the deficit is inside these kernels.
These are diagnostic figures, not a benchmark claim. The published
correctness state applies unchanged: TOKEN_GATE=FAIL, near-ties at about
0.125 nats, no deterministic denominator on this path (#2921).
Scope of a fix
A decode arm for Q4_K and Q5_K at nsb <= 32 that keeps the contiguous
superblock read #2086's sweep showed matters, rather than the warp split that
sweep rejected. gfx1151 is RDNA3.5 and has no MFMA, so #2109's prescription
does not apply here; the RDNA path is WMMA (ggml/src/ggml-cuda/mma.cuh:729 in
llama.cpp), and KQuantGemmKWmmaQ4K/KQuantGemmKWmmaQ6K already exist for
gfx1200/gfx1201 from KERNEL-QUANT-CIQ-GEMM-ROCM -- whether they build and win
on gfx1151 is unmeasured and is the first question.
Needs a spec before code. Attribution across the three arms above is a
byte-weighted estimate until #3015's rocprofv3 trace exists.
Related: #1910, #2086, #2109, #3006, #2497, #2921, #3002, #3015, #41.
Row:
KERNEL-QUANT-CIQ-GEMM-ROCM#1910 measured that
KQuantGemmKgives one warp to each output element andstrides its 32 lanes over
nsb = K/256superblocks, so lanesnsb..31executenothing. It measured
K = 4096on a 9B model. The same defect is on 61.7% ofQwen3.8-27B Q4_K_M's per-token byte budget, and #2086's fix does not reach it.
The byte budget, from the artifact
Qwen3.8-27B-Q4_K_M.gguf, sha2567e78da5d..., 866 tensors,qwen35,block_count = 65,nextn_predict_layers = 1,embedding_length = 5120,feed_forward_length = 17408,full_attention_interval = 4(48 GDN layers, 16full-attention layers). Header parsed directly; total tensor bytes computes to
17,095,778,304, which is byte-exact with the
model_sizellama-bench reports.Both engines run 64 layers (#2497's fidelity term is discharged separately), so
one token sweeps
blk.0..63plusoutput.weightplus the final norm =16,091,088,896 B.
Which arm each of those bytes takes at
m = 1Read from the tested revision,
git show c796fea41:src/vt/rocm/rocm_grouped_gemm.hip.nsb = K/256whereKis the row length, andKQuantDecodeCoopWarps(:777-788)returns 8 only for
m == 1 && nsb <= 32 && wdt == kQ6_K:nsbKQuantGemmK<OutT,0/1>, lanes idleKQuantGemmK<OutT,0/2>, lanes fullKQuantGemmKCoopQ6K<OutT,8>(#2086's fix)K = 5120givesnsb = 20, so 12 of 32 lanes idle;K = 6144givesnsb = 24, so 8 of 32 idle. Every projection in this model exceptffn_downhas
Kof 5120 or 6144.#2086 kept
Q4_KandQ5_Kon the single-warp arm deliberately and bymeasurement -- the kernel header says so at
:506-512, and the sweep found thecooperative split costs Q4_K/Q5_K 1.5x-1.8x at every width tried, because
splitting a superblock re-reads its header per warp and breaks one contiguous
128-byte read into eight 32-byte ones. That finding stands. It also means
Q4_K, which is 65.25% of this model's per-token bytes, has no decode arm at
all beyond the one #1910 diagnosed.
The second half of the mechanism: the access pattern
Lane
Lreads the whole of superblockL--w_row + L * 144for Q4_K -- andconsumes all 144 bytes itself (
:457,DotQ4Kat:225-251). Adjacent lanesare 144 bytes apart, so one memory instruction across a wave touches 32
addresses spread over 4608 bytes instead of one contiguous burst, and each lane
then walks its own block with 4-byte loads. llama.cpp's
vec_dot_q4_K_q8_1splits a superblock 16 ways so consecutive lanes read consecutive dwords. The
kernel does use the hardware
__ockl_sdot4(:78-88, landed by37255d1d5,an ancestor of the tested build), so the integer core is not the issue.
Why this is worth the row
Effective read bandwidth per token, from the figures already published in
#3002 and #2940 and the byte budget above:
b104515559679229c796fea41Both oracles land within 10% of each other near the part's published 256 GB/s
LPDDR5X ceiling. We are at 44% of llama.cpp's achieved bandwidth while holding
the board at 100% busy and a higher shader clock (#3015). The bytes are not in
dispute and the board is not idle, so the deficit is inside these kernels.
These are diagnostic figures, not a benchmark claim. The published
correctness state applies unchanged:
TOKEN_GATE=FAIL, near-ties at about0.125 nats, no deterministic denominator on this path (#2921).
Scope of a fix
A decode arm for Q4_K and Q5_K at
nsb <= 32that keeps the contiguoussuperblock read #2086's sweep showed matters, rather than the warp split that
sweep rejected. gfx1151 is RDNA3.5 and has no MFMA, so #2109's prescription
does not apply here; the RDNA path is WMMA (
ggml/src/ggml-cuda/mma.cuh:729inllama.cpp), and
KQuantGemmKWmmaQ4K/KQuantGemmKWmmaQ6Kalready exist forgfx1200/gfx1201 from
KERNEL-QUANT-CIQ-GEMM-ROCM-- whether they build and winon gfx1151 is unmeasured and is the first question.
Needs a spec before code. Attribution across the three arms above is a
byte-weighted estimate until #3015's
rocprofv3trace exists.Related: #1910, #2086, #2109, #3006, #2497, #2921, #3002, #3015, #41.