Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
19 commits
Select commit Hold shift + click to select a range
8ef6ec3
feat(KERNEL-QUANT-CIQ-GEMM-ROCM): land the W1 keep-quant providers on…
ghazni101 Aug 21, 2026
f1819b2
fix(KERNEL-QUANT-CIQ-GEMM-ROCM): keep grouped Q8_0 on the kROCM provider
mudler Sep 4, 2026
a6eb567
Revert "fix(KERNEL-QUANT-CIQ-GEMM-ROCM): keep grouped Q8_0 on the kRO…
mudler Sep 4, 2026
a338124
fix(rocm): delegate grouped Q8_0 to *Gdn kernel (#2927)
ghazni101 Sep 6, 2026
35063b2
fix(rocm): export the keep-quant providers and drop the duplicate reg…
ghazni101 Sep 8, 2026
6a3932d
spec(BACKEND-ROCM): define Qwen3.5 gfx1100 numerical gate
VikashLoomba Sep 3, 2026
9476aaa
spec(BACKEND-ROCM): repair Qwen3.5 numerical plan
VikashLoomba Sep 4, 2026
fc7c284
spec(BACKEND-ROCM): correct Qwen3.5 capture contracts
VikashLoomba Sep 4, 2026
8116aa7
perf(GFX1100-TG200): T4a MMVQ K-quant decode GEMV arm
ghazni101 Sep 3, 2026
b8ee53e
feat(GFX1100-TG200): lever-C fuses Q8_K activation quant into the Rms…
ghazni101 Sep 3, 2026
d83e7f0
T27: warp-cooperative QuantizeQ8KK for decode (+2.06%, byte-identical)
ghazni101 Aug 27, 2026
d58c923
feat(GFX1100-TG200): T24 LDS-buffered quant epilogue in RmsNormRowCoo…
ghazni101 Aug 26, 2026
aa6f4ae
feat(GFX1100-TG200): T21 keep-quant for V-head row-permuted GDN proje…
ghazni101 Aug 26, 2026
a0b7333
fix(rocm): gate the coop rmsnorm vector body to bf16 gamma, seed quan…
ghazni101 Sep 4, 2026
4960fdf
T25: keep ssm_out as Q5_K with runtime input permutation
ghazni101 Aug 26, 2026
a6be41b
fix(rocm): count the fused MMVQ route once per call
ghazni101 Sep 8, 2026
6d0d2b2
T35-r3: GDN b/a same-input GEMV merge, adjudicated bit-identical, clo…
ghazni101 Aug 29, 2026
fb55044
perf(GFX1100-TG200): merged keep-quant gate_up -- one quant GEMM per …
ghazni101 Aug 23, 2026
1263be3
T35-r3: GDN b/a same-input GEMV merge, adjudicated bit-identical, clo…
ghazni101 Aug 29, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
178 changes: 123 additions & 55 deletions .agents/specs/gfx1100-tg200.md
Original file line number Diff line number Diff line change
@@ -1,15 +1,6 @@
# Spec: GFX1100-TG200

- Original campaign issue:
[#5](https://github.com/ghazni101/vllm.cpp/issues/5) (`ghazni101/vllm.cpp`)
- Live landing owner:
[`BACKEND-ROCM` issue #2427](https://github.com/mudler/vllm.cpp/issues/2427)
- Historical landing request:
[#2164](https://github.com/mudler/vllm.cpp/issues/2164) (deleted or
unavailable; retained only as historical attribution)
- Immutable source: [`pr/1936`](https://github.com/mudler/vllm.cpp/pull/1936)
at `3a345b5ae5df7cf08f1383b6623b38db9a1335bd`
- Gate prompt: [`tools/tg200-prompt.txt`](../../tools/tg200-prompt.txt)
- Issue: [#5](https://github.com/ghazni101/vllm.cpp/issues/5) (`ghazni101/vllm.cpp`)
- Base: `019f66c1a` (upstream tip 2026-08-22; the branch carries one merge commit
pinning the base before the spec landed)
- Pull request shape: one pull request for spec and implementation per stage
Expand Down Expand Up @@ -134,9 +125,9 @@ Stage order after T1 is T1's output, not this table's.

- `tests/vt/test_rocm_quant_dot.cpp` unchanged (841 assertions, 19 cases,
fresh-build count at the issue-#9 repair) for every quant-path lever.
Provenance: the earlier "132,094 assertions" figure came from a stale
7.14-era binary whose lattice no longer matched the source. Only a fresh
configure and build in the current container is authoritative.
Provenance: the earlier "132,094 assertions" figure was read off a stale
7.14-era binary whose lattice no longer matches the source; only a fresh
configure+build in the current container is authoritative.
- Focused gate per stage: `ctest -R 'rocm|cross_device|quant'` in the 7.14
container under the gpu-ctl lock.
- The acceptance gate itself is T6's test.
Expand All @@ -160,55 +151,132 @@ Stage order after T1 is T1's output, not this table's.

## Now

Issue [#2427](https://github.com/mudler/vllm.cpp/issues/2427) owns this
records-only landing. Historical upstream issue
[#2164](https://github.com/mudler/vllm.cpp/issues/2164) is deleted or
unavailable; it remains only as attribution for the original integration
request. The campaign records come from `pr/1936` at
`3a345b5ae5df7cf08f1383b6623b38db9a1335bd`. This integration contains the
specification, 17 evidence files, and the exact
[`tools/tg200-prompt.txt`](../../tools/tg200-prompt.txt) input. It contains none
of pull request #1936's product changes. The unmerged campaign's opt-in arms,
default changes, and product changes are not reachable from this tree. The
measured position and next hypothesis that follow are historical evidence from
the source commit. They are not a current-main benchmark.

`ACTIVE`. Measured position before T21: ~103 tok/s (T18 idle-host gate
100.46 tok/s + T18 v_dot4 +2.7% matched-load). T21's measured +3.9% projects
the idle-host position to ~107 tok/s. Adopted levers: T5a shared quant-body
vectorization (+23%), T5b d128 f32-Q DecodeGqa arm (+13.5%), T6a cooperative
GDN scan (+4.6%), T6b cooperative attn preamble (+4.6%), T8 cooperative
rmsnorm row (+3.2%), T9 cooperative gated norm (+2.6%), T10 warp postconv
(+4.7%), T11 row-split scan (+3.2%, BIT-IDENTICAL), T14 row-split argmax
(−71%, BIT-IDENTICAL), T16 YTILE=4 default (+1.8% contended, +8.1% idle),
T18 v_dot4 instruction selection (+2.7%, BIT-IDENTICAL), and T21 row-permuted
GDN keep-quant (+3.9%, ADOPTED). T21's `VT_GDN_ROWPERM_KEEP_QUANT` gate is
default-enabled at 1.
`ACTIVE`. Position: **~91.2 tok/s median** (2026-08-29 three-point branch
audit, `rocm-dev:10.0.0` container = HIP 7.15 toolchain, examples/vllm-cli,
batch 1, greedy, 256-token acceptance workload, all adopted levers on).
Evidence: `docs/bench-evidence/gfx1100-tg200-t33-branch-consolidation-20260829.md`.

TOOLCHAIN BASELINE BREAK. The ROCm toolchain moved twice in the window
08-26 -> 08-29: rocm-dev:7.14 (native `/opt/rocm`, since removed from the
host) -> venv HIP 7.15 -> ROCm 10.0.0 container (HIP 7.15). Pre-swap
numbers (~103 tok/s at T18/T22) were measured under 7.14 and are NOT
comparable to post-swap measurements. The three-point audit re-measured
three branch commits under ONE container toolchain:

| commit | point | median tok/s |
|---|---|---|
| 6836c11cc (T22-era, the "~103" position) | 08-26 | 89.25 |
| b058bb752 (pre-T31/T32, post-merges) | 08-28 | 90.80 |
| 7beb76e27 (HEAD) | 08-29 | 91.24 |

Conclusions: (1) the upstream integrations (e1ea27c82, 62f37025e =
e551cf8e4) plus the fp8-KV / keep-quant / sample commits HELP: +1.7% net;
(2) T31 (device-mirror port, wash) + T32 (silu-mul+Q8_K fusion, +1.07%)
add +0.5% on top; (3) the branch did NOT regress across the merges — the
perceived 103 -> 91 drop was the toolchain swap, not branch code. HEAD's
256-token output is byte-identical to the campaign reference under the
container build; all correctness gates pass in-container.

Adopted levers (env): VT_GEMV_MMVQ, VT_SKINNY_BF16, VT_ATTN_DECODE_GQA4,
VT_GDN_SCAN_COOP, VT_ATTN_PREAMBLE_COOP, VT_NORM_QUANT_FUSED (+T32's
VT_SILU_QUANT_FUSED silu site behind the same lever), VT_RMSNORM_ROW_COOP,
VT_GDN_NORMGATED_COOP, VT_GDN_POSTCONV_COOP, VT_GDN_SCAN_SPLIT,
VT_ARGMAX_SPLIT, VT_GDN_ROWPERM_KEEP_QUANT, VT_RMSNORM_LDS_QUANT,
VT_GDN_COLPERM_KEEP_QUANT, VT_QUANT_Q8K_WARP. Plus T31's async
device-mirror port (throughput-wash on the CLI path; prerequisite for
async levers). All 17+1 verified present and wired after the merges.

Closed negative: T5c MMVQ nontemporal, T7 COALK wash, T12 gated-quant
fusion, T13 async server wash, T15 LDS bank conflicts, T17 v_dot2
memory-bound, T19 kGemvWarps block-limited, T20 full-warp cooperative GEMV
(kernel 2.4-3.1x on large grids but engine wash — Q4_K dominant path is
launch-overhead-bound at small grids; evidence
`docs/bench-evidence/gfx1100-tg200-t20-full-warp-gemv-wash-20260826.md`).
Failed-attempt ledger: 8 of 15.

Budget table (pre-T20, ~103 tok/s, ~9.7 ms/tok wall):
KQuantGemvMmvqK<Q4_K> 2.46 ms/tok (25%), wvSplitKSml 2.32 ms/tok (24%),
KQuantGemvMmvqK<Q6_K> 1.20 ms/tok (12%), RmsNormRowCoop 0.754 ms/tok (8%),
QuantizeQ8KK 0.544 ms/tok (6%), other ~1.3 ms/tok (13%), total kernel
~8.58 ms/tok (88%). Weight read floor 4.21 GB/tok = 4.38 ms/tok at 960 GB/s.
Overhead above floor: ~4.2 ms/tok — launch overhead, sync, idle gaps.

Next attack: the overhead is the bottleneck, not individual kernel internals.
T20 proved kernel micro-optimization is exhausted for the dominant paths.
The path to 200 tok/s (5.0 ms/tok) requires closing the 4.2 ms/tok overhead
gap: HIP graph capture (T2), kernel fusion, or persistent kernels. A fresh
rocprofv3 attribution capture with dispatch counts per token is the next
step to price the overhead precisely.
(evidence gfx1100-tg200-t20-full-warp-gemv-wash-20260826.md), T31 device
mirror (wash; prerequisite), T32 attempt-1 uint4 wider GEMV loads (wash —
compiler already coalesces the 4-byte memcpy pattern), T32 attempt-2
kGemvWarps 8->16 (wash/slightly worse). Both T32 attempts reverted.

Budget table (2026-08-29 T31 trace, HIP 7.15, per decode token): GEMV
5.18 ms (Q4_K 2.86 = merged gate_up 1.65 + ffn_down 0.59 + rest; Q5_K
1.12; Q6_K 1.21 incl. lm_head 0.59 at 99% of achievable BW — fixed),
RmsNorm 0.76, QuantizeQ8K 0.57 before T32 (~0.28 after the silu fusion),
PagedAttn 0.29, GdnScan 0.24, wvSplitKSml 0.13, prefill amortized ~1.03
(naive m-pass-through GEMM re-reads weight rows M times through L2 —
tiling is the unexplored lever), GPU total 8.68; host/sync ~2.2-2.3.
GEMV per-shape BW: gate_up 57%, attn_output-Q6K 26%, ssm_out 54%,
attn_gate 44% (small-N shapes GPU-underfilled; lm_head 99% — do not
touch).

Next attack (T34, 2026-08-29):
the decode step runs as ONE hipGraph replay (T2b) and the engine is
GPU-bound (`hipStreamSynchronize` 8.75 ms = the GPU step; deferring the
sync buys ~nothing). The residual splits into in-kernel time above the
byte floor (7.56 - 4.38 = 3.18 ms/tok) and ~522 device-side inter-kernel
gaps of ~4.1 us inside the graph (2.14 ms/tok, 22% of wall, diffuse).
Kernel count ~523/token; host dispatch is clean (1 graph launch, 5 eager
launches, 19 other API calls). Lever ranking priced on the capture: T35
same-input GEMV merges + remaining norm/quant epilogue folds (0.5-0.9),
T36 prefill GEMM tiling (~0.7), T37 small-N GEMV bandwidth (0.5-0.9; two
washes already, needs a new angle). Capture-infrastructure note:
rocprofv3 needs `-e HOME=<writable>` in containers or finalization aborts
and loses every buffer; `HIP_TRACE_API` is gone on this runtime.

T36 outcome (2026-08-29,
[evidence](../../docs/bench-evidence/gfx1100-tg200-t36-prefill-mtile-20260829.md)):
the m-pass-through is real but the Infinity Cache absorbs most of it.
KQuantGemmMTiledK (MT=16 activation rows per warp, bit-identical to
KQuantGemmK by construction; op gate tests/vt/test_rocm_prefill_tile.cpp
720/720 byte-identity + 1e-6 NMSE) measures +1.25% median (85.753 ->
86.825 tok/s, complete separation, 256-token outputs byte-identical).
NOTE: the prompt-honest acceptance position at the CURRENT
109-token tools/tg200-prompt.txt is ~85.8 tok/s, not 91.24 — T33's
morning median used a since-removed 71-token prompt. Closed below the
2% adoption bar; ships opt-in (VT_PREFILL_TILE=1, default OFF).
T35 outcome (2026-08-29, b/a same-input GEMV merge,
[evidence](../../docs/bench-evidence/gfx1100-tg200-t35r3-ba-merge-20260829.md)):
closed red, reverted. The b/a pair is bf16 at runtime (kTransformedWeight
under the V-row reorder, 48 wvSplitKSml launches/token), not keep-quant;
stacking it into one [64,2560] owner engaged (w_ba.bin witness) but the
N=64 merged GEMV changes the per-row reduction geometry vs two N=32
launches and the 256-token body diverged (7415e281 vs 783cea17) — a
reduction-order change that owes near-tie adjudication, out of round
scope. Lever changes reverted; tree green at the T36 commit; remaining
same-input K-quant pairs on this checkpoint (attn_q+attn_k only) price
below the bar alone. The "remaining norm/quant epilogue folds" half of
T35 is already banked: the trace's 32 residual standalone quants sit on
attention/gated-norm outputs with no producer (folding those needs new
fused kernels, not wiring).

T35-r3 outcome (2026-08-29, post-#9-fix tree,
[evidence](../../docs/bench-evidence/gfx1100-tg200-t35r3-ba-merge-20260829.md)):
the b/a expand-arm merge re-landed on the fixed kernel and adjudicated
BEFORE the A/B: tools/tg200-neartie.sh verdict=PASS divergent=0
max_gap_mnats=0.000 vs the re-minted reference a0fa1c4a (round 2's
divergence was the corrupted Q6_K arm amplifying the reduction-geometry
delta; the fixed kernel reproduces the reference bit-for-bit with the
merge ON). Clean idle-window A/B: base 85.834 vs ON 85.510 median
(-0.38%, overlapping distributions) — closed negative with numbers; the
lever ships as an adjudicated, bit-identical, default-OFF opt-in
(VT_GDN_MERGED_BA_ROCM=1) for future re-pricing. Same window, clean T36
number: VT_PREFILL_TILE=1 median 86.392 = +0.65% over base 85.834
(co-tenant-free confirmation of the earlier +1.25%, whose window
overlapped a hung container; still below the 2% bar, opt-in disposition
unchanged; both arms byte-identical to a0fa1c4a).

Owed before ANY default flip of the opt-in arms (GQA4 / GDN_SCAN_COOP /
GDN_SCAN_SPLIT / PREAMBLE_COOP / RMSNORM_ROW_COOP / GDN_NORMGATED_COOP /
GDN_POSTCONV_COOP): teacher-forced logprob-band ceremony per
`.agents/specs/rocm-m4-oracle.md`. The campaign reports into #5; each
stage lands as its own `row/GFX1100-TG200-*` branch + draft PR per the
recorded push authority.

Issue #9 (2026-08-29): the standing quant gate test_rocm_quant_dot was
found red pre-T36 (fresh e1567729e build; the green "132,094 assertions"
claims came from a stale 7.14-era binary), repaired at 80f4059f6 — the
VT_GEMV_MMVQ arm's q6_K chunk biased its 6-bit bytes with one 32-bit
subtract whose cross-byte borrow corrupted ~87% of random words. Gate
now 841/841 assertions; the two residual case-level failures are the
pre-existing CIQ owed-encoding throws, unrelated to this arm. The
campaign reference body (783cea17...) was minted under the broken arm:
post-repair ON==OFF is byte-identical end-to-end, and the 4-token
divergence vs the old reference measured 625 mnats max gap — one point
over the 500-mnat near-tie band, so the reference needs a re-mint
decision from the operator.
Loading
Loading