Skip to content

fix(sdk): prefetch the right weight chunk in batched F16 HMX matmul - #1505

Open
Tadayuki Okada (TadayukiOkada) wants to merge 1 commit into
qualcomm:mainfrom
TadayukiOkada:fix/llama-hexagon-hmx-batched-prefetch
Open

Tadayuki Okada (TadayukiOkada) wants to merge 1 commit into
qualcomm:mainfrom
TadayukiOkada:fix/llama-hexagon-hmx-batched-prefetch

Conversation

@TadayukiOkada

Copy link
Copy Markdown

Summary

  • sdk/patches/llama-hexagon-hmx-f16-batched-prefetch.patch: one-line fix in ggml-hexagon hmx_mm_f16_f32_batched. With GQA broadcast (F16 weights, r2 > 1) the second weight chunk was prefetched from weight_group + weight_stride (one row past the first chunk) instead of weight_group + n_chunk_n_cols * weight_stride. The output of the second chunk was wrong whenever the weights span more than one chunk.
  • Hit by attention prefill without flash attention at long KV (KQ from about 3072 KV x 512 tokens). Flash attention and vision encoders (r2 = 1) do not reach this path.
  • Upstream hexagon: HMX matmul with F16 activation and F16/F32 weights of any row count ggml-org/llama.cpp#29626 (open) contains the same fix; delete this patch once the submodule is bumped past it.

Test plan

  • test-backend-ops -o MUL_MAT on a v73 dev kit with a regression case (m 3072, n 512, k 128, bs [8,1], nr [2,1]): fails without the fix, passes with it (upstream llama.cpp 876c75b1f: 763/764 -> 764/764)
  • Patch applies on top of the existing patch list and the configure-time reverse check passes
  • Container preset builds on this branch (not run)

🤖 Generated with Claude Code

The batched F16 HMX matmul (GQA broadcast) prefetched the second weight
chunk from one row past the first chunk instead of from the next chunk,
so the second chunk's output was wrong whenever the weights span more
than one chunk.

Signed-off-by: Tadayuki OKADA <TadayukiOkada@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant