glm5next: avoid soft_max gridDim.y overflow in the indexer - #214
Merged
danielhanchen merged 2 commits intoSep 16, 2026
Merged
danielhanchen merged 2 commits into
danielhanchen merged 2 commits into
Conversation
Assisted-by: Cursor
Assisted-by: Cursor
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
This was referenced Sep 11, 2026
danielhanchen
added a commit
that referenced
this pull request
Sep 16, 2026
#214 landed on glm5next/upstream, so e2738e0 is no longer the head. 86ebfef is that squash on top of it: the k-pool gate logits are reshaped to 2D before ggml_soft_max so n_new_max stops mapping to gridDim.y, which is capped at 65535 and aborted the launch at n_kv >= 262144 with kpool = 4. Replayed the resolve loop on b10994: 11 clean, 2 additive, 0 hard fails.
danielhanchen
added a commit
that referenced
this pull request
Sep 16, 2026
#214 landed on glm5next/upstream after #216 was cut, so e2738e0 is no longer the head. 86ebfef is that squash on top of it: the k-pool gate logits are reshaped to 2D before ggml_soft_max so n_new_max stops mapping to gridDim.y, which CUDA caps at 65535 and which aborted the launch at n_kv >= 262144 with kpool = 4. Replayed the resolve loop on b10994 with the new pin: 11 clean, 2 additive, 0 hard fails; merge_checks clean, all 13 pins intact, llama + mtmd compile gate passed, test-llama-archs 316 rows 0 failures.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Reshape k-pool indexer gate logits before
ggml_soft_maxson_new_maxdoes not map togridDim.y, which is capped at 65535 on CUDA. During k-pool rebuild at long context (n_kv >= 262144,kpool = 4),n_new_maxreaches 65538 and the launch aborts withSOFT_MAX failed.Fold
n_new_maxandn_streamintone1viaggml_reshape_2d, then reshape back to 4D for downstreammul/sum_rows. Same pattern asqwen4exp.cpp.Stacked on ggml-org#27754.
Additional information
Related: ggml-org#27773, timkhronos#11, ggml-org#27901
Requirements