Hy4-preview merged into mainline vLLM on 08-29 (vllm-project/vllm#54160) prescribing FLASHMLA_SPARSE (SM90/100 only). On SM120 the sole alternative, FLASHINFER_MLA_SPARSE_SM120, produces degenerate output (elimination-tested; upstream issue: vllm-project/vllm#54434). So workstation Blackwell currently has no working path for this model — while its attention is the DeepSeek MLA shape (RoPE-64, head 576, top-k 2048) that B12X_MLA_SPARSE already serves for GLM-5.2/5.3.
For anyone porting, the hy_v4 deltas vs the DeepSeek path that bite silently (traced from Tencent's official llama.cpp patches in AngelSlim/Hy4-preview-GGUF and the merged vLLM code):
- Elementwise sigmoid gate (
linear_gate) multiplies the decompressed per-head attention output before o_proj — incompatible with absorbing wv_b into o_proj; absorb-style backends must unabsorb or they gate in the wrong space.
- Learnable per-head sink (softmax-denominator bias) on every layer; checkpoint magnitudes are modest, so a warned disable is tolerable short-term.
- RoPE is interleaved (
is_neox_style=False) on unpermuted weights — Tencent explicitly reverted a NEOX attempt in their llama.cpp port; transformers ≥5.15 semantics are wrong for the shipped checkpoint.
- Indexer: q comes from the compressed q path, k goes through LayerNorm-with-bias (eps aliased to rms_eps), and shared-indexer layers reuse the last preceding full layer's top-k (pattern F F F S S S…, 21 full of 78).
- MoE: sigmoid routing with
e_score_correction_bias, ungrouped (n_group=1); swiglu_limit clamp applies to routed experts only; checkpoint experts are stacked gate_up_proj (gate first).
Hy4-preview merged into mainline vLLM on 08-29 (vllm-project/vllm#54160) prescribing FLASHMLA_SPARSE (SM90/100 only). On SM120 the sole alternative,
FLASHINFER_MLA_SPARSE_SM120, produces degenerate output (elimination-tested; upstream issue: vllm-project/vllm#54434). So workstation Blackwell currently has no working path for this model — while its attention is the DeepSeek MLA shape (RoPE-64, head 576, top-k 2048) that B12X_MLA_SPARSE already serves for GLM-5.2/5.3.For anyone porting, the hy_v4 deltas vs the DeepSeek path that bite silently (traced from Tencent's official llama.cpp patches in AngelSlim/Hy4-preview-GGUF and the merged vLLM code):
linear_gate) multiplies the decompressed per-head attention output before o_proj — incompatible with absorbingwv_bintoo_proj; absorb-style backends must unabsorb or they gate in the wrong space.is_neox_style=False) on unpermuted weights — Tencent explicitly reverted a NEOX attempt in their llama.cpp port; transformers ≥5.15 semantics are wrong for the shipped checkpoint.e_score_correction_bias, ungrouped (n_group=1);swiglu_limitclamp applies to routed experts only; checkpoint experts are stackedgate_up_proj(gate first).