Skip to content

[Bug] Failed to import Triton kernels (triton_kernels.matmul_ogs) when serving DeepSeek-V4.1-Flash on 4x RTX PRO 6000 #105

Description

@mathppp

Describe the bug

I followed the unified vLLM Docker guide to serve DeepSeek-V4.1-Flash with the ds41-flash profile on 4x RTX PRO 6000 Blackwell. During model loading, the server fails to import the MXFP4 Triton kernels:

[mxfp4.py:61] Failed to import Triton kernels. Please make sure your triton version is compatible. Error: No module named 'triton_kernels.matmul_ogs'

DeepSeek V4.1's FP4 MoE weights depend on the triton_kernels.matmul_ogs path, so this looks like triton_kernels missing or version-mismatched inside the published image.

Guide / commit followed

https://github.com/local-inference-lab/rtx6kpro/blob/8a68da459c54f38b028f7afdfd9782e9f0765189/docs/unified-vllm-docker.md (commit 8a68da4)

Environment

  • Image: ghcr.io/local-inference-lab/vllm:jovian-judgement-beta-20260917-9b2a25a581e55533
  • Profile: PROFILE=ds41-flash, HARDWARE_PROFILE=rtx-pro-6000-pcie
  • GPUs: 4x NVIDIA RTX PRO 6000 Blackwell (96 GB), TP=4
  • Image runtime: CUDA 13.4.1 / PyTorch 2.14
  • Context window: 1M (--max-model-len 1048576)
  • DS4.1 launch requirements followed: --ulimit memlock=-1 --security-opt seccomp=unconfined, VLLM_PCIE_TWOSHOT_ALLREDUCE_MAX_SIZE=off

Launch configuration

PROFILE=ds41-flash
GPU_DEVICES=0,1,2,3
TP=4
PORT=<port>
SERVE_ARGS=(--mode dspark --draft-tokens 7 --engram-table-memory ram
  --no-language-model-only
  --limit-mm-per-prompt '{"image": 4}'
  --host 0.0.0.0 --port ${PORT}
  --model /models --served-model-name $Model_NAME --api-key ***}
  --max-model-len 1048576 --max-num-seqs 4
  --default-chat-template-kwargs '{"enable_thinking":"True","reasoning_effort":"max","preserve_thinking":"True"}')

Docker run block identical to the guide, with a local checkpoint mounted at /models.
Actual behavior
Model loading stops with the triton_kernels.matmul_ogs import error above.

Questions

  1. Is triton_kernels missing or version-mismatched in jovian-judgement-beta-20260917-9b2a25a581e55533?
  2. Is --max-model-len 1048576 (1M context) within the qualified scope for ds41-flash, or outside the bounded TP4/DCP1 measurements noted in the guide?
  3. Any recommended pinned triton / triton_kernels combination for this image?

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions