Describe the bug
I followed the unified vLLM Docker guide to serve DeepSeek-V4.1-Flash with the ds41-flash profile on 4x RTX PRO 6000 Blackwell. During model loading, the server fails to import the MXFP4 Triton kernels:
[mxfp4.py:61] Failed to import Triton kernels. Please make sure your triton version is compatible. Error: No module named 'triton_kernels.matmul_ogs'
DeepSeek V4.1's FP4 MoE weights depend on the triton_kernels.matmul_ogs path, so this looks like triton_kernels missing or version-mismatched inside the published image.
Guide / commit followed
https://github.com/local-inference-lab/rtx6kpro/blob/8a68da459c54f38b028f7afdfd9782e9f0765189/docs/unified-vllm-docker.md (commit 8a68da4)
Environment
- Image:
ghcr.io/local-inference-lab/vllm:jovian-judgement-beta-20260917-9b2a25a581e55533
- Profile:
PROFILE=ds41-flash, HARDWARE_PROFILE=rtx-pro-6000-pcie
- GPUs: 4x NVIDIA RTX PRO 6000 Blackwell (96 GB),
TP=4
- Image runtime: CUDA 13.4.1 / PyTorch 2.14
- Context window: 1M (
--max-model-len 1048576)
- DS4.1 launch requirements followed:
--ulimit memlock=-1 --security-opt seccomp=unconfined, VLLM_PCIE_TWOSHOT_ALLREDUCE_MAX_SIZE=off
Launch configuration
PROFILE=ds41-flash
GPU_DEVICES=0,1,2,3
TP=4
PORT=<port>
SERVE_ARGS=(--mode dspark --draft-tokens 7 --engram-table-memory ram
--no-language-model-only
--limit-mm-per-prompt '{"image": 4}'
--host 0.0.0.0 --port ${PORT}
--model /models --served-model-name $Model_NAME --api-key ***}
--max-model-len 1048576 --max-num-seqs 4
--default-chat-template-kwargs '{"enable_thinking":"True","reasoning_effort":"max","preserve_thinking":"True"}')
Docker run block identical to the guide, with a local checkpoint mounted at /models.
Actual behavior
Model loading stops with the triton_kernels.matmul_ogs import error above.
Questions
- Is triton_kernels missing or version-mismatched in jovian-judgement-beta-20260917-9b2a25a581e55533?
- Is --max-model-len 1048576 (1M context) within the qualified scope for ds41-flash, or outside the bounded TP4/DCP1 measurements noted in the guide?
- Any recommended pinned triton / triton_kernels combination for this image?
Describe the bug
I followed the unified vLLM Docker guide to serve DeepSeek-V4.1-Flash with the
ds41-flashprofile on 4x RTX PRO 6000 Blackwell. During model loading, the server fails to import the MXFP4 Triton kernels:DeepSeek V4.1's FP4 MoE weights depend on the
triton_kernels.matmul_ogspath, so this looks liketriton_kernelsmissing or version-mismatched inside the published image.Guide / commit followed
https://github.com/local-inference-lab/rtx6kpro/blob/8a68da459c54f38b028f7afdfd9782e9f0765189/docs/unified-vllm-docker.md (commit
8a68da4)Environment
ghcr.io/local-inference-lab/vllm:jovian-judgement-beta-20260917-9b2a25a581e55533PROFILE=ds41-flash,HARDWARE_PROFILE=rtx-pro-6000-pcieTP=4--max-model-len 1048576)--ulimit memlock=-1 --security-opt seccomp=unconfined,VLLM_PCIE_TWOSHOT_ALLREDUCE_MAX_SIZE=offLaunch configuration
Docker run block identical to the guide, with a local checkpoint mounted at /models.
Actual behavior
Model loading stops with the triton_kernels.matmul_ogs import error above.
Questions