using the alternative dspark-mtp0 mode (by setting MODE=dspark-mtp0) allows the service to start successfully. The failure only reproduces with the default MODE=dspark setting.
Environment
- Hardware: 8 × NVIDIA RTX 5090 (32 GB VRAM each)
- Software:
- Docker image:
voipmonitor/vllm:infernal-invocation-vllmd6cf36a-b12xf6dc512-fi1ac6942-cu133-torch213-20260827-r21
Reproduction (Docker run command)
docker run -d \
--name deepseek-v4-flash \
--gpus all \
--runtime nvidia \
--network host \
--ipc host \
--shm-size 48g \
--init \
--restart unless-stopped \
--ulimit memlock=-1 \
--ulimit stack=67108864 \
--ulimit nofile=1048576:1048576 \
-v /fast/models/DeepSeek-V4-Flash-0731:/app/model:ro \
-v /fast/cache/vllm:/cache:rw \
-v /fast/cache/tmp:/container-tmp:rw \
-e CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
-e OMP_NUM_THREADS=2 \
-e PORT=8000 \
-e SERVED_MODEL_NAME=DeepSeek-V4-Flash \
-e MODEL_PATH=/app/model \
-e MODE=dspark \
-e BACKEND=b12x-a8-dglin \
-e TP_SIZE=8 \
-e DCP_SIZE=1 \
-e ALLREDUCE_MODE=auto \
-e DSPARK_DEPTH_MODE=fixed \
-e DSPARK_TOKENS=5 \
-e DRAFT_SAMPLE_METHOD=probabilistic \
-e MAX_NUM_SEQS=16 \
-e GRAPH=auto \
-e MAX_MODEL_LEN=262144 \
-e MAX_NUM_BATCHED_TOKENS=4096 \
-e GPU_MEMORY_UTILIZATION=0.95 \
-e LMCACHE_MODE=ram \
-e LMCACHE_L1_GB=32 \
-e LMCACHE_HTTP_PORT=8099 \
-e LOAD_FORMAT=instanttensor \
-e INSTANTTENSOR_BACKEND=BUFFERED \
-e PYTHONHASHSEED=0 \
--entrypoint /usr/local/bin/lmcache-mp-wrapper.sh \
voipmonitor/vllm:infernal-invocation-vllmd6cf36a-b12xf6dc512-fi1ac6942-cu133-torch213-20260827-r21 \
/usr/local/bin/serve-ds4-flash.sh
The container script internally invokes vllm serve ... --attention-backend B12X_MLA_SPARSE ... --speculative-config ... --method dspark.
Error log (core excerpt)
(Worker_TP6 pid=1996) ERROR 08-28 16:17:44 [multiproc_executor.py:1048] WorkerProc hit an exception.
...
tvm.error.InternalError: Check failed: (ok) is false: Unsupported sparse-MLA prefill configuration: model=DSV4 num_heads=8 topk=512 page_block_size=64 topk_extra=0 extra_page_block_size=0
...
RuntimeError: Worker failed with error 'Check failed: (ok) is false: Unsupported sparse-MLA prefill configuration: model=DSV4 num_heads=8 topk=512 page_block_size=64 topk_extra=0 extra_page_block_size=0'
Full stack trace points to the CUDA graph capture phase inside self.speculator.capture() calling flashinfer.sparse_mla_sm120_paged_attention.
using the alternative dspark-mtp0 mode (by setting MODE=dspark-mtp0) allows the service to start successfully. The failure only reproduces with the default MODE=dspark setting.
Environment
voipmonitor/vllm:infernal-invocation-vllmd6cf36a-b12xf6dc512-fi1ac6942-cu133-torch213-20260827-r21Reproduction (Docker run command)
The container script internally invokes
vllm serve ... --attention-backend B12X_MLA_SPARSE ... --speculative-config ... --method dspark.Error log (core excerpt)
Full stack trace points to the CUDA graph capture phase inside
self.speculator.capture()callingflashinfer.sparse_mla_sm120_paged_attention.