Skip to content

[Bug] FlashInfer sparse MLA attention fails with "Unsupported sparse-MLA prefill configuration" during DSpark CUDA Graph capture #86

Description

@BrightMoon-000

using the alternative dspark-mtp0 mode (by setting MODE=dspark-mtp0) allows the service to start successfully. The failure only reproduces with the default MODE=dspark setting.


Environment

  • Hardware: 8 × NVIDIA RTX 5090 (32 GB VRAM each)
  • Software:
    • Docker image: voipmonitor/vllm:infernal-invocation-vllmd6cf36a-b12xf6dc512-fi1ac6942-cu133-torch213-20260827-r21

Reproduction (Docker run command)

docker run -d \
  --name deepseek-v4-flash \
  --gpus all \
  --runtime nvidia \
  --network host \
  --ipc host \
  --shm-size 48g \
  --init \
  --restart unless-stopped \
  --ulimit memlock=-1 \
  --ulimit stack=67108864 \
  --ulimit nofile=1048576:1048576 \
  -v /fast/models/DeepSeek-V4-Flash-0731:/app/model:ro \
  -v /fast/cache/vllm:/cache:rw \
  -v /fast/cache/tmp:/container-tmp:rw \
  -e CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
  -e OMP_NUM_THREADS=2 \
  -e PORT=8000 \
  -e SERVED_MODEL_NAME=DeepSeek-V4-Flash \
  -e MODEL_PATH=/app/model \
  -e MODE=dspark \
  -e BACKEND=b12x-a8-dglin \
  -e TP_SIZE=8 \
  -e DCP_SIZE=1 \
  -e ALLREDUCE_MODE=auto \
  -e DSPARK_DEPTH_MODE=fixed \
  -e DSPARK_TOKENS=5 \
  -e DRAFT_SAMPLE_METHOD=probabilistic \
  -e MAX_NUM_SEQS=16 \
  -e GRAPH=auto \
  -e MAX_MODEL_LEN=262144 \
  -e MAX_NUM_BATCHED_TOKENS=4096 \
  -e GPU_MEMORY_UTILIZATION=0.95 \
  -e LMCACHE_MODE=ram \
  -e LMCACHE_L1_GB=32 \
  -e LMCACHE_HTTP_PORT=8099 \
  -e LOAD_FORMAT=instanttensor \
  -e INSTANTTENSOR_BACKEND=BUFFERED \
  -e PYTHONHASHSEED=0 \
  --entrypoint /usr/local/bin/lmcache-mp-wrapper.sh \
  voipmonitor/vllm:infernal-invocation-vllmd6cf36a-b12xf6dc512-fi1ac6942-cu133-torch213-20260827-r21 \
  /usr/local/bin/serve-ds4-flash.sh

The container script internally invokes vllm serve ... --attention-backend B12X_MLA_SPARSE ... --speculative-config ... --method dspark.


Error log (core excerpt)

(Worker_TP6 pid=1996) ERROR 08-28 16:17:44 [multiproc_executor.py:1048] WorkerProc hit an exception.
...
tvm.error.InternalError: Check failed: (ok) is false: Unsupported sparse-MLA prefill configuration: model=DSV4 num_heads=8 topk=512 page_block_size=64 topk_extra=0 extra_page_block_size=0
...
RuntimeError: Worker failed with error 'Check failed: (ok) is false: Unsupported sparse-MLA prefill configuration: model=DSV4 num_heads=8 topk=512 page_block_size=64 topk_extra=0 extra_page_block_size=0'

Full stack trace points to the CUDA graph capture phase inside self.speculator.capture() calling flashinfer.sparse_mla_sm120_paged_attention.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions