Skip to content

Add Qwen3-Coder SGLang benchmarks - #88

Merged
haofrank merged 3 commits into
mainfrom
add-qwen3-coder-sglang-benchmarks
Sep 1, 2026
Merged

Add Qwen3-Coder SGLang benchmarks#88
haofrank merged 3 commits into
mainfrom
add-qwen3-coder-sglang-benchmarks

Conversation

@haofrank

@haofrank haofrank commented Sep 1, 2026

Copy link
Copy Markdown
Member

Title

Add Qwen3-Coder 480B SGLang benchmark configurations for MI355X

Summary

Add single-node SGLang benchmark configurations for Qwen3-Coder 480B A35B on MI355X:

  • amd/Qwen3-Coder-480B-A35B-Instruct-MXFP4
  • Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8

Both configurations use TP4 and AITER kernels. The FP8 configuration also uses FP8 KV cache with a page size of 32.

Configuration

  • GPU: 4× MI355X
  • Tensor parallelism: TP4
  • AITER: enabled
  • Workload: ISL/OSL 1024/1024, concurrency 32
  • Random range ratio: 0.8
  • FP8 KV cache: fp8_e4m3

Validation

Both configurations successfully completed 320/320 benchmark requests.

Model precision Request throughput Output throughput Mean TTFT Mean TPOT Mean E2E
MXFP4 1.93 req/s 1,783.72 tok/s 179.10 ms 17.35 ms 16.21 s
FP8 1.48 req/s 1,363.58 tok/s 229.74 ms 22.69 ms 21.18 s

@haofrank
haofrank deployed to mi355x-benchmark September 1, 2026 01:01 — with GitHub Actions Active

@ChengYao-amd ChengYao-amd left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@haofrank
haofrank merged commit a062841 into main Sep 1, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants