High-performance runtime extensions for vLLM.
-
Updated
Sep 10, 2026 - Python
High-performance runtime extensions for vLLM.
Blackwell-optimized llama.cpp bundle with DFlash2, MXFP6/MXFP8/NVFP4, TurboQuant KV, and GPU-resident speculative handoff.
To associate your repository with the mxfp6 topic, visit your repo's landing page and select "manage topics."