Triton kernels for 4-bit MoE inference: grouped NF4/MXFP4 GEMM, INT4 GEMV, FP8 paged attention, and CPU/NVMe expert streaming.
-
Updated
Sep 10, 2026 - Python
Triton kernels for 4-bit MoE inference: grouped NF4/MXFP4 GEMM, INT4 GEMV, FP8 paged attention, and CPU/NVMe expert streaming.
DeepSeek-R1 7B INT4 at 69.3 tok/s on a $300 RTX 3060. Faster than llama.cpp, vLLM, and NVIDIA TensorRT-LLM. Is one developer + Ai really better than the entire industry?
GLM-5.2 744B at 4-bit on Modal 4x H200 via vLLM, plus a static streaming chat UI.
Qwen3.8-Flash-Next (NVIDIA NVFP4 weights) served as W4A16 on 2x DGX Spark GB10 with vLLM, TP=2. Pinned engine build, 4 start-time patches, measured results, and the levers that were tried and rejected.
Argus-built, evidence-first Qwen2.5-0.5B-Instruct-AWQ W4A16 RTL with a verified 24-layer cascade and authenticated Hybrid RTL runtime.
Argus-built, evidence-first Qwen2.5-0.5B-Instruct-AWQ W4A16 RTL with a verified 24-layer cascade and authenticated Hybrid RTL runtime.
Tests worst-environment scale selection for multi-environment W4A16 quantization.
To associate your repository with the w4a16 topic, visit your repo's landing page and select "manage topics."