Run oversized sparse MoE GGUF models with bounded NVMe-to-RAM-to-VRAM caching, native multi-GPU planning, and a Linux desktop Studio.
-
Updated
Sep 7, 2026 - C++
Run oversized sparse MoE GGUF models with bounded NVMe-to-RAM-to-VRAM caching, native multi-GPU planning, and a Linux desktop Studio.
Adaptive MoE inference for Kimi K3 — beyond-memory expert streaming, reversible runtime profiles, measured optimization results, and an open NVIDIA/GPU/NPU adaptation roadmap.
Expert streaming inference engine for MoE models larger than VRAM — run 235B+ models on consumer GPUs
Experimental llama.cpp runtime for bounded DeepSeek V4 MoE streaming on dual 16 GB GPUs
To associate your repository with the expert-streaming topic, visit your repo's landing page and select "manage topics."