LLM Inference / AI Infra Engineer · 大模型推理 / AI Infra 工程师
📧 siruhe666@gmail.com · 🏠 Homepage · 🐙 GitHub
I'm SIRU HE, an LLM Inference / AI Infra engineer, currently doing a Master's in Electronic Information at SUSTech (jointly trained with SIAT-CAS).
I work on the low-level layer of AI — when models keep getting bigger and need to run fast and efficiently on phones and servers, someone has to optimize the parts closest to the hardware: GPU kernels, inference engines, distributed systems. My job is to make large models run, run fast, and run reliably on real devices.
- Kernels: CUDA / Triton operator development, memory-access and Tensor Core optimization
- Inference systems: building inference engines, KV Cache, CUDA Graph, speculative decoding
- On-device deployment: NPU optimization and streaming inference for multimodal LLMs
- Train-serve consistency: reproducible training vs. inference across distributed clusters
Tech stack: C++ · CUDA · Triton · Python · PyTorch · vLLM · llama.cpp/ggml
From kernels to systems, one complete low-level path:
CUDA Kernel → Inference Engine → On-device Deployment → Distributed Train-serve Consistency
Currently working on on-device multimodal inference engines at ModelBest, and contributing to open-source projects like RL-Kernel and vLLM-Omni.
- 📧 Email: siruhe666@gmail.com
- 🏠 Homepage / Articles: frank-2077.github.io
Open to 2026 campus recruiting · LLM Inference / AI Infra


