Repo for vLLM Hook, an vLLM plug-in for programming internal states of models deployed on vLLM
-
Updated
Sep 15, 2026 - Jupyter Notebook
Repo for vLLM Hook, an vLLM plug-in for programming internal states of models deployed on vLLM
Production-pattern Red Hat OpenShift AI 3.4.0 platform with bare-metal ESXi, GPU passthrough, KServe RawDeployment, DeepSeek R1 inference at 12–17 tok/s
An end-to-end compiler, runtime stack, and simulator for hardware-aware execution on a novel LLM inference accelerator
A privacy-first Slack bot that integrates local LLMs (Ollama/vLLM/LM Studio) with advanced tools like ComfyUI image generation, SearXNG local search, and On-Demand RAG Memory. Analyze files, execute Python code, and generate music - all while keeping your data inside your own network.
Serves cyankiwi/Qwen3.8-27B-AWQ-INT4 (dense 27B vision-language, reasoning, apache-2.0, ~21GB INT4 weights) via vLLM on a single A100-80GB. via Modal
Complete self-hosted AI server stack on the Nvidia DGX Spark (arm64)
To associate your repository with the vllm-inference topic, visit your repo's landing page and select "manage topics."