Loom is a lightweight, custom LLM inference engine built from scratch using asyncio, gRPC, and PyTorch. It implements continuous batching and KV-Cache memory management to maximize GPU throughput, serving as an educational reimplementation of systems like vLLM and TensorRT-LLM.
The system follows a centralized engine loop. The gRPC server offloads requests to the engine. The engine relies on a scheduler to manage VRAM allocation and sequence states, and a model runner to execute the forward passes.
graph TB
Client[gRPC Client] -->|GenerateRequest| Server[Loom gRPC Server]
subgraph LoomEngine["Loom Engine (Asyncio Loop)"]
Scheduler[Scheduler<br/>Queue Manager]
BlockManager[BlockManager<br/>VRAM Allocator]
ModelRunner[ModelRunner<br/>Forward Pass & KV-Cache]
end
Server -->|add_request| Scheduler
Scheduler -->|allocates blocks| BlockManager
Scheduler -->|steps sequence| ModelRunner
ModelRunner -->|yields token| Server
Server -->|stream GenerateResponse| Client
loom/
├── proto/ # gRPC protocol definitions
│ └── loom.proto
├── scripts/
│ └── client.py # CLI client to send prompts and stream responses
├── src/
│ └── loom/
│ ├── core/ # Core internals
│ │ ├── block_manager.py # Paged KV-Cache memory allocator
│ │ ├── model_runner.py # PyTorch model loading & autoregressive step
│ │ └── models.py # Sequence & Block status dataclasses
│ ├── generated/ # Auto-generated gRPC stubs (ignored by linters)
│ ├── scheduler/ # Engine scheduling logic
│ │ └── scheduler.py # Continuous batching & queue management
│ ├── engine.py # Main engine loop (weaves sequences together)
│ ├── server.py # Async gRPC servicer & stream handler
│ └── __main__.py # Server entry point
├── tests/
│ ├── unit/ # Block manager and scheduler unit tests
│ └── integration/
├── docker-compose.yml # GPU-enabled container deployment
├── Dockerfile # Multi-stage Python 3.11-slim build
└── pyproject.toml # Hatchling build, ruff, mypy config
- Continuous Batching: Dynamically injects new requests into the active GPU batch mid-generation, maximizing throughput.
- Block-Based KV-Cache: Pre-allocates GPU VRAM into fixed-size blocks. The BlockManager tracks allocation strictly to prevent Out-Of-Memory (OOM) errors.
- Custom Autoregressive Loop: Bypasses Hugging Face's
model.generate(). Managespast_key_valuesand forward passes manually for granular control. - gRPC Server-Side Streaming: Clients receive tokens over a persistent stream the millisecond they are generated.
- Strict Typing: 100%
mypy --strictcompliance and ruff linting.
- Python 3.11+ (asyncio, dataclasses)
- PyTorch (Model loading, KV-Cache management)
- Hugging Face Transformers (Tokenizer & Base Weights only)
- gRPC / Protocol Buffers (Network contract)
- Docker (Containerized GPU deployment)
Build and run the engine:
docker compose up --build -dSend a prompt to the engine:
python scripts/client.pyWhy manual past_key_values management?
To generate token N, the model needs attention over tokens 1 to N−1. Recomputing this is O(N²) compute. By managing the KV-Cache manually, we achieve O(N) compute, drastically reducing latency.
Why a custom BlockManager? Dynamically resizing tensors on the GPU causes fragmentation and crashes. Pre-allocating a pool of blocks and assigning them to sequences ensures memory safety under heavy concurrent load.
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install -e ".[dev]"
# Run strict quality gates
ruff check .
mypy src/
pytest tests/