Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Loom: Custom LLM Inference Engine

Loom is a lightweight, custom LLM inference engine built from scratch using asyncio, gRPC, and PyTorch. It implements continuous batching and KV-Cache memory management to maximize GPU throughput, serving as an educational reimplementation of systems like vLLM and TensorRT-LLM.

Architecture

The system follows a centralized engine loop. The gRPC server offloads requests to the engine. The engine relies on a scheduler to manage VRAM allocation and sequence states, and a model runner to execute the forward passes.

graph TB
    Client[gRPC Client] -->|GenerateRequest| Server[Loom gRPC Server]

    subgraph LoomEngine["Loom Engine (Asyncio Loop)"]
        Scheduler[Scheduler<br/>Queue Manager]
        BlockManager[BlockManager<br/>VRAM Allocator]
        ModelRunner[ModelRunner<br/>Forward Pass & KV-Cache]
    end

    Server -->|add_request| Scheduler
    Scheduler -->|allocates blocks| BlockManager
    Scheduler -->|steps sequence| ModelRunner
    ModelRunner -->|yields token| Server
    Server -->|stream GenerateResponse| Client
Loading

Repository Structure

loom/
├── proto/                          # gRPC protocol definitions
│   └── loom.proto
├── scripts/
│   └── client.py                   # CLI client to send prompts and stream responses
├── src/
│   └── loom/
│       ├── core/                   # Core internals
│       │   ├── block_manager.py    # Paged KV-Cache memory allocator
│       │   ├── model_runner.py     # PyTorch model loading & autoregressive step
│       │   └── models.py           # Sequence & Block status dataclasses
│       ├── generated/              # Auto-generated gRPC stubs (ignored by linters)
│       ├── scheduler/              # Engine scheduling logic
│       │   └── scheduler.py        # Continuous batching & queue management
│       ├── engine.py               # Main engine loop (weaves sequences together)
│       ├── server.py               # Async gRPC servicer & stream handler
│       └── __main__.py             # Server entry point
├── tests/
│   ├── unit/                       # Block manager and scheduler unit tests
│   └── integration/
├── docker-compose.yml              # GPU-enabled container deployment
├── Dockerfile                      # Multi-stage Python 3.11-slim build
└── pyproject.toml                  # Hatchling build, ruff, mypy config

Features

  • Continuous Batching: Dynamically injects new requests into the active GPU batch mid-generation, maximizing throughput.
  • Block-Based KV-Cache: Pre-allocates GPU VRAM into fixed-size blocks. The BlockManager tracks allocation strictly to prevent Out-Of-Memory (OOM) errors.
  • Custom Autoregressive Loop: Bypasses Hugging Face's model.generate(). Manages past_key_values and forward passes manually for granular control.
  • gRPC Server-Side Streaming: Clients receive tokens over a persistent stream the millisecond they are generated.
  • Strict Typing: 100% mypy --strict compliance and ruff linting.

Tech Stack

  • Python 3.11+ (asyncio, dataclasses)
  • PyTorch (Model loading, KV-Cache management)
  • Hugging Face Transformers (Tokenizer & Base Weights only)
  • gRPC / Protocol Buffers (Network contract)
  • Docker (Containerized GPU deployment)

Quickstart

Build and run the engine:

docker compose up --build -d

Send a prompt to the engine:

python scripts/client.py

Design Decisions

Why manual past_key_values management? To generate token N, the model needs attention over tokens 1 to N−1. Recomputing this is O(N²) compute. By managing the KV-Cache manually, we achieve O(N) compute, drastically reducing latency.

Why a custom BlockManager? Dynamically resizing tensors on the GPU causes fragmentation and crashes. Pre-allocating a pool of blocks and assigning them to sequences ensures memory safety under heavy concurrent load.

Local Development

python -m venv venv
source venv/bin/activate  # Windows: venv\Scripts\activate
pip install -e ".[dev]"

# Run strict quality gates
ruff check .
mypy src/
pytest tests/

About

Custom LLM serving infrastructure built with asyncio and gRPC. Features a paged VRAM allocator for KV-caching and a continuous batching scheduler to handle concurrent generation requests efficiently at the systems level.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages