Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
95 changes: 49 additions & 46 deletions llm-inference-optimization/README.md
Original file line number Diff line number Diff line change
@@ -1,51 +1,54 @@
# LLM Inference Optimization — Quantization, vLLM Serving, Benchmarking, Quality Evaluation

> Module of [`Agentic-AI-Engineer`](../README.md). Everything an agentic system does — planning, tool use, retrieval synthesis, reflection, multi-agent coordination — is a model call. This module is where those calls get **cheap, fast, and measurable**.
> Module of [`Agentic-AI-Engineer`](../README.md). Everything an agentic system does — planning, tool use, retrieval synthesis, reflection, and multi-agent coordination — eventually becomes a model call. This module is where those calls become **cheap, fast, reliable, and measurable**.

## How this module is organized
## Why this module matters

| File | What it is | Read when |
Inference optimization is the layer between a promising agent design and a deployable system. It answers questions such as:

- Which model and serving stack fit the hardware budget?
- Where does latency actually come from: prefill, decode, memory bandwidth, or scheduling?
- How much quality is lost when applying quantization or KV-cache compression?
- At what concurrency does the system stop meeting its latency target?
- What evidence is strong enough to recommend a deployment decision?

This module treats inference as an engineering discipline: model choice, quantization, serving, benchmarking, quality evaluation, and release certification.

## Module map

| File | Purpose | Use it when |
|---|---|---|
| [COVERAGE_GUIDE.md](COVERAGE_GUIDE.md) | Authoring spec (v3): topics, labs, tiers, constraints | Building/maintaining the module |
| [foundations.md](foundations.md) | Where inference sits in the agentic stack; prerequisites | Before Lab 1 |
| [math-foundations.md](math-foundations.md) | Every formula used, with worked examples | Alongside Labs 1, 3, 6 |
| [concepts.md](concepts.md) | Concept cards (prefill/decode, PagedAttention, prefix cache, SLO/goodput, …) | Alongside all labs |
| [diagrams.md](diagrams.md) | Mermaid diagrams (render on GitHub) | Visual reference, slides source |
| [glossary.md](glossary.md) | A–Z terms | Lookup |
| [models.md](models.md) | Pinned model registry + licenses (verification is part of lab setup) | Before any real-path lab |
| [maintenance.md](maintenance.md) | Known moving parts (tool churn watchlist) | Each maintenance pass |
| [certification/](certification/) | Per-tier end-to-end validation records (T0–T3) | Before each release |

## Learning path

```
foundations.md ─► math-foundations.md §1–2 ─► Lab 1 (memory calculator)
│
▼
concepts.md: prefill/decode, batching ─► Lab 2 (scheduler simulator)
│
▼
math §3 quantization ─► Lab 3 (NumPy quant) ─► Lab 4 (LLM Compressor, optional GPU)
│
▼
concepts: PagedAttention, prefix cache ─► Lab 5 (mock/real vLLM server + /metrics)
│
▼
math §5 SLO & queueing ─► Lab 6 (benchmark sweep, GuideLLM optional)
│
▼
math §4 eval statistics ─► Lab 7 (lm-eval, acceptance criteria)
│
▼
Lab 8 capstone: deployment decision under constraints (agent workload)
```

## Cross-links to sibling modules

- **Agentic RAG** (`../agentic-rag/`): prefix caching economics of retrieval templates → [concepts.md#prefix-caching](concepts.md#prefix-caching)
- **Agent orchestration**: routing/fallback policies consume this module's benchmark + eval evidence → [COVERAGE_GUIDE.md §8](COVERAGE_GUIDE.md#8-production-relevance-to-agentic-ai-systems)
- **Evaluation & observability**: lm-eval acceptance criteria pattern generalizes to agent evals → [concepts.md#acceptance-criteria](concepts.md#acceptance-criteria)

## Hardware tiers (pick yours, everything adapts)

T0 CPU-only (mandatory baseline, all labs runnable) · T1 consumer GPU 8–16 GB · T2 single 24–48 GB GPU (reference tier; generates all result packs) · T3 multi-GPU (stretch, demo-only). Full capability contracts: [COVERAGE_GUIDE.md → Hardware tiers](COVERAGE_GUIDE.md#hardware-tiers--capability-contracts).
| [COVERAGE_GUIDE.md](COVERAGE_GUIDE.md) | Authoring and maintenance spec: topics, labs, tiers, constraints | Extending or maintaining the module |
| [foundations.md](foundations.md) | Where inference fits in the agentic stack, prerequisites, mental model | Start here |
| [math-foundations.md](math-foundations.md) | Core formulas and worked examples: KV sizing, roofline intuition, queueing, cost, and evaluation uncertainty | Alongside Labs 1, 3, 6, 7 |
| [concepts.md](concepts.md) | Compact concept cards: prefill/decode, PagedAttention, prefix caching, goodput, SLOs, and more | Reference during all labs |
| [diagrams.md](diagrams.md) | Mermaid diagrams for the serving path, scheduler behavior, caches, and deployment decisions | Visual reference or slide source |
| [glossary.md](glossary.md) | A–Z terminology | Quick lookup |
| [models.md](models.md) | Pinned model registry, revisions, and license verification guidance | Before any real-path lab |
| [maintenance.md](maintenance.md) | Churn watchlist, version pinning, and maintenance procedure | During refresh or upgrade work |
| [certification/](certification/) | Tier-specific release validation records (T0–T3) | Before shipping or sign-off |

## Visual learning path

```mermaid
flowchart TD
A[foundations.md<br/>Inference in the agentic stack] --> B[math-foundations.md §1–2<br/>Memory sizing and throughput basics]
B --> C[Lab 1<br/>Memory calculator]

C --> D[concepts.md<br/>Prefill, decode, batching]
D --> E[Lab 2<br/>Scheduler simulator]

E --> F[math-foundations.md §3<br/>Quantization math]
F --> G[Lab 3<br/>NumPy quantization]
G --> H[Lab 4<br/>LLM Compressor<br/><i>optional GPU</i>]

H --> I[concepts.md<br/>PagedAttention and prefix caching]
I --> J[Lab 5<br/>Mock or real vLLM server<br/>plus metrics]

J --> K[math-foundations.md §5<br/>SLOs, queueing, and goodput]
K --> L[Lab 6<br/>Benchmark sweep<br/><i>GuideLLM optional</i>]

L --> M[math-foundations.md §4<br/>Evaluation uncertainty and acceptance]
M --> N[Lab 7<br/>lm-eval and quality gates]

N --> O[Lab 8 Capstone<br/>Deployment decision under constraints]
Loading