-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathglassbox.json
More file actions
135 lines (135 loc) · 4.72 KB
/
Copy pathglassbox.json
File metadata and controls
135 lines (135 loc) · 4.72 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
{
"slug": "inferenceclear",
"box": 77,
"date": "2026-09-29",
"title": "InferenceClear",
"question": "How does AI inference work?",
"hook": "Training a model happens once; using it happens billions of times. Watch a frozen network answer questions, light up a grid of multiply-adds, see why a chatbot's speed is set by memory, shrink a model to 4 bits and break it at 2, and follow one question into a 120 kW rack and back.",
"thumbText": "HOW AI ANSWERS YOU",
"field": "ai",
"minutes": 20,
"tags": [
"AI inference",
"inference",
"machine learning",
"neural network",
"LLM",
"ChatGPT",
"Gemini",
"GPU",
"TPU",
"NPU",
"matrix multiplication",
"FLOPs",
"multiply-accumulate",
"memory bandwidth",
"HBM",
"KV cache",
"batching",
"tokens per second",
"prefill",
"decode",
"time to first token",
"quantisation",
"INT8",
"INT4",
"pruning",
"distillation",
"vLLM",
"PagedAttention",
"speculative decoding",
"llama.cpp",
"on-device AI",
"edge AI",
"data centre",
"liquid cooling",
"AI energy use",
"IEA",
"Bhashini",
"UPI",
"three.js"
],
"explainer": [
{
"title": "Train once, answer forever",
"text": "Training writes a model's weights, slowly and once. Inference freezes them and runs the inputs forward to answer a question. Our spiral network cost 192 million multiply-adds to train and 304 per answer, so after about 630,000 questions answering has cost more than learning."
},
{
"title": "It is all multiply-add",
"text": "A forward pass is mostly one step: multiply an input by a weight and add it to a total. A layer is a matrix times a vector. An LLM needs about 2 × parameters FLOPs per token, and GPUs, TPUs and phone NPUs do thousands of these sums at once."
},
{
"title": "Memory is the real limit",
"text": "To write each token, a chatbot must read every weight from memory. So tokens per second ≈ memory bandwidth ÷ model size in bytes: about 200 for an 8B model in 16-bit on an H100. A growing KV cache adds more to read; batching many users shares the reads."
},
{
"title": "Fewer bits, smaller models",
"text": "Quantisation rounds weights onto a few levels: FP32 → FP16 → INT8 → INT4. Size and read time fall; accuracy holds until too few bits are left, and our net breaks at 2 bits. Pruning and distillation shrink models too, which is how phones run AI."
},
{
"title": "The journey of one question",
"text": "Your prompt travels to a data centre, is tokenised, then prefilled in one big pass. The answer is decoded one token at a time and streamed back. Racks draw around 120 kW and need liquid cooling; a median chatbot text prompt was reported at about 0.24 Wh."
},
{
"title": "Phone or cloud",
"text": "On-device AI keeps data private, works offline and has no network wait, but the model must be small. The cloud runs far bigger models but needs a connection and servers. Each answer uses little energy, yet data centres used about 1.5% of the world's electricity in 2024."
}
],
"concepts": [
{
"term": "Inference",
"def": "Using a trained model, with its weights frozen, to answer a new question."
},
{
"term": "Multiply-add (MAC)",
"def": "Multiply one input by one weight and add it to a running total; two FLOPs."
},
{
"term": "FLOP",
"def": "One floating-point operation, such as a single multiply or add."
},
{
"term": "Memory bandwidth",
"def": "How many bytes per second can move from memory to the chip; it sets LLM decoding speed."
},
{
"term": "KV cache",
"def": "Saved keys and values for every earlier token, so they need not be recomputed."
},
{
"term": "Batching",
"def": "Serving many users' next tokens with one read of the weights."
},
{
"term": "Quantisation",
"def": "Storing weights with fewer bits by rounding them onto a few allowed levels."
},
{
"term": "Prefill and decode",
"def": "Processing the whole prompt at once, then writing the answer one token at a time."
},
{
"term": "On-device AI",
"def": "Running a model on your own phone or laptop instead of in a data centre."
}
],
"links": {},
"storage": [
{
"key": "inferenceclear.v1",
"what": "Which chapters you have opened, your best quiz scores, and sound on or off."
}
],
"credits": [
{
"name": "three.js",
"license": "MIT",
"url": "https://threejs.org"
},
{
"name": "Geist, Instrument Serif",
"license": "SIL OFL 1.1",
"url": "https://openfontlicense.org"
}
]
}