This document specifies the benchmarking methodology, hardware measurement guidelines, and throughput metrics used to evaluate InferenceOS performance.
-
Prompt Ingestion Throughput (
prompt_tps):$$\text{Prompt TPS} = \frac{N_{\text{prompt_tokens}}}{T_{\text{prefill_seconds}}}$$ Measures how fast the model ingests input context. Highly dependent on the Dynamic Microbatch Scheduler setting. -
Evaluation Generation Throughput (
eval_tps):$$\text{Eval TPS} = \frac{N_{\text{gen_tokens}}}{T_{\text{decode_seconds}}}$$ Measures token generation speed during token-by-token decoding. -
Time To First Token (
TTFT):$$\text{TTFT} = T_{\text{first_token_received}} - T_{\text{request_submitted}}$$ Measures user-perceived responsiveness.
Execute the automated benchmarking suite via CLI:
inferenceos benchmark models/llama-3-8b.Q4_K_M.gguf -t 8 -g 28Or run the comparison benchmark scripts:
python run_benchmark.py
python run_ollama_comparison.py