Skip to content

Latest commit

 

History

History
36 lines (24 loc) · 1.1 KB

File metadata and controls

36 lines (24 loc) · 1.1 KB

Benchmarking Methodology & Performance Guide

This document specifies the benchmarking methodology, hardware measurement guidelines, and throughput metrics used to evaluate InferenceOS performance.


Performance Metrics

  1. Prompt Ingestion Throughput (prompt_tps): $$\text{Prompt TPS} = \frac{N_{\text{prompt_tokens}}}{T_{\text{prefill_seconds}}}$$ Measures how fast the model ingests input context. Highly dependent on the Dynamic Microbatch Scheduler setting.

  2. Evaluation Generation Throughput (eval_tps): $$\text{Eval TPS} = \frac{N_{\text{gen_tokens}}}{T_{\text{decode_seconds}}}$$ Measures token generation speed during token-by-token decoding.

  3. Time To First Token (TTFT): $$\text{TTFT} = T_{\text{first_token_received}} - T_{\text{request_submitted}}$$ Measures user-perceived responsiveness.


Running Benchmarks

Execute the automated benchmarking suite via CLI:

inferenceos benchmark models/llama-3-8b.Q4_K_M.gguf -t 8 -g 28

Or run the comparison benchmark scripts:

python run_benchmark.py
python run_ollama_comparison.py