Public performance benchmarks and trend dashboards for the Nimbus gateway.
Status: dashboards in progress. The benchmark harness and CI gates live and run in the main repo today; this repo is the public home for the published trend data and dashboards, which are being wired up. See methodology.md for what is measured and how regressions are judged.
Nimbus is a local-first AI agent whose gateway does real work on the hot path: indexing, hybrid search (BM25 + vector), embedding, and multi-step tool execution. Performance is tracked continuously rather than spot-checked. The harness lives in the main repo under packages/gateway/src/perf; this repo publishes the resulting history so anyone can see the trend without running it.
- A perf suite runs in CI across a per-OS runner matrix (Windows / macOS / Linux).
- History is recorded via
github-action-benchmark, with an advisory Bencher Cloud trend ingest running alongside (each runner registered as a Bencher testbed) for richer throughput/token charts. - An in-code
gateClasscomparator is the sole blocking gate. Bencher and the action-benchmark charts are advisory/observational — they never block a merge on their own.
Each metric is charted over time per runner. A point that crosses the configured gateClass threshold in the in-code comparator is a real regression and fails CI in the main repo; movement within noise bands is expected run-to-run variation. See methodology.md for the gate semantics and how to distinguish a regression from noise.
MIT.