Detailed LLM performance metrics for the pi coding agent — TTFT, real-time tokens/s, output token counts, context usage, cost, and per-turn statistics, rendered live in the footer.
● MiniMax-M3 · turn 3 · TTFT 597ms · 1.1s · ⚙ 1 tool
⚡ 75.5 tok/s · avg 66.1 · out 82 · think 49
ctx 12.3k/200k (6%) · in 1.5k · cached 384 · $0.0006
sess 3 turns · ↓246 ↑2.1k · cached 1.2k · $0.0019
pi install git:github.com/kurashizu/pi-token-statsOr pin a version:
pi install git:github.com/kurashizu/pi-token-stats@v1.0.0To try without installing:
pi -e git:github.com/kurashizu/pi-token-statsLoaded automatically after install. Configure with the in-chat command:
| Command | Effect |
|---|---|
/tokenstats |
Cycle off → full → brief → off |
/tokenstats off |
Hide all stats |
/tokenstats brief |
Single-line status bar (keeps default footer) |
/tokenstats full |
4-line footer (default after install) |
- Streaming indicator (
●live,✓done,○idle) - Model id
- Turn number
- TTFT (color-coded: < 500 ms green, < 1.5 s blue, ≥ 1.5 s yellow)
- Elapsed time
- Running tool count
- Stop reason on error / abort
- Real-time tokens/s (1 s sliding window)
- Average tokens/s over the turn
- Output tokens (estimated from chars/4 during stream, real from
usage.outputafter) - Reasoning tokens separately when present
- Context window usage (
tokens / limit (percentage)) - Input tokens
- Cache hit tokens (when non-zero)
- Per-turn cost
Shown only after the first turn:
- Total turns
- Total in/out tokens
- Total cached tokens
- Total session cost
- Character count is divided by 4 for live token estimates — final values come from real provider usage at
message_end. - TTFT measures the gap between
turn_startand the firsttext_delta/thinking_delta. - Multi-message turns (tool roundtrips) accumulate across messages; session totals persist.
- The extension is a no-op in non-TUI modes (RPC / JSON / print).
MIT — see LICENSE.