diff --git a/.gitignore b/.gitignore index b0051b4..118e03d 100644 --- a/.gitignore +++ b/.gitignore @@ -4,3 +4,6 @@ dist/ build/ *.egg-info/ .DS_Store + +# Claude Code session files +.claude/ diff --git a/README.md b/README.md index b6943ef..885fe0f 100644 --- a/README.md +++ b/README.md @@ -16,7 +16,7 @@
-**[Claude Code](#supported-agents) · [OpenClaw](#supported-agents) · [Hermes Agent](#supported-agents)** — one normalized core, local, read-only, zero dependencies +**[Claude Code](#supported-agents) · [Codex CLI](#supported-agents) · [Gemini CLI](#supported-agents) · [opencode](#supported-agents) · [OpenClaw](#supported-agents) · [Hermes Agent](#supported-agents)** — one normalized core, local, read-only, zero dependencies ``` uvx agentburn @@ -50,7 +50,8 @@ One command, no account, nothing leaves your computer: ```bash uvx agentburn # where it burns, and what to change -uvx agentburn limits # how fast you fill a usage window +uvx agentburn limits # how fast you fill a usage window, and how long until the wall +uvx agentburn context # what long contexts cost — and what a /clear at 150k would have saved ``` ## Two ways agents cost you, two questions @@ -70,15 +71,78 @@ Optimizing a subscription doesn't change your bill. It changes how far you get b - **Peak vs typical.** Your worst rolling 5-hour window against the median of your own active ones. The ratio is the finding: a wall is hit by the peak. - **What filled it** — by model, by source (you / subagents / scheduled work), and by kind (cache reads vs cache writes vs output). -- **Measured against your own wall.** Anthropic doesn't publish the formula behind those allowances, so agentburn refuses to invent a threshold. Tell it when you were actually cut off and the arithmetic becomes yours: - - ```bash - agentburn limits --hit "2026-08-20 14:30" - # ceiling 38.4M weighted tokens ← measured from your own cut-off - # peak window 107% of your ceiling - # last 5h 12% of your ceiling +- **Measured against your own wall — automatically.** Anthropic doesn't publish the formula behind those allowances, so agentburn refuses to invent a threshold. But Claude Code writes the cut-off into the transcript itself (*"You've hit your session limit · resets 8:30pm"*), and every one of those moments is a measured ceiling. With several, the ceiling is their median: + + ```text + YOUR MEASURED CEILING + median of 35 cut-offs Claude Code recorded itself + ceiling 146M weighted tokens + peak window 137% of your ceiling + last 5h 16% of your ceiling + TIME TO WALL 2.7 h at the pace of the last 30 min ``` + No cut-off in your logs yet? `--hit "2026-08-20 14:30"` names one by hand. A measured ceiling is remembered in `~/.agentburn/ceiling.json`, so the status line below knows it too. +- **Codex: the provider's own reading.** Codex CLI writes `rate_limits.used_percent` next to every request. agentburn pairs each reading with your weighted usage of the same window and takes the median — a ceiling from the provider's arithmetic, not from a cut-off. Treat it as an estimate: that percentage counts every device and app on the account, while your local rollouts are only part of it — and when Codex stops reporting a window (plan or client change), a later peak is flagged as measured on earlier windows, not sold as an overrun. +- **Time to wall.** Ceiling minus the current window, divided by the pace of the last half hour. The number you actually want while working. +- **The week, too.** The heaviest rolling 7-day span, how much of it this week already is, and a weekly ceiling when Claude Code recorded a weekly cut-off. +- **By project.** Sessions record their working directory; the peak window is split by it. + +### `agentburn statusline` — the wall, live, inside Claude Code + +One line, no colour, built for Claude Code's `statusLine`: + +```text +⏳ 5h 63% · wall in 47 min · week 71% +``` + +```json +{ "statusLine": { "type": "command", "command": "uvx agentburn statusline" } } +``` + +Reads only the last three days of logs (the ceiling comes from the state file), so it stays cheap enough to run on every turn. + +### `agentburn context` — what a long context costs + +Every call re-reads its whole context, and on a subscription that re-reading *is* the window: a turn at 300k costs what three turns at 100k cost. Claude Code records the exact context size of every call, so this is measured, not modelled: + +```text +📏 agentburn context — claude-code · what a long context costs + + CALLS 156,226 median context 143K · p90 316K · max 704K + + WHERE THE WINDOW GOES, BY CONTEXT SIZE + 100–200k ██████············ 35% 59,780 calls + 200–400k ████████·········· 43% 42,420 calls + >400k ██················ 11% 7,257 calls + + IF YOU HAD RESTARTED AT… + /clear at 100K → 41% of the window not spent (108,573 calls were past it) + /clear at 150K → 26% of the window not spent (73,600 calls were past it) + + WHAT A SKILL COSTS + handoff 7.96K per load × 226 = 1.8M + claude-api 33.6K per load × 14 = 470K +``` + +- **The `/clear` arithmetic** — the part of every call's context above a threshold, at the cache-read rate: the honest saving of a restart habit, assuming the same work in shorter sessions. +- **Skill costs, measured** — the context growth right after a lone `Skill` call, median of recent loads. Bundled skills never touch the disk; the transcript sees all of them. +- **By effort level** — how much of the window each `effort` setting took. +- Findings with a lever land in `agentburn fix`: the restart threshold, and the heavy skills. + +### `agentburn commits` — what a commit cost you + +Sessions record their working directory and branch; your repositories record when each commit landed. The usage between two consecutive commits is what the second one cost — read-only `git log`, nothing written: + +```text + COSTLIEST COMMITS + 124M 33_Thoforge 1f7a31a1 Aug 30 fix(ui): правки UX-аудита — раскладка, навигация + 81.2M 33_Thoforge ad19bff7 Aug 28 feat(ui): цель над деревом и развилка в карточке + + BY REPOSITORY + 33_Thoforge 1.95M median · 287 commits · 1.52B total +``` + Weighted tokens = tokens × *published* price ratios (cache read 0.1×, cache write 1.25×, output per model), normalized to one input token of the reference model. Every ratio is public; none of them is a guess about how the provider counts. ### `agentburn` — the money view @@ -111,7 +175,7 @@ Not "consider a cheaper model" but the exact file and the exact lines. Patch gen | Agent | Verified levers | |---|---| -| Claude Code | registered MCP servers (`~/.claude.json`, `.mcp.json`), always-loaded `CLAUDE.md` memory files | +| Claude Code | registered MCP servers (`~/.claude.json`, `.mcp.json`), always-loaded `CLAUDE.md` memory files, the session-restart threshold (measured), heavy skills (measured per load) | | Hermes | per-job `model` / `enabled_toolsets` (`cron/jobs.py`), per-platform toolsets (`gateway/run.py`) | | OpenClaw | `heartbeat.{every, activeHours, model, lightContext}` (`config/types.agent-defaults.ts`) | @@ -122,6 +186,7 @@ There is no `--apply` on purpose: it's your agent's config. Paste it yourself, t Token trackers quietly disagree with each other (2–91× in public issue threads). agentburn takes the opposite stance: - Numbers come from **the agent's own accounting**, read-only. No scraping, no proxies, no guessing. +- **One reply is counted once.** Claude Code writes one transcript line per content block, each carrying the same `usage`; summing lines inflates calls and tokens ~1.8×. agentburn deduplicates by `requestId` (found and fixed in 0.14.0 — earlier absolute totals from this tool were inflated by that factor; ratios were not). - Provider-billed costs are shown as-is; estimates are marked `~`; mixed data is labeled mixed. - **Where a price doesn't exist, none is invented.** Claude Code records no costs and subscription usage has no honest per-token price — so that adapter reports tokens and windows, never dollars. - Sessions with messages but **zero recorded tokens** (known accounting gaps, e.g. [hermes-agent #12023](https://github.com/NousResearch/hermes-agent/issues/12023)) are detected: totals become an explicit **lower bound**, and fixing the accounting becomes recommendation #1. @@ -158,6 +223,9 @@ Always-on agents bill you around the clock — and their built-in counters only | | **agentburn** | ccusage | codeburn | built-in `/usage` | |---|---|---|---|---| | Usage **windows** (peak vs typical, what filled them) | ✅ | — | — | current window only | +| Ceiling measured from your own recorded cut-offs · time to wall · status line | ✅ | — | — | current window % | +| The price of long contexts · what a `/clear` would have saved · skill cost per load | ✅ | — | — | — | +| Cost per git commit | ✅ | — | — | — | | Burn by *source* (cron · heartbeat · gateways · subagents) | ✅ | — | — | % only, 7 days | | 🌙 the overnight bill, isolated | ✅ | — | — | — | | Behavioral forensics (`why`: loops, retry storms, failed-run cost) | ✅ | — | — | — | @@ -176,8 +244,11 @@ One normalized model, one adapter per agent. Run `agentburn` and every agent fou | **Claude Code** | ✅ | `~/.claude/projects/**.jsonl` | tokens and **windows**, by design: no local costs, no honest per-token price for a subscription | | **OpenClaw** | ✅ | `~/.openclaw/agents/*/sessions/sessions.json` | **heartbeat is its own category** — the famous one | | **Hermes Agent** | ✅ | `~/.hermes/state.db` (+ optional request dumps) | costs from the agent's own accounting | +| **Codex CLI** | ✅ | `~/.codex/sessions/**/rollout-*.jsonl` | tokens and windows; the only agent that records the **provider's own usage %** with every request | +| **Gemini CLI** | ✅ | `~/.gemini/tmp/*/chats/session-*.json` | per-turn tokens incl. thoughts; working directory via `projects.json` | +| **opencode** | ✅ | `~/.local/share/opencode/opencode.db` | costs from the agent's own price list; free/self-hosted providers show tokens only | -Adapters are ~150 lines over a shared model. Codex CLI / opencode are natural next targets — PRs welcome. +Adapters are ~150 lines over a shared model — PRs for the next one welcome.
architecture: agent data → adapters → normalized model → report/limits/why/fix/explain/doctor/mcp
@@ -186,7 +257,7 @@ Adapters are ~150 lines over a shared model. Codex CLI / opencode are natural ne
🔌 agentburn mcp — your agent answers for its own bill -A zero-dependency MCP stdio server exposing `burn_report` / `burn_limits` / `burn_why` / `burn_card`. Register it and ask *"where do you burn my money?"* — it profiles its own database and explains. +A zero-dependency MCP stdio server exposing `burn_report` / `burn_limits` / `burn_context` / `burn_commits` / `burn_why` / `burn_card`. Register it and ask *"where do you burn my money?"* — it profiles its own database and explains. ```bash claude mcp add agentburn -- agentburn mcp diff --git a/agentburn/__init__.py b/agentburn/__init__.py index 5f029bc..13044f1 100644 --- a/agentburn/__init__.py +++ b/agentburn/__init__.py @@ -7,4 +7,4 @@ surfaced rather than hidden, and no price is invented where none exists. """ -__version__ = "0.13.3" +__version__ = "0.14.0" diff --git a/agentburn/adapters/__init__.py b/agentburn/adapters/__init__.py index a1aa8e3..883dfad 100644 --- a/agentburn/adapters/__init__.py +++ b/agentburn/adapters/__init__.py @@ -2,12 +2,15 @@ from __future__ import annotations -from . import claude_code, hermes, openclaw +from . import claude_code, codex, gemini, hermes, openclaw, opencode ADAPTERS = { "hermes": hermes, "openclaw": openclaw, "claude-code": claude_code, + "codex": codex, + "gemini": gemini, + "opencode": opencode, } diff --git a/agentburn/adapters/claude_code.py b/agentburn/adapters/claude_code.py index 727c45d..068eaf9 100644 --- a/agentburn/adapters/claude_code.py +++ b/agentburn/adapters/claude_code.py @@ -25,7 +25,16 @@ from typing import Optional from .. import cache -from ..model import BUCKET_SECONDS, ActionEvent, SessionRec, Snapshot, UsageCell +from ..model import ( + BUCKET_SECONDS, + ActionEvent, + ContextCall, + LimitHit, + SessionRec, + SkillLoad, + Snapshot, + UsageCell, +) from .hermes import salient_arg UUID_RE = re.compile(r"^[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}$", re.I) @@ -44,6 +53,54 @@ # this in place of a model name. Not a model — do not report it as one. SYNTHETIC_MODEL = "" +# The cut-off message Claude Code writes as a synthetic turn at the moment a +# usage limit is reached. Its timestamp is a measured wall. +LIMIT_HIT_RE = re.compile(r"hit your (\w+) limit", re.I) +RESET_RE = re.compile(r"resets\s+(\d{1,2})(?::(\d{2}))?\s*(am|pm)?\s*(?:\(([^)]+)\))?", re.I) +# Precision at which per-call context sizes are cached: 1k tokens is well +# below anything a `/clear` decision turns on, and keeps the cache small. +CONTEXT_STEP = 1000 + + +def _text_of(content) -> str: + if isinstance(content, str): + return content + if isinstance(content, list): + return " ".join( + b.get("text", "") for b in content if isinstance(b, dict) and isinstance(b.get("text"), str) + ) + return "" + + +def _reset_ts(text: str, at: float) -> Optional[float]: + """'resets 8:30pm (Europe/Amsterdam)' → unix seconds of that clock time, the + first occurrence at or after `at`. None when the zone can't be resolved.""" + m = RESET_RE.search(text) + if not m: + return None + hour = int(m.group(1)) + minute = int(m.group(2) or 0) + ampm = (m.group(3) or "").lower() + if ampm == "pm" and hour < 12: + hour += 12 + if ampm == "am" and hour == 12: + hour = 0 + zone = m.group(4) + try: + if zone: + from zoneinfo import ZoneInfo + + tz = ZoneInfo(zone.strip()) + else: + tz = dt.datetime.now().astimezone().tzinfo + base = dt.datetime.fromtimestamp(at, tz) + cand = base.replace(hour=hour, minute=minute, second=0, microsecond=0) + if cand.timestamp() < at: + cand += dt.timedelta(days=1) + return cand.timestamp() + except Exception: # noqa: BLE001 — no tzdata on this machine, unknown zone + return None + def _result_weight(content) -> int: """Approximate token weight of a tool result. 0 when nothing is measurable.""" @@ -60,6 +117,11 @@ def _result_weight(content) -> int: return 0 +def project_name(cwd: str) -> str: + """Last path component of a recorded working directory, on any OS.""" + return os.path.basename(cwd.rstrip("/\\")) or cwd + + def default_root() -> str: return os.path.join(os.path.expanduser("~"), ".claude", "projects") @@ -94,9 +156,16 @@ def _scan_file(path: str) -> dict: compactions, plus compact rows for events and per-bucket usage. Plain lists rather than dataclasses, because this is exactly what gets cached. + One model reply is written as SEVERAL lines — one per content block + (thinking, text, each tool_use) — and every one of them carries the SAME + `usage` object. Usage is therefore counted once per `requestId` (fallback: + `message.id`); summing rows would inflate calls and tokens ~1.8×. + Also appends ActionEvents (tool_use with salient arg; tool_result error flags linked by tool_use_id), fills `cells` — (bucket, model) → usage, the - intra-session resolution `agentburn limits` needs — and counts compactions — lines whose + intra-session resolution `agentburn limits` needs — records per-call + context sizes, skill loads (context growth after a lone `Skill` call), + limit cut-offs the agent wrote itself, and counts compactions — lines whose type/subtype mentions a compact boundary (anthropics/claude-code writes `subtype: "compact_boundary"`). Each compaction re-sends a near-full context window, so the count is a direct cost signal. @@ -108,9 +177,20 @@ def _scan_file(path: str) -> dict: lines = 0 compactions = 0 no_ts = 0 + dup_rows = 0 id_to_name = {} events: list = [] cells: dict = {} + contexts: list = [] # [ts, model, context//STEP, output, effort] + skills: list = [] # [ts, skill, tokens] + hits: list = [] # [ts, kind, reset_ts] + effort_seen: dict = {} # effort → [calls, context, output] + cwd = branch = None + seen_req: set = set() + # Skill weight = context growth between the reply that called `Skill` and + # the next reply — only when Skill was the sole tool call in that reply, + # otherwise the other tool's result lands in the delta. + pending_skill = None # [request_key, skill, context_before, tool_calls] with open(path, "r", encoding="utf-8", errors="replace") as f: for raw in f: raw = raw.strip() @@ -121,18 +201,38 @@ def _scan_file(path: str) -> dict: obj = json.loads(raw) except json.JSONDecodeError: continue + if not isinstance(obj, dict): + continue ts = _parse_ts(obj.get("timestamp")) if ts: first = ts if first is None else min(first, ts) last = ts if last is None else max(last, ts) + if cwd is None and isinstance(obj.get("cwd"), str): + cwd = obj["cwd"] + if isinstance(obj.get("gitBranch"), str) and obj["gitBranch"]: + branch = obj["gitBranch"] marker = f"{obj.get('type', '')}/{obj.get('subtype', '')}".lower() if "compact" in marker: compactions += 1 msg = obj.get("message") if not isinstance(msg, dict): continue + m_ = msg.get("model") + synthetic = m_ == SYNTHETIC_MODEL + content = msg.get("content") + if synthetic: + text = _text_of(content) + hm = LIMIT_HIT_RE.search(text) + if hm and ts: + hits.append([ts, hm.group(1).lower(), _reset_ts(text, ts)]) + req = obj.get("requestId") or msg.get("id") u = msg.get("usage") - if isinstance(u, dict): + fresh = isinstance(u, dict) and not synthetic and (req is None or req not in seen_req) + if isinstance(u, dict) and not synthetic and req is not None and req in seen_req: + dup_rows += 1 + if fresh: + if req is not None: + seen_req.add(req) calls += 1 i_ = int(u.get("input_tokens") or 0) o_ = int(u.get("output_tokens") or 0) @@ -142,12 +242,19 @@ def _scan_file(path: str) -> dict: out += o_ cw += w_ cr += r_ + ctx = i_ + r_ + w_ + effort = obj.get("effort") if isinstance(obj.get("effort"), str) else None + contexts.append([ts, m_, ctx // CONTEXT_STEP, o_, effort]) + e_ = effort_seen.setdefault(effort or "default", [0, 0, 0]) + e_[0] += 1 + e_[1] += ctx + e_[2] += o_ + if pending_skill is not None and pending_skill[0] != req: + if pending_skill[3] == 1 and ctx > pending_skill[2]: + skills.append([ts, pending_skill[1], ctx - pending_skill[2]]) + pending_skill = None if ts: - m_ = msg.get("model") - key = ( - int(ts // BUCKET_SECONDS) * BUCKET_SECONDS, - None if m_ == SYNTHETIC_MODEL else m_, - ) + key = (int(ts // BUCKET_SECONDS) * BUCKET_SECONDS, m_) c = cells.get(key) if c is None: cells[key] = [1, i_, o_, r_, w_] @@ -161,9 +268,8 @@ def _scan_file(path: str) -> dict: # No timestamp = no window. Counted so the report can say so # instead of quietly under-reporting every window. no_ts += 1 - if msg.get("model") and msg["model"] != SYNTHETIC_MODEL: - model = msg["model"] - content = msg.get("content") + if m_ and not synthetic: + model = m_ if isinstance(content, list) and len(events) < MAX_EVENTS_PER_FILE: for item in content: if not isinstance(item, dict): @@ -174,6 +280,20 @@ def _scan_file(path: str) -> dict: events.append( [ts, str(item["name"])[:40], salient_arg(item.get("input")), None, None] ) + if req is not None and isinstance(u, dict): + if pending_skill is not None and pending_skill[0] == req: + pending_skill[3] += 1 + elif item["name"] == "Skill": + skill = (item.get("input") or {}).get("skill") if isinstance(item.get("input"), dict) else None + if skill: + ctx_now = ( + int(u.get("input_tokens") or 0) + + int(u.get("cache_read_input_tokens") or 0) + + int(u.get("cache_creation_input_tokens") or 0) + ) + pending_skill = [req, str(skill)[:60], ctx_now, 1] + elif pending_skill is not None: + pending_skill[3] += 1 elif item.get("type") == "tool_result": name = id_to_name.get(item.get("tool_use_id"), "tool") events.append( @@ -192,8 +312,15 @@ def _scan_file(path: str) -> dict: "lines": lines, "compactions": compactions, "no_ts": no_ts, + "dup_rows": dup_rows, + "cwd": cwd, + "branch": branch, "events": events, "cells": [[b, m] + v for (b, m), v in cells.items()], + "contexts": contexts, + "skills": skills, + "hits": hits, + "effort": effort_seen, } @@ -221,7 +348,7 @@ def load( ] subs = glob.glob(os.path.join(root, "*", "*", "subagents", "*.jsonl")) - stats = {"parsed": 0, "cached": 0, "no_ts": 0, "truncated": 0} + stats = {"parsed": 0, "cached": 0, "no_ts": 0, "truncated": 0, "dup_rows": 0} def consider(path: str, source: str, parent: Optional[str], title: str): sid = os.path.basename(path)[:-6] @@ -245,6 +372,7 @@ def consider(path: str, source: str, parent: Optional[str], title: str): return stats["no_ts"] += scan.get("no_ts", 0) + stats["dup_rows"] += scan.get("dup_rows", 0) if len(scan["events"]) >= MAX_EVENTS_PER_FILE: stats["truncated"] += 1 for ts, name, arg_key, ok, tokens in scan["events"]: @@ -264,8 +392,30 @@ def consider(path: str, source: str, parent: Optional[str], title: str): output_tokens=o_, cache_read_tokens=r_, cache_write_tokens=w_, + session=sid, + ) + ) + for ts, m_, ctx_k, o_, effort in scan.get("contexts", []): + if days and ts is not None and ts < since: + continue + snap.context_calls.append( + ContextCall( + ts=ts, + session=sid, + model=None if m_ == SYNTHETIC_MODEL else m_, + context=ctx_k * CONTEXT_STEP, + output=o_, + effort=effort, ) ) + for ts, skill, tokens in scan.get("skills", []): + if days and ts is not None and ts < since: + continue + snap.skill_loads.append(SkillLoad(session=sid, ts=ts, skill=skill, tokens=tokens)) + for ts, kind, reset in scan.get("hits", []): + if days and ts < since: + continue + snap.limit_hits.append(LimitHit(ts=ts, kind=kind, reset_at=reset)) if scan["compactions"]: snap.compactions[sid] = scan["compactions"] snap.sessions.append( @@ -286,6 +436,8 @@ def consider(path: str, source: str, parent: Optional[str], title: str): cost_usd=None, cost_basis="unknown", message_count=scan["lines"], + project=scan.get("cwd"), + branch=scan.get("branch"), ) ) @@ -294,6 +446,10 @@ def consider(path: str, source: str, parent: Optional[str], title: str): for p in mains: project = os.path.basename(os.path.dirname(p)).strip("-").split("-")[-1] or "project" consider(p, "cli", None, f"{project}/{os.path.basename(p)[:8]}") + # A session's own record of where it ran beats the encoded directory name. + for s_ in snap.sessions: + if s_.source == "cli" and s_.project: + s_.title = f"{project_name(s_.project)}/{s_.id[:8]}" for p in subs: session_uuid = os.path.basename(os.path.dirname(os.path.dirname(p))) consider(p, "subagent", session_uuid, f"subagent {os.path.basename(p)[:18]}") @@ -321,6 +477,12 @@ def consider(path: str, source: str, parent: Optional[str], title: str): f"{stats['no_ts']:,} of {total_calls:,} API calls carry no timestamp: they are in the " "totals but not in any window, so `agentburn limits` is a lower bound." ) + if snap.limit_hits: + n_ = len(snap.limit_hits) + snap.warnings.append( + f"{n_} usage-limit cut-off(s) recorded by Claude Code itself in this window — " + "`agentburn limits` measures your ceiling from them." + ) if stats["truncated"]: snap.warnings.append( f"{stats['truncated']} transcript(s) hit the {MAX_EVENTS_PER_FILE:,}-event per-file cap; " diff --git a/agentburn/adapters/codex.py b/agentburn/adapters/codex.py new file mode 100644 index 0000000..1cbb552 --- /dev/null +++ b/agentburn/adapters/codex.py @@ -0,0 +1,257 @@ +"""Codex CLI adapter: reads ~/.codex/sessions/YYYY/MM/DD/rollout-*.jsonl (read-only). + +Layout observed in openai/codex 0.144 (September 2026), one JSONL per thread, +every line `{"timestamp", "type", "payload"}`: + session_meta once: cwd, originator ("codex_exec" CLI, "Codex Desktop"), cli_version + turn_context per turn: model, effort, cwd + event_msg payload.type == "token_count": + info.total_token_usage — cumulative for the thread + info.last_token_usage — the last request + rate_limits.primary/secondary — {used_percent, window_minutes, resets_at} + response_item function_call / function_call_output, custom_tool_call(_output) + +Usage is taken as the DELTA of the cumulative counter, so a repeated +token_count (rate-limit refreshes re-send the same totals) counts once. +`input_tokens` includes `cached_input_tokens` (total = input + output). + +Codex does not record costs locally and on a ChatGPT plan there is no honest +per-token price, so this adapter reports tokens and windows, never dollars — +same stance as the Claude Code adapter. What Codex *does* record that no other +agent does is the provider's own view of the window: `rate_limits.used_percent` +at the time of every request. Those samples are kept as RateLimitSample and +turn into a measured ceiling in `agentburn limits`. +""" + +from __future__ import annotations + +import datetime as dt +import glob +import json +import os +import time +from typing import Optional + +from .. import cache +from ..model import ( + BUCKET_SECONDS, + ActionEvent, + ContextCall, + RateLimitSample, + SessionRec, + Snapshot, + UsageCell, +) +from .hermes import salient_arg + +CHARS_PER_TOKEN = 4 +MAX_EVENTS_PER_FILE = 80_000 + + +def default_root() -> str: + return os.path.join(os.path.expanduser("~"), ".codex", "sessions") + + +def available() -> bool: + root = default_root() + return os.path.isdir(root) and bool(glob.glob(os.path.join(root, "*", "*", "*", "rollout-*.jsonl"))) + + +def _parse_ts(v) -> Optional[float]: + if isinstance(v, str): + try: + return dt.datetime.fromisoformat(v.replace("Z", "+00:00")).timestamp() + except ValueError: + return None + return None + + +def _scan_file(path: str) -> dict: + first = last = None + calls = 0 + inp = out = cr = reasoning = 0 + model = None + cwd = None + originator = None + lines = 0 + compactions = 0 + events: list = [] + cells: dict = {} + contexts: list = [] + limits: list = [] # [ts, window_minutes, used_percent, resets_at] + prev_total = None + prev = None # previous cumulative usage dict + effort = None + with open(path, "r", encoding="utf-8", errors="replace") as f: + for raw in f: + raw = raw.strip() + if not raw: + continue + lines += 1 + try: + obj = json.loads(raw) + except json.JSONDecodeError: + continue + if not isinstance(obj, dict): + continue + ts = _parse_ts(obj.get("timestamp")) + if ts: + first = ts if first is None else min(first, ts) + last = ts if last is None else max(last, ts) + kind = obj.get("type") + p = obj.get("payload") + if not isinstance(p, dict): + continue + if kind == "session_meta": + cwd = cwd or p.get("cwd") + originator = p.get("originator") + elif kind == "turn_context": + if p.get("model"): + model = p["model"] + effort = p.get("effort") if isinstance(p.get("effort"), str) else effort + cwd = p.get("cwd") or cwd + elif kind == "compacted" or (kind == "event_msg" and p.get("type") == "context_compacted"): + compactions += 1 + elif kind == "event_msg" and p.get("type") == "token_count": + info = p.get("info") or {} + tot = info.get("total_token_usage") or {} + total = int(tot.get("total_tokens") or 0) + rl = p.get("rate_limits") or {} + for key in ("primary", "secondary"): + w = rl.get(key) + if isinstance(w, dict) and w.get("used_percent") is not None and w.get("window_minutes") and ts: + limits.append([ts, int(w["window_minutes"]), float(w["used_percent"]), + w.get("resets_at")]) + if total <= 0 or total == prev_total: + continue + if prev_total is not None and total > prev_total and prev: + d_in = int(tot.get("input_tokens") or 0) - int(prev.get("input_tokens") or 0) + d_cr = int(tot.get("cached_input_tokens") or 0) - int(prev.get("cached_input_tokens") or 0) + d_out = int(tot.get("output_tokens") or 0) - int(prev.get("output_tokens") or 0) + d_rs = int(tot.get("reasoning_output_tokens") or 0) - int(prev.get("reasoning_output_tokens") or 0) + else: + # first sample, or the counter reset (a new thread after compaction) + src = info.get("last_token_usage") or tot + d_in = int(src.get("input_tokens") or 0) + d_cr = int(src.get("cached_input_tokens") or 0) + d_out = int(src.get("output_tokens") or 0) + d_rs = int(src.get("reasoning_output_tokens") or 0) + prev_total, prev = total, tot + if d_in < 0 or d_out < 0: + continue + d_cr = max(0, min(d_cr, d_in)) + calls += 1 + inp += d_in - d_cr + cr += d_cr + out += d_out + reasoning += d_rs + contexts.append([ts, model, d_in // 1000, d_out, effort]) + if ts: + key = (int(ts // BUCKET_SECONDS) * BUCKET_SECONDS, model) + c = cells.get(key) + if c is None: + cells[key] = [1, d_in - d_cr, d_out, d_cr, 0] + else: + c[0] += 1 + c[1] += d_in - d_cr + c[2] += d_out + c[3] += d_cr + elif kind == "response_item" and len(events) < MAX_EVENTS_PER_FILE: + t = p.get("type") + if t in ("function_call", "custom_tool_call"): + name = p.get("name") or "tool" + args = p.get("arguments") if t == "function_call" else p.get("input") + events.append([ts, str(name)[:40], salient_arg(args), None, None]) + elif t in ("function_call_output", "custom_tool_call_output"): + output = p.get("output") + text = output if isinstance(output, str) else json.dumps(output) if output is not None else "" + ok = None + if isinstance(output, dict) and "success" in output: + ok = bool(output.get("success")) + events.append([ts, "tool", None, ok, len(text) // CHARS_PER_TOKEN]) + return { + "first": first, "last": last, "calls": calls, "inp": inp, "out": out, "cr": cr, + "reasoning": reasoning, "model": model, "cwd": cwd, "originator": originator, + "lines": lines, "compactions": compactions, "events": events, + "cells": [[b, m] + v for (b, m), v in cells.items()], + "contexts": contexts, "limits": limits, + } + + +def load( + db_path: Optional[str] = None, + days: Optional[int] = 30, + dumps_dir: Optional[str] = None, + now: Optional[float] = None, +) -> Snapshot: + root = db_path or default_root() + if not os.path.isdir(root): + raise FileNotFoundError( + f"Codex sessions dir not found at {root}. Pass --db ~/.codex/sessions (or its actual location)." + ) + now = now or time.time() + since = now - days * 86400 if days else 0 + snap = Snapshot(agent="codex", source_path=root, generated_at=now, days=days) + files = glob.glob(os.path.join(root, "*", "*", "*", "rollout-*.jsonl")) + cache.maybe_sweep("codex") + empty = 0 + for path in files: + try: + if days and os.path.getmtime(path) < since: + continue + key = cache.stamp(path) + scan = cache.get("codex", path, key) + if scan is None: + scan = _scan_file(path) + cache.put("codex", path, key, scan) + except OSError: + continue + if scan["lines"] == 0 or (days and scan["last"] is not None and scan["last"] < since): + continue + if scan["calls"] == 0 and not scan["events"]: + empty += 1 # a thread that never got a model reply is not an accounting gap + continue + sid = os.path.basename(path)[len("rollout-"):-6] + source = "desktop" if (scan.get("originator") or "").lower().startswith("codex desktop") else "cli" + for ts, name, arg_key, ok, tokens in scan["events"]: + snap.events.append(ActionEvent(session_id=sid, ts=ts, name=name, arg_key=arg_key, ok=ok, tokens=tokens)) + for bucket, m_, calls_, i_, o_, r_, w_ in scan["cells"]: + if days and bucket < since: + continue + snap.usage_cells.append(UsageCell(start=bucket, source=source, model=m_, calls=calls_, + input_tokens=i_, output_tokens=o_, cache_read_tokens=r_, + cache_write_tokens=w_, session=sid)) + for ts, m_, ctx_k, o_, effort in scan.get("contexts", []): + if days and ts is not None and ts < since: + continue + snap.context_calls.append(ContextCall(ts=ts, session=sid, model=m_, context=ctx_k * 1000, + output=o_, effort=effort)) + for ts, wm, used, resets in scan.get("limits", []): + if days and ts < since: + continue + snap.rate_limits.append(RateLimitSample(ts=ts, window_minutes=wm, used_percent=used, + resets_at=float(resets) if resets else None)) + if scan["compactions"]: + snap.compactions[sid] = scan["compactions"] + cwd = scan.get("cwd") + title = f"{os.path.basename((cwd or '').rstrip('/')) or 'thread'}/{sid[-8:]}" + snap.sessions.append(SessionRec( + id=sid, source=source, model=scan["model"], started_at=scan["first"], ended_at=scan["last"], + parent_id=None, title=title[:80], api_calls=scan["calls"], input_tokens=scan["inp"], + output_tokens=scan["out"], cache_read_tokens=scan["cr"], cache_write_tokens=0, + reasoning_tokens=scan.get("reasoning", 0), cost_usd=None, cost_basis="unknown", + message_count=scan["lines"], project=cwd, + )) + if not snap.sessions: + raise RuntimeError( + f"Codex rollouts found but nothing with usage in the window ({empty} empty thread(s)) — try --days 0." + ) + snap.warnings.append( + "Codex does not record costs locally; a ChatGPT plan has no honest per-token price — " + "showing tokens and windows, not dollars." + ) + if snap.rate_limits: + snap.warnings.append( + f"{len(snap.rate_limits):,} rate-limit samples recorded by Codex itself — " + "`agentburn limits` measures your ceiling from them." + ) + return snap diff --git a/agentburn/adapters/gemini.py b/agentburn/adapters/gemini.py new file mode 100644 index 0000000..9d0b460 --- /dev/null +++ b/agentburn/adapters/gemini.py @@ -0,0 +1,197 @@ +"""Gemini CLI adapter: reads ~/.gemini/tmp//chats/session-*.json (read-only). + +One JSON document per session (google-gemini/gemini-cli, chatRecordingService): + {sessionId, projectHash, startTime, lastUpdated, kind, messages: [...]} +Messages of type "gemini" carry `model`, `timestamp`, `toolCalls` and + tokens: {input, output, cached, thoughts, tool, total} +already per turn (total = input + output + thoughts + tool; `input` includes +`cached`). Files are written without any telemetry configuration, unlike the +OTEL export, so they are the primary source. + +The directory under ~/.gemini/tmp is a label from ~/.gemini/projects.json +(path → label), which is how a session gets its working directory back. + +No local costs, and Gemini CLI is mostly used on a free tier or a Google +subscription: tokens and windows only, no invented dollars. +""" + +from __future__ import annotations + +import datetime as dt +import glob +import json +import os +import time +from typing import Optional + +from .. import cache +from ..model import BUCKET_SECONDS, ActionEvent, ContextCall, SessionRec, Snapshot, UsageCell +from .hermes import salient_arg + +CHARS_PER_TOKEN = 4 + + +def default_root() -> str: + return os.path.join(os.path.expanduser("~"), ".gemini", "tmp") + + +def available() -> bool: + root = default_root() + return os.path.isdir(root) and bool(glob.glob(os.path.join(root, "*", "chats", "session-*.json"))) + + +def _parse_ts(v) -> Optional[float]: + if isinstance(v, str): + try: + return dt.datetime.fromisoformat(v.replace("Z", "+00:00")).timestamp() + except ValueError: + return None + return None + + +def _projects(root: str) -> dict: + """label → path, from ~/.gemini/projects.json next to the tmp dir.""" + try: + with open(os.path.join(os.path.dirname(root), "projects.json"), "r", encoding="utf-8") as f: + data = json.load(f) + except (OSError, json.JSONDecodeError): + return {} + out = {} + for path, label in (data.get("projects") or {}).items(): + out.setdefault(label, path) + return out + + +def _scan_file(path: str) -> dict: + with open(path, "r", encoding="utf-8", errors="replace") as f: + try: + doc = json.load(f) + except json.JSONDecodeError: + return {"lines": 0} + if not isinstance(doc, dict): + return {"lines": 0} + msgs = doc.get("messages") or [] + first = _parse_ts(doc.get("startTime")) + last = _parse_ts(doc.get("lastUpdated")) + calls = inp = out = cr = thoughts = 0 + model = None + events: list = [] + cells: dict = {} + contexts: list = [] + for m in msgs: + if not isinstance(m, dict): + continue + ts = _parse_ts(m.get("timestamp")) + if ts: + first = ts if first is None else min(first, ts) + last = ts if last is None else max(last, ts) + if m.get("type") != "gemini": + continue + if m.get("model"): + model = m["model"] + t = m.get("tokens") + if isinstance(t, dict): + i_ = int(t.get("input") or 0) + c_ = max(0, min(int(t.get("cached") or 0), i_)) + o_ = int(t.get("output") or 0) + th = int(t.get("thoughts") or 0) + calls += 1 + inp += i_ - c_ + cr += c_ + out += o_ + thoughts += th + contexts.append([ts, m.get("model"), i_ // 1000, o_ + th, None]) + if ts: + key = (int(ts // BUCKET_SECONDS) * BUCKET_SECONDS, m.get("model")) + c = cells.get(key) + if c is None: + cells[key] = [1, i_ - c_, o_ + th, c_, 0] + else: + c[0] += 1 + c[1] += i_ - c_ + c[2] += o_ + th + c[3] += c_ + for tc in m.get("toolCalls") or []: + if not isinstance(tc, dict): + continue + name = str(tc.get("name") or "tool")[:40] + events.append([ts, name, salient_arg(tc.get("args")), None, None]) + res = tc.get("result") + text = json.dumps(res) if res is not None else "" + ok = None + if isinstance(tc.get("status"), str): + ok = tc["status"].lower() not in ("error", "failed", "cancelled") + events.append([ts, name, None, ok, len(text) // CHARS_PER_TOKEN]) + return { + "lines": len(msgs), "first": first, "last": last, "calls": calls, "inp": inp, "out": out, + "cr": cr, "thoughts": thoughts, "model": model, "kind": doc.get("kind"), + "session_id": doc.get("sessionId"), "events": events, + "cells": [[b, m] + v for (b, m), v in cells.items()], "contexts": contexts, + } + + +def load( + db_path: Optional[str] = None, + days: Optional[int] = 30, + dumps_dir: Optional[str] = None, + now: Optional[float] = None, +) -> Snapshot: + root = db_path or default_root() + if not os.path.isdir(root): + raise FileNotFoundError( + f"Gemini CLI dir not found at {root}. Pass --db ~/.gemini/tmp (or its actual location)." + ) + now = now or time.time() + since = now - days * 86400 if days else 0 + snap = Snapshot(agent="gemini", source_path=root, generated_at=now, days=days) + projects = _projects(root) + cache.maybe_sweep("gemini") + empty = 0 + for path in glob.glob(os.path.join(root, "*", "chats", "session-*.json")): + try: + if days and os.path.getmtime(path) < since: + continue + key = cache.stamp(path) + scan = cache.get("gemini", path, key) + if scan is None: + scan = _scan_file(path) + cache.put("gemini", path, key, scan) + except OSError: + continue + if scan["lines"] == 0 or (days and scan.get("last") is not None and scan["last"] < since): + continue + if scan["calls"] == 0 and not scan["events"]: + empty += 1 # a chat with no model reply is not an accounting gap + continue + label = os.path.basename(os.path.dirname(os.path.dirname(path))) + cwd = projects.get(label) + sid = scan.get("session_id") or os.path.basename(path)[8:-5] + source = "cli" if (scan.get("kind") or "main") == "main" else "subagent" + for ts, name, arg_key, ok, tokens in scan["events"]: + snap.events.append(ActionEvent(session_id=sid, ts=ts, name=name, arg_key=arg_key, ok=ok, tokens=tokens)) + for bucket, m_, calls_, i_, o_, r_, w_ in scan["cells"]: + if days and bucket < since: + continue + snap.usage_cells.append(UsageCell(start=bucket, source=source, model=m_, calls=calls_, + input_tokens=i_, output_tokens=o_, cache_read_tokens=r_, + cache_write_tokens=w_, session=sid)) + for ts, m_, ctx_k, o_, effort in scan.get("contexts", []): + if days and ts is not None and ts < since: + continue + snap.context_calls.append(ContextCall(ts=ts, session=sid, model=m_, context=ctx_k * 1000, + output=o_, effort=effort)) + snap.sessions.append(SessionRec( + id=sid, source=source, model=scan["model"], started_at=scan["first"], ended_at=scan["last"], + parent_id=None, title=f"{label}/{sid[:8]}", api_calls=scan["calls"], + input_tokens=scan["inp"], output_tokens=scan["out"], cache_read_tokens=scan["cr"], + cache_write_tokens=0, reasoning_tokens=scan.get("thoughts", 0), cost_usd=None, + cost_basis="unknown", message_count=scan["lines"], project=cwd, + )) + if not snap.sessions: + raise RuntimeError( + f"Gemini CLI sessions found but nothing with usage in the window ({empty} empty chat(s)) — try --days 0." + ) + snap.warnings.append( + "Gemini CLI does not record costs locally; showing tokens and windows, not dollars." + ) + return snap diff --git a/agentburn/adapters/opencode.py b/agentburn/adapters/opencode.py new file mode 100644 index 0000000..6a26c51 --- /dev/null +++ b/agentburn/adapters/opencode.py @@ -0,0 +1,169 @@ +"""opencode adapter: reads ~/.local/share/opencode/opencode.db (SQLite, read-only). + +Schema observed in sst/opencode (September 2026): + session(id, parent_id, directory, title, model, cost, tokens_*, time_created ms) + message(id, session_id, time_created ms, data JSON) — assistant rows carry + role, modelID, providerID, cost, path.cwd, tokens{input, output, reasoning, + cache{read, write}}, time{created, completed} + part(id, message_id, session_id, data JSON) — {"type": "tool", "tool", "state": + {"status", "input", "output"}} + +opencode records its own cost per message (from the provider's price list it +ships), so `cost_usd` is the agent's number, basis "actual" when non-zero. +Free/self-hosted providers leave it at 0 → basis "unknown", tokens only. +""" + +from __future__ import annotations + +import json +import os +import sqlite3 +import time +from typing import Optional + +from ..model import BUCKET_SECONDS, ActionEvent, ContextCall, SessionRec, Snapshot, UsageCell + +CHARS_PER_TOKEN = 4 + + +def default_root() -> str: + xdg = os.environ.get("XDG_DATA_HOME") or os.path.join(os.path.expanduser("~"), ".local", "share") + return os.path.join(xdg, "opencode", "opencode.db") + + +def available() -> bool: + return os.path.isfile(default_root()) + + +def _connect(path: str) -> sqlite3.Connection: + con = sqlite3.connect(f"file:{path}?mode=ro", uri=True) + con.row_factory = sqlite3.Row + return con + + +def load( + db_path: Optional[str] = None, + days: Optional[int] = 30, + dumps_dir: Optional[str] = None, + now: Optional[float] = None, +) -> Snapshot: + path = db_path or default_root() + if not os.path.isfile(path): + raise FileNotFoundError( + f"opencode database not found at {path}. Pass --db ~/.local/share/opencode/opencode.db." + ) + now = now or time.time() + since_ms = int((now - days * 86400) * 1000) if days else 0 + snap = Snapshot(agent="opencode", source_path=path, generated_at=now, days=days) + try: + con = _connect(path) + except sqlite3.Error as e: + raise RuntimeError(f"cannot open opencode database: {e}") from None + with con: + sessions = { + r["id"]: r for r in con.execute( + "SELECT id, parent_id, directory, title, model, time_created, time_updated FROM session " + "WHERE time_updated >= ?", (since_ms,) + ) + } + if not sessions: + raise RuntimeError("opencode database has no sessions in the window — try --days 0.") + agg: dict = {} + priced = 0 + for r in con.execute( + "SELECT session_id, time_created, data FROM message WHERE time_created >= ? ORDER BY time_created", + (since_ms,), + ): + sid = r["session_id"] + if sid not in sessions: + continue + try: + d = json.loads(r["data"]) + except (TypeError, json.JSONDecodeError): + continue + if not isinstance(d, dict) or d.get("role") != "assistant": + continue + t = d.get("tokens") or {} + if not isinstance(t, dict): + continue + cache_ = t.get("cache") or {} + i_ = int(t.get("input") or 0) + o_ = int(t.get("output") or 0) + rs = int(t.get("reasoning") or 0) + cr = int(cache_.get("read") or 0) + cw = int(cache_.get("write") or 0) + ts = (r["time_created"] or 0) / 1000.0 + model = d.get("modelID") + provider = d.get("providerID") + cost = float(d.get("cost") or 0) + a = agg.setdefault(sid, {"calls": 0, "inp": 0, "out": 0, "cr": 0, "cw": 0, "rs": 0, "cost": 0.0, + "model": None, "provider": None, "first": None, "last": None}) + a["calls"] += 1 + a["inp"] += i_ + a["out"] += o_ + a["cr"] += cr + a["cw"] += cw + a["rs"] += rs + a["cost"] += cost + if cost > 0: + priced += 1 + a["model"] = f"{provider}/{model}" if provider and model and "/" not in str(model) else model + a["provider"] = provider + a["first"] = ts if a["first"] is None else min(a["first"], ts) + a["last"] = ts if a["last"] is None else max(a["last"], ts) + source = "subagent" if sessions[sid]["parent_id"] else "cli" + snap.usage_cells.append(UsageCell( + start=int(ts // BUCKET_SECONDS) * BUCKET_SECONDS, source=source, model=a["model"], + calls=1, input_tokens=i_, output_tokens=o_ + rs, cache_read_tokens=cr, + cache_write_tokens=cw, session=sid, + )) + snap.context_calls.append(ContextCall(ts=ts, session=sid, model=a["model"], + context=i_ + cr + cw, output=o_ + rs)) + for r in con.execute( + "SELECT session_id, time_created, data FROM part WHERE time_created >= ?", (since_ms,) + ): + if r["session_id"] not in sessions: + continue + try: + d = json.loads(r["data"]) + except (TypeError, json.JSONDecodeError): + continue + if not isinstance(d, dict) or d.get("type") != "tool": + continue + state = d.get("state") or {} + name = str(d.get("tool") or "tool")[:40] + ts = (r["time_created"] or 0) / 1000.0 + from .hermes import salient_arg + + snap.events.append(ActionEvent(session_id=r["session_id"], ts=ts, name=name, + arg_key=salient_arg(state.get("input")))) + status = str(state.get("status") or "") + out_text = state.get("output") + snap.events.append(ActionEvent( + session_id=r["session_id"], ts=ts, name=name, + ok=None if not status else status not in ("error", "failed"), + tokens=len(out_text) // CHARS_PER_TOKEN if isinstance(out_text, str) else None, + )) + for sid, s in sessions.items(): + a = agg.get(sid) + if a is None: + continue + cost_known = a["cost"] > 0 + snap.sessions.append(SessionRec( + id=sid, source="subagent" if s["parent_id"] else "cli", model=a["model"], + started_at=a["first"] or (s["time_created"] or 0) / 1000.0, ended_at=a["last"], + parent_id=s["parent_id"], title=(s["title"] or sid)[:80], api_calls=a["calls"], + input_tokens=a["inp"], output_tokens=a["out"], cache_read_tokens=a["cr"], + cache_write_tokens=a["cw"], reasoning_tokens=a["rs"], + cost_usd=a["cost"] if cost_known else None, cost_basis="actual" if cost_known else "unknown", + message_count=a["calls"], provider=a["provider"], project=s["directory"], + )) + if not snap.sessions: + raise RuntimeError("opencode database has sessions but no assistant messages in the window.") + unpriced = sum(1 for s in snap.sessions if s.cost_basis == "unknown") + if unpriced: + snap.warnings.append( + f"{unpriced} of {len(snap.sessions)} sessions carry no cost from opencode (free or " + "self-hosted provider) — shown as tokens only." + ) + return snap diff --git a/agentburn/cache.py b/agentburn/cache.py index 60e9f55..886a330 100644 --- a/agentburn/cache.py +++ b/agentburn/cache.py @@ -21,7 +21,8 @@ import os import time -SCHEMA = 1 +# Bump when the shape of a cached parse changes: an old entry is then a miss. +SCHEMA = 2 ENV_OFF = "AGENTBURN_NO_CACHE" # Entries whose source file disappeared are swept after this long, so a machine # that churns through projects doesn't grow an unbounded cache. diff --git a/agentburn/cli.py b/agentburn/cli.py index 4e01157..9fd72a0 100644 --- a/agentburn/cli.py +++ b/agentburn/cli.py @@ -56,7 +56,10 @@ def parse_when(s: str) -> float: agentburn why behavioral forensics: loops, retry storms, idle runs agentburn why --source telegram decompose ONE source: which functions it called, loops, errors agentburn limits subscription plans bill windows, not dollars: how fast you fill one - agentburn limits --hit "2026-08-20 14:30" calibrate against the window where you actually got cut off + agentburn limits --hit "2026-08-20 14:30" calibrate by hand (Claude Code's own cut-off records are used automatically) + agentburn context the price of long contexts: what a /clear at 150k would have saved, what a skill costs + agentburn commits what each commit cost you — sessions joined to your repositories' git log + agentburn statusline one line for Claude Code's statusLine: window fill %, time to wall agentburn drift your model spend × world usage trend — are you paying for a dying model? agentburn rank you vs the community Burn Index (efficiency percentiles) agentburn --submit join the index: anonymized payload + a link YOU click @@ -74,14 +77,18 @@ def parse_when(s: str) -> float: def build_parser() -> argparse.ArgumentParser: ap = argparse.ArgumentParser( prog="agentburn", - description="Where does your AI agent burn money? Local profiler, zero deps, nothing leaves your machine.", + description="Where does your AI agent burn money or usage? Claude Code, Codex, Gemini CLI, opencode, " + "OpenClaw, Hermes. Local profiler, zero deps, nothing leaves your machine.", epilog=RECIPES, formatter_class=argparse.RawDescriptionHelpFormatter, ) ap.add_argument("command", nargs="?", - choices=["report", "doctor", "why", "limits", "explain", "mcp", "fix", "drift", "rank"], + choices=["report", "doctor", "why", "limits", "context", "commits", "statusline", + "explain", "mcp", "fix", "drift", "rank"], default="report", help="report (default) · why (forensics) · limits (subscription windows) · " + "context (price of long contexts, skill costs) · commits (cost per commit) · " + "statusline (one line for an editor status bar) · " "drift (spend × world trend) · rank (you vs the Burn Index) · " "fix (config patches) · explain (LLM) · doctor (accounting health) · mcp") ap.add_argument("--agent", default=None, choices=sorted(ADAPTERS), @@ -148,6 +155,9 @@ def pick_agents(args) -> list: " looked for: ~/.hermes/state.db (Hermes Agent)\n" " ~/.openclaw/agents/*/sessions/sessions.json (OpenClaw)\n" " ~/.claude/projects/*.jsonl (Claude Code)\n" + " ~/.codex/sessions/**/rollout-*.jsonl (Codex CLI)\n" + " ~/.gemini/tmp/*/chats/session-*.json (Gemini CLI)\n" + " ~/.local/share/opencode/opencode.db (opencode)\n" " pass --agent --db if the data lives elsewhere.", file=sys.stderr, ) @@ -226,6 +236,10 @@ def main(argv=None) -> int: args.days = 1 elif args.week: args.days = 7 + elif args.command == "statusline" and args.days == 30: + # Runs on every turn of the editor: read the last few days only. The + # ceiling comes from the state file `limits` keeps, not from history. + args.days = 3 color = sys.stdout.isatty() and not args.no_color and _ansi_ok() single_modes = (args.command in ("doctor", "explain", "fix") @@ -243,6 +257,20 @@ def load(name): snap = filter_snapshot(snap, args.source) return snap + def load_all(names): + """One adapter failing (empty window, unreadable file) must not hide the others' reports.""" + loaded, errors = [], [] + for n in names: + try: + loaded.append((n, load(n))) + except (FileNotFoundError, RuntimeError) as e: + errors.append((n, e)) + if not loaded: + raise errors[0][1] + for n, e in errors: + print(f"agentburn: {n}: {e} — skipped", file=sys.stderr) + return loaded + try: if args.command == "doctor": from .doctor import render_doctor @@ -255,7 +283,7 @@ def load(name): from .burnindex import (INDEX_URL, build_metrics, load_index, rank_against, render_rank, submit_url) - snaps = [load(n) for n in found] + snaps = [sn for _, sn in load_all(found)] analyses = [analyze(s, night_window=args.night) for s in snaps] breps = [analyze_behavior(s) for s in snaps] from .limits import build_limits @@ -282,7 +310,7 @@ def load(name): if args.command == "drift": from .drift import TRENDS_URL, build_drift, load_trends, render_drift - analyses = [analyze(load(n), night_window=args.night) for n in found] + analyses = [analyze(sn, night_window=args.night) for _, sn in load_all(found)] try: trends = load_trends(args.trends or TRENDS_URL) except RuntimeError as e: @@ -349,15 +377,21 @@ def load(name): print() return 0 - if args.command == "limits": - from .limits import build_limits, limits_json, render_limits + if args.command in ("limits", "statusline"): + from .limits import (build_limits, limits_json, load_saved_ceiling, render_limits, + save_ceiling, statusline) - reports = [ - build_limits( - load(n), hit=args.hit, window_hours=args.window or 5.0 + reports = [] + for n, sn in load_all(found): + rep = build_limits( + sn, hit=args.hit, window_hours=args.window or 5.0, + saved=load_saved_ceiling(n), ) - for n in found - ] + save_ceiling(n, rep) + reports.append(rep) + if args.command == "statusline": + print(statusline(reports[0])) + return 0 if args.json: import json as _json @@ -369,13 +403,48 @@ def load(name): print(render_limits(r, color=color)) return 0 + if args.command == "context": + from .context import build_context, context_json, render_context + + reports = [build_context(sn) for _, sn in load_all(found)] + if args.json: + import json as _json + + payloads = [context_json(r) for r in reports] + print(_json.dumps(payloads[0] if len(payloads) == 1 else payloads, + indent=2, ensure_ascii=False)) + else: + for r in reports: + print(render_context(r, color=color)) + return 0 + + if args.command == "commits": + import time as _time + + from .commits import build_commits, commits_json, render_commits + + since = _time.time() - args.days * 86400 if args.days else None + reports = [build_commits(sn, since=since) for _, sn in load_all(found)] + if args.json: + import json as _json + + payloads = [commits_json(r) for r in reports] + print(_json.dumps(payloads[0] if len(payloads) == 1 else payloads, + indent=2, ensure_ascii=False)) + else: + for r in reports: + print(render_commits(r, color=color)) + return 0 + if args.command == "why": from .behavior import analyze_behavior, behavior_json, render_behavior - if len(found) > 1 and not args.json: - print(("\033[2m" if color else "") + f"Found {len(found)} agents: {', '.join(found)} — " + loaded = load_all(found) + if len(loaded) > 1 and not args.json: + names = [n for n, _ in loaded] + print(("\033[2m" if color else "") + f"Found {len(names)} agents: {', '.join(names)} — " "one forensics report each." + ("\033[0m" if color else "")) - reports = [analyze_behavior(load(n)) for n in found] + reports = [analyze_behavior(sn) for _, sn in loaded] if args.json: import json as _json @@ -387,7 +456,9 @@ def load(name): print(render_behavior(r, color=color)) return 0 - snaps = [load(n) for n in found] + loaded = load_all(found) + found = [n for n, _ in loaded] + snaps = [sn for _, sn in loaded] analyses = [analyze(sn, night_window=args.night) for sn in snaps] except (FileNotFoundError, RuntimeError) as e: print(f"agentburn: {e}", file=sys.stderr) @@ -469,6 +540,10 @@ def _next_hints(args, color: bool, subscription: bool = False) -> None: 0, "agentburn limits → how fast you fill a 5-hour window (what a subscription actually bills)", ) + hints.insert( + 1, + "agentburn context → what long contexts cost, what a /clear at 150k would have saved", + ) if not os.path.exists(args.baseline_file or baseline.DEFAULT_PATH): hints.append("agentburn --save-baseline → snapshot now, prove your savings after you optimize") else: diff --git a/agentburn/commits.py b/agentburn/commits.py new file mode 100644 index 0000000..3421d2c --- /dev/null +++ b/agentburn/commits.py @@ -0,0 +1,238 @@ +"""`agentburn commits` — what a commit cost, in your own window. + +Claude Code records the working directory and git branch of every session. +Your repositories record when each commit landed. Joining the two attributes +every call to the commit that followed it in that repository: the weighted +usage between two consecutive commits is what the second one cost. + +Read-only on both sides: `git log` in the repositories the transcripts name, +nothing written, nothing sent. Sessions whose directory is not a git +repository (or has no commits in the window) are reported as such. +""" + +from __future__ import annotations + +import os +import statistics +import subprocess +import time +from collections import defaultdict +from dataclasses import dataclass, field +from typing import Optional + +from .limits import cell_weight +from .model import Snapshot, agent_label +from .report import fmt_tokens + + +@dataclass +class CommitCost: + repo: str + sha: str + subject: str + ts: float + weight: float + calls: int + branch: Optional[str] = None + + +@dataclass +class RepoStat: + repo: str + commits: int + median: float + total: float + uncommitted: float # usage after the last commit in the window + + +@dataclass +class CommitsReport: + agent: str + top: list = field(default_factory=list) # CommitCost, costliest first + repos: list = field(default_factory=list) # RepoStat, by total + skipped: list = field(default_factory=list) # (project, reason) + total_attributed: float = 0.0 + total_weight: float = 0.0 + unsupported: str = "" + + +def _git(args: list, cwd: str, timeout: float = 10) -> Optional[str]: + try: + r = subprocess.run( + ["git"] + args, cwd=cwd, capture_output=True, text=True, encoding="utf-8", + errors="replace", timeout=timeout, + ) + except (OSError, subprocess.SubprocessError): + return None + return r.stdout if r.returncode == 0 else None + + +def _repo_root(cwd: str) -> Optional[str]: + out = _git(["rev-parse", "--show-toplevel"], cwd) + return out.strip() if out else None + + +def _commits(root: str, since: Optional[float]) -> list: + """[(ts, sha, subject)] across all refs, oldest first.""" + args = ["log", "--all", "--no-merges", "--format=%H%x1f%ct%x1f%s"] + if since: + args.append(f"--since={int(since)}") + out = _git(args, root, timeout=30) + if not out: + return [] + rows = [] + for line in out.splitlines(): + parts = line.split("\x1f", 2) + if len(parts) == 3 and parts[1].isdigit(): + rows.append((float(parts[1]), parts[0], parts[2])) + rows.sort() + return rows + + +def build_commits(snap: Snapshot, since: Optional[float] = None, git_ok: bool = True) -> CommitsReport: + rep = CommitsReport(agent=snap.agent) + if not snap.usage_cells or not any(s.project for s in snap.sessions): + rep.unsupported = ( + f"{agent_label(snap.agent)}'s adapter does not record working directories per " + "session, so calls can't be joined to commits." + ) + return rep + if git_ok and _git(["--version"], os.getcwd()) is None: + rep.unsupported = "git is not on PATH; `agentburn commits` needs it to read your repositories." + return rep + + project_of = {s.id: s.project for s in snap.sessions} + branch_of = {s.id: s.branch for s in snap.sessions} + for s in snap.sessions: + if s.parent_id and not project_of.get(s.id): + project_of[s.id] = project_of.get(s.parent_id) + branch_of[s.id] = branch_of.get(s.parent_id) + + # cwd → repo root (one git call per distinct directory) + roots: dict = {} + for cwd in {p for p in project_of.values() if p}: + if not os.path.isdir(cwd): + rep.skipped.append((cwd, "directory no longer exists")) + roots[cwd] = None + continue + roots[cwd] = _repo_root(cwd) + if roots[cwd] is None: + rep.skipped.append((cwd, "not a git repository")) + + # repo → [(ts, weight, calls, branch)] + usage: dict = defaultdict(list) + for c in snap.usage_cells: + w = cell_weight(c) + rep.total_weight += w + if w <= 0 or not c.session: + continue + root = roots.get(project_of.get(c.session) or "") + if not root: + continue + usage[root].append((c.start, w, c.calls, branch_of.get(c.session))) + + for root, cells in usage.items(): + commits = _commits(root, since) + if not commits: + rep.skipped.append((root, "no commits in the window")) + continue + cells.sort() + costs = {sha: [0.0, 0, None] for _, sha, _ in commits} + uncommitted = 0.0 + i = 0 + for start, w, n, br in cells: + while i < len(commits) and commits[i][0] < start: + i += 1 + if i >= len(commits): + uncommitted += w + continue + entry = costs[commits[i][1]] + entry[0] += w + entry[1] += n + entry[2] = entry[2] or br + # A commit with nothing between it and the previous one (a rebase, a + # merge of someone else's work) costs nothing and is not listed. + priced = [] + for ts, sha, subj in commits: + w, n, br = costs[sha] + if w > 0: + cc = CommitCost(repo=os.path.basename(root), sha=sha[:8], subject=subj[:60], ts=ts, + weight=w, calls=n, branch=br) + priced.append(cc) + rep.top.append(cc) + if priced: + total = sum(c.weight for c in priced) + rep.total_attributed += total + rep.repos.append(RepoStat( + repo=os.path.basename(root), commits=len(priced), + median=statistics.median(c.weight for c in priced), total=total, + uncommitted=uncommitted, + )) + else: + rep.skipped.append((root, "usage recorded, but after the last commit")) + rep.top.sort(key=lambda c: -c.weight) + rep.top = rep.top[:8] + rep.repos.sort(key=lambda r: -r.total) + return rep + + +def render_commits(rep: CommitsReport, color: bool = True) -> str: + from .report import P + + p = P(color) + out = ["", p.b(f"🧾 agentburn commits — {rep.agent} · what a commit cost you")] + out.append(p.dim(" usage between two consecutive commits in a repository is what the second one cost")) + out.append("") + if rep.unsupported: + out.append(f" {rep.unsupported}") + out.append("") + return "\n".join(out) + if not rep.top: + out.append(" No commits could be priced in this window.") + for proj, why in rep.skipped[:6]: + out.append(p.dim(f" {proj}: {why}")) + out.append("") + return "\n".join(out) + + out.append(p.b(" COSTLIEST COMMITS")) + for c in rep.top: + when = time.strftime("%b %d", time.localtime(c.ts)) + out.append( + f" {fmt_tokens(c.weight):>8} {c.repo[:18]:<18} {c.sha} {when} {c.subject}" + ) + out.append("") + out.append(p.b(" BY REPOSITORY")) + out.append(p.dim(" median cost of a commit · commits · usage after the last commit")) + for r in rep.repos[:8]: + out.append( + f" {r.repo[:24]:<24} {fmt_tokens(r.median):>8} median · {r.commits:>4} commits · " + f"{fmt_tokens(r.total):>8} total" + (f" · {fmt_tokens(r.uncommitted)} uncommitted" if r.uncommitted else "") + ) + out.append("") + if rep.total_weight: + out.append( + p.dim(f" {rep.total_attributed / rep.total_weight:.0%} of weighted usage landed in a priced commit; " + f"{len(rep.skipped)} director{'y' if len(rep.skipped) == 1 else 'ies'} skipped.") + ) + out.append(p.dim(" Read-only: `git log` in the repositories your sessions name. Weighted tokens, as in `limits`.")) + out.append("") + return "\n".join(out) + + +def commits_json(rep: CommitsReport) -> dict: + return { + "agent": rep.agent, + "unsupported": rep.unsupported or None, + "top": [ + {"repo": c.repo, "sha": c.sha, "subject": c.subject, "ts": c.ts, "weight": round(c.weight), + "calls": c.calls, "branch": c.branch} + for c in rep.top + ], + "repos": [ + {"repo": r.repo, "commits": r.commits, "median": round(r.median), "total": round(r.total), + "uncommitted": round(r.uncommitted)} + for r in rep.repos + ], + "skipped": [{"path": p_, "why": w} for p_, w in rep.skipped], + "attributed_share": round(rep.total_attributed / rep.total_weight, 4) if rep.total_weight else None, + } diff --git a/agentburn/context.py b/agentburn/context.py new file mode 100644 index 0000000..a469315 --- /dev/null +++ b/agentburn/context.py @@ -0,0 +1,249 @@ +"""`agentburn context` — the price of a long context, and what a `/clear` buys. + +On a subscription the window is filled by what the model re-reads, and every +call re-reads the whole context: a turn at 300k costs the window as much as +three turns at 100k. Claude Code records the exact size of every call's +context (uncached input + cache reads + cache writes), so this is measured, +not modelled. + +Two questions, both answered from the same per-call records: +- how much of the window went to calls whose context was already past N — + the share of usage that is "long-session tax"; +- if every session had been restarted at a threshold T, how much of the + weighted window volume would not have been spent — the honest saving of + a `/clear` habit, assuming the same amount of work in shorter sessions. + +Skill costs are measured the same way: the growth of the context between the +reply that invoked a skill and the next one, when the skill was the only +tool call in that reply. Bundled skills never touch the disk, so reading the +file would miss them; the transcript sees all of them. +""" + +from __future__ import annotations + +import statistics +import time +from collections import defaultdict +from dataclasses import dataclass, field +from typing import Optional + +from .limits import token_weights +from .model import Snapshot, agent_label +from .report import fmt_tokens + +BANDS = ((0, "<50k"), (50_000, "50–100k"), (100_000, "100–200k"), (200_000, "200–400k"), (400_000, ">400k")) +THRESHOLDS = (100_000, 150_000, 200_000, 300_000) +# Median of the last K observations per skill: skills change size, a maximum +# would keep one old outlier forever, a global median lags behind a diet. +SKILL_KEEP = 5 + + +@dataclass +class Band: + label: str + calls: int = 0 + context: int = 0 # raw tokens + weight: float = 0.0 # weighted (price-ratio) tokens + + +@dataclass +class Saving: + threshold: int + share: float # weighted volume that would not have been spent + calls_over: int + + +@dataclass +class SkillCost: + skill: str + calls: int + tokens: int # median of the last SKILL_KEEP measurements + total: int # tokens × calls in this window + + +@dataclass +class ContextReport: + agent: str + calls: int = 0 + bands: list = field(default_factory=list) # Band + median_context: int = 0 + p90_context: int = 0 + max_context: int = 0 + savings: list = field(default_factory=list) # Saving + by_effort: list = field(default_factory=list) # [(effort, calls, share_of_weight)] + longest: list = field(default_factory=list) # [(session, max_context, calls)] + skills: list = field(default_factory=list) # SkillCost + total_weight: float = 0.0 + unsupported: str = "" + + +def _weight(model: Optional[str], context: int, output: int) -> float: + # A context call doesn't say which part was cache read vs uncached, so the + # whole context is weighted at the cache-read rate: the smallest ratio, + # i.e. a lower bound. Output is weighted at its real price. + w_in, w_out, w_cr, _ = token_weights(model) + return context * w_cr + output * w_out + + +def build_context(snap: Snapshot) -> ContextReport: + rep = ContextReport(agent=snap.agent) + if not snap.context_calls: + rep.unsupported = ( + f"{agent_label(snap.agent)}'s adapter does not record per-call context sizes, " + "so the price of long sessions can't be measured here." + ) + return rep + + bands = {label: Band(label) for _, label in BANDS} + sizes = [] + total_w = 0.0 + over = {t: [0, 0.0] for t in THRESHOLDS} # calls over, weight over + effort_w: dict = defaultdict(lambda: [0, 0.0]) + per_session: dict = defaultdict(lambda: [0, 0]) # max ctx, calls + for c in snap.context_calls: + w = _weight(c.model, c.context, c.output) + total_w += w + sizes.append(c.context) + label = BANDS[0][1] + for lo, lab in BANDS: + if c.context >= lo: + label = lab + b = bands[label] + b.calls += 1 + b.context += c.context + b.weight += w + for t in THRESHOLDS: + if c.context > t: + over[t][0] += 1 + # What a restart at t would have saved on this call: the part of + # the context above t, at the cache-read rate. + over[t][1] += (c.context - t) * token_weights(c.model)[2] + e = effort_w[c.effort or "default"] + e[0] += 1 + e[1] += w + ps = per_session[c.session] + ps[0] = max(ps[0], c.context) + ps[1] += 1 + + rep.calls = len(sizes) + rep.total_weight = total_w + rep.bands = [bands[lab] for _, lab in BANDS if bands[lab].calls] + sizes.sort() + rep.median_context = int(statistics.median(sizes)) + rep.p90_context = sizes[min(len(sizes) - 1, int(len(sizes) * 0.9))] + rep.max_context = sizes[-1] + if total_w > 0: + rep.savings = [ + Saving(threshold=t, share=over[t][1] / total_w, calls_over=over[t][0]) for t in THRESHOLDS + ] + rep.by_effort = sorted( + ((e, n, w / total_w) for e, (n, w) in effort_w.items()), key=lambda kv: -kv[2] + ) + titles = {s.id: s.title or s.id for s in snap.sessions} + rep.longest = sorted( + ((titles.get(sid, sid), mx, n) for sid, (mx, n) in per_session.items()), + key=lambda r: -r[1], + )[:5] + + by_skill: dict = defaultdict(list) + for sl in snap.skill_loads: + by_skill[sl.skill].append((sl.ts or 0, sl.tokens)) + costs = [] + for skill, obs in by_skill.items(): + obs.sort() + recent = [t for _, t in obs[-SKILL_KEEP:]] + med = int(statistics.median(recent)) + costs.append(SkillCost(skill=skill, calls=len(obs), tokens=med, total=med * len(obs))) + rep.skills = sorted(costs, key=lambda c: -c.total)[:8] + return rep + + +def render_context(rep: ContextReport, color: bool = True) -> str: + from .report import P + + p = P(color) + out = ["", p.b(f"📏 agentburn context — {rep.agent} · what a long context costs")] + out.append(p.dim(" every call re-reads its whole context; on a subscription that IS the window")) + out.append("") + if rep.unsupported: + out.append(f" {rep.unsupported}") + out.append("") + return "\n".join(out) + if not rep.calls: + out.append(" Nothing recorded in this window — try `--days 0`.") + out.append("") + return "\n".join(out) + + out.append( + f" {'CALLS':<25} {rep.calls:>10,} " + + p.dim(f"median context {fmt_tokens(rep.median_context)} · p90 {fmt_tokens(rep.p90_context)} · max {fmt_tokens(rep.max_context)}") + ) + out.append("") + out.append(p.b(" WHERE THE WINDOW GOES, BY CONTEXT SIZE")) + out.append(p.dim(" share of weighted volume · calls")) + for b in rep.bands: + share = b.weight / rep.total_weight if rep.total_weight else 0 + bar = "█" * max(0, min(18, round(share * 18))) + "·" * (18 - max(0, min(18, round(share * 18)))) + line = f" {b.label:<12} {bar} {share:>5.0%} {b.calls:>7,} calls" + out.append(p.red(line) if b.label in (">400k",) and share >= 0.1 else line) + out.append("") + + if rep.savings: + out.append(p.b(" IF YOU HAD RESTARTED AT…")) + out.append(p.dim(" the part of every call's context above the threshold, at the cache-read rate")) + for sv in rep.savings: + line = f" /clear at {fmt_tokens(sv.threshold):<8} → {sv.share:>5.0%} of the window not spent ({sv.calls_over:,} calls were past it)" + out.append(p.green(line) if sv.share >= 0.25 else line) + out.append(p.dim(" assumes the same work done in shorter sessions; a restart itself costs one bootstrap")) + out.append("") + + if rep.longest: + out.append(p.b(" LONGEST SESSIONS")) + for title, mx, n in rep.longest: + out.append(f" {title[:40]:<40} {fmt_tokens(mx):>8} max context · {n:,} calls") + out.append("") + + if rep.by_effort and len(rep.by_effort) > 1: + out.append(p.b(" BY EFFORT LEVEL")) + for e, n, share in rep.by_effort: + out.append(f" {e:<12} {share:>5.0%} of weighted volume · {n:,} calls") + out.append("") + + if rep.skills: + out.append(p.b(" WHAT A SKILL COSTS")) + out.append(p.dim(" measured: context growth right after a lone Skill call, median of recent loads")) + for sc in rep.skills: + out.append(f" {sc.skill[:36]:<36} {fmt_tokens(sc.tokens):>8} per load × {sc.calls:>4} = {fmt_tokens(sc.total):>8}") + out.append("") + + out.append(p.dim(" context = uncached input + cache reads + cache writes of each call, as Claude Code recorded it.")) + out.append(p.dim(" Shares are weighted by published price ratios; the context part at the cache-read rate (a lower bound).")) + out.append("") + return "\n".join(out) + + +def context_json(rep: ContextReport) -> dict: + return { + "agent": rep.agent, + "unsupported": rep.unsupported or None, + "calls": rep.calls, + "median_context": rep.median_context, + "p90_context": rep.p90_context, + "max_context": rep.max_context, + "bands": [ + {"band": b.label, "calls": b.calls, "context_tokens": b.context, + "share": round(b.weight / rep.total_weight, 4) if rep.total_weight else 0} + for b in rep.bands + ], + "savings": [ + {"clear_at": s.threshold, "share_not_spent": round(s.share, 4), "calls_over": s.calls_over} + for s in rep.savings + ], + "by_effort": [{"effort": e, "calls": n, "share": round(sh, 4)} for e, n, sh in rep.by_effort], + "longest_sessions": [{"session": t, "max_context": mx, "calls": n} for t, mx, n in rep.longest], + "skills": [ + {"skill": s.skill, "loads": s.calls, "tokens_per_load": s.tokens, "total": s.total} + for s in rep.skills + ], + "generated_at": time.time(), + } diff --git a/agentburn/doctor.py b/agentburn/doctor.py index a1367c7..a1c795f 100644 --- a/agentburn/doctor.py +++ b/agentburn/doctor.py @@ -12,7 +12,7 @@ import time from collections import Counter -from .model import Snapshot, agent_key, agent_label, agent_store +from .model import Snapshot, agent_key, agent_label, agent_store, records_costs # Upstream issue describing zero-usage streams, per agent. Only cite it where it # actually applies — pointing a Claude Code user at the Hermes tracker sends the @@ -27,12 +27,13 @@ def diagnose(snap: Snapshot) -> dict: zero_groups = Counter() unpriced_groups = Counter() zero_total = unpriced_total = 0 + priced_agent = records_costs(snap.agent) for s in snap.sessions: key = (s.provider or "unknown-provider", s.model or "unknown-model", s.source) if s.message_count > 0 and s.total_tokens == 0: zero_groups[key] += 1 zero_total += 1 - if s.total_tokens > 0 and s.cost_usd is None: + if priced_agent and s.total_tokens > 0 and s.cost_usd is None: unpriced_groups[key] += 1 unpriced_total += 1 return { @@ -41,6 +42,7 @@ def diagnose(snap: Snapshot) -> dict: "zero_groups": zero_groups.most_common(8), "unpriced_total": unpriced_total, "unpriced_groups": unpriced_groups.most_common(8), + "no_local_costs": not priced_agent, } @@ -57,9 +59,12 @@ def render_doctor(snap: Snapshot, color: bool = True) -> str: out.append( f" zero-usage sessions: {d['zero_total']} (messages exist, tokens recorded = 0)" ) - out.append( - f" unpriced sessions : {d['unpriced_total']} (tokens exist, no cost recorded)" - ) + if d["no_local_costs"]: + out.append(f" unpriced sessions : n/a ({agent_label(snap.agent)} records no prices — tokens only, by design)") + else: + out.append( + f" unpriced sessions : {d['unpriced_total']} (tokens exist, no cost recorded)" + ) out.append("") if d["zero_total"] == 0 and d["unpriced_total"] == 0: diff --git a/agentburn/fix.py b/agentburn/fix.py index fd2d4b1..3eccd41 100644 --- a/agentburn/fix.py +++ b/agentburn/fix.py @@ -11,8 +11,10 @@ - OpenClaw: `agents.defaults.heartbeat.{every, activeHours, model, lightContext}` in ~/.openclaw/openclaw.json (config/types.agent-defaults.ts). - Claude Code: registered MCP servers in ~/.claude.json / .mcp.json (documented - scopes; `claude mcp list|remove`) and the always-loaded CLAUDE.md memory - files. Both are levers the user owns; neither is a guess about pricing. + scopes; `claude mcp list|remove`), the always-loaded CLAUDE.md memory + files, skills whose measured load cost dominates (skill files the user + owns), and the session length itself (`/clear` is a documented command). + All levers the user owns; none is a guess about pricing. Findings with no verified lever stay recommendations, not patches. """ @@ -171,6 +173,63 @@ def _claude_code_fixes(a: Analysis, snap) -> list: ], ) ) + if snap is not None and getattr(snap, "context_calls", None): + from .context import build_context + + ctx = build_context(snap) + best = max(ctx.savings, key=lambda sv: sv.share) if ctx.savings else None + if best and best.share >= 0.15: + # The lowest threshold that keeps most of the best saving: cheaper + # to follow than the one with the absolute maximum. + pick = next((sv for sv in ctx.savings if sv.share >= best.share * 0.8), best) + patches.append( + Patch( + title=f"Restart sessions at ~{fmt_tokens(pick.threshold)} context (/clear)", + target="(a habit, not a file — `/clear`, or a handoff note + new session)", + target_exists=True, + why=( + f"{pick.calls_over:,} calls in this window ran past {fmt_tokens(pick.threshold)} " + f"of context; every call re-reads all of it. Median context {fmt_tokens(ctx.median_context)}, " + f"p90 {fmt_tokens(ctx.p90_context)}." + ), + impact=( + f"≈{pick.share:.0%} of the weighted window not spent, assuming the same work " + "in shorter sessions (see `agentburn context`)" + ), + proposed=( + f"when the context passes ~{fmt_tokens(pick.threshold)}: write a short handoff, /clear, continue.\n" + "long tasks → a plan file + one session per stage." + ), + notes=[ + "measured from Claude Code's own per-call usage, not modelled", + "a restart costs one bootstrap (memory files + tool definitions); the estimate ignores that", + ], + ) + ) + heavy = [sc for sc in ctx.skills if sc.tokens >= 8_000 and sc.calls >= 2] + if heavy: + total = sum(sc.total for sc in heavy) + patches.append( + Patch( + title=f"Put {len(heavy)} heavy skill(s) on a diet ({fmt_tokens(total)} loaded in this window)", + target=os.path.join(os.path.expanduser("~"), ".claude", "skills"), + target_exists=os.path.isdir(os.path.join(os.path.expanduser("~"), ".claude", "skills")), + why=( + "each load adds this much context, measured as the growth right after the call: " + + ", ".join(f"{sc.skill} {fmt_tokens(sc.tokens)}×{sc.calls}" for sc in heavy[:5]) + ), + impact="every trimmed line is paid back on every load, for the rest of the session too", + current="\n".join(f"{sc.skill:<36} {sc.tokens:>8,} tokens per load" for sc in heavy[:5]), + proposed=( + "keep the instructions that change what the agent does; move examples, catalogues " + "and reference tables into files the skill reads on demand" + ), + notes=[ + "bundled skills (not on disk) are measured the same way but can only be avoided, not trimmed", + "prove it: agentburn context → trim → agentburn context", + ], + ) + ) return patches diff --git a/agentburn/limits.py b/agentburn/limits.py index d3649af..949f23a 100644 --- a/agentburn/limits.py +++ b/agentburn/limits.py @@ -11,10 +11,16 @@ statement about how much more one token costs than another — and sums them over rolling windows. That makes YOUR windows comparable with each other: the peak, the typical one, the one running right now. -2. If you tell it when you were actually cut off (`--hit "2026-08-20 14:30"`), - the window that ended at that moment becomes a *measured* ceiling, and - everything else is reported as a percentage of it. That number is yours, - from your own wall, not a guess about the provider's arithmetic. +2. When you were actually cut off — Claude Code writes that moment into the + transcript itself ("You've hit your session limit · resets 8:30pm"), and + `--hit "2026-08-20 14:30"` lets you name one by hand — the window that + ended at that moment becomes a *measured* ceiling, and everything else is + reported as a percentage of it. That number is yours, from your own wall, + not a guess about the provider's arithmetic. With several recorded + cut-offs the ceiling is their median, so one odd window doesn't set it. +3. From the ceiling and the pace of the last half hour: how long until the + wall at this pace. That is the number you want while working, so it is + also what `agentburn statusline` prints. Unit: "weighted tokens" = tokens × price ratio, normalized so that one uncached input token of the reference model = 1. @@ -22,6 +28,8 @@ from __future__ import annotations +import json +import os import time from dataclasses import dataclass, field from typing import Optional @@ -41,6 +49,11 @@ DEFAULT_WINDOW_HOURS = 5.0 WEEK_SECONDS = 7 * 86400 +# Pace for "time to wall" is measured over this much recent usage. +PACE_SECONDS = 30 * 60 +# A measured ceiling outlives the run that found it, so that a short, fast +# `statusline` call (a few days of logs) still knows where the wall is. +STATE_PATH = os.path.join(os.path.expanduser("~"), ".agentburn", "ceiling.json") @dataclass @@ -66,7 +79,15 @@ class LimitsReport: mix: list = field(default_factory=list) # [(label, share)] over the period ceiling: Optional[float] = None ceiling_at: Optional[float] = None + ceiling_source: str = "" # "recorded" (agent wrote the cut-off) | "--hit" | "saved" + ceiling_hits: int = 0 # how many recorded cut-offs the ceiling is the median of slots_over_ceiling: int = 0 + pace: float = 0.0 # weighted tokens per second over the last PACE_SECONDS + minutes_to_wall: Optional[float] = None # at that pace; None = no ceiling or no pace + week_peak: Optional[Window] = None # heaviest rolling 7-day span + week_ceiling: Optional[float] = None # measured from a recorded weekly cut-off + provider_used: list = field(default_factory=list) # [(window_minutes, used_percent, ts)] latest reading + peak_by_project: list = field(default_factory=list) # [(project, share)] unsupported: str = "" # non-empty when the adapter can't answer this notes: list = field(default_factory=list) @@ -130,11 +151,57 @@ def _median(xs: list) -> float: return xs[mid] if len(xs) % 2 else (xs[mid - 1] + xs[mid]) / 2 +def _window_weight(series: dict, end: float, span: float, start: Optional[float] = None) -> float: + """Weight recorded in [start, end) — start defaults to end - span.""" + lo = start if start is not None else end - span + return sum(w for t, w in series.items() if lo <= t < end) + + +def load_saved_ceiling(agent: str, path: str = STATE_PATH) -> Optional[dict]: + try: + with open(path, "r", encoding="utf-8") as f: + data = json.load(f) + except (OSError, json.JSONDecodeError): + return None + entry = data.get(agent) if isinstance(data, dict) else None + return entry if isinstance(entry, dict) and entry.get("ceiling") else None + + +def save_ceiling(agent: str, rep: "LimitsReport", path: str = STATE_PATH) -> None: + """Remember a measured ceiling. Never fatal — it's a convenience for statusline.""" + if not rep.ceiling or rep.ceiling_source == "saved": + return + try: + data = {} + try: + with open(path, "r", encoding="utf-8") as f: + data = json.load(f) or {} + except (OSError, json.JSONDecodeError): + data = {} + if not isinstance(data, dict): + data = {} + data[agent] = { + "ceiling": round(rep.ceiling), + "ceiling_at": rep.ceiling_at, + "source": rep.ceiling_source, + "hits": rep.ceiling_hits, + "window_hours": rep.window_hours, + "week_ceiling": round(rep.week_ceiling) if rep.week_ceiling else None, + "saved_at": time.time(), + } + os.makedirs(os.path.dirname(path), mode=0o700, exist_ok=True) + with open(path, "w", encoding="utf-8") as f: + json.dump(data, f, indent=1) + except OSError: + pass + + def build_limits( snap: Snapshot, hit: Optional[float] = None, window_hours: float = DEFAULT_WINDOW_HOURS, now: Optional[float] = None, + saved: Optional[dict] = None, ) -> LimitsReport: now = now or snap.generated_at or time.time() span = int(window_hours * 3600) @@ -182,6 +249,29 @@ def build_limits( in_peak_source[c.source] = in_peak_source.get(c.source, 0.0) + w rep.peak_by_model = [kv for kv in _shares(in_peak_model)[:4] if kv[1] >= 0.005] rep.peak_by_source = [kv for kv in _shares(in_peak_source)[:4] if kv[1] >= 0.005] + project_of = {} + for s_ in snap.sessions: + root = s_.project + if s_.parent_id and not root: + root = None # filled below from the parent + project_of[s_.id] = root + for s_ in snap.sessions: + if s_.parent_id and not project_of.get(s_.id): + project_of[s_.id] = project_of.get(s_.parent_id) + in_peak_project: dict = {} + for c in snap.usage_cells: + if rep.peak.start <= c.start < rep.peak.end and c.session: + proj = project_of.get(c.session) + if proj: + name = os.path.basename(proj.rstrip("/\\")) or proj + in_peak_project[name] = in_peak_project.get(name, 0.0) + cell_weight(c) + if in_peak_project: + rep.peak_by_project = [kv for kv in _shares(in_peak_project)[:4] if kv[1] >= 0.005] + + week_rolling = _rolling(series, WEEK_SECONDS) + if week_rolling: + end, weight = max(week_rolling, key=lambda kv: kv[1]) + rep.week_peak = Window(start=end - WEEK_SECONDS, end=end, weight=weight) # Non-overlapping slots for "typical": a rolling maximum is by definition # unusual, and 60 overlapping views of the same busy hour would drag the @@ -203,16 +293,97 @@ def build_limits( if days: rep.busiest_day = max(days.items(), key=lambda kv: kv[1]) + # Ceiling, in order of trust: the cut-off you name by hand → cut-offs the + # agent recorded itself → one saved by an earlier run. if hit: rep.ceiling_at = hit - rep.ceiling = sum(w for t, w in series.items() if hit - span <= t < hit) - if rep.ceiling > 0: - rep.slots_over_ceiling = sum(1 for w in active if w >= rep.ceiling) - else: + rep.ceiling_source = "--hit" + rep.ceiling = _window_weight(series, hit, span) + if rep.ceiling <= 0: + rep.ceiling = None rep.notes.append( - "no usage recorded in the 5 hours before the timestamp you passed — " + f"no usage recorded in the {window_hours:g} hours before the timestamp you passed — " "check the date, or whether that window is inside --days." ) + if not rep.ceiling: + session_hits = [h for h in snap.limit_hits if h.kind not in ("weekly", "week")] + measured = [] + for h in session_hits: + # When the message said when the window resets, the window is the + # one ending there; otherwise the rolling span before the cut-off. + start = h.reset_at - span if h.reset_at and h.reset_at - span <= h.ts else None + w = _window_weight(series, h.ts, span, start) + if w > 0: + measured.append((w, h.ts)) + if measured: + measured.sort() + mid = measured[len(measured) // 2] + rep.ceiling = _median([w for w, _ in measured]) + rep.ceiling_at = mid[1] + rep.ceiling_source = "recorded" + rep.ceiling_hits = len(measured) + if not rep.ceiling and snap.rate_limits: + # The provider's own percentage next to our weighted usage of the same + # window: ceiling = weight / used_percent. One sample is noisy (the + # window may have started before our logs); the median of many is not. + est = [] + est_week = [] + latest: dict = {} + last_span_reading = None + for r in sorted(snap.rate_limits, key=lambda r: r.ts): + latest[r.window_minutes] = (r.used_percent, r.ts) + same_span = abs(r.window_minutes * 60 - span) <= BUCKET_SECONDS + if same_span: + last_span_reading = r.ts + if r.used_percent < 10: + continue + w = _window_weight(series, r.ts, r.window_minutes * 60) + if w <= 0: + continue + if same_span: + est.append((w / r.used_percent * 100, r.ts)) + elif abs(r.window_minutes * 60 - WEEK_SECONDS) <= BUCKET_SECONDS: + # exactly the week: a 30-day reading is not a weekly ceiling + est_week.append(w / r.used_percent * 100) + rep.provider_used = [(wm, used, ts) for wm, (used, ts) in sorted(latest.items())] + if est: + est.sort() + rep.ceiling = _median([w for w, _ in est]) + rep.ceiling_at = est[len(est) // 2][1] + rep.ceiling_source = "provider" + rep.ceiling_hits = len(est) + # The provider may stop reporting this window (plan change, client + # update): a peak that fell after the last reading was never + # measured against this ceiling, and the ratio would be a guess. + if rep.peak and last_span_reading is not None and rep.peak.end > last_span_reading + span: + rep.notes.append( + f"the peak window fell after the provider's last {window_hours:g}h reading " + f"({_stamp(last_span_reading)}) — the ceiling is measured on earlier windows only; " + "the % above is a comparison across periods, not a measured overrun." + ) + if est_week and not rep.week_ceiling: + rep.week_ceiling = _median(est_week) + if not rep.ceiling and saved: + rep.ceiling = float(saved["ceiling"]) + rep.ceiling_at = saved.get("ceiling_at") + rep.ceiling_source = "saved" + rep.ceiling_hits = int(saved.get("hits") or 0) + if saved.get("week_ceiling"): + rep.week_ceiling = float(saved["week_ceiling"]) + if rep.ceiling: + rep.slots_over_ceiling = sum(1 for w in active if w >= rep.ceiling) + + weekly_hits = [h for h in snap.limit_hits if h.kind in ("weekly", "week")] + wk = [_window_weight(series, h.ts, WEEK_SECONDS) for h in weekly_hits] + wk = [w for w in wk if w > 0] + if wk: + rep.week_ceiling = _median(wk) + + # Pace and time to wall: recent half hour, straight-line. + rep.pace = _window_weight(series, now, PACE_SECONDS) / PACE_SECONDS + if rep.ceiling and rep.pace > 0: + remaining = rep.ceiling - rep.current + rep.minutes_to_wall = max(0.0, remaining / rep.pace / 60.0) return rep @@ -294,15 +465,41 @@ def render_limits(rep: LimitsReport, color: bool = True) -> str: f" {'LAST 7 DAYS':<25} {_fmt(rep.week_total):>10} " + p.dim("weekly caps count this") ) + if rep.week_peak and rep.week_peak.weight > 0: + wk_bits = f"{_stamp(rep.week_peak.start)}–{_stamp(rep.week_peak.end)}" + if rep.week_total > 0: + wk_bits += f" · this week is {rep.week_total / rep.week_peak.weight:.0%} of it" + out.append( + f" {'PEAK 7 DAYS':<25} {_fmt(rep.week_peak.weight):>10} " + p.dim(wk_bits) + ) + if rep.peak_by_project: + out.append( + " " + + f"{'PEAK BY PROJECT':<25} " + + p.dim(" · ".join(f"{n} {v:.0%}" for n, v in rep.peak_by_project[:3])) + ) out.append("") if rep.ceiling: out.append(p.b(" YOUR MEASURED CEILING")) - out.append( - p.dim( - f" the window that ended {_stamp(rep.ceiling_at)}, when you say you were cut off" + if rep.ceiling_source == "recorded": + how = ( + f"median of {rep.ceiling_hits} cut-offs Claude Code recorded itself" + if rep.ceiling_hits > 1 + else f"the window that ended {_stamp(rep.ceiling_at)}, when Claude Code recorded the cut-off" ) - ) + elif rep.ceiling_source == "provider": + how = ( + f"from {rep.ceiling_hits} readings of the provider's own usage % that " + f"{agent_label(rep.agent)} recorded, against your usage in the same windows" + ) + elif rep.ceiling_source == "saved": + how = "saved by an earlier run" + ( + f" ({_stamp(rep.ceiling_at)})" if rep.ceiling_at else "" + ) + else: + how = f"the window that ended {_stamp(rep.ceiling_at)}, when you say you were cut off" + out.append(p.dim(f" {how}")) out.append(f" {'ceiling':<25} {_fmt(rep.ceiling):>10} weighted tokens") for label, val in ( ("peak window", peak.weight), @@ -319,6 +516,23 @@ def render_limits(rep: LimitsReport, color: bool = True) -> str: f" {rep.slots_over_ceiling} other slot(s) in this window reached it too" ) ) + if rep.minutes_to_wall is not None: + ttw = _fmt_minutes(rep.minutes_to_wall) + txt = f" {'TIME TO WALL':<25} {ttw:>10} at the pace of the last {PACE_SECONDS // 60} min" + out.append( + p.red(txt) if rep.minutes_to_wall < 30 else (p.yellow(txt) if rep.minutes_to_wall < 90 else txt) + ) + elif rep.pace <= 0: + out.append(p.dim(f" {'TIME TO WALL':<25} {'idle':>10} nothing in the last {PACE_SECONDS // 60} min")) + for wm, used, ts in rep.provider_used: + label = f"{wm // 60}h" if wm < 1440 else f"{wm // 1440}d" + out.append( + p.dim(f" {'provider says':<25} {used:>9.0f}% of the {label} window, as of {_stamp(ts)}") + ) + if rep.week_ceiling: + share = rep.week_total / rep.week_ceiling + txt = f" {'weekly ceiling':<25} {_fmt(rep.week_ceiling):>10} this week {share:.0%} of it" + out.append(p.red(txt) if share >= 0.9 else (p.yellow(txt) if share >= 0.6 else txt)) out.append("") if rep.mix: @@ -342,8 +556,9 @@ def render_limits(rep: LimitsReport, color: bool = True) -> str: if not rep.ceiling: out.append( p.dim( - " Been cut off before? `agentburn limits --hit \"YYYY-MM-DD HH:MM\"` turns that\n" - " moment into a ceiling measured from your own wall." + " No cut-off recorded in this window. Been cut off before? " + "`agentburn limits --hit \"YYYY-MM-DD HH:MM\"`\n" + " turns that moment into a ceiling measured from your own wall." ) ) out.append( @@ -359,6 +574,41 @@ def render_limits(rep: LimitsReport, color: bool = True) -> str: return "\n".join(out) +def _fmt_minutes(m: float) -> str: + if m < 1: + return "<1 min" + if m < 90: + return f"{m:.0f} min" + if m < 48 * 60: + return f"{m / 60:.1f} h" + return f"{m / 1440:.1f} d" + + +def statusline(rep: LimitsReport) -> str: + """One line for an editor/status bar: how full the window is, time to wall. + + Deliberately short and free of ANSI: Claude Code's `statusLine` prints + exactly what the command writes. + """ + if rep.unsupported or not rep.peak: + return "⏳ agentburn: no windows" + if rep.ceiling: + share = rep.current / rep.ceiling + bits = [f"⏳ {rep.window_hours:g}h {share:.0%}"] + if rep.minutes_to_wall is not None: + bits.append(f"wall in {_fmt_minutes(rep.minutes_to_wall)}") + elif rep.pace <= 0: + bits.append("idle") + if rep.week_ceiling and rep.week_total: + bits.append(f"week {rep.week_total / rep.week_ceiling:.0%}") + return " · ".join(bits) + bits = [f"⏳ {rep.window_hours:g}h {_fmt(rep.current)}"] + if rep.typical > 0: + bits.append(f"{rep.current / rep.typical:.1f}× typical") + bits.append("no ceiling yet") + return " · ".join(bits) + + def _tips(rep: LimitsReport) -> list: """Ranked, and only when the numbers actually support the claim.""" tips = [] @@ -430,7 +680,19 @@ def limits_json(rep: LimitsReport) -> dict: "mix": [{"kind": k, "share": round(s, 4)} for k, s in rep.mix], "ceiling": round(rep.ceiling) if rep.ceiling else None, "ceiling_at": rep.ceiling_at, + "ceiling_source": rep.ceiling_source or None, + "ceiling_hits": rep.ceiling_hits, "slots_over_ceiling": rep.slots_over_ceiling, + "pace_per_minute": round(rep.pace * 60), + "minutes_to_wall": round(rep.minutes_to_wall, 1) if rep.minutes_to_wall is not None else None, + "week_peak": ( + {"start": rep.week_peak.start, "end": rep.week_peak.end, "weight": round(rep.week_peak.weight)} + if rep.week_peak + else None + ), + "week_ceiling": round(rep.week_ceiling) if rep.week_ceiling else None, + "provider_used": [{"window_minutes": wm, "used_percent": u, "ts": ts} for wm, u, ts in rep.provider_used], + "peak_by_project": [{"project": n, "share": round(v, 4)} for n, v in rep.peak_by_project], "tips": _tips(rep), "notes": rep.notes, } diff --git a/agentburn/mcp.py b/agentburn/mcp.py index 6bb1f10..e4d2bc5 100644 --- a/agentburn/mcp.py +++ b/agentburn/mcp.py @@ -67,6 +67,24 @@ ), "inputSchema": WINDOW, }, + { + "name": "burn_context", + "description": ( + "The price of long contexts on this machine: share of the usage window spent at " + "each context size, what a /clear at 100k/150k/200k/300k would have saved, the " + "longest sessions, usage by effort level, and what each skill costs per load " + "(measured from context growth). Returns JSON." + ), + "inputSchema": WINDOW, + }, + { + "name": "burn_commits", + "description": ( + "What each git commit cost, in weighted tokens: sessions joined to the repositories " + "they ran in (read-only git log). Costliest commits, median per repository. Returns JSON." + ), + "inputSchema": WINDOW, + }, { "name": "burn_card", "description": "Anonymized shareable burn summary (plain text, safe to post).", @@ -106,6 +124,18 @@ def _call(name: str, args: dict) -> str: from .limits import build_limits, limits_json return json.dumps(limits_json(build_limits(snap)), indent=2, ensure_ascii=False) + if name == "burn_context": + from .context import build_context, context_json + + return json.dumps(context_json(build_context(snap)), indent=2, ensure_ascii=False) + if name == "burn_commits": + import time as _time + + from .commits import build_commits, commits_json + + days = (args or {}).get("days", 30) + since = _time.time() - days * 86400 if days else None + return json.dumps(commits_json(build_commits(snap, since=since)), indent=2, ensure_ascii=False) if name == "burn_card": from .share import share_text diff --git a/agentburn/model.py b/agentburn/model.py index 78ca288..7a28a42 100644 --- a/agentburn/model.py +++ b/agentburn/model.py @@ -30,6 +30,8 @@ class SessionRec: cost_basis: str # "actual" | "estimated" | "unknown" message_count: int = 0 provider: Optional[str] = None # billing provider, for doctor diagnostics + project: Optional[str] = None # working directory the session ran in, when recorded + branch: Optional[str] = None # git branch, when the agent records it @property def total_tokens(self) -> int: @@ -88,6 +90,67 @@ class UsageCell: output_tokens: int = 0 cache_read_tokens: int = 0 cache_write_tokens: int = 0 + session: str = "" # owning session id, so a window can be split by project + + +@dataclass +class LimitHit: + """The agent itself recorded that a usage limit was reached. + + Claude Code writes a synthetic assistant turn ("You've hit your session + limit · resets 8:30pm (Europe/Amsterdam)") at the moment of the cut-off. + That moment is a measured wall: the window that ended there is a ceiling + nobody had to guess. `reset_at` is the window's scheduled end when the + message stated one and it could be placed on the clock. + """ + + ts: float + kind: str # "session" (rolling 5h) | "weekly" | other wording, lowercased + reset_at: Optional[float] = None + + +@dataclass +class RateLimitSample: + """The provider's own reading of a usage window, as the agent recorded it. + + Codex writes `rate_limits.{primary,secondary}.used_percent` with every + token count. Paired with our weighted usage in the same window that is a + measured ceiling: weight_in_window / used_percent × 100. + """ + + ts: float + window_minutes: int + used_percent: float + resets_at: Optional[float] = None + + +@dataclass +class ContextCall: + """One API call's context size: what the model had to read before answering. + + On a subscription the context is the window: a call at 300k context costs + the same cache-read volume as three calls at 100k. Adapters that see + per-call usage fill these; `agentburn context` turns them into the price + of long sessions and the saving of a `/clear` at a threshold. + """ + + ts: Optional[float] + session: str + model: Optional[str] + context: int # input + cache read + cache write, i.e. everything re-read + output: int + effort: Optional[str] = None + + +@dataclass +class SkillLoad: + """A skill invocation and how much context it added (measured, not read + from disk: bundled skills never touch the disk).""" + + session: str + ts: Optional[float] + skill: str + tokens: int @dataclass @@ -104,6 +167,9 @@ class DumpComposition: "hermes": "Hermes", "openclaw": "OpenClaw", "claude-code": "Claude Code", + "codex": "Codex CLI", + "gemini": "Gemini CLI", + "opencode": "opencode", } # Storage the user would name in an upstream bug report, per agent. @@ -111,9 +177,21 @@ class DumpComposition: "hermes": "`~/.hermes/state.db`", "openclaw": "the local transcript store", "claude-code": "`~/.claude/projects/**.jsonl`", + "codex": "`~/.codex/sessions/**/rollout-*.jsonl`", + "gemini": "`~/.gemini/tmp/*/chats/session-*.json`", + "opencode": "`~/.local/share/opencode/opencode.db`", } +# Agents that record no prices at all (subscription or free tier): a session +# without a cost there is the design, not an accounting gap. +NO_LOCAL_COSTS = frozenset({"claude-code", "codex", "gemini"}) + + +def records_costs(agent: str) -> bool: + return agent_key(agent) not in NO_LOCAL_COSTS + + def agent_key(agent: str) -> str: """Bare adapter key. `behavior` may append ` · ` to Snapshot.agent.""" return agent.split(" · ", 1)[0] @@ -132,7 +210,7 @@ def agent_store(agent: str) -> str: @dataclass class Snapshot: - agent: str # "hermes" | "openclaw" | "claude-code" + agent: str # "hermes" | "openclaw" | "claude-code" | "codex" | "gemini" | "opencode" source_path: str generated_at: float days: Optional[int] @@ -150,3 +228,11 @@ class Snapshot: ) # session_id → count of context compactions # windowed usage (only adapters with per-call timestamps fill this) usage_cells: list[UsageCell] = field(default_factory=list) + # moments the agent itself recorded a limit cut-off (measured ceilings) + limit_hits: list = field(default_factory=list) # LimitHit + # per-call context sizes (only adapters with per-call usage fill this) + context_calls: list = field(default_factory=list) # ContextCall + # skill invocations with their measured context cost + skill_loads: list = field(default_factory=list) # SkillLoad + # the provider's own window readings, when the agent records them + rate_limits: list = field(default_factory=list) # RateLimitSample diff --git a/agentburn/prices.py b/agentburn/prices.py index e7c643f..1589cd1 100644 --- a/agentburn/prices.py +++ b/agentburn/prices.py @@ -30,6 +30,19 @@ "qwen/qwen3.6-plus": (0.325, 1.95), "stepfun/step-3.5-flash": (0.09, 0.3), "z-ai/glm-5-turbo": (1.2, 4.0), + # GLM (Z.ai) list prices, USD per 1M, as published 2026-09-01. The Coding + # Plan is a subscription: these ratios weigh windows, they are not a bill. + "z-ai/glm-4.5": (0.6, 2.2), + "z-ai/glm-4.5-air": (0.2, 1.1), + "z-ai/glm-4.6": (0.6, 2.2), + "z-ai/glm-4.7": (0.6, 2.2), + "z-ai/glm-5": (1.0, 3.2), + "z-ai/glm-5.1": (1.4, 4.4), + "z-ai/glm-5.2": (1.4, 4.4), + "z-ai/glm-5.3": (1.4, 4.4), + # Gemini list prices (Google AI, standard tier, ≤200k prompt). + "google/gemini-2.5-flash": (0.3, 2.5), + "google/gemini-2.5-pro": (1.25, 10.0), } CHEAP_REFERENCE = "deepseek/deepseek-chat" @@ -46,7 +59,7 @@ def _norm(model: str) -> str: # Agents that talk to one vendor log a bare model id ("claude-opus-5") where # routed setups log "anthropic/claude-opus-5". Same model, same price. -_BARE_PREFIXES = {"claude": "anthropic", "gpt": "openai", "o3": "openai"} +_BARE_PREFIXES = {"claude": "anthropic", "gpt": "openai", "o3": "openai", "glm": "z-ai", "gemini": "google"} def _with_author(m: str) -> str: diff --git a/agentburn/recommend.py b/agentburn/recommend.py index 3401f87..8fdeaf1 100644 --- a/agentburn/recommend.py +++ b/agentburn/recommend.py @@ -8,6 +8,7 @@ from __future__ import annotations from .analyze import Analysis +from .model import agent_key EXPENSIVE_HINTS = ("opus", "gpt-5", "o3", "sonnet", "pro") @@ -107,8 +108,9 @@ def recommend(a: Analysis) -> list: recs.insert( 0, f"{a.zero_token_sessions}/{a.total.sessions} sessions recorded zero tokens despite " - "having messages — fix accounting first (check provider usage reporting; " - "hermes-agent #12023), otherwise every number here is an undercount.", + "having messages — fix accounting first (check provider usage reporting" + + ("; hermes-agent #12023" if agent_key(a.agent) == "hermes" else "") + + "), otherwise every number here is an undercount.", ) return recs[:4] diff --git a/pyproject.toml b/pyproject.toml index 3f95289..1e2889d 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,14 +4,14 @@ build-backend = "setuptools.build_meta" [project] name = "agentburn" -version = "0.13.3" +version = "0.14.0" description = "Which 5-hour window took you out, and where the money goes. Local usage/cost profiler for Claude Code, OpenClaw and Hermes Agent. Zero dependencies, nothing leaves your machine." readme = "README.md" requires-python = ">=3.9" license = { text = "MIT" } authors = [{ name = "Ion" }] keywords = ["claude-code", "usage-limits", "tokens", "cost", "profiler", "ai-agents", - "hermes-agent", "openclaw", "mcp", "observability"] + "hermes-agent", "openclaw", "codex-cli", "gemini-cli", "opencode", "mcp", "observability"] classifiers = [ "Environment :: Console", "Intended Audience :: Developers", diff --git a/server.json b/server.json index 71618d7..1a9fa58 100644 --- a/server.json +++ b/server.json @@ -8,13 +8,13 @@ "source": "github" }, "websiteUrl": "https://github.com/Socialpranker/agentburn", - "version": "0.13.3", + "version": "0.14.0", "packages": [ { "registryType": "pypi", "registryBaseUrl": "https://pypi.org", "identifier": "agentburn", - "version": "0.13.3", + "version": "0.14.0", "runtimeHint": "uvx", "packageArguments": [ { diff --git a/skill/agentburn/SKILL.md b/skill/agentburn/SKILL.md index 11ce466..2004e5e 100644 --- a/skill/agentburn/SKILL.md +++ b/skill/agentburn/SKILL.md @@ -24,10 +24,24 @@ database read-only. Use it instead of guessing about costs. run `uvx agentburn limits --json` and read: `peak` (the worst rolling 5-hour window), `typical_window`, `peak_by_model`, `peak_by_source`, `mix`, `tips`. Lead with peak ÷ typical: a wall is hit by the peak. - If the user remembers when they were cut off, re-run with - `--hit "YYYY-MM-DD HH:MM"` — that turns their own cut-off into a - measured ceiling and everything else into a percentage of it. + `ceiling` is measured from cut-offs Claude Code recorded itself + (`ceiling_source: recorded`, `ceiling_hits`); `minutes_to_wall` is at + the pace of the last 30 minutes. If there is no ceiling and the user + remembers when they were cut off, re-run with `--hit "YYYY-MM-DD HH:MM"`. Never state an absolute limit: the provider's formula is not public. + Codex CLI: `ceiling_source: provider` means the ceiling comes from the + `used_percent` Codex records itself — an estimate, other devices on the + account count too; `provider_used` is the latest raw reading. + Gemini CLI and opencode: tokens only (opencode carries its own costs + when the provider is priced); no cut-offs are recorded, so `ceiling` + is absent unless `--hit` names one. +4b. For "why is my context so big / should I /clear / what does skill X + cost": run `uvx agentburn context --json` and read `bands`, `savings` + (`clear_at` → `share_not_spent`), `longest_sessions`, `by_effort`, + `skills` (`tokens_per_load`, measured). Recommend the lowest `clear_at` + that keeps most of the best saving. +4c. For "what did that commit / PR cost": run `uvx agentburn commits --json` + and read `top` (costliest commits) and `repos` (median per repository). 5. For one channel ("what did you do in telegram?"): add `--source telegram` (or cron / heartbeat / subagent / cli). 6. Answer in the user's language, lead with the verdict diff --git a/tests/selftest.py b/tests/selftest.py index 5a87d8a..fa1aff4 100644 --- a/tests/selftest.py +++ b/tests/selftest.py @@ -658,9 +658,9 @@ def log_message(self, *a): lines = [json.loads(l) for l in r_mcp.stdout.strip().splitlines()] byid = {l.get("id"): l for l in lines} ok("mcp: initialize → serverInfo", byid[1]["result"]["serverInfo"]["name"] == "agentburn") - ok("mcp: tools/list → 4 tools", + ok("mcp: tools/list → 6 tools", {t["name"] for t in byid[2]["result"]["tools"]} - == {"burn_report", "burn_why", "burn_limits", "burn_card"}) + == {"burn_report", "burn_why", "burn_limits", "burn_card", "burn_context", "burn_commits"}) body0 = json.loads(byid[3]["result"]["content"][0]["text"]) ok("mcp: tools/call burn_report returns the report JSON", byid[3]["result"]["isError"] is False and body0["agentburn"] == 1 and body0["total"]["sessions"] > 0) @@ -1073,6 +1073,432 @@ def digest(sn): ok("card svg: same window line", "peak" in svg and svg.startswith("", "usage": {"input_tokens": 0, "output_tokens": 0}, + "content": [{"type": "text", "text": "You've hit your session limit · resets 8:30pm (Europe/Amsterdam)"}]}}), + ] + with open(os.path.join(dd_proj, "22222222-2222-4222-8222-222222222222.jsonl"), "w", encoding="utf-8") as f: + f.write("\n".join(lines) + "\n") + dd = cc.load(db_path=dd_root, days=30, now=now) + rec = dd.sessions[0] + ok("dedup: three rows of one reply count as ONE call", rec.api_calls == 3, str(rec.api_calls)) + ok("dedup: usage summed once per requestId", + rec.cache_read_tokens == 360_000 and rec.input_tokens == 2000, str(rec.cache_read_tokens)) + ok("dedup: synthetic cut-off row is not an API call", all(c.model != "" for c in dd.usage_cells)) + ok("dedup: cells agree with the session total", + sum(c.cache_read_tokens for c in dd.usage_cells) == rec.cache_read_tokens) + ok("adapter: cwd and branch recorded on the session", rec.project == repo_dir and rec.branch == "main") + ok("adapter: title uses the recorded working directory", + rec.title.startswith(os.path.basename(repo_dir))) + ok("hits: the recorded cut-off is a LimitHit with kind + reset", + len(dd.limit_hits) == 1 and dd.limit_hits[0].kind == "session" + and abs(dd.limit_hits[0].ts - (t0 + 1260)) < 1) + ok("hits: warning names the recorded cut-off", any("cut-off" in w for w in dd.warnings)) + lim_auto = build_limits(dd, now=now) + ok("limits: ceiling measured from the recorded cut-off, no --hit needed", + lim_auto.ceiling is not None and lim_auto.ceiling > 0 and lim_auto.ceiling_source == "recorded") + ok("limits: week peak found", lim_auto.week_peak is not None and lim_auto.week_peak.weight > 0) + ok("limits: peak split by project uses the recorded cwd", + lim_auto.peak_by_project and lim_auto.peak_by_project[0][0] == os.path.basename(repo_dir)) + ok("limits: --hit by hand still wins over the recorded one", + build_limits(dd, hit=t0 + 1230, now=now).ceiling_source == "--hit") + r_auto = render_limits(lim_auto, color=False) + ok("limits render: says the cut-off was recorded by Claude Code", "recorded" in r_auto) + ok("limits render: time to wall line present", "TIME TO WALL" in r_auto) + lim_busy = build_limits(dd, now=t0 + 1300) + ok("limits: pace over the last 30 min gives minutes to wall", + lim_busy.pace > 0 and lim_busy.minutes_to_wall is not None) + sl = statusline(lim_busy) + ok("statusline: one line, percent of ceiling, no ANSI", + "\n" not in sl and "%" in sl and "\033" not in sl, sl) + state = os.path.join(tempfile.mkdtemp(), "ceiling.json") + save_ceiling("claude-code", lim_auto, path=state) + saved = load_saved_ceiling("claude-code", path=state) + ok("ceiling state: saved and reloaded", saved is not None and saved["ceiling"] == round(lim_auto.ceiling)) + dd_nohit = cc.load(db_path=dd_root, days=30, now=now) + dd_nohit.limit_hits = [] + lim_saved = build_limits(dd_nohit, now=now, saved=saved) + ok("ceiling state: a run without cut-offs falls back to the saved ceiling", + lim_saved.ceiling_source == "saved" and lim_saved.ceiling == float(saved["ceiling"])) + ok("statusline: no ceiling → says so instead of inventing one", + "no ceiling" in statusline(build_limits(dd_nohit, now=now))) + ok("limits json: new fields", "minutes_to_wall" in limits_json(lim_busy) and "week_peak" in limits_json(lim_busy)) + + ctx = build_context(dd) + ok("context: one record per deduplicated call", ctx.calls == 3) + ok("context: max context is the long call", ctx.max_context == 250_000) + ok("context: saving at 200k counts the one call past it", + any(sv.threshold == 200_000 and sv.calls_over == 1 and sv.share > 0 for sv in ctx.savings)) + ok("context: effort levels split", {e for e, _, _ in ctx.by_effort} == {"high", "max"}) + ok("context: skill cost measured from the growth after a lone Skill call", + len(ctx.skills) == 1 and ctx.skills[0].skill == "deploy-verify" and ctx.skills[0].tokens == 10_000, + str([(s_.skill, s_.tokens) for s_ in ctx.skills])) + r_ctx = render_context(ctx, color=False) + ok("context render: bands, savings, skills", "BY CONTEXT SIZE" in r_ctx and "/clear at" in r_ctx and "deploy-verify" in r_ctx) + ok("context json: shape", context_json(ctx)["skills"][0]["tokens_per_load"] == 10_000) + ok("context: adapters without per-call usage say so", + bool(build_context(hermes.load(db_path=env_db, days=30)).unsupported)) + r_ctx_cli = subprocess.run([sys.executable, "-m", "agentburn.cli", "context", "--agent", "claude-code", + "--db", dd_root, "--no-color"], capture_output=True, text=True, encoding="utf-8", errors="replace") + ok("cli context: end-to-end", r_ctx_cli.returncode == 0 and "WHERE THE WINDOW GOES" in r_ctx_cli.stdout, r_ctx_cli.stderr[-300:]) + r_sl = subprocess.run([sys.executable, "-m", "agentburn.cli", "statusline", "--agent", "claude-code", + "--db", dd_root], capture_output=True, text=True, encoding="utf-8", errors="replace") + ok("cli statusline: one line", r_sl.returncode == 0 and r_sl.stdout.count("\n") == 1 and "⏳" in r_sl.stdout, r_sl.stdout) + + # fix: the /clear lever and heavy skills come from the same measurements + from agentburn.fix import build_fixes as _bf # noqa: E402 + a_dd = analyze(dd) + dd_heavy = cc.load(db_path=dd_root, days=30, now=now) + from agentburn.model import SkillLoad # noqa: E402 + dd_heavy.skill_loads += [SkillLoad(session="s", ts=now, skill="fat-skill", tokens=20_000) for _ in range(3)] + fx = _bf("claude-code", dd_root, a_dd, None, dd_heavy) + fx_titles = " | ".join(p_.title for p_ in fx) + ok("fix: /clear lever proposed from measured context", "Restart sessions" in fx_titles, fx_titles) + ok("fix: heavy skill flagged with its measured size", "fat-skill" in " ".join(p_.why for p_ in fx), fx_titles) + + # commits: join sessions to the repository's git log + from agentburn.commits import build_commits, render_commits, commits_json # noqa: E402 + git_env = dict(os.environ, GIT_AUTHOR_NAME="t", GIT_AUTHOR_EMAIL="t@t", GIT_COMMITTER_NAME="t", + GIT_COMMITTER_EMAIL="t@t", GIT_CONFIG_NOSYSTEM="1") + have_git = subprocess.run(["git", "--version"], capture_output=True).returncode == 0 + if have_git: + subprocess.run(["git", "init", "-q", repo_dir], check=True, env=git_env) + def commit(msg, ts): + with open(os.path.join(repo_dir, "f.txt"), "a", encoding="utf-8") as f: + f.write(msg + "\n") + subprocess.run(["git", "-C", repo_dir, "add", "f.txt"], check=True, env=git_env) + stamp = str(int(ts)) + subprocess.run(["git", "-C", repo_dir, "commit", "-q", "-m", msg], check=True, + env=dict(git_env, GIT_AUTHOR_DATE=stamp, GIT_COMMITTER_DATE=stamp)) + commit("first", t0 + 600) # after req-1/req-2 (t0, t0+60) → costs those + commit("second", t0 + 1800) # after req-3 (t0+1200) → costs that one + cm = build_commits(dd, since=now - 30 * 86400) + ok("commits: both commits priced", len(cm.top) == 2, str([(c.subject, c.weight) for c in cm.top])) + first = next(c for c in cm.top if c.subject == "first") + second = next(c for c in cm.top if c.subject == "second") + ok("commits: the long-context call lands on the commit that followed it", + second.weight > first.weight and second.calls == 1 and first.calls == 2) + ok("commits: everything attributed", abs(commits_json(cm)["attributed_share"] - 1.0) < 1e-6) + ok("commits render", "COSTLIEST COMMITS" in render_commits(cm, color=False) and "second" in render_commits(cm, color=False)) + dd_norepo = cc.load(db_path=dd_root, days=30, now=now) + dd_norepo.sessions[0].project = tempfile.mkdtemp() + ok("commits: a session outside any repo is skipped with a reason", + any("not a git" in why for _, why in build_commits(dd_norepo, since=None).skipped)) + else: + print(" (git not found — commits checks skipped)") + ok("commits: adapters without cwd say so", bool(build_commits(hermes.load(db_path=env_db, days=30)).unsupported)) + + # ------------------------------------------------ codex / gemini / opencode + print("codex: cumulative token_count delta, rate_limits, tools:") + from agentburn.adapters import ADAPTERS, codex, gemini, opencode # noqa: E402 + ok("registry: six adapters in order", + list(ADAPTERS) == ["hermes", "openclaw", "claude-code", "codex", "gemini", "opencode"], str(list(ADAPTERS))) + + def iso(ts): + return time.strftime("%Y-%m-%dT%H:%M:%S", time.gmtime(ts)) + ".000Z" + + cx_root = os.path.join(tempfile.mkdtemp(), "sessions") + cx_dir = os.path.join(cx_root, "2026", "09", "01") + os.makedirs(cx_dir) + cx_t0 = now - 3 * 3600 + cx_cwd = os.path.join(tempfile.gettempdir(), "codexproj") + + def cx_line(ts, kind, payload): + return json.dumps({"timestamp": iso(ts), "type": kind, "payload": payload}) + + def token_count(ts, inp, cached, out, reasoning, used=50.0): + tot = {"input_tokens": inp, "cached_input_tokens": cached, "output_tokens": out, + "reasoning_output_tokens": reasoning, "total_tokens": inp + out} + return cx_line(ts, "event_msg", { + "type": "token_count", + "info": {"total_token_usage": tot, "last_token_usage": tot}, + "rate_limits": {"primary": {"used_percent": used, "window_minutes": 300, "resets_at": int(ts) + 3600}, + "secondary": {"used_percent": 12.0, "window_minutes": 10080, "resets_at": None}}, + }) + + cx_lines = [ + cx_line(cx_t0, "session_meta", {"cwd": cx_cwd, "originator": "codex_exec", "cli_version": "0.144.0"}), + cx_line(cx_t0, "turn_context", {"model": "gpt-5.5", "effort": "high", "cwd": cx_cwd}), + # cumulative counter: 10k (2k cached) → 30k (12k cached) → repeat → 45k + token_count(cx_t0 + 10, 10_000, 2_000, 500, 100), + token_count(cx_t0 + 400, 30_000, 12_000, 1_500, 300), + token_count(cx_t0 + 401, 30_000, 12_000, 1_500, 300), # rate-limit refresh re-sends the same totals + token_count(cx_t0 + 800, 45_000, 20_000, 2_500, 500, used=60.0), + cx_line(cx_t0 + 20, "response_item", {"type": "function_call", "name": "shell", + "arguments": json.dumps({"command": "ls -la"}), "call_id": "c1"}), + cx_line(cx_t0 + 21, "response_item", {"type": "function_call_output", "call_id": "c1", + "output": {"output": "x" * 400, "success": True}}), + cx_line(cx_t0 + 30, "response_item", {"type": "custom_tool_call", "name": "apply_patch", "input": "*** Begin"}), + cx_line(cx_t0 + 31, "response_item", {"type": "custom_tool_call_output", "output": "Done"}), + cx_line(cx_t0 + 900, "event_msg", {"type": "context_compacted"}), + ] + with open(os.path.join(cx_dir, "rollout-2026-09-01T10-00-00-abcdef.jsonl"), "w", encoding="utf-8") as f: + f.write("\n".join(cx_lines) + "\n") + # a thread that never got a reply: not a session, not an error + with open(os.path.join(cx_dir, "rollout-2026-09-01T11-00-00-empty.jsonl"), "w", encoding="utf-8") as f: + f.write(cx_line(cx_t0, "session_meta", {"cwd": cx_cwd, "originator": "codex_exec"}) + "\n") + cxs = codex.load(db_path=cx_root, days=30, now=now) + ok("codex: one session, the empty thread skipped", len(cxs.sessions) == 1 and cxs.agent == "codex") + cr = cxs.sessions[0] + ok("codex: repeated token_count is not a call", cr.api_calls == 3, str(cr.api_calls)) + ok("codex: deltas of the cumulative counter, cached split out of input", + cr.input_tokens == 25_000 and cr.cache_read_tokens == 20_000 and cr.output_tokens == 2_500 + and cr.reasoning_tokens == 500, f"{cr.input_tokens} {cr.cache_read_tokens} {cr.output_tokens}") + ok("codex: model from turn_context, cwd from session_meta, cli source", + cr.model == "gpt-5.5" and cr.project == cx_cwd and cr.source == "cli") + ok("codex: title from the working directory", cr.title.startswith("codexproj/")) + ok("codex: no dollars", cr.cost_usd is None and cr.cost_basis == "unknown") + ok("codex: cells agree with the session", sum(c.cache_read_tokens for c in cxs.usage_cells) == 20_000 + and sum(c.calls for c in cxs.usage_cells) == 3) + ok("codex: compaction counted", cxs.compactions.get(cr.id) == 1) + ok("codex: rate-limit samples kept for both windows", + len(cxs.rate_limits) == 8 and {r.window_minutes for r in cxs.rate_limits} == {300, 10080}) + ok("codex: context per call with effort", + len(cxs.context_calls) == 3 and cxs.context_calls[0].effort == "high" and cxs.context_calls[1].context == 20_000) + names = [e.name for e in cxs.events] + ok("codex: tool calls and outputs as events", names == ["shell", "tool", "apply_patch", "tool"], str(names)) + ok("codex: shell arg grouped on the command, output priced in tokens", + cxs.events[0].arg_key and "ls" in cxs.events[0].arg_key and cxs.events[1].ok is True and cxs.events[1].tokens == 107) + ok("codex: warning about no local prices", any("dollars" in w for w in cxs.warnings)) + lim_cx = build_limits(cxs, now=now) + ok("codex limits: ceiling from the provider's used_percent", + lim_cx.ceiling_source == "provider" and lim_cx.ceiling and lim_cx.ceiling > 0, lim_cx.ceiling_source) + ok("codex limits: latest provider reading per window", + [(wm, u) for wm, u, _ in lim_cx.provider_used] == [(300, 60.0), (10080, 12.0)], str(lim_cx.provider_used)) + ok("codex limits render: provider line", "provider says" in render_limits(lim_cx, color=False)) + ok("codex limits: week ceiling from the weekly window", lim_cx.week_ceiling is not None and lim_cx.week_ceiling > 0) + ok("codex limits: peak inside the readings' coverage → no staleness note", not lim_cx.notes, str(lim_cx.notes)) + # readings stop, then a bigger peak happens: the ratio must be flagged, not sold as an overrun + from agentburn.model import RateLimitSample as _RLS, UsageCell as _UC # noqa: E402 + cx_stale = codex.load(db_path=cx_root, days=30, now=now) + cx_stale.usage_cells.append(_UC(start=int((cx_t0 + 7 * 3600) // 300) * 300, source="desktop", model="gpt-5.5", + calls=5, input_tokens=900_000, output_tokens=50_000, cache_read_tokens=0, + cache_write_tokens=0, session="later")) + lim_stale = build_limits(cx_stale, now=now) + ok("codex limits: peak after the last 5h reading is flagged as unmeasured", + lim_stale.ceiling_source == "provider" and any("last 5h reading" in n_ for n_ in lim_stale.notes), str(lim_stale.notes)) + # a 30-day reading is not a weekly ceiling + cx_30d = codex.load(db_path=cx_root, days=30, now=now) + cx_30d.rate_limits = [r for r in cx_30d.rate_limits if r.window_minutes == 300] + cx_30d.rate_limits.append(_RLS(ts=cx_t0 + 800, window_minutes=43_200, used_percent=90.0, resets_at=None)) + ok("codex limits: a 30-day reading does not become the weekly ceiling", + build_limits(cx_30d, now=now).week_ceiling is None) + try: + codex.load(db_path=cx_root, days=1, now=now + 10 * 86400) + ok("codex: empty window raises", False) + except RuntimeError as e: + ok("codex: empty window raises with a hint", "--days 0" in str(e)) + r_cx = subprocess.run([sys.executable, "-m", "agentburn.cli", "--agent", "codex", "--db", cx_root, "--no-color"], + capture_output=True, text=True, encoding="utf-8", errors="replace") + ok("cli codex: end-to-end", r_cx.returncode == 0 and "gpt-5.5" in r_cx.stdout, (r_cx.stdout + r_cx.stderr)[-400:]) + + print("gemini: per-turn tokens, projects.json label→cwd, toolCalls:") + gm_home = tempfile.mkdtemp() + gm_root = os.path.join(gm_home, "tmp") + gm_cwd = os.path.join(tempfile.gettempdir(), "geminiproj") + os.makedirs(os.path.join(gm_root, "proj", "chats")) + os.makedirs(os.path.join(gm_root, "orphan", "chats")) + with open(os.path.join(gm_home, "projects.json"), "w", encoding="utf-8") as f: + json.dump({"projects": {gm_cwd: "proj"}}, f) + gm_t0 = now - 2 * 3600 + + def gm_msg(ts, model, inp, cached, out, thoughts, tool_calls=None): + return {"type": "gemini", "model": model, "timestamp": iso(ts), "content": "…", + "tokens": {"input": inp, "output": out, "cached": cached, "thoughts": thoughts, "tool": 0, + "total": inp + out + thoughts}, + "toolCalls": tool_calls or []} + + gm_doc = {"sessionId": "11111111-aaaa-4bbb-8ccc-000000000001", "projectHash": "h", "kind": "main", + "startTime": iso(gm_t0), "lastUpdated": iso(gm_t0 + 700), + "messages": [ + {"type": "user", "timestamp": iso(gm_t0), "content": "hi"}, + gm_msg(gm_t0 + 5, "gemini-2.5-pro", 8_000, 3_000, 400, 200, + [{"name": "read_file", "args": {"path": "/x/y.py"}, "status": "success", "result": {"o": "z" * 200}}]), + gm_msg(gm_t0 + 400, "gemini-2.5-pro", 20_000, 15_000, 600, 100, + [{"name": "run_shell_command", "args": {"command": "pytest"}, "status": "error", "result": "boom"}]), + ]} + with open(os.path.join(gm_root, "proj", "chats", "session-2026-09-01T10-00-00-abc.json"), "w", encoding="utf-8") as f: + json.dump(gm_doc, f) + with open(os.path.join(gm_root, "orphan", "chats", "session-2026-09-01T12-00-00-def.json"), "w", encoding="utf-8") as f: + json.dump({"sessionId": "22222222-aaaa-4bbb-8ccc-000000000002", "kind": "subagent", "startTime": iso(gm_t0), + "lastUpdated": iso(gm_t0 + 5), + "messages": [gm_msg(gm_t0 + 5, "gemini-2.5-flash", 1_000, 0, 50, 0)]}, f) + with open(os.path.join(gm_root, "proj", "chats", "session-2026-09-01T13-00-00-nil.json"), "w", encoding="utf-8") as f: + json.dump({"sessionId": "3", "kind": "main", "messages": [{"type": "user", "timestamp": iso(gm_t0), "content": "?"}]}, f) + gms = gemini.load(db_path=gm_root, days=30, now=now) + ok("gemini: two sessions with usage, the reply-less chat skipped", len(gms.sessions) == 2 and gms.agent == "gemini") + gr = next(s_ for s_ in gms.sessions if s_.id.endswith("0001")) + go = next(s_ for s_ in gms.sessions if s_.id.endswith("0002")) + ok("gemini: cwd resolved through projects.json by label", gr.project == gm_cwd, str(gr.project)) + ok("gemini: unknown label → no project, not a crash", go.project is None) + ok("gemini: per-turn tokens summed, cached split out of input", + gr.api_calls == 2 and gr.input_tokens == 10_000 and gr.cache_read_tokens == 18_000 + and gr.output_tokens == 1_000 and gr.reasoning_tokens == 300, + f"{gr.input_tokens} {gr.cache_read_tokens} {gr.output_tokens} {gr.reasoning_tokens}") + ok("gemini: model and title", gr.model == "gemini-2.5-pro" and gr.title.startswith("proj/")) + ok("gemini: kind main → cli, other kinds → subagent", gr.source == "cli" and go.source == "subagent") + ok("gemini: time span from the messages", gr.started_at is not None and gr.ended_at - gr.started_at >= 700) + ok("gemini: cells carry thoughts as output", + sum(c.output_tokens for c in gms.usage_cells if c.session == gr.id) == 1_300) + ok("gemini: context per call in whole input", any(c.context == 20_000 for c in gms.context_calls)) + ev = [e for e in gms.events if e.session_id == gr.id] + ok("gemini: two events per tool call — the call and its result", + [e.name for e in ev] == ["read_file", "read_file", "run_shell_command", "run_shell_command"], str([e.name for e in ev])) + ok("gemini: status → ok, result sized in tokens", + ev[1].ok is True and ev[1].tokens and ev[1].tokens > 40 and ev[3].ok is False) + ok("gemini: no dollars", gr.cost_usd is None and any("dollars" in w for w in gms.warnings)) + try: + gemini.load(db_path=gm_root, days=1, now=now + 10 * 86400) + ok("gemini: empty window raises", False) + except RuntimeError as e: + ok("gemini: empty window raises with a hint", "--days 0" in str(e)) + r_gm = subprocess.run([sys.executable, "-m", "agentburn.cli", "--agent", "gemini", "--db", gm_root, "--no-color"], + capture_output=True, text=True, encoding="utf-8", errors="replace") + ok("cli gemini: end-to-end", r_gm.returncode == 0 and "gemini-2.5-pro" in r_gm.stdout, (r_gm.stdout + r_gm.stderr)[-400:]) + + print("opencode: sqlite session/message/part, agent-priced cost:") + oc_path = os.path.join(tempfile.mkdtemp(), "opencode.db") + oc_t0_ms = int((now - 3600) * 1000) + ocon = sqlite3.connect(oc_path) + ocon.executescript( + """ + CREATE TABLE session(id TEXT PRIMARY KEY, parent_id TEXT, directory TEXT, title TEXT, model TEXT, + cost REAL, time_created INTEGER, time_updated INTEGER); + CREATE TABLE message(id TEXT PRIMARY KEY, session_id TEXT, time_created INTEGER, data TEXT); + CREATE TABLE part(id TEXT PRIMARY KEY, message_id TEXT, session_id TEXT, time_created INTEGER, data TEXT); + """ + ) + + def oc_msg(role, i_, o_, rs, cr, cw, cost, provider="anthropic", model="claude-sonnet-5"): + return json.dumps({"role": role, "modelID": model, "providerID": provider, "cost": cost, + "tokens": {"input": i_, "output": o_, "reasoning": rs, "cache": {"read": cr, "write": cw}}}) + + ocon.executemany("INSERT INTO session VALUES (?,?,?,?,?,?,?,?)", [ + ("ses_main", None, "/w/repo", "Fix the build", "anthropic/claude-sonnet-5", 0.5, oc_t0_ms, oc_t0_ms + 900_000), + ("ses_sub", "ses_main", "/w/repo", "explore", "ollama/llama", 0.0, oc_t0_ms, oc_t0_ms + 900_000), + ("ses_old", None, "/w/old", "ancient", None, 0.0, oc_t0_ms - 90 * 86400_000, oc_t0_ms - 90 * 86400_000), + ("ses_bare", None, "/w/repo", "no reply yet", None, 0.0, oc_t0_ms, oc_t0_ms), + ]) + ocon.executemany("INSERT INTO message VALUES (?,?,?,?)", [ + ("m1", "ses_main", oc_t0_ms, json.dumps({"role": "user"})), + ("m2", "ses_main", oc_t0_ms + 1_000, oc_msg("assistant", 5_000, 300, 50, 20_000, 1_000, 0.10)), + ("m3", "ses_main", oc_t0_ms + 400_000, oc_msg("assistant", 6_000, 700, 0, 25_000, 0, 0.15)), + ("m4", "ses_sub", oc_t0_ms + 2_000, oc_msg("assistant", 1_000, 100, 0, 0, 0, 0.0, "ollama", "llama3")), + ("m5", "ses_old", oc_t0_ms - 90 * 86400_000, oc_msg("assistant", 9, 9, 0, 0, 0, 9.0)), + ]) + ocon.executemany("INSERT INTO part VALUES (?,?,?,?,?)", [ + ("p1", "m2", "ses_main", oc_t0_ms + 1_500, json.dumps({"type": "tool", "tool": "bash", + "state": {"status": "completed", "input": {"command": "make"}, + "output": "o" * 800}})), + ("p2", "m2", "ses_main", oc_t0_ms + 1_600, json.dumps({"type": "text", "text": "…"})), + ("p3", "m3", "ses_main", oc_t0_ms + 400_500, json.dumps({"type": "tool", "tool": "read", + "state": {"status": "error", "input": {"filePath": "/w/x"}}})), + ]) + ocon.commit() + ocon.close() + ocs = opencode.load(db_path=oc_path, days=30, now=now) + ok("opencode: sessions in the window with assistant replies only", + sorted(s_.id for s_ in ocs.sessions) == ["ses_main", "ses_sub"], str([s_.id for s_ in ocs.sessions])) + om = next(s_ for s_ in ocs.sessions if s_.id == "ses_main") + osb = next(s_ for s_ in ocs.sessions if s_.id == "ses_sub") + ok("opencode: tokens summed, cache read/write kept apart", + om.api_calls == 2 and om.input_tokens == 11_000 and om.output_tokens == 1_000 and om.reasoning_tokens == 50 + and om.cache_read_tokens == 45_000 and om.cache_write_tokens == 1_000, f"{om.input_tokens} {om.cache_read_tokens}") + ok("opencode: cost is the agent's own, basis actual", abs(om.cost_usd - 0.25) < 1e-9 and om.cost_basis == "actual") + ok("opencode: zero-cost provider → tokens only, basis unknown", osb.cost_usd is None and osb.cost_basis == "unknown") + ok("opencode: model qualified with provider", om.model == "anthropic/claude-sonnet-5" and om.provider == "anthropic") + ok("opencode: parent_id → subagent, directory → project", + osb.source == "subagent" and osb.parent_id == "ses_main" and om.project == "/w/repo" and om.title == "Fix the build") + ok("opencode: one cell per assistant message, reasoning inside output", + sum(c.calls for c in ocs.usage_cells) == 3 and sum(c.output_tokens for c in ocs.usage_cells if c.session == om.id) == 1_050) + ok("opencode: context = input + cache read + cache write", any(c.context == 26_000 for c in ocs.context_calls)) + oev = [e for e in ocs.events if e.session_id == om.id] + ok("opencode: tool parts → events, text parts ignored", + [e.name for e in oev] == ["bash", "bash", "read", "read"], str([e.name for e in oev])) + ok("opencode: status and output size on the result event", + oev[1].ok is True and oev[1].tokens == 200 and oev[3].ok is False and oev[3].tokens is None) + ok("opencode: unpriced sessions warned", any("no cost" in w for w in ocs.warnings)) + ok("opencode: read-only — the db file is untouched", not os.path.exists(oc_path + "-journal")) + try: + opencode.load(db_path=oc_path, days=1, now=now + 10 * 86400) + ok("opencode: empty window raises", False) + except RuntimeError as e: + ok("opencode: empty window raises with a hint", "--days 0" in str(e)) + a_oc = analyze(ocs) + ok("opencode analyze: dollars flow through, basis mixed", a_oc.cost_basis in ("actual", "mixed") and a_oc.daily_cost) + r_oc = subprocess.run([sys.executable, "-m", "agentburn.cli", "--agent", "opencode", "--db", oc_path, "--no-color"], + capture_output=True, text=True, encoding="utf-8", errors="replace") + ok("cli opencode: end-to-end", r_oc.returncode == 0 and "Fix the build" in r_oc.stdout, (r_oc.stdout + r_oc.stderr)[-400:]) + + print("cli: one empty adapter does not sink the others:") + import shutil + mh = tempfile.mkdtemp() + shutil.copytree(dd_root, os.path.join(mh, ".claude", "projects")) + mh_cx = os.path.join(mh, ".codex", "sessions", "2026", "09", "01") + os.makedirs(mh_cx) + with open(os.path.join(mh_cx, "rollout-2026-09-01T09-00-00-bare.jsonl"), "w", encoding="utf-8") as f: + f.write(cx_line(now - 60, "session_meta", {"cwd": cx_cwd, "originator": "codex_exec"}) + "\n") + r_multi = subprocess.run([sys.executable, "-m", "agentburn.cli", "--no-color"], capture_output=True, text=True, + encoding="utf-8", errors="replace", env=fake_home(mh)) + ok("cli multi: claude-code report printed although codex had nothing, header counts only the loaded", + r_multi.returncode == 0 and "Found" not in r_multi.stdout and "codex" in r_multi.stderr + and "skipped" in r_multi.stderr, (r_multi.stdout + r_multi.stderr)[-400:]) + r_multi_lim = subprocess.run([sys.executable, "-m", "agentburn.cli", "limits", "--no-color"], capture_output=True, + text=True, encoding="utf-8", errors="replace", env=fake_home(mh)) + ok("cli multi limits: same tolerance", r_multi_lim.returncode == 0 and "codex" in r_multi_lim.stderr, + (r_multi_lim.stdout + r_multi_lim.stderr)[-400:]) + r_single = subprocess.run([sys.executable, "-m", "agentburn.cli", "--agent", "codex", "--no-color"], + capture_output=True, text=True, encoding="utf-8", errors="replace", env=fake_home(mh)) + ok("cli --agent codex alone: still an error with the hint", r_single.returncode == 2 and "--days 0" in r_single.stderr, + r_single.stderr[-300:]) + + print("doctor: agents without local prices are not 'unpriced':") + from agentburn.doctor import render_doctor as _rd # noqa: E402 + for name_, snap_ in (("codex", cxs), ("gemini", gms)): + doc_ = _rd(snap_, color=False) + ok(f"doctor {name_}: healthy, no fake pricing gap, no issue template", + "healthy" in doc_ and "by design" in doc_ and "GitHub issue" not in doc_, doc_[-300:]) + doc_oc = _rd(ocs, color=False) + ok("doctor opencode: a zero-cost provider IS reported as unpriced", + "1 × ollama" in doc_oc and "unpriced sessions : 1" in doc_oc, doc_oc[-400:]) + # ------------------------------------------------- fix for claude code print("fix (claude-code levers):") from agentburn.fix import build_fixes, render_fixes # noqa: E402