AI Charts plots published AI benchmark scores against cost and tokens per task, marking the best score at every budget. A local collector measures your own agents' token use.
The repository also contains a local collector that measures your own coding agents' token use. It is in development and has no packaged release yet, so build it from source. Reports you open on the site stay in your browser tab. Account sync needs a collector enrolled on a Mac and works only while the site's account service is enabled.
The homepage leads with an interactive Pareto frontier: compare model capability against output tokens or cost, then inspect the configuration behind each point. The original coding-agent charts have a focused home at /coding. The separate /benchmarks library covers coding, reasoning, research, memory, images, video, audio, and world models. Charted results, source guides, and emerging evaluations are labeled separately, and older research cohorts are labeled with their dates.
The site header links Charts (/), Benchmarks (/benchmarks), Usage (/usage), Dashboard (/dashboard), Leaderboard (/leaderboard), and Notes (/blog). Coding comparisons, model pages, and source data are linked where relevant. Legacy root chart and atlas links resolve to the matching workspace; new shares use that workspace’s URL. Each canonical page also has a Markdown representation through Accept: text/markdown.
The site is a static-data Next.js and TypeScript application. Its benchmark snapshots are committed to the repository, validated at build time, and refreshed automatically. Production never depends on an upstream data source being available during a request.
- Use the official, version-pinned Terminal-Bench 4.0 snapshot as the current coding standard, with exact model, agent, effort, trial, uncertainty, cost, token, duration, and source metadata.
- Explore native score rankings, cost trade-offs where reported, tables, and up to three selected configurations within one benchmark. Inspect exact values, uncertainty, source dates, and evaluation conditions, then share the configured view.
- Compare checked ARC-AGI-2 and ARC-AGI-3 tracks, selected DeepResearch Bench II systems, and LongMemEval-V2 Small and Medium baselines without combining their scores.
- Explore four separate Arena preference-rating cohorts for image generation, image editing, text-to-video, and image-to-video, with source dates, intervals, and vote counts. Compare selected open-weight speech configurations on the Open ASR AMI-Cleaned English test; lower word-error rate is better.
- Explore WISE Verified image generation, GEditBench v2 editing, VideoPhy2 human evaluation, OmniDocBench v1.6_full document parsing, and WorldScore's historical author cohort. Other evaluations have source guides until their data imports are qualified.
- Keep a checked, version-pinned Terminal-Bench-Science 0.1 owner snapshot, with per-domain results, uncertainty, cost, and token use, separate from GDPval-AA v2, OSWorld 2.0, and Humanity's Last Exam instead of collapsing the families into one composite score.
- Treat CursorBench 3.2 as supplemental closed evidence for the model-plus-Cursor system, not an independently reproducible coding standard.
- Compare Artificial Analysis Intelligence Index v4.3.2 with native output-only tokens and cost per Index task in the leading model-level Pareto chart. The historical v4.1.1 dataset remains separately accessible; scores are never relabeled across versions.
- Compare Artificial Analysis Coding Agent Index v1.5, DeepSWE v1.1, Terminal-Bench 4, and SWE-Atlas-QnA results.
- Plot each result against cost, duration, or total token use.
- Pin a model to see its nearby performance cohort, or pin a provider to inspect its range.
- Explore the cost/performance Pareto frontier and per-provider score ranges.
- Follow a checked timeline of newly detected models, settings, and material benchmark changes.
- Share the current axes and selection as a link or export a full-resolution PNG.
- Open a profile-specific benchmark card for each model, then share its branded image through the native share sheet, download it, or post its URL from
/models. Cataloged identities and settings use stable canonical routes; newly observed identities or profile settings use deterministic provisional routes until reviewed. Every curated model carries a checked first-party release date and source shared by all of its profiles. - Read sourced analysis at
/blog, including the current AA Index versus cost snapshot analysis, open models on coding-agent benchmarks, how GPT-5.6 Luna changed one daily news-page cost, Terminal-Bench-Science, and why a high score still needs a holdout. - Inspect the current snapshot, methodology, provenance, full configuration table, and machine-readable distribution at
/data.
Each checked snapshot records its source, named version, retrieval time, and available revision or content fingerprint. The atlas catalog JSON links one compact JSON distribution per measured cohort; /data provides benchmark definitions, comparison rules, limitations, and richer source snapshots. Source-only guides have no invented observations or dataset download. AI Charts is an independent visualization and is not affiliated with the benchmark owners or model providers represented in the data.
AI Charts uses Bun 1.3.14 and Node.js 24.
bun install --frozen-lockfile
bun run devOpen http://localhost:3000. Analytics are disabled outside production on aicharts.io, so local development does not send PostHog events.
Run the complete local gate before opening a pull request:
bun run checkThe complete gate also requires Rust 1.97.1 with rustfmt and Clippy, pinned in rust-toolchain.toml. It validates generated files and the checked data contract, checks the local usage workspace and AI Charts skill helper, runs strict TypeScript and ESLint, executes example and property tests, and verifies the production build in a browser. The website's ordinary development/build commands do not invoke Rust.
The detailed usage dashboard adds UTC trends,
client/provider/model drilldowns, token composition, separate cost bases, source
coverage, and exact numeric exports. The aicharts stats command imports the
known local formats of the 55 sources in the pinned Tokscale parser registry.
--all reads 54 of them; it leaves out the 9router selector because Gajae-Code
already includes that channel. Run aicharts stats --list-clients for the list.
See detailed usage reports for commands, acquisition
requirements, and what parser support does and does not cover. The scheduled publisher guide covers
native refresh profiles, dry runs, failure recovery and a reversible macOS cutover.
The Rust workspace contains a local-only Codex/Claude Code usage reader, a closed numeric wire format, a private numeric SQLite ledger and matching TypeScript validation/rollups. Explicit collection can retain measurements across restarts with atomic source checkpoints and a pending-queue preview. The separate inspect command reads retained numeric totals without source scanning, writes or SQLite recovery. The inspect command does not enable sign-in, uploads, a public leaderboard or background collection. See the local usage guide for explicit source selection, private namespace keys and current measurement/recovery limitations.
The session usage view opens local numeric reports for session totals, model mix, and time breakdowns with their source coverage. Reports stay in the browser tab. See session collection and timing for the command, concurrent-work denominators and source coverage.
The website includes Hraness Accounts sign-in and device enrollment, each
switched on separately in production. The identity design
and activation runbook record dated production
evidence, the exact-deployment health check (bun run usage:deployment:verify)
and what still needs testing against real providers. The
claim inventory lists every support, platform and
live-status statement in this README and the usage guides with its evidence.
The AI Charts menu bar shows whether usage collection is working: when the last pass ran, when your usage last synced, and any failures in plain words, with the collector's error log one click away. It also opens the usage dashboard and lists your two newest outputs. Build it explicitly, then run it:
bun run menubar:build
bun run menubar:install
bun run menubarmenubar:install copies the release-built binary to
~/Library/Application Support/AI Charts/bin/aicharts-menubar atomically. The
launcher never compiles on startup: it uses that installed copy when present,
or a prebuilt checkout binary otherwise. A second copy exits with status 3 and
says AI Charts is already in your menu bar. Use bun run menubar:uninstall to
remove the installed copy.
To open it at login, run aicharts menubar install (or turn on "Open at
login" in its menu). macOS then shows a notice that aicharts-menubar can open
at login. aicharts menubar status says whether it is running and opens at
login, aicharts menubar start opens it now, and aicharts menubar uninstall
removes the login item. Nothing here runs launchctl; the login item takes
effect at your next login.
The menu reads two files from ~/.aicharts (or AICHARTS_HOME):
collector-status.json, which aicharts daemon --status-file ~/.aicharts/collector-status.json writes after each pass, and autosubmit-runtime/last-cycle.json, which each publishing
cycle writes. Both hold times, results and fixed error codes, never account
IDs, paths or session content.
The build uses the committed Cargo lockfile. Installation stages the replacement in a private temporary directory and refuses symlinked or externally writable managed directories. Launch also refuses symlinked or externally writable executables. The menu bar is not an updater or a privileged service.
The portable AI Charts skill retrieves one public benchmark cohort at a time with its version, units, configuration, source and coverage intact. Its dependency-free Node.js helper makes anonymous requests only to aicharts.io, bounds response sizes and pagination, and never opens local usage data. A separate local mode can explain a supplied numeric summary or use an already available reviewed aicharts inspect binary for an explicitly selected ledger. It never collects source logs or uploads usage.
The skill lives at skills/aicharts; load or install that directory through your agent's supported skill workflow. The source command is documented in the local usage guide; this repository does not yet distribute a signed native installer. The skill does not install a runtime, obtain keys or enable background collection.
The standalone benchmark helper and native CLI source offer free AI Charts product updates and optional paid development support after useful completed reads. These use the shared Hraness invitation preferences; they never change benchmark or numeric JSON stdout, send usage data, or open checkout automatically. Imported benchmark calls, CI, background collection, authentication, uploads and control commands stay quiet. The native change is source support, not a claim that an installed binary has been released.
Use node skills/aicharts/scripts/atlas.mjs support protocol --json for the local machine-readable handoff, or aicharts support protocol --json with a verified build exposing that command. The Node helper includes its reviewed shared runtime and needs no dependency installation. An agent presents a due invitation once at task closeout and records it only after persistent human-visible output. Git email is an optional unverified suggestion; signup requires explicit authorization and inbox confirmation, and payment requires human checkout approval. See the skill handoff contract. Set HRANESS_SUPPORT_AUDIENCE=off for a delegated child, HRANESS_SUPPORT=off to suppress incidental invitations, or use support dismiss to save the suite preference.
The data-refresh.yml workflow checks first-party release sources and OpenRouter discovery hourly, off the top of the hour. It checks Terminal-Bench 4, Terminal-Bench-Science 0.1, direct DeepSWE evidence, the lightweight Artificial Analysis Intelligence model snapshot, and the reasoning and multimodal atlas imports every four hours, then adds the heavier Artificial Analysis coding-agent import to one daily full run at 10:43 UTC. Manual runs can select release-only, benchmark-only, or full refreshes; the legacy discovery mode remains a combined non-AAI alias. Poll-metadata-only checks remain visible in Actions without creating a data pull request. It treats each importer as a separate failure domain:
- The registry has 30 lab-owned release sources from 24 labs: Anthropic, OpenAI, Google DeepMind, Meta, xAI, Mistral AI, Cohere, DeepSeek, Z.ai, Moonshot AI/Kimi, Alibaba/Qwen, MiniMax, ByteDance Seed, Microsoft AI, NVIDIA Nemotron, Amazon Nova, Baidu ERNIE, Tencent Hunyuan, Xiaomi MiMo, AI21 Labs, IBM Granite, Ai2, StepFun, and Cognition. Conservative URL parsing retains every parsed model in a multi-model release; unknown versioned model-family shapes enter the review queue as unresolved instead of disappearing, while reviewed irregular launches use exact route mappings. New canonical URLs drive discovery. Mutable source timestamps remain secondary evidence, and discovery never writes the reviewed official-date ledger.
- OpenRouter's public models API supplies a bounded 90-day identity radar and the model-ID catalog used to identify direct benchmark observations. Its listing timestamp is discovery metadata, not a claimed release date.
- Harbor Framework's official Terminal-Bench leaderboard repository supplies the current 4.0.0 snapshot. The importer pins one immutable commit, asserts the 4.0 leaderboard definition, and rejects duplicate configurations, incomplete trials, score arithmetic errors, version changes, or unsafe row loss.
- Terminal-Bench-Science's official 0.1 owner API supplies a separate scientific-workflow snapshot pinned to release
v0.1.0, its immutable commit, and exact Harbor dataset version. The importer retains five-domain results, validates resolution-rate and binomial-error arithmetic, records unpublished protocol fields as null, and preserves aggregate and domain costs without inventing reconciliation. - Artificial Analysis's public model-page Flight payload supplies the current v4.3.2 Intelligence Index efficiency snapshot, including native weighted output-only token and cost components per Index task. The importer cross-checks the page's exact ten-evaluation roster and every public JSON-LD leaderboard score against the payload. Future index versions require separate admission; historical v4.1.1 remains frozen and checked offline.
- Artificial Analysis supplies the AA Index, DeepSWE, Terminal-Bench v2.1, and SWE-Atlas-QnA observations in the separate coding-agent interactive chart and cards. Terminal-Bench v2.1 remains labeled and separate from 4.0; this heavier import refreshes daily.
- DataCurve's official DeepSWE v1.1 artifact supplies early harness-specific pass@1 evidence. Ambiguous model matches fail closed, unmatched models remain explicit, and every observation retains its harness, effort, run count, attempts, and source provenance.
- The first-party and OpenRouter radars are discovery-only, and direct DeepSWE observations are early-evidence-only. Missing values remain missing. None of these sources can invent another source's score, chart point, model card, or official release date.
- The reasoning and multimodal atlas importers retain explicit reviewed cohorts and exact evaluation versions. Mutable publisher endpoints have checked content fingerprints, not invented commit pins. New ARC model families, changed paper cohorts, and revised evaluation protocols require a reviewed admission change; a new release is not automatically a new chart point. See
docs/benchmark-sourcing-protocol.mdfor the precise source and selection contracts. - Arena media refreshes four separate overall cohorts from one owner-released CC BY 4.0 dataset revision; pagination, source dates, confidence intervals, and vote counts are checked. The Open ASR importer rechecks a reviewed immutable cohort, not an automatically expanding audio leaderboard. Both have independent failure reporting.
- The workflow can atomically update only its eleven allowlisted snapshots: first-party and OpenRouter release radars; Terminal-Bench and Terminal-Bench-Science; current Intelligence v4.3.2 and coding-agent data; direct DeepSWE evidence; reasoning and multimodal atlas data; Arena media; and Open ASR audio. Historical Intelligence v4.1.1 is not automation-owned. These snapshots pass the full project check before a dedicated pull request, required CI on its exact head, and a protected squash merge. Failed importers leave their last-known-good production snapshots in place.
- Official card dates live in the manually reviewed
data/model-release-dates.jsonledger, keyed by stable canonical model ID. Marketplace and sitemap timestamps never populate it or appear as official release dates.
The repository keeps default workflow-token permissions read-only and grants write capabilities only inside this workflow. GitHub's repository-level “Allow GitHub Actions to create and approve pull requests” setting must remain enabled so that the scoped token can open its data PR; the workflow never submits reviews. First-party sources refresh independently, so one outage retains that source's last-known-good slice while healthy sources continue. A new candidate updates the durable release-review issue before benchmark refreshes run. Dependency installation is retried, and any unhealthy run creates or updates a separate automation-health issue. Source-shape changes, suspicious data loss, failed required CI, and unmerged update PRs still fail closed for publication, leaving the last-known-good production snapshot in place.
Publication is bound to the CI run for the update PR's exact head commit. The workflow names that run, waits up to 25 minutes for its Required result, reconciles the branch once if main moves, and lets GitHub squash-merge when the check passes. Pull-request CI tests the merge with current main, so a red main fails the update PR without any fault in the refreshed data; the health issue then quotes the CI run and states which case applies, and the PR is closed so the next scheduled run rebuilds from current main. To retry sooner than the schedule, run gh workflow run data-refresh.yml -f mode=full (or benchmarks or releases) and read the run's step summary; the health issue closes itself after a healthy full run.
To refresh locally:
bun run releases:refresh
bun run first-party-releases:refresh
bun run terminal-bench:refresh
bun run terminal-bench-science:refresh
bun run aa-intelligence-v4-3:refresh
bun run data:refresh
bun run releases:reconcile
bun run deepswe:refresh
bun run atlas:reasoning:refresh
bun run atlas:multimodal:refresh
bun run atlas:arena-media:refresh
bun run atlas:audio:refresh
bun run checkReview the resulting data diff before committing it. The corresponding *:check commands are network-free validations of the committed snapshots. bun run releases:reconcile updates only OpenRouter radar benchmark statuses from the checked Artificial Analysis snapshot. Official-date and first-party candidate review statuses remain reviewed edits; scheduled discovery preserves them.
PostHog is initialized only in production on the canonical AI Charts domains. The browser configuration is cookieless and privacy constrained:
- no person profiles, persistent identifiers, autocapture, session replay, surveys, heatmaps, or feature flags;
- memory-only persistence, Do Not Track support, masked text and element attributes;
- page-view, page-leave, Core Web Vitals, and typed allowlisted product events only.
Product events cover chart and model-card interaction, delegated public-link clicks, and footer signup requests. They contain controlled enum-like properties, never raw URLs, query strings, hashes, link text, visitor-entered values, or model-level user data. Every event receives a grouped page classification plus a bounded public article or model-card content ID. The complete event and privacy contract lives in docs/analytics-instrumentation.md.
The optional AI Charts mailing list is separate from product use and every
other Hraness audience. Its footer sends the entered email address, the
aicharts audience, the form source, and a short-lived Cloudflare Turnstile
proof to Hraness Accounts at account.hraness.com. Cloudflare verifies the
anti-abuse proof. Hraness Accounts records dated consent, and Resend sends the
confirmation and subscribed messages from news.hraness.com. The address is
not subscribed until its confirmation link is used. After confirmation, each
newsletter message includes an AI Charts-specific unsubscribe link, which does
not change subscriptions to other Hraness products.
The durable positioning, search-intent map, technical invariants, event schema, baseline, and review cadence live in docs/seo-strategy.md. Search Console measures impressions, queries, clicks, click-through rate, and search position. PostHog measures acquisition and qualified engagement after a visitor arrives.
Copy .env.example to .env.local to exercise configuration. The public project token and ingest host are safe browser variables. POSTHOG_API_KEY is a private build credential used only to upload production source maps; never expose it through a NEXT_PUBLIC_ variable.
The production site is deployed from main with Vercel. The repository-level vercel.json pins the Bun install and build commands. Generated automation/model-data-refresh-<run>-<attempt> branches skip their disposable Vercel Preview: the data-refresh workflow runs the complete application check before publishing the branch, then requires CI on that commit before merge. Production and ordinary branch previews still build. Missing or unrecognized Vercel system values also continue the build.
To inspect a hosted Preview for a generated refresh, redeploy it and uncheck Use project's Ignore Build Step in Vercel's redeploy dialog. This is Vercel's documented manual override.
Configure these environment variables in Vercel:
| Variable | Scope | Purpose |
|---|---|---|
NEXT_PUBLIC_POSTHOG_KEY |
Production, Preview | Public PostHog project token |
NEXT_PUBLIC_POSTHOG_HOST |
Production, Preview | Regional PostHog ingest host |
NEXT_PUBLIC_HRANESS_MAILING_TURNSTILE_SITEKEY |
Production, Preview | Public hostname-restricted Turnstile widget key for the aicharts audience |
POSTHOG_API_KEY |
Production | Private key for source-map upload |
POSTHOG_PROJECT_ID |
Production | Numeric PostHog project ID |
POSTHOG_UI_HOST |
Production | https://us.posthog.com or https://eu.posthog.com |
Production source maps are uploaded only when all private build settings and Vercel's commit SHA are present, then removed from the deployment output.
app/contains the App Router chart, sourced benchmark notes, metadata, error states, and product styling.components/contains the interactive chart, model cards, update timeline, linked summaries, sharing, export, and local UI primitives.lib/contains strict data and model-card boundaries, chart math, deterministic layout, analytics events, and property tests.data/contains the checked benchmark, discovery radar, official model-release ledger, and model-card catalog snapshots.scripts/contains the guarded benchmark and release-radar refreshes and deterministic color generator.styles/contains the portable plain-publication styles used by the benchmark notes.docs/contains the current search, measurement, and engineering strategy..github/workflows/contains CI, hourly release discovery, four-hour benchmark checks, and the daily full refresh.
The application code is available under the MIT License. The repository also contains normalized public facts sourced from Harbor Framework's Terminal-Bench, Terminal-Bench-Science, Artificial Analysis model and coding-agent leaderboards, the OpenRouter Models API, the DataCurve DeepSWE leaderboard, and provider-owned release sources. The MIT license does not grant rights to third-party data, names, logos, or trademarks; see NOTICE.md.
GitHub can generate a citation from CITATION.cff. Cite AI Charts for this software or its visualization method and cite the named benchmark owner for measurements. Cite OpenRouter only for discovery and model-identity metadata, and cite the linked provider source for an official release date. Include the source URL, version, retrieval date, metric, harness, effort, and trial policy when a claim depends on a snapshot.
Contributions are welcome. Start with CONTRIBUTING.md, and report security issues through the process in SECURITY.md.