Skip to content

linny006/agent-eval-harness

Repository files navigation

Agent Eval Harness

Live, open-source benchmark for comparing AI coding agents on real GitHub issues

Stars Last Commit Items Updated

⭐ Star this repo to bookmark — fresh data every 15 minutes

English · 中文 · 日本語 · 한국어 · Español · Português


💡 What is this?

A standardized benchmark suite that runs coding agents against live, real-world GitHub issues with reproduction steps. Unlike static academic benchmarks, it outputs a weekly-updated public leaderboard, enabling developers to compare agents like OpenCode, Codex, and Claude Code in realistic scenarios.

This list is auto-updated every 15 minutes by a GitHub Actions cron. Each commit reflects a real change in the upstream data source — new items added, expired items removed — so you can rely on what you see being current.


📋 Current Items

⏰ Last updated: 2026-07-26 03:30 UTC

Data source: GitHub Search API

The table below is rewritten on every cron tick. Star the repo to bookmark.

# Name Lang Updated Description
1 Arize-ai/phoenix 10730 Python 2026-07-26 AI Observability & Evaluation
2 saddled-panicattack529/idea-evaluation-pipeline 0 2026-07-26 Streamline research idea evaluation for finance and economics to reach top journal quality using an iterative, AI-assist
3 Kondwani10/Origin-Continuum 0 2026-07-26 🌐 Define and explore the Origin ↔ Continuum framework, ensuring proper attribution and continuity in dependency relation
4 Sans-cell-art/-Project-Phoenix-The-E-Waste-Supercomputer- 0 2026-07-26 ♻️ Transform e-waste into a powerful, low-cost cloud operating system, unlocking computing potential and promoting resou
5 bhavya7995/AI_governance 2 PowerShell 2026-07-26 🤖 Streamline AI-assisted development with a governance kit for rules, enforcement, and decision-making, ensuring speed a
6 promptfoo/promptfoo 23598 TypeScript 2026-07-26 Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, C
7 sammyjdev/gnomon-eval 0 Python 2026-07-26 Honest RAG evaluation harness: judge metrics with confidence intervals, cost and latency first-class, offline-first.
8 IonDen/mlx-quant-fidelity 2 Python 2026-07-25 Measure MLX quantization quality loss — KL divergence, perplexity, top-token agreement for KV cache and weights
9 vitorwilher/copom-rag-service 0 Python 2026-07-25 Serviço de RAG sobre atas do Copom + boletim Focus, servido como API (FastAPI+Docker), com eval harness (golden set, LLM
10 tkarim45/multi-agent-eval 0 Python 2026-07-25 Does multi-agent beat single-agent? Benchmarks a planner→workers→critic system vs single-agent on quality/cost/latency —
11 tkarim45/agent-eval-harness 0 Python 2026-07-25 Agent eval harness — measure task success, tool-call accuracy, step efficiency, and cost for tool-using LLM agents (Clau
12 UiPath/coder_eval 107 Python 2026-07-24 Evaluate & benchmark AI coding agents and Claude Code skills — sandboxed, reproducible YAML eval suites for Claude Code,
13 truera/trulens 3459 Python 2026-07-24 Evaluation and Tracking for LLM Experiments and AI Agents
14 ipezygj/evalgate 2 Python 2026-07-24 Dependency-free statistical checks for AI eval claims: multiple-comparisons correction, LLM-judge bias tests, and leave-
15 NoesisVision/nasde-toolkit 11 Python 2026-07-24 CLI for benchmarks & evals of AI coding agents — on tasks you already understand, using your Claude / Codex / Gemini ind
16 10div10/agent-trajectory-evaluator 0 Python 2026-07-24 LLM-as-judge evaluator for scoring AI agent trajectories on task completion, tool correctness, and efficiency
17 jeremylongshore/j-rig-skill-binary-eval 0 TypeScript 2026-07-24 Binary-criteria evaluation harness for Claude skills with planned extension to plugins, agents, and MCP servers. Score e
18 verifywise-ai/verifywise 322 TypeScript 2026-07-26 Complete AI governance and LLM Evals platform with support for EU AI Act, ISO 42001, NIST AI RMF and 20+ more AI framewo
19 valbaudo/awf 1 Go 2026-07-23 Run agents you don't babysit, and trust the result. awf runs agentic workflows with independent gates that check every s
20 iZenDeveloper/auditai 0 Python 2026-07-23 Developer-first LLM/RAG safety audits for CI/CD — faithfulness, relevancy, prompt injection (BYOK OpenAI/xAI). pip insta
21 lftherios/session-link 0 Go 2026-07-23 A local-first CLI that turns any LLM session into a permanent URL you can inspect, share, and revisit.
22 Giskard-AI/giskard-oss 5713 Python 2026-07-23 🐢 Open-Source Evaluation & Testing library for LLM Agents
23 Eval-core/evalcore 14 Rust 2026-07-26 Snapshot testing for LLM apps and agents, built to run locally and block regressions in CI.
24 cklxx/ckl-bench 1 HTML 2026-07-22 ckl's personal benchmark for doc writing, infra code, and paper reading — one-click evaluation of the latest models via
25 monospaceai/evaldata 2 Python 2026-07-22 Evaluate AI-generated SQL with pytest.
26 RudrenduPaul/memtrust 0 Python 2026-07-22 Agent memory backends each publish their own benchmark numbers, on different tests, measured different ways. memtrust ru
27 homemade-software-inc/completion-kit 1 Ruby 2026-07-21 Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and c
28 lehigh-university-libraries/htr 2 Go 2026-07-21 Handwritten Text Recognition llm eval tool
29 ozlar34/job-match-radar 1 Python 2026-07-21 Self-hosted n8n + Supabase pipeline that scrapes LinkedIn and a watchlist of company ATS endpoints, scores listings agai
30 mohanish3/eval-atlas 0 TypeScript 2026-07-21 LLM eval dashboard for multiple choice and open ended prompts. Upload JSON or JSONL eval sets, run against OpenAI, Anthr
31 SFX-TECH/sfx-lead-intelligence 0 2026-07-20 SFX Lead Intelligence Command Center: local-LLM hub plus lead dashboard, quality lifted 61 to 99 percent via a ground-tr
32 AshC2004/llm-eval-harness 0 Python 2026-07-18 Automated groundedness scoring, hallucination detection, and regression tracking for RAG and agent outputs
33 pdxlab/trustmodel-mcp-server 0 TypeScript 2026-07-25 TrustModel MCP Server — trust evaluation, red-team, and governance for AI agents via the Model Context Protocol. npm: @t
34 sahuno/confounded 0 R 2026-07-15 A gym for scientific judgment. AI agents render fatally flawed analyses at publication quality — we measured whether any
35 vijayarjun7/flipcheck 0 TypeScript 2026-07-14 LLM self-consistency and sycophancy diagnostic — catches contradictions and false-premise acceptance hiding behind confi
36 QuesmaOrg/BinaryAudit 97 Shell 2026-07-14 An open-source benchmark for evaluating AI agents' ability to find backdoors hidden in compiled binaries.
37 harnexa/nexa-gauge 40 Python 2026-07-20 An graph-eval framework for LLM's
38 multivon-ai/multivon-eval 8 Python 2026-07-13 Practical LLM evaluation for teams that ship to production. Deterministic + LLM-as-judge evaluators, dataset support, CI
39 attogram/ollama-multirun 16 Shell 2026-07-12 Run a prompt against all, or some, of your models running on Ollama. Creates web pages with the output, performance stat
40 nishant20/testgen-eval 0 TypeScript 2026-07-10 An agent-eval harness that gates and judges testgen's generated test suites via MCP — deterministic checks plus an LLM-j
41 ahwurm/localshift 3 Python 2026-07-09 Migrate headless Claude/AI workloads to local LLMs with a derived, per-workload quality eval — cron job in, zero-margina
42 coffee-converter/onthemoney 0 Python 2026-07-14 An AI agent that answers questions about U.S. campaign money against real FEC data. Source-cited, graded on a public eva
43 MihirBindu/dynamo-log-report-fix 0 Python 2026-07-09 A broken Terminal-Bench 2 (Harbor) task, repaired: reproducible env, solution-leak removed, gameable verifier replaced w
44 christianmacion26/judge-harness 0 Python 2026-07-08 LLM-as-judge validated vs humans (Cohen's κ), with position-bias exposed.
45 plwslpld-arch/loopward 1 TypeScript 2026-07-07 Stress-test the tool-routing decision inside an agent loop: audit confusable tools, red-team routing robustness, compare
46 ContextJet-ai/skillvitals 1 Python 2026-07-06 Vital signs for your agent skills: measure whether a skill triggers when it should and whether it actually helps, on a c
47 nikolas-sapa/sigeval 1 Python 2026-07-05 Statistically rigorous LLM evaluation framework for Python — pytest for LLMs that isn't flaky. Treats every eval as a pr
48 Sushant-Dagar/agent-eval-harness 0 Python 2026-07-04 CI-integrated eval harness for LLM agents — intent accuracy, retrieval precision/recall/MRR, and hallucination detection
49 lokesh75-kank/agenteval 0 TypeScript 2026-07-03 Reliability and audit-evidence testing for LLM agents - wrap any agent, assert behavior, measure determinism, check grou
50 Ruthwik-Data/finrag-eval 0 Python 2026-07-02 Local RAG eval on real SEC 10-Ks that catches confident financial hallucinations — and surfaced a metric bug now merged
51 hydrangeas20/safetylens 0 Jupyter Notebook 2026-07-01 Empirical study of evaluation robustness in large language models. Compares benchmark style evaluation prompts with real
52 kilocommits/campaign-eval-harness 0 Python 2026-06-30 An LLM-as-judge harness that scores AI-generated campaign phone scripts against a weighted quality rubric with a real Ha
53 Merchantlee99/myrealtrip-cancel-recovery-eval 0 Python 2026-06-30 Codex eval plugin for auditing cancel-recovery replacement-product exposure policies.
54 reaatech/agent-eval-harness 0 TypeScript 2026-07-20 End-to-end agent evaluation — trajectory eval, tool-use correctness, cost-per-task, latency budgets, regression suites w
55 TheAnacondA57/BidAgent 1 Python 2026-06-29 RAG agentique sur des documents de concession télécom publique (DSP/RIP), pensé eval-first et contrôlé en CI.
56 chquandogong/mission-spec 0 TypeScript 2026-06-29 Mission Spec — AI 에이전트 워크플로를 위한 task contract layer
57 G59-Toneli/dataset-eval-skill 1 JavaScript 2026-06-25 A Claude skill for building golden sets to test AI systems — matching, RAG, LLM-as-judge — without false greens.
58 thewonderofyou777z-dot/tjoe-reviewkit 0 Python 2026-06-25 TjoeReviewKit:tjoe 的本地离线工作流复盘检查工具;不运行任务、不联网、不接管工具调用、不采集生产日志
59 RaphaelFakhri/reagent 0 Python 2026-06-24 Tool-using ReAct + RAG agent (enterprise assistant) with a built-in evaluation harness scoring accuracy, tool selection,
60 melody-ling-L/eval-resume 0 HTML 2026-06-24 中文 LLM 简历改写诚实度 benchmark:20 脱敏简历 × 3 模型 × 4 维度 · promptfoo + LLM-as-judge · 含在线报告
61 gashel01/evalmcp 0 Python 2026-06-22 Evaluation for AI agents — judge-based scoring and native RAG metrics (faithfulness, relevancy, context precision/recall
62 anejakartik/evalstack 0 TypeScript 2026-06-22 Open-source LLM evaluation framework — drop-in SDK + CI plugin. LLM-as-judge, regression detection, free + self-hostable
63 jmpei/nl2sql-agents 0 Python 2026-06-21 NL→SQL multi-agent pipeline (LangGraph + Claude) with deterministic SQL-injection guardrails and golden-set eval.
64 TeracAI/svg-arena 0 TypeScript 2026-06-20 A forkable example of the human-in-the-loop model-improvement loop: AI generates, humans judge via the Terac MCP, you im
65 Ayubjon/refusal-radar 1 JavaScript 2026-06-20 Zero-dependency detector and classifier for LLM refusals, deflections, and capability disclaimers — CLI + library with s
66 melody-ling-L/judgebuddy 0 HTML 2026-06-20 Single-file labeling tool for LLM-as-judge calibration. Three-pane comparison + multi-dim scoring. Zero deployment.
67 ramenprotokol/hallucination-hunter 0 Python 2026-06-20 Detect & score LLM hallucinations by groundedness — labeled data, precision/recall/F1, runs offline with no API key. Plu
68 gititya/Quality-Agency-support 0 Python 2026-07-09 Five local QA judges that review B2B and B2C customer-support replies, catch the risky parts, and explain what to fix.
69 tushariitr-19/assay 4 Go 2026-06-17 Framework-agnostic evaluation harness for Go — test your MCP servers and AI agents with scored, CI-ready checks.
70 jedobe/skill-evaluator 0 Python 2026-06-17 Score any Claude Code skill against a research-backed rubric derived from the top 9 most-starred skill repos on GitHub
71 ALEX-nlp/OpenSkillEval 12 Python 2026-06-15 OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents
72 mpuodziukas-labs/eval-harness-template 0 Python 2026-06-14 Eval harness template for LLM systems: golden regression, LLM-as-judge, invariants
73 mizcausevic-dev/agent-eval-arena 0 TypeScript 2026-07-20 Agent and LLM evaluation harness — golden datasets, multi-scorer execution, regression detection across model versions,
74 ejentum/eval 4 Python 2026-06-11 A/B evaluate any LLM task with and without Ejentum cognitive injection. n8n workflow + TypeScript module.
75 akanjilal-work/agent-eval-harness 0 Python 2026-06-10 A lightweight harness to test agent behaviour (tool-call correctness, injection refusal, cost ceilings) before deploymen
76 karlmehta/trustmodel-mcp 0 TypeScript 2026-06-10 TrustModel MCP Server — trust evaluation, red-team & governance for AI agents via the Model Context Protocol. Public can
77 alyssadata/continuity-keys 1 2026-06-08 Continuity Keys: tests for “same someone” returns. Behavioral identity consistency under pressure. Origin (Alyssa Solen)
78 reaatech/classifier-evals 0 TypeScript 2026-07-22 Offline classifier evaluation harness — dataset loader, confusion matrices, LLM-as-judge with cost accounting, regressio
79 reaatech/rag-eval-pack 0 TypeScript 2026-07-20 RAG evaluation toolkit — faithfulness, answer relevance, context precision/recall, cost accounting, CI gates. Pairs with
80 Juanllenato/llm-eval-harness 0 Python 2026-06-03 A small, production-minded evaluation and observability harness for LLM/RAG features. Runs offline or live, gates CI on
81 Victor-David-Medina/llm-eval-harness 0 Python 2026-06-03 LLM evaluation harness that gates quality in CI: golden datasets, regression detection, grounding and faithfulness check
82 thestio/thest-eval 0 Python 2026-06-02 The CI regression gate and governance-evidence layer for LLM systems — zero-dependency, vendor-neutral, offline.
83 monkeyin92/voice-agent-testops 0 TypeScript 2026-06-01 Regression testing for voice agents: scripted conversations, safety assertions, CI-ready reports.
84 fastxyz/skill-optimizer 71 TypeScript 2026-05-28 Benchmark, evaluate, and optimize skills to ensure reliable performance across all LLMs
85 ajmeese7/local-llms 1 Python 2026-05-27 Use local Large Language Models for production use cases, and perform benchmarking for task-specific performance evaluat
86 rogue-socket/focusgroup 0 Python 2026-07-13 Persona-driven dynamic testing for conversational AI products. Focus groups for your agents.
87 sanya2025/edututor-eval 0 Python 2026-05-21 A lightweight evaluation framework for AI tutoring responses, built for education-focused LLM systems
88 Alexanderk30/context-override-resistance 0 Python 2026-05-19 RL-style eval measuring intent/action divergence in frontier agents: model acknowledges a correction, then acts on the s
89 GiuseppeSp/n8n-customer-interview-synthesizer 0 2026-05-19 Multi-agent customer-interview synthesis pipeline in n8n with LLM-as-judge eval, Slack human-in-the-loop approval, and d
90 gmitt98/fieldtest 0 Python 2026-05-16 LLM evaluation framework — define what correct, well-formed, and safe means before you measure
91 verifywise-ai/plugin-marketplace 3 TypeScript 2026-05-15 VerifyWise AI Governance Plugin Marketplace
92 AI-QL/tuui 1153 TypeScript 2026-05-14 A desktop MCP client designed as a tool unitary utility integration, accelerating AI adoption through the Model Context
93 prompt-foundry/typescript-sdk 6 TypeScript 2026-05-13 The prompt engineering, prompt management, and prompt evaluation tool for TypeScript, JavaScript, and NodeJS.
94 prompt-foundry/python-sdk 8 Python 2026-05-13 The prompt engineering, prompt management, and prompt evaluation tool for Python
95 Ruthwik-Data/mechanictrust 0 2026-05-11 AI product case study for trust, pricing transparency, and explainable diagnosis in auto repair.
96 SAY-5/eval-observability 0 Python 2026-05-10 Python LLM eval framework with full OTel tracing, structured logs, and daily Welch's-t-test regression detection persist
97 Ruthwik-Data/self-improving-prompt-agent 0 Python 2026-05-10 Automated prompt-optimization loop (edit→evaluate→keep) with an LLM judge — the RLHF/AutoML pattern in ~100 lines. Score
98 SAY-5/genai-eval 0 Python 2026-05-07 Multilingual GenAI evaluation service across 5 task types and 3 languages, with regression-trend dashboard
99 advitrocks9/java2kotlin-eval 0 Kotlin 2026-05-05 headless Java→Kotlin converter pipeline -- IntelliJ plugin + Kotlin eval harness
100 HumphreySun98/repoagentbench 30 Python 2026-04-30 SWE-bench for your codebase — mine your merged PRs into local, contamination-free coding-agent benchmarks. Adapters: cla

🔍 How it works

Every 15 minutes, a GitHub Action runs tracker.py. That script:

  1. Fetches the latest state from GitHub Search API.
  2. Diffs against data/items.json (the previous snapshot).
  3. Rewrites the table above between the <!-- TRACKER_TABLE_* --> markers.
  4. Commits feat: +N added, -M removed (timestamp) if anything changed.

No external services. No paid APIs. Just a public data source and a free GitHub Action.


🤝 Contributing

See CONTRIBUTING.md — usually you don't need to: the tracker keeps itself current. If you spot a data-source bug or want to suggest a new column for the table, open an issue.


🔗 Related live trackers

If you find this useful, you might also like these other auto-updated trackers from the same maintainer — same mechanism, different upstream:


📜 License

MIT — see LICENSE.

More from linny006

  • Awesome Agent Skills — Curated, auto-updated awesome-list of vetted AI agent skills with quality ratings for Claude, GPT, and open-source agents (⭐ 0)

  • Agent Skills Daily Tracker — Real-time tracking of every new GitHub 'skills' repo to capture the AI agent skill ecosystem trend (⭐ 0)

  • Agent Eval Harness — Live, open-source benchmark for comparing AI coding agents on real GitHub issues (⭐ 0)

  • Prompt Tools Live — Live-updating tracker of prompt engineering tools, libraries, and techniques — refreshed every 15 minutes (⭐ 0)

  • LLMOps Radar — Live index of the newest LLMOps tooling — track what's shipping in LLM observability and deployment (⭐ 0)