Turn any git repository into a verifiable RL environment for coding agents.
Website · Documentation · Architecture · Quickstart · Why
Frontier coding agents are now trained with reinforcement learning on real software-engineering work: fix this bug, add this feature, refactor this module, make this function faster. The models are not the bottleneck any more. The environments are. Every task needs a reproducible repository snapshot, an executable oracle, a golden solution that proves it is solvable, protection against reward hacking, and a gym-style interface, and today most teams hand-build all of that per project.
repogym makes it a task.yaml and one command.
git clone https://github.com/shi1720/repogym.git
cd repogym
python3 -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e ".[dev]"
repogym validate tasks --repeat 2 # prove every task is sound and deterministic
repogym run tasks --agent golden --report report.htmlfrom repogym import Task, RepoEnv, Read, Edit, Submit
env = RepoEnv(Task.load("tasks/bugfix-median")) # from a checkout of this repo
obs = env.reset() # problem statement + file tree
obs = env.step(Read("stats.py"))
obs = env.step(Edit("stats.py",
" return ordered[mid]\n return ordered[mid]\n",
" return ordered[mid]\n return (ordered[mid - 1] + ordered[mid]) / 2\n"))
obs = env.step(Submit())
print(obs.reward, obs.done) # 1.0 True| Declarative tasks | A task is a folder: task.yaml + repo/ snapshot (or git URL + commit) + optional hidden/ tests + solution.patch. Four task types: bugfix, feature, refactor, perf. |
| Verified by construction | repogym validate proves hidden tests are not leaked, fail_to_pass tests fail at baseline, pass_to_pass tests pass at baseline, doing nothing scores 0, the golden solution scores 1, and grading is deterministic. |
| Gym-style environment | reset() / step(action) with Read, Write, Edit, Run, Test, ListFiles, Submit. Sparse reward in [0, 1] on submit, optional dense shaping from visible tests, step limits, command allowlist, every episode recorded as a JSON trajectory. |
| Composable graders | TestGrader (JUnit XML from any runner), PerfGrader (benchmark speed-up, log-scaled partial credit), QualityGrader (cyclomatic complexity, function length), ConstraintGrader (scope, forbidden/required patterns, lines changed). Objectives earn weighted credit, gates zero the score when they fail; a submission passes only if every component passes. |
| Anti reward-hacking | Protected test files are restored before grading, hidden tests are injected only at grade time, benchmarks are timed externally, git metadata lives outside the agent's tree and grading fails closed if it is destroyed, an empty submission always scores 0, and secrets are scrubbed from the sandbox environment. |
| Task mining | repogym mine <repo> turns real bug-fix commits into verified tasks. repogym mutate <repo> synthesises verified bug-fix tasks by injecting AST-guided faults that the test-suite catches. |
| Any agent | Built-in golden, noop, ShellAgent (plug in Claude Code, Aider, Codex, OpenHands, any CLI), ClaudeAgent (Anthropic SDK tool loop), or subclass Agent in ten lines. |
| Sandboxes | LocalSandbox (zero setup) or DockerSandbox (pinned image, network off, CPU/memory limits). |
| Reports & export | Rich terminal tables, leaderboards, a standalone HTML report, JSONL export of tasks (SWE-bench-compatible fields) and trajectories (chat-style messages for SFT/RL pipelines). |
| Language-agnostic | Anything that emits JUnit XML works. Ships example tasks in Python and JavaScript. |
Install from GitHub. This release is not published on PyPI. The checkout includes the example tasks; their Python tests need pytest (included in the dev extra), and the JavaScript example needs Node.js. Use local execution only for trusted code.
git clone https://github.com/shi1720/repogym.git
cd repogym
python3 -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e ".[dev]"
repogym doctor # which toolchains are available
repogym list tasks # the five example tasks
repogym validate tasks -v # every check, per task
repogym run tasks --agent noop # baseline: every task scores 0
repogym run tasks --agent golden -w 4 --record trajectories -o results.json --report report.htmlPlug in a real agent:
# any CLI agent, {prompt} is replaced with the shell-quoted problem statement
repogym run tasks --agent "shell:claude -p {prompt} --allowedTools Edit,Write,Bash"
repogym run tasks --agent "shell:aider --message {prompt} --yes"
# the built-in Anthropic tool-loop agent
python -m pip install -e ".[anthropic]" && export ANTHROPIC_API_KEY=...
repogym run tasks --agent claude --record trajectories
# your own agent class
repogym run tasks --agent my_package.agents:MyAgentGenerate tasks instead of writing them:
repogym mutate ./my-project -n 20 --out tasks/mutants # synthetic, verified bug-fix tasks
repogym mine ./my-project --limit 500 -n 30 --out tasks/mined # real fixes from git history
repogym validate tasks/mined # they come out validated, but check anyway
repogym export tasks -o tasks.jsonl # SWE-bench-style datasettasks/bugfix-median/
├── task.yaml # spec: statement, tests, constraints, grading, env limits
├── repo/ # repository snapshot the agent sees (or git url + commit)
├── hidden/ # tests injected only at grading time
└── solution.patch # golden solution (or solution/ directory)
id: bugfix-median
title: "median() returns the wrong value for even-length input"
type: bugfix # bugfix | feature | refactor | perf
difficulty: easy
language: python
problem_statement: |
`stats.median()` should return the mean of the two middle values ...
tests:
command: "python -m pytest -q --junitxml={junit}" # any JUnit-emitting runner
hidden: [tests/test_hidden_median.py]
fail_to_pass: [tests/test_stats.py::test_median_even, tests/test_hidden_median.py::test_median_two_values]
pass_to_pass: [tests/test_stats.py::test_mean, tests/test_stats.py::test_median_odd]
constraints:
allowed_files: ["stats.py"]
forbidden_patterns: ["import statistics"]
grading:
weights: {tests: 1.0}
solution:
patch: solution.patch
env:
max_steps: 20Refactor and performance tasks are just as declarative:
# refactor: behaviour preserved (tests are a gate), complexity limits are the objective
constraints:
paths: ["orders.py"]
max_cyclomatic_complexity: 8
max_function_length: 30
# perf: benchmark timed externally; reward is log-scaled up to the target speed-up
perf:
command: "python bench.py"
min_speedup: 3.0
runs: 3submit ──▶ snapshot patch ──▶ ConstraintGrader ──▶ TestGrader ──▶ QualityGrader ──▶ PerfGrader ──▶ score
gate: scope, restore protected complexity / wall-clock objectives weighted,
patterns, size tests, inject function length speed-up any failed gate = 0
hidden, JUnit XML
testsscore = fraction offail_to_passtests passing (or 0/1 without partial credit); any brokenpass_to_passtest zeroes it.perfscore = log-scaled speed-up between a noise floor andmin_speedup, measured with interleaved baseline/candidate runs timed from outside the process so the agent cannot fake it.qualityscore (refactor tasks) = fraction of complexity/length limits met.- Components are objectives (weighted credit) or gates (no credit, zero the score when they fail): scope/pattern constraints and regression-only tests are gates, so an empty submission always scores 0.
from repogym import Task, RepoEnv, load_tasks, run_suite, GoldenAgent, validate_task
# validate
report = validate_task(Task.load("tasks/bugfix-median"), repeat=3)
assert report.passed, report.failures
# your own agent
from repogym import Agent, Submit
class MyAgent(Agent):
name = "mine"
def solve(self, env, obs):
while not obs.done:
action = my_policy(obs.text) # -> Action / dict
obs = env.step(action)
return obs
results = run_suite(load_tasks("tasks"), MyAgent, workers=4, record_dir="trajectories")Trajectories are the training-data product. Each episode is a JSON file with every
(action, observation, reward) tuple plus the final grade and patch; repogym export --trajectories trajectories -o sft.jsonl --only-passed turns them into chat-style
records.
Existing options are either benchmarks or platforms. SWE-bench, SWE-Gym and
R2E-Gym are datasets with bespoke harnesses; environment platforms (Prime Intellect,
HUD, Mechanize and friends) are hosted products. What was missing is the small,
open, local-first library layer: the FastAPI-shaped thing you pip install,
point at your own repository, and get a validated environment out of. repogym is
that layer:
- Verification is a first-class command, not a notebook you run once. Tasks
rot as repositories move;
repogym validatein CI catches it. - The reward is defensible. Hidden tests, protected files, scope rules and complexity limits are declared next to the task and enforced on every grade.
- Task supply scales. Mining git history and mutation-based synthesis turn one well-tested repository into hundreds of executable tasks, each with a golden patch, with no LLM in the loop.
- It is agent- and language-agnostic. JUnit XML as the contract means pytest,
node --test, cargo-nextest, jest, gradle and go test all work, and the agent can be a Python function, an LLM tool loop or an external CLI.
- Website · Docs
- Architecture: components, data flow, design decisions
- Authoring tasks · task.yaml reference · Grading · Environment API · Agents · Mining · CLI
- Contributing · Security · Changelog
0.1.0. The core API (Task, RepoEnv, graders, agents, miners, CLI) is
covered by 125 tests at ~94% statement coverage, and the shipped tasks are validated in CI
on Linux and macOS across Python 3.9 to 3.13. Roadmap: OpenEnv/Gymnasium adapters,
an async batched environment server, Rust/Go/Java example tasks, LLM-authored
problem statements for mined tasks, and per-task Docker image builds.
MIT © Shivam Gupta