Skip to content
shi1720Public

About

Turn any git repository into a verifiable RL environment for coding agents. Declarative tasks, gym-style env, composable graders, task validation and mining.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

repogym

repogym

Turn any git repository into a verifiable RL environment for coding agents.

CI Python MIT Docs

Website · Documentation · Architecture · Quickstart · Why


Frontier coding agents are now trained with reinforcement learning on real software-engineering work: fix this bug, add this feature, refactor this module, make this function faster. The models are not the bottleneck any more. The environments are. Every task needs a reproducible repository snapshot, an executable oracle, a golden solution that proves it is solvable, protection against reward hacking, and a gym-style interface, and today most teams hand-build all of that per project.

repogym makes it a task.yaml and one command.

git clone https://github.com/shi1720/repogym.git
cd repogym
python3 -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e ".[dev]"
repogym validate tasks --repeat 2   # prove every task is sound and deterministic
repogym run tasks --agent golden --report report.html
from repogym import Task, RepoEnv, Read, Edit, Submit

env = RepoEnv(Task.load("tasks/bugfix-median"))  # from a checkout of this repo
obs = env.reset()                                   # problem statement + file tree
obs = env.step(Read("stats.py"))
obs = env.step(Edit("stats.py",
                    "        return ordered[mid]\n    return ordered[mid]\n",
                    "        return ordered[mid]\n    return (ordered[mid - 1] + ordered[mid]) / 2\n"))
obs = env.step(Submit())
print(obs.reward, obs.done)                         # 1.0 True

What you get

Declarative tasks A task is a folder: task.yaml + repo/ snapshot (or git URL + commit) + optional hidden/ tests + solution.patch. Four task types: bugfix, feature, refactor, perf.
Verified by construction repogym validate proves hidden tests are not leaked, fail_to_pass tests fail at baseline, pass_to_pass tests pass at baseline, doing nothing scores 0, the golden solution scores 1, and grading is deterministic.
Gym-style environment reset() / step(action) with Read, Write, Edit, Run, Test, ListFiles, Submit. Sparse reward in [0, 1] on submit, optional dense shaping from visible tests, step limits, command allowlist, every episode recorded as a JSON trajectory.
Composable graders TestGrader (JUnit XML from any runner), PerfGrader (benchmark speed-up, log-scaled partial credit), QualityGrader (cyclomatic complexity, function length), ConstraintGrader (scope, forbidden/required patterns, lines changed). Objectives earn weighted credit, gates zero the score when they fail; a submission passes only if every component passes.
Anti reward-hacking Protected test files are restored before grading, hidden tests are injected only at grade time, benchmarks are timed externally, git metadata lives outside the agent's tree and grading fails closed if it is destroyed, an empty submission always scores 0, and secrets are scrubbed from the sandbox environment.
Task mining repogym mine <repo> turns real bug-fix commits into verified tasks. repogym mutate <repo> synthesises verified bug-fix tasks by injecting AST-guided faults that the test-suite catches.
Any agent Built-in golden, noop, ShellAgent (plug in Claude Code, Aider, Codex, OpenHands, any CLI), ClaudeAgent (Anthropic SDK tool loop), or subclass Agent in ten lines.
Sandboxes LocalSandbox (zero setup) or DockerSandbox (pinned image, network off, CPU/memory limits).
Reports & export Rich terminal tables, leaderboards, a standalone HTML report, JSONL export of tasks (SWE-bench-compatible fields) and trajectories (chat-style messages for SFT/RL pipelines).
Language-agnostic Anything that emits JUnit XML works. Ships example tasks in Python and JavaScript.

Quickstart

Install from GitHub. This release is not published on PyPI. The checkout includes the example tasks; their Python tests need pytest (included in the dev extra), and the JavaScript example needs Node.js. Use local execution only for trusted code.

git clone https://github.com/shi1720/repogym.git
cd repogym
python3 -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e ".[dev]"

repogym doctor                 # which toolchains are available
repogym list tasks             # the five example tasks
repogym validate tasks -v      # every check, per task
repogym run tasks --agent noop # baseline: every task scores 0
repogym run tasks --agent golden -w 4 --record trajectories -o results.json --report report.html

Plug in a real agent:

# any CLI agent, {prompt} is replaced with the shell-quoted problem statement
repogym run tasks --agent "shell:claude -p {prompt} --allowedTools Edit,Write,Bash"
repogym run tasks --agent "shell:aider --message {prompt} --yes"

# the built-in Anthropic tool-loop agent
python -m pip install -e ".[anthropic]" && export ANTHROPIC_API_KEY=...
repogym run tasks --agent claude --record trajectories

# your own agent class
repogym run tasks --agent my_package.agents:MyAgent

Generate tasks instead of writing them:

repogym mutate ./my-project -n 20 --out tasks/mutants   # synthetic, verified bug-fix tasks
repogym mine ./my-project --limit 500 -n 30 --out tasks/mined   # real fixes from git history
repogym validate tasks/mined                             # they come out validated, but check anyway
repogym export tasks -o tasks.jsonl                      # SWE-bench-style dataset

Anatomy of a task

tasks/bugfix-median/
├── task.yaml            # spec: statement, tests, constraints, grading, env limits
├── repo/                # repository snapshot the agent sees (or git url + commit)
├── hidden/              # tests injected only at grading time
└── solution.patch       # golden solution (or solution/ directory)
id: bugfix-median
title: "median() returns the wrong value for even-length input"
type: bugfix                  # bugfix | feature | refactor | perf
difficulty: easy
language: python
problem_statement: |
  `stats.median()` should return the mean of the two middle values ...
tests:
  command: "python -m pytest -q --junitxml={junit}"   # any JUnit-emitting runner
  hidden: [tests/test_hidden_median.py]
  fail_to_pass: [tests/test_stats.py::test_median_even, tests/test_hidden_median.py::test_median_two_values]
  pass_to_pass: [tests/test_stats.py::test_mean, tests/test_stats.py::test_median_odd]
constraints:
  allowed_files: ["stats.py"]
  forbidden_patterns: ["import statistics"]
grading:
  weights: {tests: 1.0}
solution:
  patch: solution.patch
env:
  max_steps: 20

Refactor and performance tasks are just as declarative:

# refactor: behaviour preserved (tests are a gate), complexity limits are the objective
constraints:
  paths: ["orders.py"]
  max_cyclomatic_complexity: 8
  max_function_length: 30

# perf: benchmark timed externally; reward is log-scaled up to the target speed-up
perf:
  command: "python bench.py"
  min_speedup: 3.0
  runs: 3

How grading works

submit ──▶ snapshot patch ──▶ ConstraintGrader ──▶ TestGrader ──▶ QualityGrader ──▶ PerfGrader ──▶ score
                              gate: scope,          restore protected   complexity /     wall-clock      objectives weighted,
                              patterns, size        tests, inject       function length  speed-up        any failed gate = 0
                                                    hidden, JUnit XML
  • tests score = fraction of fail_to_pass tests passing (or 0/1 without partial credit); any broken pass_to_pass test zeroes it.
  • perf score = log-scaled speed-up between a noise floor and min_speedup, measured with interleaved baseline/candidate runs timed from outside the process so the agent cannot fake it.
  • quality score (refactor tasks) = fraction of complexity/length limits met.
  • Components are objectives (weighted credit) or gates (no credit, zero the score when they fail): scope/pattern constraints and regression-only tests are gates, so an empty submission always scores 0.

Use it from Python

from repogym import Task, RepoEnv, load_tasks, run_suite, GoldenAgent, validate_task

# validate
report = validate_task(Task.load("tasks/bugfix-median"), repeat=3)
assert report.passed, report.failures

# your own agent
from repogym import Agent, Submit

class MyAgent(Agent):
    name = "mine"
    def solve(self, env, obs):
        while not obs.done:
            action = my_policy(obs.text)      # -> Action / dict
            obs = env.step(action)
        return obs

results = run_suite(load_tasks("tasks"), MyAgent, workers=4, record_dir="trajectories")

Trajectories are the training-data product. Each episode is a JSON file with every (action, observation, reward) tuple plus the final grade and patch; repogym export --trajectories trajectories -o sft.jsonl --only-passed turns them into chat-style records.

Why repogym

Existing options are either benchmarks or platforms. SWE-bench, SWE-Gym and R2E-Gym are datasets with bespoke harnesses; environment platforms (Prime Intellect, HUD, Mechanize and friends) are hosted products. What was missing is the small, open, local-first library layer: the FastAPI-shaped thing you pip install, point at your own repository, and get a validated environment out of. repogym is that layer:

  • Verification is a first-class command, not a notebook you run once. Tasks rot as repositories move; repogym validate in CI catches it.
  • The reward is defensible. Hidden tests, protected files, scope rules and complexity limits are declared next to the task and enforced on every grade.
  • Task supply scales. Mining git history and mutation-based synthesis turn one well-tested repository into hundreds of executable tasks, each with a golden patch, with no LLM in the loop.
  • It is agent- and language-agnostic. JUnit XML as the contract means pytest, node --test, cargo-nextest, jest, gradle and go test all work, and the agent can be a Python function, an LLM tool loop or an external CLI.

Documentation

Project status

0.1.0. The core API (Task, RepoEnv, graders, agents, miners, CLI) is covered by 125 tests at ~94% statement coverage, and the shipped tasks are validated in CI on Linux and macOS across Python 3.9 to 3.13. Roadmap: OpenEnv/Gymnasium adapters, an async batched environment server, Rust/Go/Java example tasks, LLM-authored problem statements for mined tasks, and per-task Docker image builds.

License

MIT © Shivam Gupta

About

Turn any git repository into a verifiable RL environment for coding agents. Declarative tasks, gym-style env, composable graders, task validation and mining.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages