Skip to content

Repository files navigation

OR-Space logo

OR-Space

A full-lifecycle workspace benchmark for industrial optimization agents.

Dataset License Benchmark

OR-Space evaluates whether language-model agents can perform reliable operations research work in executable, multi-file workspaces. Each instance separates business requirements, structured data, code artifacts, solver state, and evaluation targets instead of flattening the problem into one prompt.

Overview of the OR-Space Build, Revise, and Explain benchmark

Benchmark

OR-Space contains 100 industrial optimization topologies, each represented by three task views of the same underlying problem.

Task Participant-visible artifacts Evaluation
Build Business documents, tabular data, and an empty src/ directory Execute generated code and compare its objective with the oracle
Revise Revised documents and data plus the correct Build implementation Execute the revised program and compare its objective with the revised oracle
Explain Correct original and revised workspaces, execution logs, and solver records Score checklist coverage, reasoning, evidence grounding, answer quality, and unsupported claims

Build and Revise require an Optimal status and an objective within 1% relative error of the reference value. Explain uses instance-specific checklists and the released rubric. The paper's default track uses the Filesystem interface, Revise-code context, and Gurobi. Solver tracks require complete programs written against the corresponding solver API.

Repository contents

This repository contains participant workspaces, evaluation references and programs, and benchmark construction code.

.
├── 01_build/ ... 03_revise_business/  Benchmark construction code
├── Workspace_OR/                   Build, Revise, and Explain workspaces
├── evaluation/                    Task references and Explain rubrics
├── evaluation_programs/           Build, Revise, and Explain evaluators
├── supporting_files/              Metadata, task splits, and visual assets
├── tools/                         Staging and release-validation utilities
└── tests/                         Evaluator tests

Quick start

git clone https://github.com/0xzhouchenyu/OR-Space.git
cd OR-Space
python evaluation_programs/validate_dataset.py
import pandas as pd

index = pd.read_csv("supporting_files/metadata/workspace_index.csv")
print(index.groupby(["task_type", "difficulty"]).size())

The benchmark release is organized as:

.
├── Workspace_OR/
│   ├── build_workspaces/
│   ├── revise_workspaces/
│   └── explain_workspaces/
├── evaluation/
│   ├── build_evaluation/
│   ├── revise_evaluation/
│   └── explain_evaluation/
├── evaluation_programs/
└── supporting_files/

Empirical difficulty

Difficulty labels are derived from real benchmark outcomes under a fixed evaluation panel, separately for each task. Build and Revise use empirical executable pass rates; Explain uses the empirical mean rubric score. Boundaries approximate tertiles without splitting tied scores. The released distributions are Build 35/32/33, Revise 39/30/31, and Explain 33/33/34 for Easy/Medium/Hard. See supporting_files/metadata/difficulty_methodology.md and supporting_files/metadata/empirical_difficulty.csv.

Evaluation

evaluation_programs/ provides runnable scorers, while evaluation/ contains task references and Explain rubrics. The Explain release includes normalized exact checks, semantic checklist judgments, evidence verification, the judge prompt and schema, and the final 35/35/20/10 rubric with an unsupported-claim penalty of up to 20 points.

Validation

python tools/validate_public_release.py
python evaluation_programs/validate_dataset.py
python -m unittest discover -s tests

The validator checks difficulty metadata, workspace completeness, and accidental credentials.

Citation

@misc{zhou2026orspace,
  title = {OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents},
  author = {Zhou, Chenyu and Lu, Xinyun and Zhao, Jiangyue and Lin, Jianghao and Ge, Dongdong and Ye, Yinyu},
  year = {2026},
  note = {Dataset: https://huggingface.co/datasets/Chenyu-Zhou/OR-Space}
}

License

The release is for non-commercial research use under CC BY-NC 4.0-compatible terms, following the inherited license constraints of the IndustryOR seed topologies. Proprietary solver binaries, commercial credentials, and third-party model services are not redistributed.

About

OR-Space: a full-lifecycle workspace benchmark for industrial optimization agents

Topics

Resources

Stars

10 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages