New here? Start with one defect
Before committing to a benchmark, try the isolated backend first trial. It uses a disposable playground with one project-local skill; no global skill installation or production access is needed.
Stop after that backend exercise. Report whether setup worked, whether a regression test failed before the fix and passed afterward, and one unclear or missing instruction. A blocked attempt is useful too. Use the workflow feedback form, with task class Playground first trial.
This is an onboarding check, not evidence of token savings. The controlled replication protocol below remains the path for efficiency claims. No star, fork, endorsement, or full transcript is required. Remove private code, credentials, personal paths, and account details from feedback.
GPT-6 Astra replication
The Astra compatibility update is merged. Its changes clarify existing authorization, limit redundant verification, and describe useful authorized delegation. Efficiency on Astra is still unmeasured.
For this experiment, compare these three conditions with gpt-6-astra held constant:
- Engineering-loop disabled.
- Previous skill at
da88c13.
- Revised skill at
fbac8f3.
Use fresh sessions and identical task starting states, reasoning effort, tools, permissions, acceptance checks, and other available skills. Include the whole skill directory, since the references also changed. The control must actually disable engineering-loop; leaving its name out of a prompt does not prevent automatic selection.
Report acceptance, human corrections, unnecessary approval pauses, checks repeated without new evidence, elapsed time, and available token totals across all participating agents. Include failed and blocked runs; leave unavailable metrics blank. Use distinct condition labels and exact revisions in your receipt, and keep these Astra results separate from the historical GPT-5.6 experiment below. Repeat paired comparisons and alternate run order before drawing conclusions.
See the Astra comparison guidance for the rationale.
Why this replication matters
The repository currently has two controlled GPT-5.6-sol task sets:
- on a small backend boundary fix, all variants passed and the
no-repository-skill control used the fewest reported tokens;
- on a medium 2048 browser game, all variants passed and the lean
engineering-loop used 31.2% fewer reported tokens than v0.2.0 and 54.0%
fewer than the control.
That mixed result is more useful than a universal claim, but two seeded tasks
are not enough to guide team policy. We want independent, representative runs,
including results where the skill provides no advantage.
What to run
Follow the
three-way measurement protocol
from three identical starting copies:
no_skill: do not install or name a repository skill;
full_skill: install and invoke engineering-loop from v0.2.0;
lean_skill: install and invoke the current engineering-loop.
Keep the model, reasoning effort, starting commit, Codex surface, sandbox,
tool access, repository instructions, dependency state, task contract, and
acceptance criteria fixed.
You may use:
What to report
Quality comes before efficiency:
- acceptance without human code correction;
- evidence completeness;
- focused and required checks;
- retries and human corrections;
- elapsed time;
- input and output tokens when exposed by the Codex surface;
- model, reasoning effort, runtime, and skill revision;
- limitations, failed environment checks, and anything unverified.
Use the
CSV schema
or attach an equivalent table. Remove private code, prompts, paths, credentials,
customer information, and proprietary measurements before posting.
Ways to contribute
- Reply here with an anonymized result.
- Open a pull request adding auditable CSV rows and a short interpretation.
- Fork the benchmark into a stack-specific or translated edition.
- Review the protocol and identify a counting rule or validity threat that
would change the conclusion.
Maintainer disclosure: I maintain Codex How To. The goal is to collect
falsifiable evidence, not favorable testimonials.
New here? Start with one defect
Before committing to a benchmark, try the isolated backend first trial. It uses a disposable playground with one project-local skill; no global skill installation or production access is needed.
Stop after that backend exercise. Report whether setup worked, whether a regression test failed before the fix and passed afterward, and one unclear or missing instruction. A blocked attempt is useful too. Use the workflow feedback form, with task class Playground first trial.
This is an onboarding check, not evidence of token savings. The controlled replication protocol below remains the path for efficiency claims. No star, fork, endorsement, or full transcript is required. Remove private code, credentials, personal paths, and account details from feedback.
GPT-6 Astra replication
The Astra compatibility update is merged. Its changes clarify existing authorization, limit redundant verification, and describe useful authorized delegation. Efficiency on Astra is still unmeasured.
For this experiment, compare these three conditions with
gpt-6-astraheld constant:da88c13.fbac8f3.Use fresh sessions and identical task starting states, reasoning effort, tools, permissions, acceptance checks, and other available skills. Include the whole skill directory, since the references also changed. The control must actually disable engineering-loop; leaving its name out of a prompt does not prevent automatic selection.
Report acceptance, human corrections, unnecessary approval pauses, checks repeated without new evidence, elapsed time, and available token totals across all participating agents. Include failed and blocked runs; leave unavailable metrics blank. Use distinct condition labels and exact revisions in your receipt, and keep these Astra results separate from the historical GPT-5.6 experiment below. Repeat paired comparisons and alternate run order before drawing conclusions.
See the Astra comparison guidance for the rationale.
Why this replication matters
The repository currently has two controlled GPT-5.6-sol task sets:
no-repository-skill control used the fewest reported tokens;
engineering-loopused 31.2% fewer reported tokens than v0.2.0 and 54.0%fewer than the control.
That mixed result is more useful than a universal claim, but two seeded tasks
are not enough to guide team policy. We want independent, representative runs,
including results where the skill provides no advantage.
What to run
Follow the
three-way measurement protocol
from three identical starting copies:
no_skill: do not install or name a repository skill;full_skill: install and invokeengineering-loopfromv0.2.0;lean_skill: install and invoke the currentengineering-loop.Keep the model, reasoning effort, starting commit, Codex surface, sandbox,
tool access, repository instructions, dependency state, task contract, and
acceptance criteria fixed.
You may use:
backend playground;
2048 benchmark;
or
What to report
Quality comes before efficiency:
Use the
CSV schema
or attach an equivalent table. Remove private code, prompts, paths, credentials,
customer information, and proprietary measurements before posting.
Ways to contribute
would change the conclusion.
Maintainer disclosure: I maintain Codex How To. The goal is to collect
falsifiable evidence, not favorable testimonials.