Skip to content

Replication wanted: measure engineering-loop on a real task #23

Description

@Phelan164

New here? Start with one defect

Before committing to a benchmark, try the isolated backend first trial. It uses a disposable playground with one project-local skill; no global skill installation or production access is needed.

Stop after that backend exercise. Report whether setup worked, whether a regression test failed before the fix and passed afterward, and one unclear or missing instruction. A blocked attempt is useful too. Use the workflow feedback form, with task class Playground first trial.

This is an onboarding check, not evidence of token savings. The controlled replication protocol below remains the path for efficiency claims. No star, fork, endorsement, or full transcript is required. Remove private code, credentials, personal paths, and account details from feedback.

GPT-6 Astra replication

The Astra compatibility update is merged. Its changes clarify existing authorization, limit redundant verification, and describe useful authorized delegation. Efficiency on Astra is still unmeasured.

For this experiment, compare these three conditions with gpt-6-astra held constant:

  1. Engineering-loop disabled.
  2. Previous skill at da88c13.
  3. Revised skill at fbac8f3.

Use fresh sessions and identical task starting states, reasoning effort, tools, permissions, acceptance checks, and other available skills. Include the whole skill directory, since the references also changed. The control must actually disable engineering-loop; leaving its name out of a prompt does not prevent automatic selection.

Report acceptance, human corrections, unnecessary approval pauses, checks repeated without new evidence, elapsed time, and available token totals across all participating agents. Include failed and blocked runs; leave unavailable metrics blank. Use distinct condition labels and exact revisions in your receipt, and keep these Astra results separate from the historical GPT-5.6 experiment below. Repeat paired comparisons and alternate run order before drawing conclusions.

See the Astra comparison guidance for the rationale.

Why this replication matters

The repository currently has two controlled GPT-5.6-sol task sets:

  • on a small backend boundary fix, all variants passed and the
    no-repository-skill control used the fewest reported tokens;
  • on a medium 2048 browser game, all variants passed and the lean
    engineering-loop used 31.2% fewer reported tokens than v0.2.0 and 54.0%
    fewer than the control.

That mixed result is more useful than a universal claim, but two seeded tasks
are not enough to guide team policy. We want independent, representative runs,
including results where the skill provides no advantage.

What to run

Follow the
three-way measurement protocol
from three identical starting copies:

  1. no_skill: do not install or name a repository skill;
  2. full_skill: install and invoke engineering-loop from v0.2.0;
  3. lean_skill: install and invoke the current engineering-loop.

Keep the model, reasoning effort, starting commit, Codex surface, sandbox,
tool access, repository instructions, dependency state, task contract, and
acceptance criteria fixed.

You may use:

What to report

Quality comes before efficiency:

  • acceptance without human code correction;
  • evidence completeness;
  • focused and required checks;
  • retries and human corrections;
  • elapsed time;
  • input and output tokens when exposed by the Codex surface;
  • model, reasoning effort, runtime, and skill revision;
  • limitations, failed environment checks, and anything unverified.

Use the
CSV schema
or attach an equivalent table. Remove private code, prompts, paths, credentials,
customer information, and proprietary measurements before posting.

Ways to contribute

  • Reply here with an anonymized result.
  • Open a pull request adding auditable CSV rows and a short interpretation.
  • Fork the benchmark into a stack-specific or translated edition.
  • Review the protocol and identify a counting rule or validity threat that
    would change the conclusion.

Maintainer disclosure: I maintain Codex How To. The goal is to collect
falsifiable evidence, not favorable testimonials.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    experimentMeasured workflow or knowledge-maintenance experimentfeedbackReproducible workflow feedback from usershelp wantedExtra attention is neededworkflowChanges or feedback involving an engineering workflow

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions