The scored benchmark cannot be the place where we discover that fixture reset, blinding, or redaction is broken.
Freeze the private Ditto Proof v1 task pack, then pass a separate six-execution, non-scored pilot before any paid or scored run.
Acceptance criteria
- Create primary and held-out variants for work/done, design, and writing.
- Freeze every task as a clean Git commit plus a deterministic content hash before the pilot.
- Keep scored fixture content and the private operator rubric outside Git until the final evidence package is approved.
- Use separate
pilot-work, pilot-design, and pilot-write fixtures that can never enter v1 aggregates.
- Run each pilot fixture once in cold and
+Ditto conditions: six isolated executions total.
- Capture every required field without manually patching records.
- Prove deterministic reset, clean condition separation, opaque review IDs, artifact hashing, and append-only storage.
- Catch seeded secret, username, local-path, profile-text, and receipt canaries.
- Produce a sanitized pilot package labelled
non-scored and non-comparable.
- Record the observed provider cost and quota, then get explicit approval before any later scored execution.
Failure policy
A failed pilot is diagnosed and rerun from clean pilot fixtures after the harness is fixed. Pilot outputs never become public performance evidence.
The scored benchmark cannot be the place where we discover that fixture reset, blinding, or redaction is broken.
Freeze the private Ditto Proof v1 task pack, then pass a separate six-execution, non-scored pilot before any paid or scored run.
Acceptance criteria
pilot-work,pilot-design, andpilot-writefixtures that can never enter v1 aggregates.+Dittoconditions: six isolated executions total.non-scoredandnon-comparable.Failure policy
A failed pilot is diagnosed and rerun from clean pilot fixtures after the harness is fixed. Pilot outputs never become public performance evidence.