Skip to content

Ditto Proof v1: reproducible paired benchmark tracker #17

Description

@ohad6k

Ditto has launch traffic, installs, and user feedback. What it does not have yet is controlled, inspectable evidence showing where the personal layer changes the result.

Ditto Proof v1 is the narrow evidence milestone. It compares the same complete system under a clean-host cold start and with frozen Ditto context enabled.

Frozen v1 shape

  • Two preselected general-purpose systems: one Codex-host and one Claude-host.
  • Three task families: work/done, design, and writing.
  • Primary and held-out variants for each family.
  • Two independent trials per system, family, and variant.
  • Cold and +Ditto conditions in fresh isolated cells.
  • 24 paired comparisons and 48 cell executions.
  • Ditto frozen at v0.3.7 for the complete scored run.
  • Public label: small-n, directional only.

Workstream

  • Build the disabled-by-default Ditto Proof v1 harness #14
  • Freeze scored fixtures and pass the six-execution non-scored pilot #15
  • Capture two live system identities, host versions, selection evidence, quota, and expected cost before execution.
  • Get explicit approval for the exact manifest and expected provider cost.
  • Capture independent third-party blind verdicts. Ohad cannot cast them.
  • Recalculate win/tie/loss denominators, Wilson 95% intervals where defined, and hard failures from sanitized records.
  • Pass automated redaction, seeded-canary checks, and manual privacy review.
  • Complete all 48 valid cells or use a clearly named incomplete non-v1 pilot label.
  • Get a separate evidence-digest ship approval before publishing claims, clips, or a benchmark release.

Public outcomes

The primary public outcomes are valid paired blind preference and hard-failure counts by condition. Writing preference is omitted when the reviewer is familiar with the operator or the voice makes blinding invalid. Profile-rubric adherence stays labelled as mechanism validation, not standalone proof of value.

No significance language, universal model ranking, traffic forecast, or 1000x claim belongs in this release. If the result is neutral, mixed, invalid, or negative, it becomes product evidence instead of a positive launch claim.

Community dependency

#16 improves the contribution path around this work, but it is not allowed to weaken the privacy or ship gates above.

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestproof-v1Reproducible, privacy-safe Ditto Proof v1 work

    Projects

    No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions