Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 4 additions & 3 deletions docs/specs/2026-08-22-offline-e2e-release.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,14 +12,15 @@ The fixture uses no external provider account or production credentials. A loopb
Langfuse-compatible traces
→ revision-attributed collection
→ Mine admission and immutable snapshots
→ evidence-bound diagnosis and clusters
→ evidence-bound diagnosis, human-confirmed clusters
→ leakage-safe training/eval/selection/admission exports
→ controlled tool-file candidate
→ paired baseline/candidate gates and one-shot admission
→ repeated paired baseline/candidate gates with exact evidence
→ holdout-safe Fit experience index and one-shot admission
→ durable scheduler PROMOTE job
→ isolated Git commit, review PR reference, and reverse patch
```

The planted fixture contains one verified-good trace and four failures assigned to frontier, regression, selection, and admission partitions. The candidate fixes frontier and sealed holdouts while preserving the regression case. A real heartbeat materializes the seven-job DAG; the scheduler restarts before promotion and preserves its budget ledger. The test proves one winner, one PR, no deploy, no Langfuse write, and a non-empty rollback artifact.
The planted fixture contains one verified-good trace and four failures assigned to frontier, regression, selection, and admission partitions. A content-bound human review confirms every failure cluster before export. The candidate fixes frontier and sealed holdouts while preserving the regression case. Five paired developer repeats produce five discordant frontier wins and an exact one-sided probability of 0.03125. The Fit experience index preserves developer trace lineage and dense feedback while excluding selection/admission trace identities. A real heartbeat materializes the seven-job DAG; the scheduler restarts before promotion and preserves its budget ledger. The test proves one winner, one PR, no deploy, no Langfuse write, and a non-empty rollback artifact.

This is the local-v0 release boundary. Distributed scheduling, cloud workspaces, provider-backed candidate generation, and production deployment remain explicit adapters rather than hidden behavior in the offline proof.
31 changes: 28 additions & 3 deletions tests/test_e2e_release.py
Original file line number Diff line number Diff line change
Expand Up @@ -73,6 +73,7 @@
Worker,
WorkerId,
ofw,
read_fit_experience,
)
from ofw.contracts import AssetAccess
from ofw.diagnosis import ClusterId, DiagnosisResult, DiagnosisRun
Expand Down Expand Up @@ -200,7 +201,9 @@ def _harness(tmp_path: Path, base_url: str, monkeypatch: pytest.MonkeyPatch) ->
" target = any(name in output for name in ('frontier', 'selection', 'admission'))\n"
" passed = ('FIXED' in output) if target else ('BROKEN' not in output)\n"
" verdict = VerifierVerdict.PASS if passed else VerifierVerdict.FAIL\n"
" return VerifierResult(verdict, 1.0 if passed else 0.0, 'offline')\n",
" return VerifierResult(\n"
" verdict, 1.0 if passed else 0.0, 'offline:' + output\n"
" )\n",
encoding="utf-8",
)
(root / "diagnoser.py").write_text(
Expand Down Expand Up @@ -424,17 +427,39 @@ def test_offline_trace_to_review_release(
0.0,
1.0,
1.0,
PairedEvidencePolicy(StatisticalGateMode.EFFECT_SIZE_ONLY, 0, 1.0),
PairedEvidencePolicy(StatisticalGateMode.EXACT_SIGN_TEST, 5, 0.05),
)
campaign = ofw.fit(
harness,
bundle,
(candidate,),
benchmark_policy=BenchmarkPolicy(1, 10, 0, 0.25),
benchmark_policy=BenchmarkPolicy(5, 12, 0, 0.25),
policy=fit_policy,
)
fit_result = campaign.wait()
assert fit_result.winner_id == candidate.candidate.id
outcome = fit_result.outcomes[0]
frontier_evidence = next(
evidence
for evidence in outcome.paired_evidence
if evidence.partition is ExportPartition.FRONTIER
)
assert frontier_evidence.wins == 5
assert frontier_evidence.losses == 0
assert frontier_evidence.exact_one_sided_probability == pytest.approx(0.03125)
experience = read_fit_experience(fit_result)
assert tuple(
case.trace_id for case in experience.candidates[0].developer_cases
) == (TraceId("frontier"), TraceId("regression"))
assert all(
verifier.feedback.startswith("offline:")
for case in experience.candidates[0].developer_cases
for attempt in case.attempts
for verifier in attempt.candidate_verifiers
)
experience_payload = fit_result.experience_path.read_bytes()
assert b'"trace_id":{"value":"selection"}' not in experience_payload
assert b'"trace_id":{"value":"admission"}' not in experience_payload

remote = tmp_path / "review.git"
subprocess.run(("git", "init", "--bare", "-q", str(remote)), check=True)
Expand Down