Skip to content

Commit 66a0fbd

Browse files
feat(efficacy): implement #69 rulings a1,b1,c1,d1,e2,f1 — protocol v2.1 (#75)
Owner ruled the six open D1 questions (#69) as **a1, b1, c1, d1, e2, f1**. This PR lands the whole ruling in one change, per the issue's done-condition ("rulings recorded here, protocol amended to v2.1 with examples updated"). ## What's in it - **`vexometer/docs/EFFICACY-PROTOCOL.adoc` → v2.1** — every ruled semantic written into the normative text, all JSON examples regenerated from the binary (doc == tool is enforced by exact-value equality in the test suite): - **a1** — a target metric with `B_m = 0` gets `G_m := 0`, so it cannot improve and the verdict is `reject_null`; the report lists offenders in a diagnosability warning. Zero-baseline *collateral* metrics stay fully protected. - **b1** — per-probe identity gate is normative when `probes.results` exists in both measurements (at most one baseline-passing probe may fail after; newly-passing probes buy nothing back); aggregate pass-rate is the degraded fallback. `capability.probes_regressed` records which gate applied. - **c1** — all-targets rule: every declared target must improve or the verdict is `reject_null`. - **d1** — plural `frontier_records`, one per-metric record per target in target order; length mismatch is a hard error; the singular `frontier_record` key fails validation. - **e2** — new `lift` subcommand: mechanical v1→v2.1 lift, v1 fields verbatim, missing evidence as explicit `null` (never synthesised), `lifted_from` marker, `verdict: "unverified"` reserved for lifted reports. Normative spec + worked example in the doc. - **f1** — held-out scenario-set registry (`vexometer/data/scenario_sets/registry.json`, ships empty — no corpus exists yet, no invented hashes); `validate --scenario-registry` rejects scoring on a tuning partition and unknown scenario sets, covering both efficacy reports and frontier attempts. Registering the first partition is a precondition for the first satellite evaluation (D6). - **`vexometer-efficacy`** — every `AwaitingRuling` refusal replaced with the ruled semantics; exit code 2 (the "open D1 question" refusal) retired, codes now 0/1/3; the renamed `--frontier-record` flag gets an explicit error pointing at `--frontier-records`. - **Justfile** — `efficacy-lift` recipe; the protocol's stale "these recipes do not yet exist" sentence replaced with the real tooling list. - **README/ROADMAP** updated; **trust manifests regenerated in-PR** per the manifest contract. ## Verification - `cargo test`: 29/29 (protocol examples are the fixtures; report + lift roundtrips assert exact JSON equality with the doc) - `cargo clippy --all-targets -- -D warnings` clean, `cargo fmt --check` clean - `just must-all` green, `just trust-manifest-verify` green, `just --evaluate` parses Closes #69 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
1 parent 1ac7e41 commit 66a0fbd

15 files changed

Lines changed: 1184 additions & 359 deletions

File tree

‎Justfile‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -69,6 +69,10 @@ efficacy-report *ARGS:
6969
efficacy-attempt *ARGS:
7070
cd vexometer-efficacy && cargo run --release --quiet -- attempt "$@"
7171

72+
# Mechanically lift a vexometer-efficacy-v1 report to v2.1 shape (ruling e2)
73+
efficacy-lift *ARGS:
74+
cd vexometer-efficacy && cargo run --release --quiet -- lift "$@"
75+
7276
# Validate efficacy reports and frontier records by recomputation
7377
efficacy-validate *ARGS:
7478
cd vexometer-efficacy && cargo run --release --quiet -- validate "$@"

‎lazy-eliminator/.trust/trust-manifest.sha256‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,6 @@
11
# trust-manifest v1
22
# component=lazy-eliminator
3-
# generated_at=2026-09-01T14:22:53Z
3+
# generated_at=2026-09-01T23:29:45Z
44
339d25795fa89149354d4101c533492f0e5bbe39fc953ac02be17a225f2fd270 README.adoc
55
6772e621da4e50257728886f568bd652b8e9c58a02aec325112e57def73da46a ROADMAP.adoc
66
504199ed09a9acbd183fe9c37a8330ec254f33f5612e0f36f2f5e7be170a57ff SECURITY.adoc

‎satellite-template/.trust/trust-manifest.sha256‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,6 @@
11
# trust-manifest v1
22
# component=satellite-template
3-
# generated_at=2026-09-01T14:22:53Z
3+
# generated_at=2026-09-01T23:29:45Z
44
e90437cd512f3ac6824b42e733394bc69dfa4bcab046d779842bc876e11f870e README.adoc
55
5a62f5611eecfafa43d931b4d7e0917fa1b0f0275fdfd74c61415ab635e72805 ROADMAP.adoc
66
38ccfdc1a04c12616acfb030522358383702184480f52507a5f72fccccbe76b9 SECURITY.adoc
Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -1,8 +1,8 @@
11
# trust-manifest v1
22
# component=vexometer-efficacy
3-
# generated_at=2026-09-01T15:55:49Z
4-
3ea7341c2a55bea766ffa7c34879168001f701982c3427c8f3ff874b4b907c3e README.adoc
5-
99bf8c708656fee9beba0c4812aac55a6fd3b4fdaaa989a9b6a13b7dc3c4b5ba ROADMAP.adoc
3+
# generated_at=2026-09-01T23:29:45Z
4+
df1c5e511fc5e2e7cbcd094fea57728ce193ee267a267037f4043ae5edbf3279 README.adoc
5+
272d7670bb46837c46db61d519d849c60314fecaaf3cd807ee9fdd18ec76ba35 ROADMAP.adoc
66
b1245e468709a6c75e530412da6480943bf53c836df0ca108aaf39843886e6cb SECURITY.adoc
77
9c80ff2e60fdb772a0479b46b140e0ce08e4e37bc39e6d7e257aa3d5d1281d18 contractiles/must/Mustfile
88
3ac4606620454d844d8f0d0580fe32072a8a0b6821c93a74df64c3ed597e3640 contractiles/trust/Trustfile.a2ml

‎vexometer-efficacy/README.adoc‎

Lines changed: 66 additions & 45 deletions
Original file line numberDiff line numberDiff line change
@@ -11,52 +11,65 @@ applies the six-verdict acceptance rule with its precedence order, emits
1111
records under the monotone-frontier invariant, and validates both
1212
document shapes by recomputing every derived number.
1313

14-
== Design rule: refuse where the protocol is undecided
14+
== The ruled semantics (protocol v2.1)
1515

16-
Six normative questions are open in
17-
https://github.com/hyperpolymath/vexometer/issues/69[issue #69] (debt
18-
item D1). Where one of them bites, this tool *refuses with an explicit
19-
error* naming the question rather than silently picking a semantic:
16+
The six normative questions this tool originally refused to guess at
17+
(https://github.com/hyperpolymath/vexometer/issues/69[issue #69],
18+
D1a-D1f) were ruled `a1, b1, c1, d1, e2, f1` on 2026-09-01 and are
19+
implemented here:
2020

21-
[cols="1,4,2",options="header"]
21+
[cols="1,4,3",options="header"]
2222
|===
23-
|Question |When it bites |Behaviour
24-
25-
|D1a
26-
|A declared target metric has baseline `B_m = 0` (division by zero in
27-
`G_m`)
28-
|Hard refusal, exit code 2
29-
30-
|D1b
31-
|Per-probe results are supplied for both measurements and the aggregate
32-
pass-rate gate disagrees with the per-probe identity gate
33-
|Hard refusal, exit code 2
34-
35-
|D1c
36-
|Multiple targets are declared and some improved while others did not
37-
|Hard refusal, exit code 2
38-
39-
|D1d
40-
|Multiple targets with the singular `frontier_record` field
41-
|Report is emitted, with a warning on stderr
42-
43-
|D1e
23+
|Ruling |When it bites |Behaviour
24+
25+
|a1
26+
|A declared target metric has baseline `B_m = 0`
27+
|`G_m := 0`, so the target cannot improve and the verdict is
28+
`reject_null`; the report lists the metric in a diagnosability warning.
29+
Zero-baseline *collateral* metrics stay fully protected.
30+
31+
|b1
32+
|Per-probe results exist in both measurements
33+
|The per-probe identity gate is normative: at most one baseline-passing
34+
probe may fail after; newly-passing probes buy nothing back. The
35+
aggregate pass-rate gate is the degraded fallback when per-probe results
36+
are absent.
37+
38+
|c1
39+
|Multiple targets are declared
40+
|All-targets rule: every declared target must improve, or the verdict is
41+
`reject_null`.
42+
43+
|d1
44+
|Frontier references
45+
|Plural `frontier_records`, one per-metric record per target, in target
46+
order; a length mismatch is a hard error and the pre-ruling singular
47+
`frontier_record` key fails validation.
48+
49+
|e2
4450
|v1→v2 lifting
45-
|Unimplemented — no `lift` subcommand exists
51+
|The `lift` subcommand: v1 fields carried verbatim, missing v2 evidence
52+
as explicit `null`, `lifted_from` marker, `verdict: "unverified"`
53+
(reserved for lifted reports).
54+
55+
|f1
56+
|Scenario-set provenance
57+
|`validate --scenario-registry FILE` checks every scored `scenario_set`
58+
against the held-out partition registry at
59+
`../vexometer/data/scenario_sets/registry.json`.
4660
|===
4761

48-
After the rulings land and the protocol is amended to v2.1, these
49-
refusals are replaced by the ruled semantics.
50-
5162
== The protocol's examples are the test fixtures
5263

5364
The integration tests read `../vexometer/docs/EFFICACY-PROTOCOL.adoc`
5465
at build time, extract its example JSON blocks, and require that the
55-
validator accepts both and that the evaluator reproduces the efficacy
66+
validator accepts them all, that the evaluator reproduces the efficacy
5667
example value-for-value from raw inputs (including `D_ISA = -2.71`
5768
under the default category weights in
58-
link:../vexometer/docs/METRICS.adoc[METRICS.adoc]). If the protocol and
59-
this implementation drift apart, `cargo test` fails loudly.
69+
link:../vexometer/docs/METRICS.adoc[METRICS.adoc]), and that `lift`
70+
reproduces the protocol's lifted example from its v1 example. If the
71+
protocol and this implementation drift apart, `cargo test` fails
72+
loudly.
6073

6174
== CLI
6275

@@ -66,7 +79,8 @@ $ vexometer-efficacy report --baseline baseline.json --after after.json \
6679
--targets LPS,TII --satellite vex-verbosity-compressor \
6780
--sample-size 500 --output report.json \
6881
[--methodology "A/B testing with vexometer validation"] \
69-
[--notes "..."] [--frontier-record frontier/LPS-....json] \
82+
[--notes "..."] \
83+
[--frontier-records frontier/LPS-....json]... \
7084
[--traces-available true|false] [--date YYYY-MM-DD] [--scenario-set SHA]
7185
7286
$ vexometer-efficacy attempt --frontier frontier/LPS-2026-09-01.json \
@@ -75,18 +89,24 @@ $ vexometer-efficacy attempt --frontier frontier/LPS-2026-09-01.json \
7589
[--model-profile STR] [--timestamp ISO8601] [--scenario-set SHA] \
7690
[--baseline-isa 4.63] # required when creating a new frontier record
7791
78-
$ vexometer-efficacy validate report.json frontier.json ...
92+
$ vexometer-efficacy lift --input v1-report.json --output lifted.json
93+
94+
$ vexometer-efficacy validate report.json frontier.json ... \
95+
[--scenario-registry ../vexometer/data/scenario_sets/registry.json]
7996
----
8097

81-
Bare `validate` arguments are routed by each document's own `version`
82-
field; `--efficacy FILE` / `--frontier FILE` force a kind when a
83-
document lacks one. The same commands are exposed at the monorepo root
84-
as `just efficacy-report`, `just efficacy-attempt`, and
85-
`just efficacy-validate`.
98+
Pass `--frontier-records` once per target metric, in target order
99+
(ruling d1). Bare `validate` arguments are routed by each document's own
100+
`version` field; `--efficacy FILE` / `--frontier FILE` force a kind when
101+
a document lacks one, and `--scenario-registry` enforces ruling f1
102+
against every scored scenario set. The same commands are exposed at the
103+
monorepo root as `just efficacy-report`, `just efficacy-attempt`,
104+
`just efficacy-lift`, and `just efficacy-validate`.
86105

87106
Exit codes: `0` success (any verdict, including rejections — a computed
88-
rejection is a successful evaluation), `1` usage or data error, `2` open
89-
D1 ruling required, `3` validation failed.
107+
rejection is a successful evaluation), `1` usage or data error, `3`
108+
validation failed. (Exit code `2`, the pre-ruling "open D1 question"
109+
refusal, is retired.)
90110

91111
== Measurement input format
92112

@@ -115,8 +135,9 @@ pass over one content-addressed scenario set:
115135
`{score, std_dev, confidence, p_value}` object are both accepted;
116136
statistics are carried into the report when present.
117137
* `probes.results` (per-probe outcomes) is optional; when both
118-
measurements carry it, the per-probe identity gate is cross-checked
119-
against the aggregate gate (see D1b above).
138+
measurements carry it, the per-probe identity gate is normative
139+
(ruling b1) and the report records any regressed probes in
140+
`capability.probes_regressed`.
120141
* `scenario_set` must match between baseline and after — tuning against
121142
a different set than you score on is exactly what the protocol's
122143
audit trail exists to catch.

‎vexometer-efficacy/ROADMAP.adoc‎

Lines changed: 8 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -11,13 +11,15 @@
1111
* [x] Protocol examples as live test fixtures (drift fails `cargo test`)
1212
* [x] Explicit refusals on open D1 questions (issue #69)
1313

14-
== After the D1 rulings (v0.2, blocked on issue #69)
14+
== The D1 rulings (v0.2, shipped — issue #69 ruled `a1,b1,c1,d1,e2,f1`)
1515

16-
* [ ] Replace each D1a–D1d refusal with the ruled semantic
17-
* [ ] `frontier_records` plurality per ruling (d)
18-
* [ ] v1→v2 lifting: implement or formally drop per ruling (e)
19-
* [ ] Held-out scenario-set support per ruling (f)
20-
* [ ] Track the protocol's v2.1 text (same PR as the amendment)
16+
* [x] Replace each D1a–D1d refusal with the ruled semantic (a1 zero
17+
baseline, b1 per-probe identity gate, c1 all-targets rule)
18+
* [x] `frontier_records` plurality per ruling d1
19+
* [x] v1→v2 lifting: `lift` subcommand per ruling e2
20+
* [x] Held-out scenario-set registry + `--scenario-registry` per
21+
ruling f1
22+
* [x] Track the protocol's v2.1 text (same PR as the amendment)
2123

2224
== Later
2325

0 commit comments

Comments
 (0)