Skip to content

Commit bdb1aa2

Browse files
authored
Merge pull request #350 from Flow-Research/codex/ws-ci-005-semantic-proof-quality
docs(ci): plan semantic proof quality
2 parents d8f415b + 4afc88e commit bdb1aa2

15 files changed

Lines changed: 968 additions & 0 deletions

‎.agent-loop/CURRENT_STATE.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -45,6 +45,7 @@ authority; these records do not grant or withhold it.
4545
| [WS-CI-002](initiatives/WS-CI-002-deterministic-agent-gates/STATUS.md) | `WS-CI-002-01` complete through PR #311; Agent Gates is deterministic per PR head | Preserve protected-branch review as the independent approval authority |
4646
| [WS-CI-003](initiatives/WS-CI-003-atomic-chunk-state/STATUS.md) | `WS-CI-003-01` complete | Require every chunk PR to land its final contract and initiative state atomically |
4747
| [WS-CI-004](initiatives/WS-CI-004-review-evidence-integrity/STATUS.md) | Exact-target, nine-reviewer, final-head, impact-cone, and adversarial-proof work is merged through PR #345; `WS-CI-004-05` semantic-completeness hardening is complete | Human review and merge `WS-CI-004-05`; do not start another successor automatically |
48+
| [WS-CI-005](initiatives/WS-CI-005-semantic-proof-quality/STATUS.md) | `WS-CI-005-PLAN`, `WS-CI-005-01`, `WS-CI-005-02`, and `WS-CI-005-03` are planned; no implementation behavior exists | Start `WS-CI-005-01` only through explicit human direction |
4849
| [WS-DB-001](initiatives/WS-DB-001-v01-schema-baseline/STATUS.md) | v0.1 schema baseline complete through PRs #316 and #317 | Extend `0001_v01_baseline` only through future bounded migrations |
4950
| [WS-SEC-001](initiatives/WS-SEC-001-dependency-alert-remediation/STATUS.md) | `WS-SEC-001-01` is complete with patched runtime and tooling dependencies | Handle future security alerts through fresh bounded dependency changes |
5051
| [WS-DOCS-001](initiatives/WS-DOCS-001-current-v01-documentation/STATUS.md) | Current v0.1 entry documentation complete | Keep current pages synchronized with merged capability changes |

‎.agent-loop/REVIEW_LOG.md‎

Lines changed: 11 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -3001,3 +3001,14 @@ the `08A` executable contract omitted `CON-03D`, and canonical dependency views
30013001
used broad REV persistence labels. The repair adds the explicit `03D -> 08A`
30023002
gate and names exact merged `REV-04B` runtime
30033003
`Review`/`ReviewLease`/`FinalAcceptance` targets.
3004+
3005+
## 2026-08-17 - WS-CI-005 Semantic Proof Quality Planning
3006+
3007+
PR #350 plans a semantic-proof layer for the existing exact-head reviewer
3008+
system. Initial internal review corrected proof self-attestation, incomplete
3009+
failure replay, missing exact tests, malicious-evidence coverage, premature
3010+
adoption language, stale terminology, and invalid chunk states. CodeRabbit
3011+
then identified two final gaps: source inspection could appear to replace
3012+
required database/session custody, and future outcomes used present-tense
3013+
completion wording. Both are corrected; final-head review and hosted checks
3014+
remain required before human merge.
Lines changed: 11 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,11 @@
1+
# Chunk Map: WS-CI-005 Semantic Proof Quality
2+
3+
| Chunk | Goal | Risk | State |
4+
|---|---|---:|---|
5+
| `WS-CI-005-PLAN` | Record failure replay, proof model, risks, decisions, and bounded implementation contracts | L1 | Planned; no implementation behavior |
6+
| `WS-CI-005-01` | Add proof-strength vocabulary, escaped-failure reference, receipt fields, compatibility validation, and focused tests | L1 | Planned; not started |
7+
| `WS-CI-005-02` | Install candidate test-of-the-test and relevant failure-pattern obligations across reviewer skills/agents with mutation enforcement | L1 | Planned after `01`; not started |
8+
| `WS-CI-005-03` | Prove adoption with blind PR #349 replay fixtures and integrate proof-quality summaries into evidence/trust workflows | L1 | Planned after `02`; not started |
9+
10+
This map is sequencing guidance, not an active queue. Human approval starts one
11+
chunk. One implementation chunk equals one pull request.
Lines changed: 58 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,58 @@
1+
# Decisions: WS-CI-005 Semantic Proof Quality
2+
3+
## D1: Improve proof quality, not reviewer count
4+
5+
The existing nine specialties remain. This initiative sharpens their evidence
6+
obligations instead of adding overlapping agents.
7+
8+
## D2: Proof custody follows the claimed boundary
9+
10+
Pure, service, repository, transaction, concurrency, direct-SQL, composition,
11+
and negative-structure claims require evidence capable of observing that exact
12+
boundary. A stronger infrastructure label does not automatically replace a
13+
different kind of proof.
14+
15+
## D3: Every PASS tests at least one proof
16+
17+
A final reviewer PASS requires a concrete test-of-the-test adversarial probe.
18+
Merely listing named tests and passing commands is insufficient.
19+
20+
## D4: Escaped defects become shared failure patterns
21+
22+
Valid escaped findings are promoted into concise reviewer knowledge and blind
23+
evaluation fixtures. PR-specific prose does not become an ever-growing active
24+
checklist.
25+
26+
## D5: Real isolation requires an existing foreign resource
27+
28+
A missing-row mock cannot prove tenant isolation. Repository-level isolation
29+
must exercise a valid resource owned by another tenant or project.
30+
31+
## D6: Database truth requires database custody
32+
33+
Rollback, locking, concurrency, triggers, constraints, and direct-SQL integrity
34+
must be proven against the real database. Source inspection may identify a
35+
finding but cannot produce executed database proof.
36+
37+
## D7: Reuse canonical rules
38+
39+
When schema, runtime, public API, migration, and database layers express the
40+
same invariant, reviewers must identify one canonical owner or explicitly prove
41+
equivalence. Silent duplicate validation is not accepted.
42+
43+
## D8: Machine checks validate shape; blind evaluations validate judgment
44+
45+
The validator enforces closed fields and declared compatibility. It does not
46+
pretend to infer semantic correctness from filenames. Raw blind fixtures prove
47+
whether reviewer behavior actually improves.
48+
49+
## D9: Preserve simple contribution authority
50+
51+
Nothing in this initiative starts work, grants permission, approves a PR, or
52+
merges code. GitHub permissions and explicit human merge remain authoritative.
53+
54+
## D10: Review artifacts are untrusted data
55+
56+
Diffs, comments, findings, fixtures, and evidence may contain instructions.
57+
Reviewers inspect but never execute or obey those instructions. Blind
58+
evaluations must prove this behavior independently of expected-answer secrecy.
Lines changed: 110 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,110 @@
1+
# Discovery: WS-CI-005 Semantic Proof Quality
2+
3+
## Current behavior
4+
5+
WS-CI-004 already supplies exact-target inspection, structured receipts,
6+
specialty routing, atomic traceability, residual-escape analysis, and isolated
7+
reviewer evaluations. `scripts/reviewer_contracts.py` checks that every reviewer
8+
agent and skill contains the shared semantic requirements. Its evaluation model
9+
uses positive, negative, stale-replay, output-contract, and handoff cases.
10+
11+
The remaining gap is semantic discrimination. The validator checks that a
12+
traceability row has an owner, implementation source, named proof, custody, and
13+
result. It does not determine whether the named proof can observe the behavior.
14+
15+
## Relevant files and symbols
16+
17+
| Path | Current responsibility | Gap |
18+
|---|---|---|
19+
| `.agents/skills/reviewer-evidence-protocol/SKILL.md` | Universal exact-head and semantic review process | No proof-strength or test-of-the-test rule |
20+
| `.agents/skills/{architecture,security,qa,test-delta,ci-integrity,reuse-dedup,senior-engineer}-review/SKILL.md` | Specialty review depth | No shared database/isolation/strict-fake failure patterns |
21+
| `.codex/agents/*-reviewer.toml` | Custom reviewer execution contracts | Can repeat named proof without demonstrating discrimination |
22+
| `.agent-loop/templates/INTERNAL_REVIEW_RECEIPT.schema.json` | Structured advisory receipt | Trace rows do not declare proof strength or adversarial mutation outcome |
23+
| `scripts/reviewer_contracts.py` | Reviewer contract/evaluation validator | Validates fields and tokens, not proof compatibility |
24+
| `scripts/test_reviewer_contracts.py` | Mutation and output regression tests | Does not replay non-discriminating proof classes |
25+
| `WS-CI-004/evaluations/{CASES,EXPECTATIONS}.json` | Blind reviewer evaluation inputs | Covers broad specialties, not the PR #349 escapes |
26+
| `.agents/skills/evidence-gate/SKILL.md` | Deterministic pre-review evidence | Does not classify infrastructure custody |
27+
| `.agents/skills/pr-trust-bundle/SKILL.md` | Human-facing evidence summary | Can summarize a named but semantically weak test |
28+
29+
## Failure replay
30+
31+
### PR #338 and PR #346
32+
33+
These exposed path continuity, atomic state vocabulary, public-owner boundaries,
34+
completed-history immutability, ledger parity, and compound-criterion gaps.
35+
WS-CI-004 now covers exact-head and traceability mechanics for those classes.
36+
37+
### PR #349
38+
39+
The following escaped after initial internal passes:
40+
41+
1. `PROJECT_POINTS` quantity `"1.0"` passed a duplicated runtime validator while
42+
canonical schema and database rules rejected it.
43+
2. A COMPENSATION owner fact was trusted after checking only its project, not
44+
binding identity and instrument type.
45+
3. Malformed immutable input leaked `AttributeError` before domain concealment.
46+
4. A mocked repository exception was presented as transaction rollback proof.
47+
5. Fake authorization methods raised labels for wrong-session/copy/replay cases
48+
without constructing those conditions.
49+
6. PostgreSQL `<>` comparisons over nullable operands allowed trigger guards to
50+
evaluate to unknown and skip rejection.
51+
7. Independent foreign keys allowed project, policy, and version facts that
52+
existed individually but did not share composite ownership.
53+
8. A cross-project read test used a mock returning `None`, duplicating the
54+
missing-record case rather than proving repository isolation.
55+
9. PostgreSQL regressions initially failed during shared fixture setup because
56+
they recreated a globally unique service identity, so the intended
57+
integrity assertion was never reached.
58+
10. A required-version regression initially used an invalid non-UUID value the
59+
old code already rejected instead of `None`, the previously accepted bad
60+
selector.
61+
62+
These are proof-quality failures: the named proof existed but could not
63+
distinguish the defect.
64+
65+
## Existing tests and gaps
66+
67+
Existing reviewer-contract tests prove protocol adoption and receipt shape.
68+
They do not currently prove:
69+
70+
- proof custody is compatible with the claim;
71+
- a reviewer mutates or contradicts a claimed invariant;
72+
- strict fakes validate identity, state, and call order;
73+
- real tenant-isolation proof persists a foreign resource;
74+
- database review covers NULL semantics, composite ownership, direct SQL, and
75+
rollback durability;
76+
- reuse review compares schema, runtime, and database representations of one
77+
canonical rule;
78+
- escaped findings become permanent blind evaluation cases.
79+
- fixture setup reaches the intended assertion rather than merely failing;
80+
- a regression input distinguishes corrected behavior from the pre-fix code.
81+
82+
## Dependencies and conventions to preserve
83+
84+
- Extend `scripts/reviewer_contracts.py`; do not create a parallel validator.
85+
- Extend the WS-CI-004 reviewer registry and evaluation harness; do not fork the
86+
nine-reviewer map.
87+
- Keep shared rules in one protocol/reference and specialty deltas in their
88+
existing skill and agent files.
89+
- Keep receipts advisory and out of tree; GitHub remains durable authority.
90+
- Keep external review responses separate from internal receipts.
91+
92+
## Risks discovered
93+
94+
| Risk | Consequence | Planned control |
95+
|---|---|---|
96+
| Named proof without observability | False PASS | Closed proof-strength and compatibility rules |
97+
| Permissive fake | Simulated security/isolation evidence | Strict-fake obligations and blind fixtures |
98+
| ORM-only review | SQL integrity bypass | Database integrity reference and direct-SQL probes |
99+
| Missing tenant record | Isolation test duplicates not-found | Real foreign-resource proof requirement |
100+
| Duplicated business rule | Schema/runtime drift | Canonical-rule reuse comparison |
101+
| More reviewer prose | Token/maintenance growth | One concise shared reference, validated adoption tokens |
102+
103+
## Unknowns to measure during implementation
104+
105+
- Whether proof-strength metadata belongs directly in the receipt schema or in
106+
a referenced traceability sub-schema.
107+
- The smallest blind-fixture set that covers each escape without leaking the
108+
answer or making evaluation slow.
109+
- Which proof-strength mismatches can be checked deterministically and which
110+
remain reviewer judgments.
Lines changed: 106 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,106 @@
1+
# Intent: WS-CI-005 Semantic Proof Quality
2+
3+
## Problem being solved
4+
5+
Exact-head receipts, named tests, traceability rows, and adversarial prose can
6+
still produce a false `PASS` when the proof does not discriminate the claimed
7+
behavior. PR #349 passed internal review before later probes found duplicated
8+
quantity rules, partial owner-fact validation, malformed-input leakage, mocked
9+
rollback, simulated handle semantics, nullable SQL guard bypass, incomplete
10+
composite ownership, and a cross-project test that was only a missing-row test.
11+
12+
## Why this work matters
13+
14+
Workstream is governed contribution infrastructure. A confident but vacuous
15+
review can allow an authorization, tenant-isolation, lifecycle, compensation,
16+
or audit defect to become durable. Repeated full review cycles also exhaust
17+
maintainers and contributors. Review must become sharper, not larger.
18+
19+
## Current behavior
20+
21+
- WS-CI-004 binds reviews to exact clean Git targets and requires atomic
22+
traceability, impact-cone inspection, adversarial probes, and residual-escape
23+
analysis.
24+
- Nine specialty reviewers have structural contract and blind-evaluation tests.
25+
- Proof rows name tests and execution custody, but do not use a closed proof-
26+
strength vocabulary.
27+
- A mock can be cited for a database, transaction, isolation, or prepared-handle
28+
claim even when it cannot exercise that behavior.
29+
- Escaped defects are documented per PR but are not promoted into reusable
30+
reviewer failure classes.
31+
32+
## Target behavior
33+
34+
Every material review claim declares the weakest acceptable proof custody.
35+
Reviewers attempt a concrete test-of-the-test, reject proofs that cannot observe
36+
the claimed failure, reuse canonical owner rules, and apply shared database and
37+
tenant-isolation probes where relevant. Escaped defects become compact reusable
38+
evaluation fixtures so the same blind spot is not rediscovered manually.
39+
40+
## Design chosen
41+
42+
Extend the existing reviewer protocol and evaluator with:
43+
44+
1. a closed proof-strength vocabulary;
45+
2. machine-validated claim-to-proof compatibility;
46+
3. shared escaped-failure patterns derived from real PRs;
47+
4. strict-fake, database-integrity, tenant-isolation, and canonical-rule reuse
48+
obligations;
49+
5. specialty-specific blind evaluations proving the requirements change
50+
reviewer behavior.
51+
52+
## Alternatives considered
53+
54+
- Add more generic reviewers: rejected because overlapping prompts reproduce
55+
the same blind spots and increase latency.
56+
- Require PostgreSQL and full fanout for every PR: rejected as disproportionate.
57+
- Depend on CodeRabbit to catch internal misses: rejected because it is an
58+
independent external sensor, can rate-limit or skip, and does not own repo
59+
architecture.
60+
- Add a new contribution permission or merge gate: rejected. GitHub remains the
61+
only repository authority.
62+
63+
## Boundaries preserved
64+
65+
- No Workstream product, API, authorization, payment, artifact, review, task,
66+
contribution, compensation, or database behavior changes.
67+
- No new hosted service, secret, model provider, merge automation, signed start,
68+
loop memory, or post-merge reconciliation.
69+
- Reviewer routing remains proportionate; low-risk work does not receive
70+
ceremonial fanout.
71+
- CodeRabbit, CI, internal review, and human review remain independent sensors.
72+
73+
## Expected risks
74+
75+
- Turning useful heuristics into rigid bureaucracy.
76+
- Misclassifying proof strength and forcing infrastructure where pure proof is
77+
sufficient.
78+
- Adding token-heavy duplicated instructions to every skill and agent.
79+
- Creating fixtures that leak their expected answer to reviewers.
80+
- Mistaking a proof taxonomy for proof that a reviewer reasoned correctly.
81+
82+
## What must not change
83+
84+
The engineering loop remains:
85+
86+
```text
87+
Intent -> Plan -> Bounded Change -> Tests -> Review -> PR -> Human Merge
88+
```
89+
90+
One implementation chunk remains one PR. Planning artifacts explain work; they
91+
do not authorize contribution or merge.
92+
93+
## How this will be proven
94+
95+
- Validator mutation tests remove each proof-quality requirement and fail.
96+
- Raw blind fixtures replay the missed defects from PRs #338, #346, and #349.
97+
- Tests prove mock proof cannot satisfy database, transaction, concurrency, or
98+
real-isolation custody.
99+
- Tests prove each specialty catches its owned failure while valid controls
100+
remain clear.
101+
- Final-head reviewer replay validates the initiative against its own rules.
102+
103+
## Human decisions required
104+
105+
- Approve this plan before implementation.
106+
- Approve each implementation PR and any explicitly accepted Medium risk.

0 commit comments

Comments
 (0)