diff --git a/.agent-loop/LOOP_STATE.md b/.agent-loop/LOOP_STATE.md index 38edc02cd..4621387ca 100644 --- a/.agent-loop/LOOP_STATE.md +++ b/.agent-loop/LOOP_STATE.md @@ -4,21 +4,17 @@ - Active initiative: `WS-POL-001` - Submission Artifact Policy Foundation - Active planning chunk: none -- Active implementation chunk: `WS-POL-001-15` - Agent Derivation Policy - Conflict Hardening -- Branch: `codex/ws-pol-001-15-agent-derivation-hardening` -- Status: The accepted no-DB Terminal Benchmark live API drill from `main` - reached project/guide/snapshot/sufficiency, then failed during - agent-derived submission artifact policy creation because the agent emitted a - required artifact and forbidden artifact pattern that matched each other. -- Last merged implementation SHA: `ebf9d1d` -- Last merge commit: `53a57c3` -- Current gate: write internal review evidence, run the evidence gate, and open - a PR for human review after the derivation contract hardening, full project - tests, internal reviewers, and accepted no-DB Terminal Benchmark live API - drill passed on this branch. -- Next chunk: inactive until this corrective chunk is merged and the user - explicitly starts it. +- Active implementation chunk: none +- Branch: `main` +- Status: `WS-POL-001-15` merged through PR #81. The project setup derivation + prompt now explicitly prevents required/forbidden artifact self-conflicts, + keeps derivation project-scoped, and the accepted no-DB Terminal Benchmark + live API drill passes after hardening. +- Last merged implementation SHA: `b72a5b9` +- Last merge commit: `b1a9851` +- Current gate: post-merge memory update for PR #81, then stop for the user's + next explicit implementation chunk. +- Next chunk: inactive until the user explicitly starts it. ## Operating Rule diff --git a/.agent-loop/REVIEW_LOG.md b/.agent-loop/REVIEW_LOG.md index 73ba8f967..a00448435 100644 --- a/.agent-loop/REVIEW_LOG.md +++ b/.agent-loop/REVIEW_LOG.md @@ -418,8 +418,29 @@ External review response: `.agent-loop/initiatives/WS-POL-001-submission-artifac External review status: CodeRabbit comments triaged; valid findings fixed; GitHub checks and CodeRabbit passed before merge. -Next gate update: the accepted no-DB Terminal Benchmark live API drill from -`main` exposed an agent-derived submission artifact policy self-conflict during -policy derivation. Corrective chunk `WS-POL-001-15` is now active to harden the -derivation contract before any next implementation chunk starts. The corrective -branch has rerun that live API drill successfully after hardening. +Next gate update before PR #81: the accepted no-DB Terminal Benchmark live API +drill from `main` exposed an agent-derived submission artifact policy +self-conflict during policy derivation. Corrective chunk `WS-POL-001-15` +hardened the derivation contract and reran that live API drill successfully. + +## 2026-07-08 - WS-POL-001-15 Merged + +PR #81 merged into `main` as `b1a9851a5fe00580b704fe42bdeb511638dfe688`. + +Result: PASS after internal review, CodeRabbit, Agent Gates, and Backend checks. + +Scope: hardened the OpenAI Agents SDK submission artifact policy derivation +prompt so it produces a project-level worker submission contract, avoids +required/forbidden self-conflicts, avoids secret-like required fields, and +requires exact safe relative artifact paths. Server-side default forbidden +artifact validation remains fail-closed. + +Evidence: `.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-15-internal-review-evidence.md` + +External review response: `.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-15-external-review-response.md` + +External review status: CodeRabbit generated no actionable code comments; the +description warning was fixed before merge. + +Next gate: no active implementation chunk. Wait for the user's next explicit +chunk start signal. diff --git a/.agent-loop/WORK_QUEUE.md b/.agent-loop/WORK_QUEUE.md index f1b73473b..9ea7fc571 100644 --- a/.agent-loop/WORK_QUEUE.md +++ b/.agent-loop/WORK_QUEUE.md @@ -4,7 +4,7 @@ | Chunk | Title | Risk | Status | |---|---|---:|---| -| `WS-POL-001-15` | Agent Derivation Policy Conflict Hardening | L1 | Active on `codex/ws-pol-001-15-agent-derivation-hardening` | +| none | none | - | Waiting for user to explicitly start the next chunk | ## Completed @@ -27,13 +27,12 @@ | `WS-POL-001-12` | Project Setup And Policy Visibility APIs | L1 | Merged through PR #76 as `46e74de` | | `WS-POL-001-13` | Task Context And Submission Requirement APIs | L1 | Merged through PR #77 as `b567bac` on 2026-07-08 | | `WS-POL-001-14` | Submission Finalize And No-DB Terminal Benchmark Proof | L1 | Merged through PR #79 as `53a57c3` on 2026-07-08 | +| `WS-POL-001-15` | Agent Derivation Policy Conflict Hardening | L1 | Merged through PR #81 as `b1a9851` on 2026-07-08 | ## Proposed Next -Complete `WS-POL-001-15` after the accepted no-DB Terminal Benchmark live API -drill exposed an agent-derived submission artifact policy self-conflict. Do not -start the next implementation chunk until this corrective PR is merged and the -user explicitly approves the next chunk. +Stop after the PR #81 post-merge memory update. Do not start the next +implementation chunk until the user explicitly starts it. ## Blocked diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/STATUS.md b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/STATUS.md index 3e3677a40..6e6ac3fc2 100644 --- a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/STATUS.md +++ b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/STATUS.md @@ -2,21 +2,19 @@ ## Current Status -`WS-POL-001-01` through `WS-POL-001-14` are merged to `main`. +`WS-POL-001-01` through `WS-POL-001-15` are merged to `main`. The post-actor-registry Terminal Benchmark live API drill passed through real HTTP calls, and task context visibility is now exposed through APIs. `WS-POL-001-14` replaced public submission lock wording with finalization, defined system actor audit semantics, and merged PR #79's HTTP-visible Terminal Benchmark proof evidence. The accepted post-merge no-DB Terminal Benchmark drill from `main` exposed a self-conflicting agent-derived submission artifact -policy, now tracked as corrective chunk `WS-POL-001-15`. -The corrective branch now reruns that accepted drill successfully after -hardening the derivation contract. +policy. Corrective chunk `WS-POL-001-15` hardened the derivation contract and +reran that accepted drill successfully before merging through PR #81. ## Active Chunk -`WS-POL-001-15` on branch -`codex/ws-pol-001-15-agent-derivation-hardening`. +None. Waiting for the user's next explicit implementation chunk. ## Chunk Status @@ -36,7 +34,7 @@ hardening the derivation contract. | `WS-POL-001-12` | Merged | `codex/ws-pol-001-12-project-setup-policy-visibility` | 76 | Adds project setup-run and project policy visibility APIs for setup runs, sufficiency reports, submission artifact policies, effective policy, and compiled project pre-submit checker policy. | | `WS-POL-001-13` | Merged | `codex/ws-pol-001-13-task-context-apis` | 77 | Adds task work-context, worker submission-requirements, and operator-only locked-context APIs. | | `WS-POL-001-14` | Merged | `codex/ws-pol-001-14-submission-finalize` | 79 | Replaces public submission lock with finalize, defines system actor audit semantics, scopes operator visibility, and proves the Terminal Benchmark flow through HTTP-visible lifecycle responses. | -| `WS-POL-001-15` | Active | `codex/ws-pol-001-15-agent-derivation-hardening` | - | Hardens agent-derived submission artifact policy instructions after the no-DB Terminal Benchmark drill exposed a required-artifact/forbidden-pattern self-conflict. | +| `WS-POL-001-15` | Merged | `codex/ws-pol-001-15-agent-derivation-hardening` | 81 | Hardens agent-derived submission artifact policy instructions after the no-DB Terminal Benchmark drill exposed a required-artifact/forbidden-pattern self-conflict. | ## Blockers diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-15-post-merge-memory-external-review-response.md b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-15-post-merge-memory-external-review-response.md new file mode 100644 index 000000000..9e20bd4cd --- /dev/null +++ b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-15-post-merge-memory-external-review-response.md @@ -0,0 +1,65 @@ +# External Review Response: WS-POL-001-15 Post-Merge Memory + +## Comments Addressed + +CodeRabbit review on PR #82 reported one pre-merge description warning: + +- PR description covered summary and validation, but omitted template sections + such as Chunk, Goal, Human-Approved Intent, What Changed, Why It Changed, + Design Chosen, Scope Control, Evidence, Acceptance Criteria, Test Delta, and + reviewer details. + +Resolution: + +- Updated PR #82 description to include the missing template sections. + +CodeRabbit later posted three actionable wording comments: + +- The internal review evidence claimed only the evidence file changed after the + reviewed SHA. +- The internal review evidence used future-tense external review wording even + though the PR and checks now exist. +- The work queue used `approves` where the loop state uses `starts` for the + explicit user signal that begins the next chunk. + +Resolution: + +- Updated the evidence file to state that the reviewed SHA contains the + WS-POL-001-15 memory updates and that post-review edits are limited to review + evidence artifacts for this memory-only chunk. +- Updated the external review separation section to describe the current PR + tracking state. +- Updated the work queue wording to say the next chunk waits until the user + explicitly starts it. + +Internal reviewer repair pass also found one low documentation cleanup: + +- `docs/roadmap_status.md` listed Chunk 15 before Chunks 11-14. + +Resolution: + +- Moved the Chunk 15 completed item after Chunk 14. + +## Comments Deferred + +None. + +## Human Decisions Needed + +None from external review. Human still decides whether PR #82 is acceptable to +merge. + +## Commands Rerun + +No runtime code changed. Existing checks before this response: + +```bash +INTERNAL_REVIEW_CHUNK_ID=WS-POL-001-15-post-merge-memory python3 scripts/check_internal_review_evidence.py +python3 scripts/check_stale_workstream_wording.py +python3 scripts/check_markdown_links.py +git diff --check +``` + +## Remaining Risks + +None from the CodeRabbit description warning. diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-15-post-merge-memory-internal-review-evidence.md b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-15-post-merge-memory-internal-review-evidence.md new file mode 100644 index 000000000..f895c1af8 --- /dev/null +++ b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-15-post-merge-memory-internal-review-evidence.md @@ -0,0 +1,76 @@ +# Internal Review Evidence: WS-POL-001-15 Post-Merge Memory + +## Chunk + +WS-POL-001-15-post-merge-memory + +open sub-agent sessions: none + +valid findings addressed: yes + +## Reviewed Revision + +Reviewed code SHA: 60e70e5508b77e58f5cb97b13a9b67141f7769ac + +Reviewed at: 2026-07-08T19:37:52Z + +Reviewer run IDs: senior-engineering-019f4339-308b-7553-8a34-8f2ab3c88531, qa-test-019f4339-35a7-7120-8477-07a2c401d60f, security-auth-019f4339-41ab-75a3-a8bf-522adb985cb8, product-ops-019f4339-56a7-7fc3-a21f-a13a4005e6a2, docs-019f4339-650b-7572-a227-576d76c70205, architecture-019f4339-734c-74d2-bcaf-6774182fa15a + +The reviewed SHA contains the loop state, work queue, review log, initiative +status, review response, and roadmap memory updates for `WS-POL-001-15`. +Post-review edits are limited to review evidence artifacts for this same +memory-only chunk. + +## Reviewed Change + +Scope: + +- Marks `WS-POL-001-15` as merged through PR #81. +- Records PR #81 merge commit `b1a9851a5fe00580b704fe42bdeb511638dfe688`. +- Records implementation SHA `b72a5b9`. +- Clears active implementation chunk state. +- Sets next chunk to inactive until the user explicitly starts it. +- Moves `WS-POL-001-15` into completed work queue/status/roadmap memory. +- Records that the accepted no-DB Terminal Benchmark live API drill now passes after derivation hardening. + +## Reviewer Results + +| Reviewer | Result | Blocking findings | Notes | +|---|---:|---|---| +| senior engineering | PASS WITH LOW RISKS | None | Confirmed memory/docs content is safe for evidence refresh; no maintainability finding beyond the expected stale evidence before this file update. | +| QA/test | PASS WITH LOW RISKS | None | Confirmed active chunk none, next chunk inactive until user start, PR #81 details recorded, roadmap order fixed, and review separation clear. | +| security/auth | PASS WITH LOW RISKS | None | Confirmed no auth, payment, secrets, tenant boundary, runtime, CI, migration, or policy-enforcement change. | +| product/ops | PASS WITH LOW RISKS | None | Confirmed operational state is correct, no product lifecycle confusion, and no next chunk is active. | +| architecture | PASS WITH LOW RISKS | None | Confirmed no product implementation drift and evidence can be refreshed after this review pass. | +| docs | PASS | None | Confirmed internal/external review separation, roadmap ordering, and stale wording/link/whitespace checks. | + +## Valid Findings Addressed + +- Added the missing completed `WS-POL-001-15` roadmap bullet so `docs/roadmap_status.md` does not end the completed list at Chunk 14 while later mentioning Chunk 15. +- Added this post-merge memory internal review evidence file so engineering-loop state changes satisfy the internal evidence gate. +- Updated CodeRabbit wording follow-ups: external review separation now describes current PR tracking, work queue uses `starts` for the explicit user signal, and evidence wording no longer claims only the evidence file changed after the old review. +- Moved the Chunk 15 roadmap item after Chunk 14 so the completed list reads chronologically. + +## Commands Run + +```bash +python3 scripts/check_stale_workstream_wording.py +python3 scripts/check_markdown_links.py +git diff --check +``` + +Results: + +- Stale wording check: passed. +- Markdown link check: passed for 5 changed Markdown files. +- Diff whitespace check: passed. + +## External Review Separation + +External review is tracked separately from this internal reviewer evidence. +CodeRabbit comments and GitHub checks are recorded in the PR and the external +review response artifact. + +## Remaining Risks + +None known. This branch is memory-only and does not change runtime behavior. diff --git a/docs/roadmap_status.md b/docs/roadmap_status.md index 1f7377d57..69985bd38 100644 --- a/docs/roadmap_status.md +++ b/docs/roadmap_status.md @@ -47,6 +47,7 @@ Current phase: Week 3 review and revision preparation. - Chunk 12 project setup-run and project policy visibility APIs for setup runs, sufficiency reports, submission artifact policies, effective policy, and compiled project pre-submit checker policy. - Chunk 13 task work-context, worker submission-requirements, and operator-only locked-context APIs. - Chunk 14 submission finalization, system actor pre-review gate audit semantics, scoped operator visibility, and HTTP-visible Terminal Benchmark proof. +- Chunk 15 agent-derivation hardening after the accepted no-DB Terminal Benchmark drill exposed a required/forbidden self-conflict. ## Review Tracks Closed @@ -66,9 +67,9 @@ Current phase: Week 3 review and revision preparation. - Week 3 must keep review decisions canonical: `accept`, `needs_revision`, and `reject`. - `needs_revision` from human review must carry `outcome_source = human_review` and a review decision id; checker-caused `needs_revision` keeps `outcome_source = auto_checker`. - Review findings, revision replay, and reviewer-quality metrics are the next backend contracts to lock. -- The accepted no-DB Terminal Benchmark drill exposed an agent-derived - submission artifact policy self-conflict from `main`; `WS-POL-001-15` hardens - the derivation contract before the next implementation chunk starts. +- `WS-POL-001-15` hardened the agent-derived submission artifact policy + contract after the accepted no-DB Terminal Benchmark drill exposed a + required/forbidden self-conflict; the drill now passes after hardening. ## Pending Before Pilot