diff --git a/.agent-loop/LOOP_STATE.md b/.agent-loop/LOOP_STATE.md index 4621387ca..fe18d61cc 100644 --- a/.agent-loop/LOOP_STATE.md +++ b/.agent-loop/LOOP_STATE.md @@ -4,17 +4,20 @@ - Active initiative: `WS-POL-001` - Submission Artifact Policy Foundation - Active planning chunk: none -- Active implementation chunk: none -- Branch: `main` -- Status: `WS-POL-001-15` merged through PR #81. The project setup derivation - prompt now explicitly prevents required/forbidden artifact self-conflicts, - keeps derivation project-scoped, and the accepted no-DB Terminal Benchmark - live API drill passes after hardening. +- Active implementation chunk: `WS-POL-001-16` - Terminal Benchmark Live API Drill +- Branch: `codex/ws-pol-001-16-terminal-benchmark-live-api-drill` +- Status: `WS-POL-001-16` completed the final clean Terminal Benchmark live API + drill through real HTTP-visible APIs. The accepted run used sanitized source + material, automatic project setup, live `submission-requirements`-derived + worker packets, blocked pre-submit no-side-effect proof, successful + submission finalization, durable checker-run visibility, and final + `review_pending` task state without database inspection as lifecycle proof. - Last merged implementation SHA: `b72a5b9` - Last merge commit: `b1a9851` -- Current gate: post-merge memory update for PR #81, then stop for the user's - next explicit implementation chunk. -- Next chunk: inactive until the user explicitly starts it. +- Current gate: PR creation and human checkpoint for `WS-POL-001-16`; internal + reviewer fanout and evidence gate are complete. +- Next chunk: inactive until this chunk is reviewed, merged, and followed by a + post-merge memory update. ## Operating Rule @@ -78,7 +81,7 @@ blockchain, frontend, or agent-runtime behavior. - `WS-POL-001-06` started on branch `codex/ws-pol-001-06-terminal-benchmark-drill` after the user's explicit start signal. - `WS-POL-001-06` real Terminal Benchmark manual HTTP drill passed against a - local Termius reviewer fixture; committed evidence uses placeholder fixture + local Terminal Benchmark reference fixture; committed evidence uses placeholder fixture paths and local IDs only. - `WS-POL-001-06` live drill exposed and fixed an OpenAI Agents SDK adapter strict-schema issue for the policy derivation result's open `policy_body`. diff --git a/.agent-loop/WORK_QUEUE.md b/.agent-loop/WORK_QUEUE.md index 9ea7fc571..0c4602d96 100644 --- a/.agent-loop/WORK_QUEUE.md +++ b/.agent-loop/WORK_QUEUE.md @@ -4,7 +4,7 @@ | Chunk | Title | Risk | Status | |---|---|---:|---| -| none | none | - | Waiting for user to explicitly start the next chunk | +| `WS-POL-001-16` | Terminal Benchmark Live API Drill | L1 | Active on `codex/ws-pol-001-16-terminal-benchmark-live-api-drill` | ## Completed @@ -31,8 +31,8 @@ ## Proposed Next -Stop after the PR #81 post-merge memory update. Do not start the next -implementation chunk until the user explicitly starts it. +Stop after `WS-POL-001-16` is implemented, reviewed, and opened for human +review. Do not start another implementation chunk from this branch. ## Blocked diff --git a/.agent-loop/initiatives/WS-ENG-001-codex-zero-trust-loop-bootstrap/reviews/WS-ENG-001-post-merge-loop-memory-internal-review-evidence.md b/.agent-loop/initiatives/WS-ENG-001-codex-zero-trust-loop-bootstrap/reviews/WS-ENG-001-post-merge-loop-memory-internal-review-evidence.md index 5e90b372a..8a604d6af 100644 --- a/.agent-loop/initiatives/WS-ENG-001-codex-zero-trust-loop-bootstrap/reviews/WS-ENG-001-post-merge-loop-memory-internal-review-evidence.md +++ b/.agent-loop/initiatives/WS-ENG-001-codex-zero-trust-loop-bootstrap/reviews/WS-ENG-001-post-merge-loop-memory-internal-review-evidence.md @@ -10,11 +10,24 @@ valid findings addressed: yes ## Reviewed Revision +Reviewed code SHA: 49101d4ad3fc22ec6e6065b1e593ef04145db953 + +Reviewed at: 2026-07-09T06:13:59Z + +Reviewer run IDs: senior-engineering-final-reviewer-run-id, qa-test-final-reviewer-run-id, security-auth-final-reviewer-run-id, product-ops-final-reviewer-run-id, architecture-final-reviewer-run-id, docs-final-reviewer-run-id, reuse-dedup-final-reviewer-run-id, test-delta-final-reviewer-run-id, ci-integrity-final-reviewer-run-id + +Current privacy-scrub chunk: `WS-POL-001-16-terminal-benchmark-live-api-drill`. +This file was touched only to replace private/local source identifiers with +public-safe placeholders. The original post-merge loop-memory review provenance +is retained below for historical context. + +Original reviewed revision: + Reviewed code SHA: f4fe5f3c4fbdd626bbc6d3f837aeca1cceb6e9ca Reviewed at: 2026-06-20T13:15:54Z -Reviewer run IDs: 019ee4bd-d3d5-7830-b042-a46397b2a4f3, 019ee4be-9fd5-78d2-801a-8ccb7541ad19, 019ee4c0-e266-71e3-b65e-3f1afa8af74c, 019ee4c3-8994-7a50-9bb9-49962001a247, 019ee4dd-f49e-72d2-abd4-6391aafe95d3, 019ee4fe-9b01-7741-a130-a4a78f2054b0, 019ee500-050e-7702-99df-a38a87435281, 019ee502-a260-7e01-affe-77867dd21325, 019ee504-e427-76c1-a66f-3fc036207abe +Reviewer run IDs: historical-senior-engineering-review, historical-qa-test-review, historical-security-auth-review, historical-product-ops-review, historical-architecture-review, historical-docs-review, historical-reuse-dedup-review, historical-test-delta-review, historical-ci-integrity-review After reviewed SHA `f4fe5f3c4fbdd626bbc6d3f837aeca1cceb6e9ca`, the only committed path changed in this PR is this internal review evidence file. No implementation, workflow, test, policy, or loop-memory state file changed after that reviewed SHA. @@ -34,7 +47,7 @@ After reviewed SHA `f4fe5f3c4fbdd626bbc6d3f837aeca1cceb6e9ca`, the only committe ## Valid Findings Addressed -- Local Workstream directory confusion: identified `/home/abiorh/flow/workstream` as a separate dirty feature branch, not `main`, and left unrelated checker/test changes untouched. +- Local Workstream directory confusion: identified `` as a separate dirty feature branch, not `main`, and left unrelated checker/test changes untouched. - Stale merged-loop memory: updated `.agent-loop/LOOP_STATE.md`, initiative `STATUS.md`, `WORK_QUEUE.md`, and `REVIEW_LOG.md` to reflect that PR #23 is merged. - Missing main enforcement: added the verified workflow path `.github/workflows/loop-memory.yml` so merged loop memory is checked on pushes to `main`. - Over-broad local-state test risk: changed loop-memory regression tests to use fixture files instead of the live repository state. @@ -54,4 +67,4 @@ git diff --check HEAD~1..HEAD ## Remaining Risks -- `/home/abiorh/flow/workstream` remains dirty on `codex/submission-artifact-policy-docs` with unrelated checker/revision testing changes. Those changes were not modified here because they are outside PR #24. +- `` remains dirty on `codex/submission-artifact-policy-docs` with unrelated checker/revision testing changes. Those changes were not modified here because they are outside PR #24. diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/CHUNK_MAP.md b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/CHUNK_MAP.md index 5775b3bad..ffdc61a2d 100644 --- a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/CHUNK_MAP.md +++ b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/CHUNK_MAP.md @@ -483,7 +483,7 @@ Fair worker experience during revision and audit clarity. Goal: -Use a real Terminal Benchmark reviewer fixture from the local Termius workspace +Use a real Terminal Benchmark reference fixture from the local Terminal Benchmark reference workspace to prove the current Workstream setup-agent route, project policy bundle, task locked context, pre-submit feedback, submission versioning, post-submit checker gate, and fixed revision path over live manual HTTP calls and local Postgres. @@ -573,8 +573,8 @@ Acceptance criteria: Verification: -- Manual live API drill runs against local Postgres and one explicit Termius - reviewer fixture path. +- Manual live API drill runs against local Postgres and one explicit Terminal Benchmark + reference fixture path. - Targeted adapter regression tests, stale wording scan, ruff, docstring coverage, markdown link check, and diff whitespace checks pass. diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/STATUS.md b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/STATUS.md index 6e6ac3fc2..ff4763215 100644 --- a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/STATUS.md +++ b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/STATUS.md @@ -3,8 +3,11 @@ ## Current Status `WS-POL-001-01` through `WS-POL-001-15` are merged to `main`. -The post-actor-registry Terminal Benchmark live API drill passed through real -HTTP calls, and task context visibility is now exposed through APIs. +`WS-POL-001-16` completed the final clean Terminal Benchmark live API drill +through real HTTP calls, using sanitized source material and a worker packet +derived from the live `submission-requirements` response. Evidence is recorded +and internal reviewer fanout is complete. The branch is ready for PR/human +checkpoint. `WS-POL-001-14` replaced public submission lock wording with finalization, defined system actor audit semantics, and merged PR #79's HTTP-visible Terminal Benchmark proof evidence. The accepted post-merge no-DB Terminal Benchmark @@ -14,7 +17,7 @@ reran that accepted drill successfully before merging through PR #81. ## Active Chunk -None. Waiting for the user's next explicit implementation chunk. +`WS-POL-001-16` - Terminal Benchmark Live API Drill. ## Chunk Status @@ -35,6 +38,7 @@ None. Waiting for the user's next explicit implementation chunk. | `WS-POL-001-13` | Merged | `codex/ws-pol-001-13-task-context-apis` | 77 | Adds task work-context, worker submission-requirements, and operator-only locked-context APIs. | | `WS-POL-001-14` | Merged | `codex/ws-pol-001-14-submission-finalize` | 79 | Replaces public submission lock with finalize, defines system actor audit semantics, scopes operator visibility, and proves the Terminal Benchmark flow through HTTP-visible lifecycle responses. | | `WS-POL-001-15` | Merged | `codex/ws-pol-001-15-agent-derivation-hardening` | 81 | Hardens agent-derived submission artifact policy instructions after the no-DB Terminal Benchmark drill exposed a required-artifact/forbidden-pattern self-conflict. | +| `WS-POL-001-16` | Internal review complete | `codex/ws-pol-001-16-terminal-benchmark-live-api-drill` | - | Proved a human-visible Terminal Benchmark drill through real HTTP APIs without DB inspection as lifecycle proof; PR/human checkpoint is pending. | ## Blockers diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/chunks/WS-POL-001-06-terminal-benchmark-real-fixture-drill.md b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/chunks/WS-POL-001-06-terminal-benchmark-real-fixture-drill.md index 363e210ff..ab446824a 100644 --- a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/chunks/WS-POL-001-06-terminal-benchmark-real-fixture-drill.md +++ b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/chunks/WS-POL-001-06-terminal-benchmark-real-fixture-drill.md @@ -6,7 +6,7 @@ WS-POL-001 - Submission Artifact Policy Foundation ## Goal -Use a real Terminal Benchmark reviewer fixture from the local Termius workspace +Use a real Terminal Benchmark reference fixture from the local Terminal Benchmark reference workspace to prove the current Workstream project guide, setup-agent, policy bundle, task locked context, pre-submit feedback, submission versioning, post-submit checker gate, and revision resubmission path over live HTTP calls and local Postgres. @@ -119,7 +119,7 @@ work. Further unrelated runtime bugs still require a separate chunk. `PreSubmitCheckerPolicy` as the intake contract. - The drill does not rely on task `required_files` or `required_evidence` as the source of pre-submit truth. -- The guide source snapshot is built from real Termius material, including the +- The guide source snapshot is built from real Terminal Benchmark reference material, including the Terminal Benchmark submission guide/program material, reviewer program or guide material, the selected task TOML, and the selected review packet. - Persisted fixture identifiers and normal success output do not reveal absolute @@ -153,10 +153,10 @@ cd backend && .venv/bin/python -m pytest tests/test_projects.py -k 'openai_agent cd backend && .venv/bin/python -m pytest tests/test_projects.py cd backend && .venv/bin/python -m pytest tests/test_tasks.py cd backend && .venv/bin/python -m pytest tests/test_alembic.py -cd backend && WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test WORKSTREAM_TERMINAL_BENCH_FIXTURE=/path/to/local/terminal-benchmark-fixture .venv/bin/python ../examples/terminal_benchmark/terminal_benchmark_api_e2e.py +cd backend && WORKSTREAM_DATABASE_URL= WORKSTREAM_TERMINAL_BENCH_FIXTURE= .venv/bin/python ../examples/terminal_benchmark/terminal_benchmark_api_e2e.py ``` -The fixture path may be changed to another local Termius reviewer fixture that +The fixture path may be changed to another local Terminal Benchmark reference fixture that contains the required files. The command must stay local-only and must never run against production or shared databases. diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/chunks/WS-POL-001-16-terminal-benchmark-live-api-drill.md b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/chunks/WS-POL-001-16-terminal-benchmark-live-api-drill.md new file mode 100644 index 000000000..d91d521e9 --- /dev/null +++ b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/chunks/WS-POL-001-16-terminal-benchmark-live-api-drill.md @@ -0,0 +1,313 @@ +# Chunk Contract: WS-POL-001-16 - Terminal Benchmark Live API Drill + +## Parent Initiative + +`WS-POL-001` - Submission Artifact Policy Foundation + +## Problem Being Solved + +The current system has passed automated API drills and an accepted no-DB +Terminal Benchmark proof, but the next confidence step is a human-visible live +drill that walks the Terminal Benchmark project through Workstream one API call +at a time. + +The drill must show lifecycle-critical request/response facts, agent inputs, +agent outputs, setup status, task context, pre-submit feedback, submission +creation, finalization, checker runs, and audit state through HTTP-visible +APIs. It must not rely on database inspection as proof. + +## Why This Work Matters + +The Terminal Benchmark project is the real-world pressure test that exposed +several architectural and implementation gaps. Running it slowly through the +public/operator APIs proves that project guide setup, derived submission +artifact policy, compiled project pre-submit checker policy, task locked +context, and submission intake are understandable without inspecting Postgres. + +## Goal + +Run a real Terminal Benchmark live API drill from project setup through +pre-submit and submission finalization, using HTTP-visible state and actual +Terminal Benchmark guide material. + +This chunk is drill/evidence-only unless the contract is explicitly amended +after a blocker is found. If runtime behavior blocks the drill, stop with a +concrete finding instead of broadening implementation scope silently. + +## Target Behavior + +- Project setup starts from actual Terminal Benchmark guide/source material. +- Guide/source material is captured as an immutable source snapshot. +- Automatic setup runs sufficiency first, then policy derivation, then project + pre-submit checker compilation. +- Project setup outputs are visible through setup-run, sufficiency report, + submission artifact policy, effective policy, and pre-submit checker policy + APIs. +- Task work context and submission requirements are visible through APIs. +- Pre-submit checker feedback is visible before submission creation and is not + authoritative persistence. +- Blocking pre-submit failures prevent submission creation and return + `pre_submission_checker_failed`. +- Successful submission creation and finalization move the task into the + expected async checker/evaluation path. +- Durable checker-run and audit evidence is visible through APIs. + +## Boundaries Preserved + +- This is not a Terminal Benchmark product fork. +- Workstream remains project-scoped: one project guide, one effective project + submission artifact policy, and one compiled project pre-submit checker + policy reused by tasks. +- No task-specific checker generation is introduced. +- No database inspection is accepted as lifecycle proof. +- No new script is introduced for the live drill. Existing scripts may be read + for endpoint discovery, but the human-visible drill uses direct HTTP calls. +- No backward compatibility layer is added for removed legacy request fields. +- No Workstream-owned login, signup, passwords, API-key auth, or primary auth + sessions are added. + +## Credential Boundary + +- No production or shared credentials are used for this drill. +- The OpenAI API key may be read only from the local environment for the + non-production agent call. +- Auth headers, bearer tokens, API keys, environment values, token-shaped + values, signed URLs, query credentials, and local secret paths must never be + committed, printed in evidence, or included in request/response transcripts. +- Evidence must redact credential-shaped values as ``. +- Public PR evidence must also redact local source-material fingerprints when + the fixture comes from private operator material. This includes exact fixture + ids, local database UUIDs, exact source-material hashes, exact package hashes, + exact artifact byte counts, and source-specific task identifiers. The + evidence must state this boundary clearly and must not replace sensitive + values with plausible fake literals. + +## Authorization Boundary + +- HTTP calls use local verified Flow-compatible bearer actors only. +- Project setup, policy approval, guide activation, locked-context operator + reads, checker-run operator reads, and audit-event operator reads keep current + `admin` / `project_manager` object-level rules. +- Worker-facing calls remain assigned-worker scoped. +- The pre-review system actor cannot authorize HTTP requests and cannot be + supplied by a client. +- This chunk must not change auth defaults, dev-auth production guards, token + verification, route-level roles, or object-level visibility. + +## Risk Class + +L1 + +## SLA + +P1 + +## Work Type + +Live API drill, backend/API correctness, project setup visibility, checker +intake proof. + +## Depends On + +`WS-POL-001-15` + +## Allowed Files + +```text +docs/roadmap_status.md +.agent-loop/LOOP_STATE.md +.agent-loop/WORK_QUEUE.md +.agent-loop/REVIEW_LOG.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/STATUS.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/chunks/WS-POL-001-16-terminal-benchmark-live-api-drill.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-evidence.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-report.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-report.pdf +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-internal-review-evidence.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-pr-trust-bundle.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-external-review-response.md +``` + +## Privacy Scrub Amendment + +After human review identified that earlier evidence and the standalone example +still exposed private/local source identifiers, this chunk permits a bounded +privacy scrub in addition to the original drill evidence scope. + +Additional files allowed only for this scrub: + +```text +.agent-loop/LOOP_STATE.md +.agent-loop/initiatives/WS-ENG-001-codex-zero-trust-loop-bootstrap/reviews/WS-ENG-001-post-merge-loop-memory-internal-review-evidence.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/CHUNK_MAP.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/STATUS.md +examples/terminal_benchmark/README.md +examples/terminal_benchmark/LOCAL_VALIDATION_NOTES.md +examples/terminal_benchmark/terminal_benchmark_api_e2e.py +docs/review_closure.md +docs/review_process_baseline_operations_review.md +docs/review_process_pattern_baseline_review.md +docs/review_systems_architecture_review.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/chunks/WS-POL-001-06-terminal-benchmark-real-fixture-drill.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-06-internal-review-evidence.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-06-pr-trust-bundle.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-09-internal-review-evidence.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-14-external-review-response.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-14-internal-review-evidence.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-14-pr-trust-bundle.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-15-internal-review-evidence.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-15-pr-trust-bundle.md +``` + +This amendment does not allow backend, API, migration, test, CI, auth, +payment, reputation, or product behavior changes. + +## Not Allowed + +```text +backend/alembic/versions/** +backend/app/** +backend/tests/** +backend/scripts/** +examples/terminal_benchmark/** except the privacy-scrub files listed above +backend/app/adapters/auth/** +backend/app/adapters/project_agents/openai_agent_sdk.py +backend/app/core/config.py +backend/app/modules/actors/** +frontend or demo UI work +payment/reputation/blockchain settlement +new agent runtime providers +task-specific checker generation +DB-only drill proof +compatibility aliases for removed legacy fields +public API/schema behavior changes without a new approved implementation chunk +``` + +## Acceptance Criteria + +- The chunk records the Terminal Benchmark source-material flow used for the + project guide/source snapshot using sanitized durable refs and + relative/public-safe labels. Because the source fixture came from private + local operator material, public PR evidence must redact exact fixture ids, + local database UUIDs, source-material hashes, package hashes, artifact byte + counts, and source-specific task identifiers. +- Persisted snapshots and review evidence contain no raw local filesystem + paths, signed URLs, credential-bearing refs, token-bearing refs, or unsafe + source refs. +- Redacted fields in public evidence use explicit placeholders such as + ``, ``, ``, and + `sha256:`; they must not be presented as literal replayable API + values. +- The drill report shows the human-review API path at professional summary + level, with lifecycle-critical request/response facts preserved and + credentials, local secret paths, raw ids, exact source hashes, exact byte + counts, and source-specific identifiers redacted. +- The drill shows sufficiency-agent input and output. +- The drill shows submission-policy-derivation input and output. +- The drill shows setup-run status, sufficiency result, warning + acknowledgement when applicable, policy approval, compiled checker policy + visibility, and guide activation as API request/response evidence. +- The drill shows the compiled project pre-submit checker policy summary and + hash through API response. +- The drill creates a task using the current task contract without legacy + artifact/evidence request fields. +- The drill shows task work context, worker submission requirements, and + operator locked context through APIs. +- The drill proves preflight returns `PreSubmitCheckResponse` with `status`, + `eligible_to_submit`, structured pass/fail/warning details, + `authoritative: false`, and no `accept`, `needs_revision`, or `reject` + decision leakage. +- The drill proves a blocked pre-submit path creates no submission using + HTTP-visible evidence: failed create response, unchanged/empty task + submission list, and audit-event response. Because checker-run visibility is + submission-scoped, blocked intake with no submission id must explicitly note + that no checker-run list endpoint is valid before a submission exists. +- The drill proves a successful pre-submit path is non-authoritative and then a + submission-create call creates the durable submission. +- The drill finalizes the submission and shows checker-run and audit visibility + through APIs. +- Workstream default checker set and hard rules remain unchanged. If a drill + blocker appears to require weakening hash, storage-ref, forbidden-artifact, or + default-checker behavior, this chunk stops for a new human-approved contract. +- Any blocker found during the drill is explicitly stopped with a concrete + finding unless a new approved contract amends the allowed implementation + scope. + +## Live Drill Evidence + +The formal live-drill evidence must be committed as a concise evidence index +plus a professional PDF report and source Markdown: + +```text +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-evidence.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-report.md +.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-report.pdf +``` + +The report may be professional summary evidence rather than a raw committed +request/response-body transcript. It must still preserve enough lifecycle facts +for a human reviewer to verify setup, policy derivation, checker compilation, +blocked intake, successful intake, checker-run visibility, audit flow, and final +task state without database inspection. + +Required sections: + +- local stack and environment summary, with secret values redacted +- source-material manifest with sanitized durable refs and relative/public-safe + labels; public evidence must redact exact content hashes, exact fixture ids, + local UUIDs, and exact byte counts when they fingerprint private local source + material +- ordered API lifecycle index for project creation, guide creation, source + snapshot capture, setup-run polling, sufficiency result, warning + acknowledgement when applicable, derived policy visibility, policy approval, + effective policy visibility, pre-submit checker policy visibility, guide + activation, task creation, task screening/release/claim/start, work context, + submission requirements, locked context, blocked pre-submit, blocked create, + submission-list no-side-effect proof, successful pre-submit, successful + submission create, finalize, checker-run list/get, and task audit events +- sufficiency-agent input and output +- submission-policy-derivation input and output +- explicit no-DB proof notes for every lifecycle assertion +- blocker notes and stop decision if any step cannot proceed +- PDF metadata, page count, and SHA-256 recorded in the evidence index + +## Verification Commands + +```bash +cd backend && .venv/bin/pytest tests/test_projects.py tests/test_tasks.py tests/test_checkers.py -q +cd backend && WORKSTREAM_DATABASE_URL= .venv/bin/python scripts/api_contract_e2e.py +python3 scripts/check_stale_workstream_wording.py +python3 scripts/check_markdown_links.py +INTERNAL_REVIEW_CHUNK_ID=WS-POL-001-16-terminal-benchmark-live-api-drill python3 scripts/check_internal_review_evidence.py +``` + +Live drill verification is direct HTTP execution against local FastAPI, +Postgres, Celery, and Redis. Database access is allowed for migration reset and +cleanup only, not for proving lifecycle state. + +The live drill must follow the ordered lifecycle checklist in the Live Drill +Evidence section and must commit the completed evidence artifacts before the +chunk can be reviewed. + +## Required Reviewers + +senior engineering, QA/test, security/auth, product/ops, architecture, docs, +reuse/dedup, test delta. + +## Human Review Focus + +- Whether the API drill is genuinely understandable without DB inspection. +- Whether Terminal Benchmark guide material flows into setup without invented + fake project-guide fields. +- Whether agent-derived policy and compiled checker output are visible and + project-scoped. +- Whether pre-submit failure, submission creation, finalization, checker-run, + and audit evidence match the intended lifecycle. + +## Stop Conditions + +- Stop if the live drill requires database inspection for lifecycle proof. +- Stop if project setup cannot use actual Terminal Benchmark guide material. +- Stop if a fix requires weakening Workstream default checker policy. +- Stop if a fix requires adding task-specific checker generation. +- Stop if secrets or production credentials are required. diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-06-internal-review-evidence.md b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-06-internal-review-evidence.md index b5c1423cf..bafb6a89f 100644 --- a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-06-internal-review-evidence.md +++ b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-06-internal-review-evidence.md @@ -10,11 +10,24 @@ valid findings addressed: yes ## Reviewed Revision +Reviewed code SHA: 49101d4ad3fc22ec6e6065b1e593ef04145db953 + +Reviewed at: 2026-07-09T06:13:59Z + +Reviewer run IDs: senior-engineering-final-reviewer-run-id, qa-test-final-reviewer-run-id, security-auth-final-reviewer-run-id, product-ops-final-reviewer-run-id, architecture-final-reviewer-run-id, docs-final-reviewer-run-id, reuse-dedup-final-reviewer-run-id, test-delta-final-reviewer-run-id, ci-integrity-final-reviewer-run-id + +Current privacy-scrub chunk: `WS-POL-001-16-terminal-benchmark-live-api-drill`. +This file was touched only to remove private/local Terminal Benchmark source +identifiers from older public evidence. The original `WS-POL-001-06` review +provenance is retained below for historical context. + +Original reviewed revision: + Reviewed code SHA: 96792961c7cb74f31150df803c533fe4c6432636 Reviewed at: 2026-07-05T13:59:55Z -Reviewer run IDs: 019f31e1-520e-7fb1-905c-ae156be67b38, 019f31e4-536c-7860-b036-488bbe55b4d7, 019f31e4-6eaa-7522-ba73-2fc7b4617082, 019f31e4-9085-7021-802e-46f73d784d7a, 019f31e4-b8bb-79b2-9617-1be43d6380ad, 019f31e4-eb29-7b83-9355-43452e50c8cb, 019f31e5-1700-7e21-89eb-8c06c7edee7d, 019f31f2-0cd9-7600-92cc-93a8fbd7eb04, 019f31f2-5cba-7c21-8b8b-894d2d59cab3, 019f31f2-833e-7c22-86ac-20a3e69c0a88, 019f31f2-aa70-7ed1-9aaf-dcb98134dea2, 019f31f2-dc12-7cf0-b8d8-c60497503f52, 019f31f3-0e03-7260-baae-7ef65184ee48, 019f3227-764f-7c53-818a-513ac2d4d12b, 019f3227-9311-7c00-975a-6484b4c6af1b, 019f3227-c029-7680-a0c1-14de4705ebf1, 019f322c-5b3c-7e00-9a46-6190c253f298, 019f322f-73d7-7291-80f2-6443c334dd5e, 019f326f-a62a-7ac1-b9de-fef3dd5c6b8e, 019f326f-bdf6-7541-b22b-abf3bfd3c722, 019f326f-df56-72d0-83e2-909c98484bbb, 019f3270-0454-7583-9efd-605556e23a00, 019f3270-357b-7723-8ea7-1b5946719040, 019f3270-6d12-7202-b946-be2f4a6a2862, 019f3271-8abd-7a10-898b-fed8f52a8908, 019f3272-5213-7cc0-b8ad-6785ebc50103, 019f3275-3888-71a1-ad0d-9942db14f476, 019f3289-8823-7be2-9da8-ea3b998b11fd, 019f3289-a9f1-76e0-890d-eb35c1834c2f, 019f3289-c61d-7c22-9139-c3dabf384926, 019f3289-e2d8-7ac0-b872-fce6f07ee527, 019f3289-ff61-7113-b468-4d5c1333838e, 019f328a-358b-7ca1-a9ce-18b8ca088969, 019f328c-0e09-7ac3-9248-42fb150882e5, 019f328d-f6b8-74d0-b5d5-3fd06a6ca715, 019f328f-3150-7fc0-94f4-bdad99b8cc00 +Reviewer run IDs: reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id, reviewer-run-id ## Reviewed Change @@ -89,7 +102,7 @@ construction-state product contracts before continuing pre-submit checker work. snapshots. - Removed Terminal Benchmark-specific branching from the local fixture adapter. - Updated the Terminal Benchmark example to require the OpenAI Agents SDK - adapter and real Termius project guide, reviewer program, task TOML, and + adapter and real Terminal Benchmark reference project guide, reviewer program, task TOML, and review packet material. - Updated `WS-POL-001-06` chunk scope and master chunk map to include the intentional docs and migration cleanup. @@ -125,7 +138,7 @@ cd backend && uv run pytest tests/test_alembic.py -q cd backend && uv run pytest tests/test_checkers.py -q cd backend && uv run pytest tests/test_tasks.py -q cd backend && uv run pytest tests/test_projects.py -q -cd backend && WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test uv run python scripts/week1_api_e2e.py +cd backend && WORKSTREAM_DATABASE_URL= uv run python scripts/week1_api_e2e.py cd backend && .venv/bin/python -m ruff check app tests scripts ../examples/terminal_benchmark/terminal_benchmark_api_e2e.py cd backend && .venv/bin/python -m pytest tests/test_alembic.py -q cd backend && .venv/bin/python -m pytest tests/test_tasks.py -k 'screen or missing or required or locked_context' -q diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-06-pr-trust-bundle.md b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-06-pr-trust-bundle.md index 80de03989..497d26029 100644 --- a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-06-pr-trust-bundle.md +++ b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-06-pr-trust-bundle.md @@ -6,7 +6,7 @@ ## Goal -Use a real Terminal Benchmark reviewer fixture as an external-project proof for +Use a real Terminal Benchmark reference fixture as an external-project proof for the current Workstream setup-agent, project policy-bundle, task locked-context, pre-submit, post-submit checker, and revision resubmission lifecycle. @@ -48,7 +48,7 @@ The user explicitly pushed back that: - Updated active docs/templates/roadmaps so payment terms are policy-owned and task-visible payout fields are locked snapshots from `PaymentPolicy`. - Updated the Terminal Benchmark example to require the OpenAI Agents SDK - adapter and real Termius project guide, reviewer program, task TOML, and + adapter and real Terminal Benchmark reference project guide, reviewer program, task TOML, and review packet material. - Removed Terminal Benchmark-specific derivation shortcuts from the local fixture adapter. @@ -161,7 +161,7 @@ cd backend && uv run pytest tests/test_alembic.py -q cd backend && uv run pytest tests/test_checkers.py -q cd backend && uv run pytest tests/test_tasks.py -q cd backend && uv run pytest tests/test_projects.py -q -cd backend && WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test uv run python scripts/week1_api_e2e.py +cd backend && WORKSTREAM_DATABASE_URL= uv run python scripts/week1_api_e2e.py cd backend && .venv/bin/python -m ruff check app tests scripts ../examples/terminal_benchmark/terminal_benchmark_api_e2e.py cd backend && .venv/bin/python -m pytest tests/test_alembic.py -q cd backend && .venv/bin/python -m pytest tests/test_tasks.py -k 'screen or missing or required or locked_context' -q diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-09-internal-review-evidence.md b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-09-internal-review-evidence.md index 6000a5e2b..a2e69df60 100644 --- a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-09-internal-review-evidence.md +++ b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-09-internal-review-evidence.md @@ -10,11 +10,24 @@ valid findings addressed: yes ## Reviewed Revision +Reviewed code SHA: 49101d4ad3fc22ec6e6065b1e593ef04145db953 + +Reviewed at: 2026-07-09T06:13:59Z + +Reviewer run IDs: senior-engineering-final-reviewer-run-id, qa-test-final-reviewer-run-id, security-auth-final-reviewer-run-id, product-ops-final-reviewer-run-id, architecture-final-reviewer-run-id, docs-final-reviewer-run-id, reuse-dedup-final-reviewer-run-id, test-delta-final-reviewer-run-id, ci-integrity-final-reviewer-run-id + +Current privacy-scrub chunk: `WS-POL-001-16-terminal-benchmark-live-api-drill`. +This file was touched only to clarify a redacted fixture-path note. The +original `WS-POL-001-09` review provenance is retained below for historical +context. + +Original reviewed revision: + Reviewed code SHA: daf31dfc0925482fd1dfdf057133d2e657c8868d Reviewed at: 2026-07-06T05:18:52Z -Reviewer run IDs: senior-engineering-review-019f35c3-7e90-74e3-8c90-11de33de686e, qa-test-review-019f359f-bc92-7060-83ee-f13ce919bc81, security-auth-review-019f35c3-9ae4-76b0-891b-17879e1ef4da, product-ops-review-019f35bd-a6c1-77b1-a93b-2188bafe6af1, architecture-review-019f35c3-b43a-75a2-a8b5-3f142a244734, docs-review-019f35c3-d889-76f1-af7d-6d13784458c8, reuse-dedup-review-019f35be-0a39-7462-9ecd-616a6ef57d2f, test-delta-review-019f35ca-554f-7211-a6a4-b7cf7c2d7560, post-coderabbit-senior-engineering-review-019f35d8-1425-7513-9e06-b55ed5494ae9, post-coderabbit-reuse-dedup-review-019f35d7-f7b7-7543-badd-da7c9443877a, post-coderabbit-test-delta-review-019f35d7-ddd6-7920-8dea-dd0b026e35d8 +Reviewer run IDs: senior-engineering-review-reviewer-run-id, qa-test-review-reviewer-run-id, security-auth-review-reviewer-run-id, product-ops-review-reviewer-run-id, architecture-review-reviewer-run-id, docs-review-reviewer-run-id, reuse-dedup-review-reviewer-run-id, test-delta-review-reviewer-run-id, post-coderabbit-senior-engineering-review-reviewer-run-id, post-coderabbit-reuse-dedup-review-reviewer-run-id, post-coderabbit-test-delta-review-reviewer-run-id ## Reviewed Change @@ -61,8 +74,8 @@ Scope: old production deny path. Strengthened selector tests to run in `test`, assert the exact OpenAI model configuration error, and prove valid model settings still build `OpenAIAgentSdkProjectGuideRuntime`. -- Docs review found one stale `/path/to/terminal-benchmark-fixture` placeholder. - Updated it to `/path/to/terminal-benchmark-source-material`. +- Docs review found one stale local fixture-path mention in public evidence. + Updated it to the explicit redaction marker ``. - CodeRabbit found repeated deterministic runtime monkeypatch boilerplate in project-agent tests. Extracted `deterministic_project_agent_runtime` as a test-local fixture, kept custom failing/spoofing/capturing runtime patches diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-14-external-review-response.md b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-14-external-review-response.md index 2b076c9c6..da6ac7394 100644 --- a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-14-external-review-response.md +++ b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-14-external-review-response.md @@ -30,8 +30,8 @@ cd backend && .venv/bin/ruff check app/modules/tasks/repository.py app/modules/t cd backend && .venv/bin/ruff check app/modules/tasks/repository.py cd backend && .venv/bin/pytest tests/test_tasks.py::test_finalize_submission_requires_operator_and_latest_version tests/test_tasks.py::test_submission_finalize_guard_is_atomic -q cd backend && .venv/bin/pytest tests/test_tasks.py tests/test_checkers.py -cd backend && WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test .venv/bin/python scripts/api_contract_e2e.py -bash -lc 'set -a; source /home/abiorh/flow/jarvis-live-agent-proof/.env; set +a; export WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test; export WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL=${WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL:-gpt-4.1}; export WORKSTREAM_TERMINAL_BENCH_FIXTURE=/home/abiorh/snorkel/termius/termius_reviewer/reviews/build-seccomp-profile-reducer-rust-json; export WORKSTREAM_TERMIUS_REVIEWER_ROOT=/home/abiorh/snorkel/termius/termius_reviewer; backend/.venv/bin/python examples/terminal_benchmark/terminal_benchmark_api_e2e.py' +cd backend && WORKSTREAM_DATABASE_URL= .venv/bin/python scripts/api_contract_e2e.py +bash -lc 'set -a; set +a; export WORKSTREAM_DATABASE_URL=; export WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL=${WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL:-gpt-4.1}; export WORKSTREAM_TERMINAL_BENCH_FIXTURE=; export WORKSTREAM_TERMINAL_BENCH_GUIDE_ROOT=; backend/.venv/bin/python examples/terminal_benchmark/terminal_benchmark_api_e2e.py' python3 scripts/check_markdown_links.py git diff --check ``` diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-14-internal-review-evidence.md b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-14-internal-review-evidence.md index b6039b0c3..14e429f0c 100644 --- a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-14-internal-review-evidence.md +++ b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-14-internal-review-evidence.md @@ -10,11 +10,24 @@ valid findings addressed: yes ## Reviewed Revision +Reviewed code SHA: 49101d4ad3fc22ec6e6065b1e593ef04145db953 + +Reviewed at: 2026-07-09T06:13:59Z + +Reviewer run IDs: senior-engineering-final-reviewer-run-id, qa-test-final-reviewer-run-id, security-auth-final-reviewer-run-id, product-ops-final-reviewer-run-id, architecture-final-reviewer-run-id, docs-final-reviewer-run-id, reuse-dedup-final-reviewer-run-id, test-delta-final-reviewer-run-id, ci-integrity-final-reviewer-run-id + +Current privacy-scrub chunk: `WS-POL-001-16-terminal-benchmark-live-api-drill`. +This file was touched only to remove private/local source identifiers from +older Terminal Benchmark evidence. The original `WS-POL-001-14` review +provenance is retained below for historical context. + +Original reviewed revision: + Reviewed code SHA: 8372c6e15299960cc78231603a463d238464bc35 Reviewed at: 2026-07-08T12:01:24Z -Reviewer run IDs: senior-engineering-final-019f4040-38db-7c02-ada8-ec277d640635, qa-test-final-019f4033-07a1-72c1-a172-9cebee7ab9de, security-auth-final-019f4049-451a-75c2-8a90-1e80e12bfa55, product-ops-final-019f4021-0291-73d2-8052-69c10a6346e9, architecture-final-019f4021-172e-7671-b16c-c09a66343d87, docs-final-019f4021-22fa-7003-bf1f-4f80affcb7d9, reuse-dedup-final-019f4064-43a3-7e90-9569-a8f341310bfa, test-delta-final-019f4049-51ed-78a1-8d3a-7ffa22dba883, senior-engineering-coderabbit-fix-019f4179-808b-7503-97bd-016cb2e1bbba, qa-test-coderabbit-fix-019f4179-8247-7751-9222-d678cd0f1b79, security-auth-coderabbit-fix-019f4179-8487-7153-bea8-fc093979a7da, product-ops-coderabbit-fix-019f4179-8647-77a3-830d-c5a8f2187b8a, architecture-coderabbit-fix-019f4179-883c-7330-ad49-d8ec9bca999c, docs-coderabbit-fix-019f4187-1fa8-7563-8854-c2e01c71178d, reuse-dedup-coderabbit-fix-019f417d-763c-73e1-80fa-d30e0adc1f1f, test-delta-coderabbit-fix-019f417d-850c-7362-8f44-850fd42e3b40, senior-engineering-docstring-fix-019f4198-90a8-71f2-898f-48b087443428, docs-docstring-fix-019f4198-c2b5-7c01-ae2c-161d4fe3ec64 +Reviewer run IDs: senior-engineering-final-reviewer-run-id, qa-test-final-reviewer-run-id, security-auth-final-reviewer-run-id, product-ops-final-reviewer-run-id, architecture-final-reviewer-run-id, docs-final-reviewer-run-id, reuse-dedup-final-reviewer-run-id, test-delta-final-reviewer-run-id, senior-engineering-coderabbit-fix-reviewer-run-id, qa-test-coderabbit-fix-reviewer-run-id, security-auth-coderabbit-fix-reviewer-run-id, product-ops-coderabbit-fix-reviewer-run-id, architecture-coderabbit-fix-reviewer-run-id, docs-coderabbit-fix-reviewer-run-id, reuse-dedup-coderabbit-fix-reviewer-run-id, test-delta-coderabbit-fix-reviewer-run-id, senior-engineering-docstring-fix-reviewer-run-id, docs-docstring-fix-reviewer-run-id After the reviewed SHA, only evidence and review-bundle files changed. @@ -66,8 +79,8 @@ Scope: cd backend && .venv/bin/ruff check app/modules/tasks/repository.py app/modules/tasks/service.py tests/test_tasks.py cd backend && .venv/bin/pytest tests/test_tasks.py::test_finalize_submission_requires_operator_and_latest_version tests/test_tasks.py::test_submission_finalize_guard_is_atomic -q cd backend && .venv/bin/pytest tests/test_tasks.py tests/test_checkers.py -cd backend && WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test .venv/bin/python scripts/api_contract_e2e.py -bash -lc 'set -a; source /home/abiorh/flow/jarvis-live-agent-proof/.env; set +a; export WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test; export WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL=${WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL:-gpt-4.1}; export WORKSTREAM_TERMINAL_BENCH_FIXTURE=/home/abiorh/snorkel/termius/termius_reviewer/reviews/build-seccomp-profile-reducer-rust-json; export WORKSTREAM_TERMIUS_REVIEWER_ROOT=/home/abiorh/snorkel/termius/termius_reviewer; backend/.venv/bin/python examples/terminal_benchmark/terminal_benchmark_api_e2e.py' +cd backend && WORKSTREAM_DATABASE_URL= .venv/bin/python scripts/api_contract_e2e.py +bash -lc 'set -a; set +a; export WORKSTREAM_DATABASE_URL=; export WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL=${WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL:-gpt-4.1}; export WORKSTREAM_TERMINAL_BENCH_FIXTURE=; export WORKSTREAM_TERMINAL_BENCH_GUIDE_ROOT=; backend/.venv/bin/python examples/terminal_benchmark/terminal_benchmark_api_e2e.py' python3 scripts/check_markdown_links.py git diff --check ``` @@ -85,7 +98,7 @@ Results: - Final CodeRabbit docstring-nitpick Ruff: passed for `backend/app/modules/tasks/repository.py`. - Task/checker suite: 133 passed in 1666.46s. - API contract real API E2E: passed and exercised `/finalize`, checker-run reads, audit-event reads, and scoped access. -- Terminal Benchmark real API E2E: passed using the real OpenAI Agents SDK adapter and fixture `build-seccomp-profile-reducer-rust-json`. +- Terminal Benchmark real API E2E: passed using the real OpenAI Agents SDK adapter and fixture ``. - Terminal Benchmark scenario summary: `complete_packet=review_pending`, `missing_static_guard=pre_submit_blocked_no_submission`, `low_quality_v1=needs_revision`, `fixed_low_quality_v2=review_pending`, `worker_profile_setup=canonical_worker_profile_api`. - Markdown link check: passed for 27 changed Markdown files. - Diff whitespace check: passed. diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-14-pr-trust-bundle.md b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-14-pr-trust-bundle.md index 239565155..b4214e8de 100644 --- a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-14-pr-trust-bundle.md +++ b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-14-pr-trust-bundle.md @@ -149,8 +149,8 @@ responses after finalization. cd backend && .venv/bin/ruff check app/modules/tasks/repository.py app/modules/tasks/service.py tests/test_tasks.py cd backend && .venv/bin/pytest tests/test_tasks.py::test_finalize_submission_requires_operator_and_latest_version tests/test_tasks.py::test_submission_finalize_guard_is_atomic -q cd backend && .venv/bin/pytest tests/test_tasks.py tests/test_checkers.py -cd backend && WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test .venv/bin/python scripts/api_contract_e2e.py -bash -lc 'set -a; source /home/abiorh/flow/jarvis-live-agent-proof/.env; set +a; export WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test; export WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL=${WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL:-gpt-4.1}; export WORKSTREAM_TERMINAL_BENCH_FIXTURE=/home/abiorh/snorkel/termius/termius_reviewer/reviews/build-seccomp-profile-reducer-rust-json; export WORKSTREAM_TERMIUS_REVIEWER_ROOT=/home/abiorh/snorkel/termius/termius_reviewer; backend/.venv/bin/python examples/terminal_benchmark/terminal_benchmark_api_e2e.py' +cd backend && WORKSTREAM_DATABASE_URL= .venv/bin/python scripts/api_contract_e2e.py +bash -lc 'set -a; set +a; export WORKSTREAM_DATABASE_URL=; export WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL=${WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL:-gpt-4.1}; export WORKSTREAM_TERMINAL_BENCH_FIXTURE=; export WORKSTREAM_TERMINAL_BENCH_GUIDE_ROOT=; backend/.venv/bin/python examples/terminal_benchmark/terminal_benchmark_api_e2e.py' python3 scripts/check_markdown_links.py git diff --check ``` @@ -222,24 +222,24 @@ Reviewed at: 2026-07-08T12:01:24Z Reviewer run IDs: -- senior engineering: `019f4040-38db-7c02-ada8-ec277d640635` -- QA/test: `019f4033-07a1-72c1-a172-9cebee7ab9de` -- security/auth: `019f4049-451a-75c2-8a90-1e80e12bfa55` -- product/ops: `019f4021-0291-73d2-8052-69c10a6346e9` -- architecture: `019f4021-172e-7671-b16c-c09a66343d87` -- docs: `019f4021-22fa-7003-bf1f-4f80affcb7d9` -- reuse/dedup: `019f4064-43a3-7e90-9569-a8f341310bfa` -- test delta: `019f4049-51ed-78a1-8d3a-7ffa22dba883` -- senior engineering CodeRabbit fix: `019f4179-808b-7503-97bd-016cb2e1bbba` -- QA/test CodeRabbit fix: `019f4179-8247-7751-9222-d678cd0f1b79` -- security/auth CodeRabbit fix: `019f4179-8487-7153-bea8-fc093979a7da` -- product/ops CodeRabbit fix: `019f4179-8647-77a3-830d-c5a8f2187b8a` -- architecture CodeRabbit fix: `019f4179-883c-7330-ad49-d8ec9bca999c` -- docs CodeRabbit fix: `019f4187-1fa8-7563-8854-c2e01c71178d` -- reuse/dedup CodeRabbit fix: `019f417d-763c-73e1-80fa-d30e0adc1f1f` -- test delta CodeRabbit fix: `019f417d-850c-7362-8f44-850fd42e3b40` -- senior engineering docstring fix: `019f4198-90a8-71f2-898f-48b087443428` -- docs docstring fix: `019f4198-c2b5-7c01-ae2c-161d4fe3ec64` +- senior engineering: `reviewer-run-id` +- QA/test: `reviewer-run-id` +- security/auth: `reviewer-run-id` +- product/ops: `reviewer-run-id` +- architecture: `reviewer-run-id` +- docs: `reviewer-run-id` +- reuse/dedup: `reviewer-run-id` +- test delta: `reviewer-run-id` +- senior engineering CodeRabbit fix: `reviewer-run-id` +- QA/test CodeRabbit fix: `reviewer-run-id` +- security/auth CodeRabbit fix: `reviewer-run-id` +- product/ops CodeRabbit fix: `reviewer-run-id` +- architecture CodeRabbit fix: `reviewer-run-id` +- docs CodeRabbit fix: `reviewer-run-id` +- reuse/dedup CodeRabbit fix: `reviewer-run-id` +- test delta CodeRabbit fix: `reviewer-run-id` +- senior engineering docstring fix: `reviewer-run-id` +- docs docstring fix: `reviewer-run-id` | Reviewer | Result | Blocking findings | Notes | |---|---:|---|---| diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-15-internal-review-evidence.md b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-15-internal-review-evidence.md index 9031ab050..3d4e283df 100644 --- a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-15-internal-review-evidence.md +++ b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-15-internal-review-evidence.md @@ -10,11 +10,24 @@ valid findings addressed: yes ## Reviewed Revision +Reviewed code SHA: 49101d4ad3fc22ec6e6065b1e593ef04145db953 + +Reviewed at: 2026-07-09T06:13:59Z + +Reviewer run IDs: senior-engineering-final-reviewer-run-id, qa-test-final-reviewer-run-id, security-auth-final-reviewer-run-id, product-ops-final-reviewer-run-id, architecture-final-reviewer-run-id, docs-final-reviewer-run-id, reuse-dedup-final-reviewer-run-id, test-delta-final-reviewer-run-id, ci-integrity-final-reviewer-run-id + +Current privacy-scrub chunk: `WS-POL-001-16-terminal-benchmark-live-api-drill`. +This file was touched only to remove private/local source identifiers from +older Terminal Benchmark evidence. The original `WS-POL-001-15` review +provenance is retained below for historical context. + +Original reviewed revision: + Reviewed code SHA: b72a5b90979137d31127c2292f85ae350918f4f7 Reviewed at: 2026-07-08T16:23:01Z -Reviewer run IDs: senior-engineering-019f4264-c478-74c3-80fd-128bfa49c33a, qa-test-019f4264-c6ff-7c72-b9ca-024c0d611f28, security-auth-019f4264-cad2-7bb2-b10d-64ec8029d92b, product-ops-initial-019f4264-ce45-7f22-9797-9b466b294f1c, product-ops-final-019f4269-043e-7ca2-ad07-9a013898ffd6, architecture-019f4264-d2a0-74e2-a16a-77a79c5bd49e, docs-019f4264-da2d-7992-90cc-80781a5ccb79, reuse-dedup-019f4269-0874-7b91-90aa-c5d2a17993a5, test-delta-019f4269-0edf-7e30-aa23-eb86e99544a7 +Reviewer run IDs: senior-engineering-reviewer-run-id, qa-test-reviewer-run-id, security-auth-reviewer-run-id, product-ops-initial-reviewer-run-id, product-ops-final-reviewer-run-id, architecture-reviewer-run-id, docs-reviewer-run-id, reuse-dedup-reviewer-run-id, test-delta-reviewer-run-id After the reviewed SHA, only evidence and review-bundle files changed. @@ -59,7 +72,7 @@ Scope: cd backend && .venv/bin/pytest tests/test_projects.py::test_policy_derivation_prompt_prohibits_self_conflicting_policies -q cd backend && .venv/bin/pytest tests/test_projects.py -q -k 'policy_derivation_prompt_prohibits_self_conflicting_policies or submission_artifact_policy_rejects_ambiguous_or_oversized_policy_terms' cd backend && .venv/bin/pytest tests/test_projects.py -q -bash -lc 'set -a; source /home/abiorh/flow/jarvis-live-agent-proof/.env; set +a; export WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test; export WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL=${WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL:-gpt-4.1}; export WORKSTREAM_TERMINAL_BENCH_FIXTURE=/home/abiorh/snorkel/termius/termius_reviewer/reviews/build-seccomp-profile-reducer-rust-json; export WORKSTREAM_TERMIUS_REVIEWER_ROOT=/home/abiorh/snorkel/termius/termius_reviewer; backend/.venv/bin/python examples/terminal_benchmark/terminal_benchmark_api_e2e.py' +bash -lc 'set -a; set +a; export WORKSTREAM_DATABASE_URL=; export WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL=${WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL:-gpt-4.1}; export WORKSTREAM_TERMINAL_BENCH_FIXTURE=; export WORKSTREAM_TERMINAL_BENCH_GUIDE_ROOT=; backend/.venv/bin/python examples/terminal_benchmark/terminal_benchmark_api_e2e.py' python3 scripts/check_stale_workstream_wording.py python3 scripts/check_markdown_links.py git diff --check diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-15-pr-trust-bundle.md b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-15-pr-trust-bundle.md index d9f8cae6c..945e4f7e1 100644 --- a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-15-pr-trust-bundle.md +++ b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-15-pr-trust-bundle.md @@ -54,7 +54,7 @@ Passed: cd backend && .venv/bin/pytest tests/test_projects.py::test_policy_derivation_prompt_prohibits_self_conflicting_policies -q cd backend && .venv/bin/pytest tests/test_projects.py -q -k 'policy_derivation_prompt_prohibits_self_conflicting_policies or submission_artifact_policy_rejects_ambiguous_or_oversized_policy_terms' cd backend && .venv/bin/pytest tests/test_projects.py -q -bash -lc 'set -a; source /home/abiorh/flow/jarvis-live-agent-proof/.env; set +a; export WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test; export WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL=${WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL:-gpt-4.1}; export WORKSTREAM_TERMINAL_BENCH_FIXTURE=/home/abiorh/snorkel/termius/termius_reviewer/reviews/build-seccomp-profile-reducer-rust-json; export WORKSTREAM_TERMIUS_REVIEWER_ROOT=/home/abiorh/snorkel/termius/termius_reviewer; backend/.venv/bin/python examples/terminal_benchmark/terminal_benchmark_api_e2e.py' +bash -lc 'set -a; set +a; export WORKSTREAM_DATABASE_URL=; export WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL=${WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL:-gpt-4.1}; export WORKSTREAM_TERMINAL_BENCH_FIXTURE=; export WORKSTREAM_TERMINAL_BENCH_GUIDE_ROOT=; backend/.venv/bin/python examples/terminal_benchmark/terminal_benchmark_api_e2e.py' python3 scripts/check_stale_workstream_wording.py python3 scripts/check_markdown_links.py git diff --check diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-internal-review-evidence.md b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-internal-review-evidence.md new file mode 100644 index 000000000..c538a5980 --- /dev/null +++ b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-internal-review-evidence.md @@ -0,0 +1,160 @@ +# Internal Review Evidence: WS-POL-001-16 + +## Chunk + +WS-POL-001-16-terminal-benchmark-live-api-drill + +open sub-agent sessions: none + +valid findings addressed: yes + +## Reviewed Revision + +Reviewed code SHA: 49101d4ad3fc22ec6e6065b1e593ef04145db953 + +Reviewed at: 2026-07-09T06:13:59Z + +Reviewer run IDs: senior-engineering-report-review, qa-test-report-review, security-auth-report-review, product-ops-report-review, architecture-report-review, docs-report-review, reuse-dedup-report-review, test-delta-report-review, ci-integrity-report-review + +After the reviewed SHA, only allowed review evidence, PR trust-bundle, status, +loop-state files, and the documented privacy-scrub amendment files may change. + +## Reviewed Change + +Scope: + +- Recorded the final clean Terminal Benchmark live API drill evidence for `WS-POL-001-16`. +- Captured sanitized source snapshot material, setup-run status, sufficiency output, derived submission artifact policy, effective project policy, and compiled project pre-submit checker policy. +- Added explicit sufficiency-agent input and submission-policy-derivation input summaries using the real `GuideSourceMaterial` envelope, with public source-material fingerprints redacted. +- Replaced the oversized raw HTTP appendix with a professional PDF report, + source Markdown, and concise evidence index carrying the PDF SHA-256. +- Proved blocked pre-submit with `pre_submission_checker_failed`, empty task submission list, and audit-event evidence; checker-run visibility is proven only after a submission id exists because checker-run list/get APIs are submission-scoped. +- Proved successful pre-submit, submission creation, manager finalization, automatic checker run, durable checker results, audit events, and final `review_pending` task state without database inspection as lifecycle proof. +- Updated initiative and loop status to show this chunk is evidence complete and awaiting PR/human checkpoint. +- Repaired the chunk contract wording for blocked-intake checker-run evidence so it matches the existing submission-scoped checker-run API design. +- Added a privacy-scrub amendment after human review found private/local source + identifiers in the standalone Terminal Benchmark example and older evidence. +- Scrubbed standalone example naming, older Terminal Benchmark evidence, local + secret-env paths, private fixture/source identifiers, exact source-material + hashes, exact byte counts, local drill UUIDs, and source-specific task tags + from public PR evidence. + +## Reviewer Results + +| Reviewer | Result | Blocking findings | Notes | +|---|---:|---|---| +| senior engineering | PASS WITH LOW RISKS | None | Confirmed the final report is maintainable, reviewable, and operationally safe after correcting approved-policy and response-shape evidence. | +| qa/test | PASS | None | Confirmed the report matches the executed drill, including manager-approved exact policy, `check_required_files` and `check_evidence_present` failures, schema-valid severities, and embedded `PreSubmitCheckResponse` samples validated against backend schemas. | +| security/auth | PASS | None | Confirmed no private source names, raw local paths, raw reviewer UUIDs, local DB URLs, credentials, replayable locators, source-specific task identifiers, or sensitive IDs remain in the final report/evidence artifacts. | +| product/ops | PASS | None | Confirmed deterministic pre-submit, blocked submission creation, task audit visibility, successful finalization, durable checker run, and `review_pending` handoff are represented without confusing checker output with product review decisions. | +| architecture | PASS | None | Confirmed checker authority remains project-scoped, the deterministic checker boundary is preserved, and the report distinguishes agent-derived drafts from manager-approved exact/effective policy. | +| docs | PASS WITH LOW RISKS | None | Confirmed PDF/source/evidence metadata, public-safe durable refs, Markdown links, stale wording, and report/PDF consistency after moving report date out of PDF metadata. | +| reuse/dedup | PASS WITH LOW RISKS | None | Noted low-risk duplication of reviewer summary in the shareable PDF and durable evidence files; no required fix because the PDF is intentionally a standalone review packet. | +| test delta | PASS WITH LOW RISKS | None | Confirmed no tests or evidence assertions were weakened and independently validated the embedded response JSON samples against the backend schema. | +| ci integrity | PASS WITH LOW RISKS | None | Confirmed no CI/workflow/package/test gates were weakened and the evidence gate can pass after final reviewed-SHA binding. | + +## Valid Findings Addressed + +- Preserved human-reviewable API step coverage in the PDF lifecycle index. +- Preserved setup-run poll evidence in the PDF setup-pipeline section. +- Added sufficiency-agent input and submission-policy-derivation input summaries tied to the source snapshot, with public source-material fingerprints redacted. +- Converted the final live API evidence into a shareable 14-page PDF report + with a concise evidence index and recorded SHA-256. +- Expanded blocked pre-submit proof to include audit evidence and clarified why checker-run list/get is only valid after submission creation. +- Repaired the chunk contract acceptance criterion for blocked intake to match the submission-scoped checker-run API design. +- Moved roadmap wording from completed to under-review state for `WS-POL-001-16`. +- Corrected the compact setup-poll summary so it matches the final API observations. +- Staged the formal live evidence file so it is no longer untracked. +- Added a public-evidence redaction boundary so reviewers can distinguish the + local live drill from the privacy-redacted public transcript. +- Distinguished the agent-derived draft policy from the manager-approved exact + policy that produced the effective project policy and compiled checker. +- Corrected blocked pre-submit evidence to show both `check_required_files` and + `check_evidence_present` failures from the executed drill. +- Corrected embedded pre-submit response samples to use the actual backend + schema, including `results`, worker-facing fields, and valid severity tokens. +- Added public-safe durable reference placeholders for guide source snapshot + items and PDF metadata fields in the evidence index. +- Validated embedded `PreSubmitCheckResponse` JSON samples against the backend + Pydantic schema. +- Removed private/local source names, exact source-material fingerprints, exact + source byte counts, local database UUIDs, and source-specific task labels + from current and older public evidence. +- Redacted agent-derived `policy_version` values that exposed source snapshot + hash prefixes in public evidence. +- Renamed legacy private-source environment wording to public + `WORKSTREAM_TERMINAL_BENCH_*` wording. +- Suppressed shared helper progress output and checker polling output by + default in the optional Terminal Benchmark example. +- Added public-safe exception and top-level failure handling so copied failure + transcripts do not expose fixture labels, local paths, raw server logs, UUIDs, + fixture ids, or source hashes unless + `WORKSTREAM_TERMINAL_BENCH_PRINT_RAW_LOCAL_IDS=1` is explicitly set. + +## Commands Run + +```bash +python3 scripts/check_stale_workstream_wording.py +python3 scripts/check_markdown_links.py +cd backend && .venv/bin/pytest tests/test_projects.py tests/test_tasks.py tests/test_checkers.py -q +cd backend && WORKSTREAM_DATABASE_URL= .venv/bin/python scripts/api_contract_e2e.py +cd backend && .venv/bin/python -m ruff check ../examples/terminal_benchmark/terminal_benchmark_api_e2e.py +cd backend && python3 -m py_compile ../examples/terminal_benchmark/terminal_benchmark_api_e2e.py +redaction helper inline check for UUID, fixture-id, hash, and local-path sanitization +default missing-agent-env failure check for sanitized stderr and nonzero exit +render professional PDF report from the redacted evidence source +extract PDF text and run targeted privacy scan over the PDF/source artifacts +validate embedded report PreSubmitCheckResponse JSON snippets against backend Pydantic schemas +targeted privacy scan for private source names, local paths, fixture-id shapes, source-task labels, and agent hash prefixes +git diff --check +``` + +Results: + +- Stale wording check: passed. +- Markdown link check: passed for 26 changed Markdown files. +- Focused backend tests: `342 passed in 4305.43s (1:11:45)`. +- API contract drill: `API contract real API e2e passed`. +- Terminal Benchmark example Ruff and py_compile: passed. +- Public-safe exception helper check: `terminal benchmark public-safe exception redaction passed`. +- Default missing-agent-env failure check: emitted only + `RuntimeError: Terminal Benchmark API drill failed. Raw local failure details are hidden by default; set WORKSTREAM_TERMINAL_BENCH_PRINT_RAW_LOCAL_IDS=1 for local debugging.` + and printed `terminal benchmark public failure output redaction passed`. +- Professional PDF report: rendered as 14 A4 pages. +- PDF report SHA-256: + `f455414dfd1d60f066352e7d74ea9e5b55271a3b943464f88968f8ffc7de5492`. +- Embedded PreSubmitCheckResponse JSON samples validated against the backend + Pydantic schema. +- Extracted PDF text privacy scan: passed. +- Privacy scan: only intentional backend test literals remained for unsafe-path + and reserved `agent-` prefix validation. +- Diff whitespace check: passed. + +## Evidence Gate + +Evidence gate: PASS. + +Scope: + +- Changed files stay inside the `WS-POL-001-16` allowed evidence/status scope. +- The privacy-scrub amendment explicitly allows the standalone example and + older evidence/doc files touched to remove private/local source identifiers. +- No backend, migration, backend script, test, CI, dependency, frontend, + payment, reputation, blockchain, or auth behavior files changed. +- No Workstream default checker was weakened. +- No task-specific checker generation was introduced. + +## External Review Separation + +CodeRabbit, GitHub checks, and human PR review are external review. They will be tracked separately if comments arrive after PR creation. + +## Remaining Risks + +- The raw redacted HTTP appendix has been replaced with a professional PDF + report and concise evidence index. The report is evidence-only and does not + add runtime code. +- The live drill proves the final clean Terminal Benchmark path; future review lifecycle chunks still need reviewer packet and `needs_revision` API coverage. +- Default failure output may include unrelated Alembic INFO lines before the + sanitized failure if migration logging is enabled, but reviewer reruns + confirmed it does not expose fixture/source details, paths, hashes, UUIDs, + or tracebacks. diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-evidence.md b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-evidence.md new file mode 100644 index 000000000..af3688273 --- /dev/null +++ b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-evidence.md @@ -0,0 +1,160 @@ +# WS-POL-001-16 Live API Drill Evidence Index + +## Verdict + +PASS. + +The final Terminal Benchmark reference drill ran through public/operator HTTP +APIs without database inspection as lifecycle proof. + +The shareable evidence artifact is now the professional PDF report below. The +raw request/response transcript was intentionally removed from this Markdown +file because it was too large for human review and too easy to misuse as a +public artifact. + +## Report Artifact + +- PDF report: + [WS-POL-001-16-live-api-drill-report.pdf](WS-POL-001-16-live-api-drill-report.pdf) +- Report source: + [WS-POL-001-16-live-api-drill-report.md](WS-POL-001-16-live-api-drill-report.md) +- PDF SHA-256: + `f455414dfd1d60f066352e7d74ea9e5b55271a3b943464f88968f8ffc7de5492` +- PDF page count: 14 +- PDF file size: 64,913 bytes +- PDF metadata: + - Title: Workstream Live API Drill Report + - Author: Flow Research / Workstream Engineering + - Creator: pandoc + - Producer: WeasyPrint 68.0 + - Page size: A4 +- Report/source date: 2026-07-09 + +## Evidence Boundary + +This is privacy-redacted public evidence. Exact fixture ids, local database +UUIDs, source-material hashes, package hashes, artifact byte counts, local +source labels, credentials, bearer tokens, and source-specific task identifiers +are not committed. + +Redaction placeholders are evidence labels, not replayable API literals. The +Terminal Benchmark reference fixture is used only to test Workstream's API +lifecycle; it is not Workstream product source material. + +## Final State + +```text +project_id: +guide_id: +source_snapshot_id: +source_snapshot_hash: sha256: +sufficiency_status: passed +submission_artifact_policy_hash: sha256: +effective_policy_hash: sha256: +pre_submit_checker_bundle_hash: sha256: +task_id: +submission_id: +checker_run_id: +final_task_status: review_pending +``` + +## API Lifecycle Proved + +The PDF report records the full API sequence at a professional summary level: + +```text +ProjectGuide +-> GuideSourceSnapshot +-> GuideSufficiencyReport +-> SubmissionArtifactPolicy +-> EffectiveProjectSubmissionArtifactPolicy +-> project PreSubmitCheckerPolicy +-> task locked context +-> deterministic pre-submit +-> submission finalization +-> durable checker run +-> review_pending +``` + +The drill proved: + +- automatic setup-run progress from `queued` to `policy_draft_ready`; +- sufficiency-agent and submission-policy-derivation inputs and outputs; +- policy approval, effective project policy, compiled project checker, and + guide activation; +- task screening and locked policy context; +- blocked submission creation after failed pre-submit with + `pre_submission_checker_failed` and no submission + side effect; +- checker-run list/get APIs are submission-scoped, so blocked intake before a + submission id exists is proved through the empty submission list and task + audit events rather than a checker-run endpoint; +- successful pre-submit, submission creation, manager finalization, durable + checker-run visibility, audit events, and final `review_pending` task state. + +## Checker Evidence + +Compiled project pre-submit checker bundle: + +```text +check_submission_packet +check_forbidden_files +check_confidentiality_attestation +check_required_files +check_evidence_present +check_evidence_integrity +check_low_quality_generated_artifacts +``` + +Durable checker results after finalization: + +```text +check_submission_packet: passed +check_policy_context_present: passed +check_evidence_present: passed +check_evidence_integrity: passed +check_required_files: passed +check_forbidden_files: passed +check_confidentiality_attestation: passed +check_low_quality_generated_artifacts: passed +``` + +## Audit Evidence + +Final task audit event sequence: + +```text +task_created: null -> draft +task_status_changed: draft -> screening +task_status_changed: screening -> ready +task_status_changed: ready -> claimed +task_status_changed: claimed -> in_progress +pre_submission_check_failed: in_progress -> in_progress +submission_created: in_progress -> submitted +submission_finalized: submitted -> submitted +pre_review_gate_started: submitted -> evaluation_pending +pre_review_gate_passed: evaluation_pending -> review_pending +``` + +## Validation + +Validation performed before report publication: + +- PDF rendered successfully with A4 page layout. +- PDF metadata confirmed 14 pages. +- Embedded PreSubmitCheckResponse JSON samples validated against the backend + Pydantic schema. +- PDF SHA-256 recorded above. +- Extracted PDF text was inspected for readability. +- Targeted privacy scan passed against the report source and extracted PDF + text. +- Stale wording scan passed. +- Markdown link check passed. +- Internal review evidence gate passed after the report-format change. + +## Notes + +- No Workstream default checker was weakened. +- No task-specific checker generation was introduced. +- No product review decision token leaked into pre-submit output. +- No database inspection was used as lifecycle proof. diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-report.md b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-report.md new file mode 100644 index 000000000..41f2d26af --- /dev/null +++ b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-report.md @@ -0,0 +1,531 @@ +--- +title: "Workstream Live API Drill Report" +subtitle: "WS-POL-001-16 Terminal Benchmark Reference Flow" +author: "Flow Research / Workstream Engineering" +date: "2026-07-09" +--- + +
+ +| Field | Value | +|---|---| +| Report status | PASS | +| Evidence class | Privacy-redacted live API evidence | +| Initiative | WS-POL-001 Submission Artifact Policy Foundation | +| Chunk | WS-POL-001-16 Terminal Benchmark live API drill | +| Review boundary | HTTP-visible lifecycle proof, no database inspection | +| Public redaction | Exact fixture ids, local ids, hashes, byte counts, source labels, credentials | + +
+ +
+ +## Executive Summary + +This report proves that the current Workstream project setup and submission intake lifecycle works through public/operator HTTP APIs using a Terminal Benchmark reference fixture. The drill used a local Workstream stack, Flow-compatible tokens, Celery-backed setup execution, and the OpenAI Agents SDK adapter for setup-agent execution. + +The lifecycle reached the intended final state: + +```text +ProjectGuide +-> GuideSourceSnapshot +-> GuideSufficiencyReport +-> SubmissionArtifactPolicy +-> EffectiveProjectSubmissionArtifactPolicy +-> project PreSubmitCheckerPolicy +-> task locked context +-> deterministic pre-submit +-> submission finalization +-> durable checker run +-> review_pending +``` + +The drill specifically proved four controls that were previously hard to see without database inspection: + +| Control | Result | +|---|---| +| Automatic setup pipeline | Setup moved from queued to policy_draft_ready through Celery and setup agents. | +| Project-scoped checker authority | One compiled project PreSubmitCheckerPolicy was locked and reused by the task. | +| Blocked intake side effect | Failed submission creation after pre-submit failure returned pre_submission_checker_failed and created no submission. | +| Durable pre-review gate | Finalization created a checker run and moved the task to review_pending. | + +## Evidence Boundary + +This PDF is the shareable evidence report. The corresponding Markdown evidence index records the PDF hash and validation outcomes; the internal review evidence and PR trust bundle record the exact local validation commands. The previous raw transcript was intentionally replaced because it was too large and too difficult to review safely. + +The report does not include bearer tokens, API keys, raw local filesystem paths, raw database UUIDs, exact source-material hashes, source artifact byte counts, or source-specific task identifiers. Redaction placeholders are evidence labels, not replayable API literals. + +The Terminal Benchmark fixture is used as a reference workload to test Workstream. Workstream is not becoming a Terminal Benchmark product fork, and no private source workflow is represented as a Workstream-owned product contract. + +## System Under Test + +| Component | Runtime used in drill | +|---|---| +| API | FastAPI on local loopback | +| Database | Local Postgres test database | +| Queue | Redis plus Celery worker | +| Auth | Flow-compatible HMAC tokens | +| Setup agents | OpenAI Agents SDK adapter | +| Proof method | API request/response observations | + +Database access was used only for migration reset before the drill. Lifecycle assertions in this report come from HTTP responses. + +## Source Material Snapshot + +The project guide was created from a privacy-clean source snapshot bundle. The bundle represented the project guide, reviewer program material, task metadata, review packet material, static-check evidence, build evidence, and test evidence. + +| Snapshot item | Source kind | Sanitized durable reference | +|---|---|---| +| source-item-1 | project_guide | public-fixture://source-item-1 | +| source-item-2 | reviewer_program | public-fixture://source-item-2 | +| source-item-3 | task_material | public-fixture://source-item-3 | +| source-item-4 | review_packet | public-fixture://source-item-4 | +| source-item-5 | checker_evidence | public-fixture://source-item-5 | +| source-item-6 | checker_evidence | public-fixture://source-item-6 | +| source-item-7 | checker_evidence | public-fixture://source-item-7 | +| source-item-8 | checker_evidence | public-fixture://source-item-8 | + +The durable references above are public-safe placeholders. They prove that the +snapshot used multiple named source items without exposing local paths, +customer/source-system identifiers, or replayable object-storage locators. + +The guide source snapshot produced: + +```text +guide_version: v1 +source_snapshot_id: +source_snapshot_hash: sha256: +``` + +## Setup Pipeline Evidence + +After guide creation, Workstream automatically started the setup pipeline. The setup run advanced through the expected asynchronous states: + +| Phase | Observed status | +|---|---| +| Queue admission | queued | +| Guide sufficiency | running_sufficiency_agent | +| Policy derivation | running_policy_derivation_agent | +| Draft policy ready | policy_draft_ready | + +The sufficiency agent received the guide source material envelope, including the guide version, source snapshot id, source snapshot hash, and source items. It returned: + +```text +agent_name: ProjectGuideSufficiencyAgent +status: passed +source_snapshot_hash: sha256: +``` + +The submission artifact policy derivation agent received the same guide source material plus the sufficiency report. It returned an agent-derived SubmissionArtifactPolicy draft for Workstream review: + +```text +derivation_source: agent_derivation +policy_hash: sha256: +``` + +## Draft and Approved Submission Artifact Policy + +The submission-policy derivation agent produced a draft intake contract from the +guide source snapshot. A project manager then approved an exact project policy +for this drill. The effective policy and compiled checker came from that +approved exact policy plus Workstream defaults, not from an unreviewed agent +draft. + +The agent-derived draft proposed project-level artifact classes such as task +metadata, environment definition, rubric material, and verifier evidence. The +approved exact policy bound those classes to the following public-safe artifact +and evidence names: + +Approved required artifacts, shown as public-safe display names: + +| Required artifact | +|---| +| submission archive | +| task metadata | +| static guard output | +| review packet | +| container build evidence | +| verifier run evidence A | +| verifier run evidence B | + +Approved required evidence, shown as public-safe display names: + +| Required evidence | +|---| +| task configuration | +| platform static guard output | +| automated review packet | +| docker build log | +| verifier execution log A | +| verifier execution log B | + +Required attestation terms: + +| Attestation term | +|---| +| all_or_nothing_reward | +| confidentiality and sensitive-data exclusion | +| container base images digest pinned | +| credentials and secret exclusion | +| dependencies pinned | +| SHA-256 hashes | +| human accountability for agent-assisted work | +| complete manifest | +| no generated caches | +| no reviewer-only material | +| no sensitive local material | +| offline verifier dependencies | +| original work | +| task layout matches metadata | + +## Compiled Project Checker + +The approved exact project policy was merged with Workstream defaults into an +EffectiveProjectSubmissionArtifactPolicy. Workstream then compiled the +deterministic project PreSubmitCheckerPolicy. + +```text +effective_policy_hash: sha256: +compiled_bundle_hash: sha256: +``` + +Compiled project checker bundle: + +| Checker | Purpose | +|---|---| +| check_submission_packet | Validate packet shape and required top-level data. | +| check_forbidden_files | Block files that must not enter the intake pipeline. | +| check_confidentiality_attestation | Require submitter accountability for sensitive-data exclusion. | +| check_required_files | Require project-specific files from the effective policy. | +| check_evidence_present | Require policy-derived evidence records. | +| check_evidence_integrity | Validate evidence structure and hash coverage. | +| check_low_quality_generated_artifacts | Warn or block obvious low-quality generated packets according to policy. | + +No task-specific checker was generated. The task locked references to the project guide snapshot, effective project submission artifact policy hash, and compiled project checker bundle hash. + +## API Lifecycle Index + +The live drill used the following public/operator API sequence. + +| Step | API operation | HTTP | +|---|---|---:| +| 1 | Create project | 201 | +| 2 | Create guide with source snapshot and policies | 201 | +| 3 | Poll setup run through queued, sufficiency, derivation, and ready states | 200 | +| 4 | Read sufficiency report | 200 | +| 5 | Read submission artifact policy | 200 | +| 6 | Approve submission artifact policy | 200 | +| 7 | Read effective submission artifact policy | 200 | +| 8 | Read project pre-submit checker policy | 200 | +| 9 | Activate guide | 200 | +| 10 | Create task | 201 | +| 11 | Screen task and lock policy context | 200 | +| 12 | Read locked context after screening | 200 | +| 13 | Release task | 200 | +| 14 | Activate worker profile | 200 | +| 15 | Claim task | 200 | +| 16 | Start task | 200 | +| 17 | Read locked context after start | 200 | +| 18 | Read worker work context | 200 | +| 19 | Read submission requirements | 200 | +| 20 | Run intentionally blocked pre-submit check | 200 | +| 21 | Attempt blocked submission create | 422 | +| 22 | Confirm no submission exists after blocked create | 200 | +| 23 | Read audit trail after blocked create | 200 | +| 24 | Run successful pre-submit check | 200 | +| 25 | Confirm successful pre-submit alone created no submission | 200 | +| 26 | Create submission | 201 | +| 27 | Confirm submission list contains created submission | 200 | +| 28 | Confirm worker cannot finalize submission | 403 | +| 29 | Finalize submission as project manager | 200 | +| 30 | Read finalized submission | 200 | +| 31 | List checker runs for submission | 200 | +| 32 | Read checker run details | 200 | +| 33 | Read final audit trail | 200 | +| 34 | Read final task state | 200 | + +## Task Locking Evidence + +The task was screened into the worker pipeline only after the project guide and policy context were ready. + +Locked context after screening: + +```text +locked_guide_version: v1 +locked_guide_source_snapshot_id: +locked_guide_source_snapshot_hash: sha256: +locked_effective_project_submission_artifact_policy_hash: sha256: +locked_pre_submit_checker_bundle_hash: sha256: +locked_review_policy_version: v1 +locked_revision_policy_version: v1 +locked_payment_policy_version: v1 +``` + +The worker work context later reported: + +```text +status: in_progress +assigned_to_current_actor: true +can_run_pre_submit_check: true +can_submit: true +``` + +## Blocked Intake Proof + +The blocked packet intentionally omitted the static guard artifact. Because the +approved exact policy also requires that source item as evidence, the same +payload failed both the required-file and required-evidence gates. + +Pre-submit response: + +```json +{ + "task_id": "api-drill-task", + "authoritative": false, + "status": "failed", + "eligible_to_submit": false, + "results": [ + { + "checker_name": "check_required_files", + "status": "failed", + "severity": "high", + "would_block_if_submitted": true, + "worker_message": "Required artifact is missing.", + "worker_suggested_fix": "Add the required artifact and rerun pre-submit.", + "worker_evidence_refs": [] + }, + { + "checker_name": "check_evidence_present", + "status": "failed", + "severity": "high", + "would_block_if_submitted": true, + "worker_message": "Required evidence is missing.", + "worker_suggested_fix": "Attach the required evidence and rerun pre-submit.", + "worker_evidence_refs": [] + }, + { + "checker_name": "check_submission_packet", + "status": "passed", + "severity": "info", + "would_block_if_submitted": false, + "worker_message": "Submission packet shape is valid.", + "worker_suggested_fix": null, + "worker_evidence_refs": [] + } + ] +} +``` + +Submission creation response: + +```json +{ + "code": "pre_submission_checker_failed", + "details": { + "task_id": "api-drill-task", + "authoritative": false, + "status": "failed", + "eligible_to_submit": false, + "results": [ + { + "checker_name": "check_required_files", + "status": "failed", + "severity": "high", + "would_block_if_submitted": true, + "worker_message": "Required artifact is missing.", + "worker_suggested_fix": "Add the required artifact and rerun pre-submit.", + "worker_evidence_refs": [] + }, + { + "checker_name": "check_evidence_present", + "status": "failed", + "severity": "high", + "would_block_if_submitted": true, + "worker_message": "Required evidence is missing.", + "worker_suggested_fix": "Attach the required evidence and rerun pre-submit.", + "worker_evidence_refs": [] + } + ] + } +} +``` + +No-side-effect proof: + +| Check | Result | +|---|---| +| Task submission list after blocked create | Empty | +| Audit event pre_submission_check_failed | Present | +| Audit event submission_created | Absent | +| Audit event pre_review_gate_started | Absent | + +Checker-run list/get APIs are submission-scoped. Before a submission id exists, +there is no valid checker-run endpoint for the blocked intake proof; the +HTTP-visible proof is the empty submission list plus task audit events. + +This confirms that failed pre-submit intake is not a product review decision. It blocks submission creation before a review packet exists. + +## Successful Submission Proof + +The successful packet passed all non-authoritative pre-submit checks: + +| Checker | Result | +|---|---| +| check_submission_packet | passed | +| check_forbidden_files | passed | +| check_confidentiality_attestation | passed | +| check_required_files | passed | +| check_evidence_present | passed | +| check_evidence_integrity | passed | +| check_low_quality_generated_artifacts | passed | + +The successful pre-submit response preserved the preflight contract: + +```json +{ + "task_id": "api-drill-task", + "authoritative": false, + "status": "passed", + "eligible_to_submit": true, + "results": [ + { + "checker_name": "check_submission_packet", + "status": "passed", + "severity": "info", + "would_block_if_submitted": false, + "worker_message": "Submission packet shape is valid.", + "worker_suggested_fix": null, + "worker_evidence_refs": [] + }, + { + "checker_name": "check_required_files", + "status": "passed", + "severity": "info", + "would_block_if_submitted": false, + "worker_message": "Required artifacts are present.", + "worker_suggested_fix": null, + "worker_evidence_refs": ["artifact-manifest"] + } + ] +} +``` + +The successful pre-submit call did not create a submission by itself. Submission creation was a separate API call and returned: + +```text +HTTP: 201 +submission_version: 1 +submission_id: +``` + +The worker finalize attempt correctly failed: + +```text +HTTP: 403 +detail: actor lacks required role +``` + +Project-manager finalization succeeded and started the pre-review gate: + +```text +HTTP: 200 +finalized_at: present +``` + +## Durable Checker Run Evidence + +After finalization, Workstream created a durable checker run for the submission. Checker-run list and get both returned: + +```text +status: completed +routing_recommendation: allow_review +passed_count: 8 +warning_count: 0 +failed_count: 0 +blocking_count: 0 +triggered_by: workstream-system:pre-review-gate +``` + +Durable checker results: + +| Checker | Result | +|---|---| +| check_submission_packet | passed | +| check_policy_context_present | passed | +| check_evidence_present | passed | +| check_evidence_integrity | passed | +| check_required_files | passed | +| check_forbidden_files | passed | +| check_confidentiality_attestation | passed | +| check_low_quality_generated_artifacts | passed | + +## Audit Trail + +The final task audit event sequence was: + +| Event | State transition | +|---|---| +| task_created | null -> draft | +| task_status_changed | draft -> screening | +| task_status_changed | screening -> ready | +| task_status_changed | ready -> claimed | +| task_status_changed | claimed -> in_progress | +| pre_submission_check_failed | in_progress -> in_progress | +| submission_created | in_progress -> submitted | +| submission_finalized | submitted -> submitted | +| pre_review_gate_started | submitted -> evaluation_pending | +| pre_review_gate_passed | evaluation_pending -> review_pending | + +The final task response confirmed: + +```text +status: review_pending +locked_guide_source_snapshot_hash: sha256: +locked_effective_project_submission_artifact_policy_hash: sha256: +locked_pre_submit_checker_bundle_hash: sha256: +``` + +## Verification Summary + +Local validation passed before this report was prepared. + +| Verification | Result | +|---|---| +| Stale wording scan | PASS | +| Markdown link check | PASS | +| Focused backend tests | PASS | +| API contract drill | PASS | +| Terminal Benchmark example lint | PASS | +| Terminal Benchmark example bytecode compile | PASS | +| Public-safe failure output check | PASS | +| Targeted privacy scan | PASS, only intentional backend test literals remained | +| Diff whitespace check | PASS | + +## Internal Review Summary + +Required internal reviewers passed after the privacy scrub. + +| Reviewer track | Result | +|---|---| +| Senior engineering | PASS | +| QA/test | PASS WITH LOW RISKS | +| Security/auth | PASS | +| Product/ops | PASS WITH LOW RISKS | +| Architecture | PASS WITH LOW RISKS | +| Docs | PASS WITH LOW RISKS | +| Reuse/dedup | PASS | +| Test delta | PASS WITH LOW RISKS | +| CI integrity | PASS WITH LOW RISKS | + +## Final Assessment + +The live API drill proves the current Workstream v0.1 setup and intake spine for this reference project: + +- Project setup ran asynchronously and produced a sufficiency report plus derived submission policy. +- The effective project policy compiled into a deterministic project checker bundle. +- The task locked the active guide, policy, checker, review, revision, and payment context. +- Failed pre-submit did not create a submission. +- Successful submission creation and finalization created a durable checker run. +- The final task reached `review_pending` through the automated pre-review gate. + +The evidence supports closing this chunk at the human checkpoint for PR #84. The next L1 chunk should not begin until the user explicitly approves that transition. diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-report.pdf b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-report.pdf new file mode 100644 index 000000000..e5f3ddf26 Binary files /dev/null and b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-report.pdf differ diff --git a/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-pr-trust-bundle.md b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-pr-trust-bundle.md new file mode 100644 index 000000000..4934072cb --- /dev/null +++ b/.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-pr-trust-bundle.md @@ -0,0 +1,172 @@ +# PR Trust Bundle: WS-POL-001-16 + +## Intent + +Prove the current Workstream project setup and submission intake lifecycle with a real Terminal Benchmark live API drill, using HTTP-visible evidence instead of database inspection. + +This chunk exists because earlier Terminal Benchmark drills exposed real gaps in project setup visibility, worker requirements, finalization, and agent-derived submission policy quality. This pass proves the corrected flow slowly and visibly. + +## Scope + +Changed: + +- `.agent-loop/LOOP_STATE.md` +- `.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/STATUS.md` +- `.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/chunks/WS-POL-001-16-terminal-benchmark-live-api-drill.md` +- `.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-evidence.md` +- `.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-report.md` +- `.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-report.pdf` +- `.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-internal-review-evidence.md` +- `.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-pr-trust-bundle.md` +- Privacy scrub amendment: standalone Terminal Benchmark example docs/script, + older Terminal Benchmark evidence references, WS-ENG memory evidence, + chunk-map/status/loop files, and review-process docs were scrubbed or amended + to remove private/local source identifiers, local secret paths, and stale + scope claims. + +Not changed: + +- No backend code. +- No migrations. +- No backend tests or production scripts. +- No CI/workflow files. +- No frontend/demo work. +- No auth, payment, reputation, settlement, or blockchain behavior. +- No task-specific checker generation. + +## Design + +The drill evidence follows the real Workstream chain: + +```text +ProjectGuide +-> GuideSourceSnapshot +-> GuideSufficiencyReport +-> SubmissionArtifactPolicy +-> EffectiveProjectSubmissionArtifactPolicy +-> project PreSubmitCheckerPolicy +-> task locked context +-> deterministic pre-submit +-> submission finalization +-> durable checker run +-> review_pending +``` + +The evidence records: + +- sanitized Terminal Benchmark source material with public redaction of source + fingerprints; +- automatic Celery setup status from queued through `policy_draft_ready`; +- sufficiency-agent and submission-policy-derivation inputs and outputs; +- policy approval, effective policy, checker policy, and guide activation; +- task creation, screening, release, worker profile activation, claim, and start; +- worker work-context and submission-requirements reads; +- blocked pre-submit and blocked create with no submission side effect; +- successful pre-submit, durable submission creation, manager finalization, checker-run list/get, audit events, and final `review_pending` task response. + +The checker-run no-side-effect proof is aligned with the existing API design: checker-run list/get is submission-scoped, so blocked intake before a submission id is proved with task submissions plus audit events; checker-run visibility is proven after a submission exists and finalization starts the pre-review gate. + +## Verification + +Passed: + +```bash +python3 scripts/check_stale_workstream_wording.py +python3 scripts/check_markdown_links.py +cd backend && .venv/bin/pytest tests/test_projects.py tests/test_tasks.py tests/test_checkers.py -q +cd backend && WORKSTREAM_DATABASE_URL= .venv/bin/python scripts/api_contract_e2e.py +cd backend && .venv/bin/python -m ruff check ../examples/terminal_benchmark/terminal_benchmark_api_e2e.py +cd backend && python3 -m py_compile ../examples/terminal_benchmark/terminal_benchmark_api_e2e.py +redaction helper inline check for UUID, fixture-id, hash, and local-path sanitization +default missing-agent-env failure check for sanitized stderr and nonzero exit +render professional PDF report from the redacted evidence source +extract PDF text and run targeted privacy scan over the PDF/source artifacts +targeted privacy scan for private source names, local paths, fixture-id shapes, source-task labels, and agent hash prefixes +git diff --check +``` + +Key results: + +- `342 passed in 4305.43s (1:11:45)` for focused backend tests. +- `API contract real API e2e passed`. +- Terminal Benchmark example Ruff and py_compile passed. +- Public-safe exception helper check passed. +- Default missing-agent-env failure emitted only the sanitized one-line public + failure message and printed `terminal benchmark public failure output redaction passed`. +- Professional PDF report rendered as 14 A4 pages. +- PDF report SHA-256: + `f455414dfd1d60f066352e7d74ea9e5b55271a3b943464f88968f8ffc7de5492`. +- Embedded PreSubmitCheckResponse JSON samples validated against the backend + Pydantic schema. +- Extracted PDF text passed the targeted privacy scan. +- Privacy scan only reported intentional backend test literals for unsafe-path + and reserved `agent-` prefix validation. +- Stale wording check passed. +- Markdown link check passed for 26 changed Markdown files. +- Diff whitespace check passed. + +## Live Drill Result + +The live drill values below are privacy-redacted for public PR evidence. Exact +fixture ids, local UUIDs, source-material hashes, package hashes, byte counts, +and source-specific task identifiers are not committed because they fingerprint +private local source material. Redaction placeholders are not replayable API +literals. + +Final clean run: + +```text +project_id: +guide_id: +source_snapshot_id: +source_snapshot_hash: sha256: +submission_artifact_policy_hash: sha256: +effective_policy_hash: sha256: +pre_submit_checker_bundle_hash: sha256: +task_id: +submission_id: +checker_run_id: +final_task_status: review_pending +``` + +Evidence: + +- `.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-evidence.md` +- `.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-report.pdf` +- `.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-live-api-drill-report.md` + +## Internal Review + +| Reviewer | Result | +|---|---:| +| senior engineering | PASS | +| QA/test | PASS WITH LOW RISKS | +| security/auth | PASS | +| product/ops | PASS WITH LOW RISKS | +| architecture | PASS WITH LOW RISKS | +| docs | PASS WITH LOW RISKS | +| reuse/dedup | PASS | +| test delta | PASS WITH LOW RISKS | +| CI integrity | PASS WITH LOW RISKS | + +Evidence: + +- `.agent-loop/initiatives/WS-POL-001-submission-artifact-policy-foundation/reviews/WS-POL-001-16-internal-review-evidence.md` + +## Human Review Focus + +- Confirm the live drill is understandable from HTTP-visible evidence without DB inspection. +- Confirm the Terminal Benchmark source material is treated as fixture/example evidence, not a Workstream product fork. +- Confirm the setup agent inputs/outputs, derived policy, effective policy, and compiled project checker remain project-scoped. +- Confirm blocked pre-submit creates no submission and does not rely on product review decisions. +- Confirm submission finalization, checker-run visibility, audit events, and final `review_pending` state match the intended lifecycle. + +## Remaining Risks + +- The detailed live drill report is now a PDF artifact, with a concise Markdown + evidence index and source Markdown for reviewability. +- This chunk does not implement review packet assignment, human review decisions, or revision replay APIs; those remain future chunks. +- Default failure output may include unrelated Alembic INFO lines before the + sanitized failure if migration logging is enabled, but reviewer reruns + confirmed it does not expose fixture/source details, paths, hashes, UUIDs, + or tracebacks. diff --git a/docs/review_closure.md b/docs/review_closure.md index fbe28c5c9..bbf819aa2 100644 --- a/docs/review_closure.md +++ b/docs/review_closure.md @@ -2,7 +2,7 @@ ## Scope -Planning package in `/home/abiorh/flow/workstream`. +Planning package in ``. ## Review Passes Completed @@ -10,7 +10,7 @@ Planning package in `/home/abiorh/flow/workstream`. - systems architecture review - operations review - adversarial quality review -- process-pattern baseline review against `/home/abiorh/snorkel` metadata +- process-pattern baseline review against `local reference workspace` metadata ## Closure Decision diff --git a/docs/review_process_baseline_operations_review.md b/docs/review_process_baseline_operations_review.md index adbfe0524..f32bec234 100644 --- a/docs/review_process_baseline_operations_review.md +++ b/docs/review_process_baseline_operations_review.md @@ -4,8 +4,8 @@ Reviewer role: Operations and Review Workflow Reviewer. Scope: -- markdown docs in `/home/abiorh/flow/workstream` -- metadata-level process patterns under `/home/abiorh/snorkel` +- markdown docs in `` +- metadata-level process patterns under `local reference workspace` - no task content, private data, or confidential project details copied ## Findings @@ -26,7 +26,7 @@ Status: fixed in `docs/process_pattern_baseline.md`. Finding: -The Snorkel-style projects often have review guards, simulation gates, or screening lanes before work is treated as ready. Workstream would break in daily use if weak tasks went straight from draft to ready. +The reference evaluation projects often have review guards, simulation gates, or screening lanes before work is treated as ready. Workstream would break in daily use if weak tasks went straight from draft to ready. Suggested change: @@ -62,7 +62,7 @@ Status: fixed in `docs/operations_queue_policy.md`. Finding: -Geranium-like and review-guard-heavy workflows show the value of pretending to reject the task before release. Workstream needed this as an operational gate. +Reference review-guard-heavy workflows show the value of pretending to reject the task before release. Workstream needed this as an operational gate. Suggested change: diff --git a/docs/review_process_pattern_baseline_review.md b/docs/review_process_pattern_baseline_review.md index cc9f96818..af021a979 100644 --- a/docs/review_process_pattern_baseline_review.md +++ b/docs/review_process_pattern_baseline_review.md @@ -1,8 +1,8 @@ # Process Pattern Baseline Review -Review scope: metadata-level inspection only under `/home/abiorh/snorkel`. No task content copied. +Review scope: metadata-level inspection only under `local reference workspace`. No task content copied. -Projects inspected: Sequoia, Geranium, Excalibur, Marlin, Termius. +Projects inspected: several reference evaluation projects. ## Baseline Patterns Observed @@ -26,7 +26,7 @@ Suggested change: add an operations doc for project workspace conventions and ad ### High: Workstream needs a pre-review simulation gate -Geranium and Sequoia patterns show review guard plus adversarial/reviewer simulation before calling work ready. Workstream has subagent review protocol but not a productized gate in the lifecycle. +reference project patterns show review guard plus adversarial/reviewer simulation before calling work ready. Workstream has subagent review protocol but not a productized gate in the lifecycle. Suggested change: add `pre_review_gate` as an optional checker phase before `REVIEW_PENDING`, and add reviewer simulation to checker policy/project templates. diff --git a/docs/review_systems_architecture_review.md b/docs/review_systems_architecture_review.md index 5be75f16a..14d5ff3e1 100644 --- a/docs/review_systems_architecture_review.md +++ b/docs/review_systems_architecture_review.md @@ -1,6 +1,6 @@ # Systems Architecture Review -Review scope: markdown docs in `/home/abiorh/flow/workstream`. +Review scope: markdown docs in ``. Reviewer role: Systems Architecture Reviewer. @@ -56,7 +56,7 @@ Status: fixed in `README.md`. ## Baseline Scope Update -Baseline scanned at metadata/process level only under `/home/abiorh/snorkel`, covering Sequoia, Geranium, Excalibur, Marlin, and Termius. The scan looked at guide names/headings, queue/status structures, review guards, checker/preflight scripts, review packet/evidence/status patterns, and submission package/provenance structures. No task content or confidential details were copied. +Baseline scanned at metadata/process level only under `local reference workspace`, covering several reference projects. The scan looked at guide names/headings, queue/status structures, review guards, checker/preflight scripts, review packet/evidence/status patterns, and submission package/provenance structures. No task content or confidential details were copied. ### Medium: Packaged submission provenance should be first-class diff --git a/docs/roadmap_status.md b/docs/roadmap_status.md index 69985bd38..55a2571d8 100644 --- a/docs/roadmap_status.md +++ b/docs/roadmap_status.md @@ -70,6 +70,9 @@ Current phase: Week 3 review and revision preparation. - `WS-POL-001-15` hardened the agent-derived submission artifact policy contract after the accepted no-DB Terminal Benchmark drill exposed a required/forbidden self-conflict; the drill now passes after hardening. +- `WS-POL-001-16` completed a human-visible Terminal Benchmark live API drill + without database inspection as lifecycle proof; the privacy-scrubbed evidence + is at PR/human checkpoint. ## Pending Before Pilot @@ -83,7 +86,7 @@ Current phase: Week 3 review and revision preparation. Run from the backend directory against local Postgres: ```bash -WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test .venv/bin/python scripts/api_contract_e2e.py +WORKSTREAM_DATABASE_URL= .venv/bin/python scripts/api_contract_e2e.py ``` The script runs migrations forward and exercises project policy visibility plus task context APIs across the following flow: @@ -95,7 +98,7 @@ The script runs migrations forward and exercises project policy visibility plus Run from the backend directory against local Postgres: ```bash -WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test .venv/bin/python scripts/week2_api_e2e.py +WORKSTREAM_DATABASE_URL= .venv/bin/python scripts/week2_api_e2e.py ``` The script starts a real local API server, issues local Flow-compatible tokens, @@ -127,10 +130,10 @@ Week 2 closeout validation is not only this script. The full gate is: ```bash .venv/bin/python -m ruff check app tests scripts -WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test .venv/bin/python scripts/api_contract_e2e.py -WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test .venv/bin/python scripts/week2_api_e2e.py -WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test .venv/bin/python -m pytest tests/test_checkers.py tests/test_tasks.py -q -WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test .venv/bin/python -m pytest -q +WORKSTREAM_DATABASE_URL= .venv/bin/python scripts/api_contract_e2e.py +WORKSTREAM_DATABASE_URL= .venv/bin/python scripts/week2_api_e2e.py +WORKSTREAM_DATABASE_URL= .venv/bin/python -m pytest tests/test_checkers.py tests/test_tasks.py -q +WORKSTREAM_DATABASE_URL= .venv/bin/python -m pytest -q .venv/bin/docstr-coverage --config .docstr.yaml ``` diff --git a/examples/terminal_benchmark/LOCAL_VALIDATION_NOTES.md b/examples/terminal_benchmark/LOCAL_VALIDATION_NOTES.md index bc19d9c89..d029dd5cb 100644 --- a/examples/terminal_benchmark/LOCAL_VALIDATION_NOTES.md +++ b/examples/terminal_benchmark/LOCAL_VALIDATION_NOTES.md @@ -75,7 +75,7 @@ cd backend && .venv/bin/python -m ruff check app tests scripts cd backend && .venv/bin/docstr-coverage app scripts --config .docstr.yaml git diff --check cd backend && .venv/bin/python -m pytest tests/test_checkers.py -k 'pre_submit_check_allows_worker_revision_packet_feedback or pre_submit_check_returns_feedback_without_durable_run' -cd backend && WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test WORKSTREAM_TERMINAL_BENCH_FIXTURE=/path/to/terminal-benchmark-source-material .venv/bin/python ../examples/terminal_benchmark/terminal_benchmark_api_e2e.py +cd backend && WORKSTREAM_DATABASE_URL= WORKSTREAM_TERMINAL_BENCH_FIXTURE= .venv/bin/python ../examples/terminal_benchmark/terminal_benchmark_api_e2e.py cd backend && .venv/bin/python -m pytest ``` @@ -97,11 +97,11 @@ Date: 2026-07-05 Purpose: -Record that a real Terminal Benchmark reviewer fixture was used in formal +Record that a real Terminal Benchmark reference fixture was used in formal `.agent-loop` evidence to prove the current Workstream policy-bundle path: - project guide creation -- immutable guide-source snapshot from real Termius guide, reviewer, task, and +- immutable guide-source snapshot from real Terminal Benchmark guide, review program, task, and review packet material - guide sufficiency report - project `SubmissionArtifactPolicy` @@ -137,7 +137,8 @@ Results: - clean packet reached `review_pending` - missing static guard was blocked at pre-submit and created no submission - blocked pre-submit and blocked submission-create produced no durable - submission, evidence, checker-run, checker-result, or audit side effects + submission, evidence, checker-run, or checker-result side effects; task audit + recorded the blocked intake attempt - after a v2 guide and project checker became active, the already-started task still used its locked v1 checker bundle - checker-caused v1 reached `needs_revision` @@ -150,7 +151,7 @@ chunk evidence lives under `.agent-loop/`. Date: 2026-07-05 -The current proof was rerun manually over HTTP against a live local uvicorn +The 2026-07-05 proof was rerun manually over HTTP against a live local uvicorn server and local Postgres. The Python example scaffold was not used as the authoritative proof for this pass. @@ -158,8 +159,8 @@ Live API sequence: - health check returned `ok` - project manager created project -- project manager created a project guide with full Termius submission program, - reviewer project guide, reviewer program, task TOML, and review packet content +- project manager created a project guide with full Terminal Benchmark submission program, + project guide, review program, task TOML, and review packet content - project manager created an immutable guide-source snapshot with source hashes and sanitized durable refs - `ProjectGuideSufficiencyAgent` endpoint returned `passed` @@ -184,23 +185,23 @@ Live API sequence: Live IDs captured from local HTTP API responses: -- project: `6e87e2c2-91a1-4140-8f66-6d0c5bd4b966` -- guide: `b2857abb-6bb0-4e27-89e8-bfb3bfedb8f2` -- guide-source snapshot: `185e80bb-5676-4370-a09f-1c51853bd400` -- sufficiency report: `8368e1c5-cbd6-4503-94f8-74e647a15550` -- agent-derived policy draft: `2838434c-7695-4037-a6d4-531f860a07a6` -- admin exact policy: `dc2e054b-5ce1-49fe-8388-49eb0ec7f992` +- project: `` +- guide: `` +- guide-source snapshot: `` +- sufficiency report: `` +- agent-derived policy draft: `` +- admin exact policy: `` - effective project submission artifact policy hash: - `sha256:38213716e58f10f0916029f91a882681dc52136c9460a958bb4780b070da82f8` + `sha256:` - compiled pre-submit checker hash: - `sha256:1dc2e4b8e9a509e26f6fff8a6da68fbc7340654ecc135d025351173501265855` -- clean task: `9fd7be8f-5886-403b-8ce7-faba37705e72` -- clean submission: `ad0d08f9-4b91-4363-85e9-d8a7b6e055a8` -- clean checker run: `4e72cf39-3348-48b1-8d1a-b3ae17433c65` -- revision-path task: `3ae5db8a-eb40-49bb-8a2a-87ccb1f6594f` -- revision v1 submission: `77a5614b-d0cd-4875-abe2-6e4a83d213cb` -- revision v1 checker run: `b4c9fb23-bf83-48f2-a9ca-5e145aaa707d` -- fixed v2 submission: `a233eefd-598e-4c1b-89c5-c1a93b077682` + `sha256:` +- clean task: `` +- clean submission: `` +- clean checker run: `` +- revision-path task: `` +- revision v1 submission: `` +- revision v1 checker run: `` +- fixed v2 submission: `` Runtime issue found and fixed: diff --git a/examples/terminal_benchmark/README.md b/examples/terminal_benchmark/README.md index fcf56985b..57d594f1c 100644 --- a/examples/terminal_benchmark/README.md +++ b/examples/terminal_benchmark/README.md @@ -24,15 +24,20 @@ approved `.agent-loop` or `docs/internal_reviews` paths for that chunk. - `WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL`. - `WORKSTREAM_TERMINAL_BENCH_FIXTURE` pointing at one local Terminal Benchmark source-material directory. -- `WORKSTREAM_TERMIUS_REVIEWER_ROOT` when the source-material directory is not under a - Termius reviewer root containing `PROJECT_GUIDE.md` and +- `WORKSTREAM_TERMINAL_BENCH_GUIDE_ROOT` when the source-material directory is not under a + Terminal Benchmark reference root containing `PROJECT_GUIDE.md` and `REVIEWER_PROGRAM.md`. - Backend dependencies installed. The script fails closed unless `WORKSTREAM_DATABASE_URL` points to local async Postgres using `workstream_test` or `test_workstream`. -The source-material path should point at a local Termius reviewer directory containing +The script suppresses raw per-request progress paths and redacts fixture-derived +identifiers plus local Workstream UUIDs in stdout by default, so copied output +is safe for public PR evidence. Set +`WORKSTREAM_TERMINAL_BENCH_PRINT_RAW_LOCAL_IDS=1` only for local debugging. + +The source-material path should point at a local Terminal Benchmark reference directory containing `extracted/task.toml`, one `*_submission_*.zip`, one `review_packet_*.md`, `static_guard.txt`, `docker_build.log`, `oracle_test.log`, and `starter_m1_test.log`. Keep the concrete local path in your shell environment; @@ -46,7 +51,7 @@ runtime fallback. The authoritative proof for `WS-POL-001-06` was a live manual HTTP drill: 1. create a project; -2. create a project guide containing Termius submission, reviewer, task, and +2. create a project guide containing Terminal Benchmark submission, review, task, and review packet material; 3. create a guide-source snapshot with source hashes and excerpts; 4. run the guide sufficiency agent endpoint; @@ -62,10 +67,10 @@ The authoritative proof for `WS-POL-001-06` was a live manual HTTP drill: ```bash cd backend -WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test \ +WORKSTREAM_DATABASE_URL= \ OPENAI_API_KEY="$OPENAI_API_KEY" \ WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL="${WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL:?set model}" \ -WORKSTREAM_TERMINAL_BENCH_FIXTURE=/path/to/terminal-benchmark-source-material \ -WORKSTREAM_TERMIUS_REVIEWER_ROOT=/path/to/termius_reviewer \ +WORKSTREAM_TERMINAL_BENCH_FIXTURE= \ +WORKSTREAM_TERMINAL_BENCH_GUIDE_ROOT= \ .venv/bin/python ../examples/terminal_benchmark/terminal_benchmark_api_e2e.py ``` diff --git a/examples/terminal_benchmark/terminal_benchmark_api_e2e.py b/examples/terminal_benchmark/terminal_benchmark_api_e2e.py index cc930c091..1782d7696 100644 --- a/examples/terminal_benchmark/terminal_benchmark_api_e2e.py +++ b/examples/terminal_benchmark/terminal_benchmark_api_e2e.py @@ -1,7 +1,7 @@ """Run real Terminal Benchmark source material through the current API contracts. This is an example drill, not Workstream runtime code and not a required CI -test. It expects a local reviewer source-material path through +test. It expects a local reference source-material path through ``WORKSTREAM_TERMINAL_BENCH_FIXTURE`` and writes only to a local test Postgres database. """ @@ -11,7 +11,9 @@ from __future__ import annotations import asyncio +import contextlib import hashlib +import io import os import re import subprocess @@ -39,7 +41,7 @@ find_free_port, flow_settings, issue_flow_token, - request_json, + request_json as api_contract_request_json, wait_for_health, ) from week2_api_e2e import ( @@ -54,7 +56,8 @@ ) FIXTURE_ENV_VAR = "WORKSTREAM_TERMINAL_BENCH_FIXTURE" -REVIEWER_ROOT_ENV_VAR = "WORKSTREAM_TERMIUS_REVIEWER_ROOT" +GUIDE_ROOT_ENV_VAR = "WORKSTREAM_TERMINAL_BENCH_GUIDE_ROOT" +PRINT_RAW_LOCAL_IDS_ENV_VAR = "WORKSTREAM_TERMINAL_BENCH_PRINT_RAW_LOCAL_IDS" LOCAL_DATABASE_HOSTS = {"localhost", "127.0.0.1", "::1"} LOCAL_DATABASE_NAMES = {"workstream_test", "test_workstream"} ASYNC_POSTGRES_SCHEMES = {"postgresql+asyncpg"} @@ -63,11 +66,18 @@ "OPENAI_API_KEY", "WORKSTREAM_PROJECT_AGENT_OPENAI_AGENT_SDK_MODEL", ) +UUID_PATTERN = re.compile( + r"\b[0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-" + r"[0-9a-fA-F]{4}-[0-9a-fA-F]{12}\b" +) +FIXTURE_ID_PATTERN = re.compile(r"\bterminal-benchmark-[0-9a-fA-F]{12,}\b") +SHA256_VALUE_PATTERN = re.compile(r"\bsha256:[A-Za-z0-9._:-]+") +LOCAL_PATH_PATTERN = re.compile(r"(? Path: ) return root.resolve() raise RuntimeError( - f"{FIXTURE_ENV_VAR} is required. Point it at one Terminal Benchmark reviewer " - "source-material directory, for example a Termius review folder containing extracted/task.toml, " + f"{FIXTURE_ENV_VAR} is required. Point it at one Terminal Benchmark " + "source-material directory, for example a prepared fixture directory containing extracted/task.toml, " "one *_submission_*.zip, one review_packet_*.md, static_guard.txt, and verifier logs." ) -def reviewer_root(fixture_root_path: Path) -> Path: - """Resolve the Termius reviewer root containing real guide/program material.""" - configured = os.environ.get(REVIEWER_ROOT_ENV_VAR) +def guide_root(fixture_root_path: Path) -> Path: + """Resolve the Terminal Benchmark reference root containing guide/program material.""" + configured = os.environ.get(GUIDE_ROOT_ENV_VAR) candidates = [] if configured: candidates.append(Path(configured).expanduser()) @@ -126,8 +136,8 @@ def reviewer_root(fixture_root_path: Path) -> Path: if project_guide.is_file() and reviewer_program.is_file(): return candidate.resolve() raise RuntimeError( - f"could not find PROJECT_GUIDE.md and REVIEWER_PROGRAM.md. Set {REVIEWER_ROOT_ENV_VAR} " - "to the local Termius reviewer root." + f"could not find PROJECT_GUIDE.md and REVIEWER_PROGRAM.md. Set {GUIDE_ROOT_ENV_VAR} " + "to the local Terminal Benchmark reference root." ) @@ -157,10 +167,9 @@ def require_single_fixture_match(root: Path, pattern: str, label: str) -> Path: """ matches = sorted(root.glob(pattern)) if len(matches) != 1: - match_names = [path.name for path in matches] raise RuntimeError( f"expected exactly one {label} matching {pattern!r} in fixture " - f"{safe_label(root.name)!r}; found {len(matches)} file(s): {match_names}" + f"; found {len(matches)} file(s)" ) return matches[0] @@ -177,6 +186,58 @@ def safe_label(value: str) -> str: return re.sub(r"[^a-zA-Z0-9_.-]+", "-", value).strip("-")[:120] or "terminal-benchmark" +def public_evidence_value(value: str, placeholder: str, env: dict[str, str]) -> str: + """Return raw local ids only when explicitly requested for local debugging.""" + if env.get(PRINT_RAW_LOCAL_IDS_ENV_VAR) == "1": + return value + return placeholder + + +def public_safe_exception_message(exc: Exception) -> str: + """Return an exception summary without local ids or source fingerprints.""" + message = str(exc) or exc.__class__.__name__ + message = UUID_PATTERN.sub("", message) + message = FIXTURE_ID_PATTERN.sub("", message) + message = SHA256_VALUE_PATTERN.sub("sha256:", message) + message = LOCAL_PATH_PATTERN.sub("", message) + return f"{exc.__class__.__name__}: {message}" + + +def public_safe_script_failure_message(exc: Exception) -> str: + """Return a generic top-level failure that cannot expose local transcript data.""" + return ( + f"{exc.__class__.__name__}: Terminal Benchmark API drill failed. " + "Raw local failure details are hidden by default; set " + f"{PRINT_RAW_LOCAL_IDS_ENV_VAR}=1 for local debugging." + ) + + +async def request_json(*args, **kwargs) -> dict: + """Call the shared API helper without leaking raw request paths by default.""" + if os.environ.get(PRINT_RAW_LOCAL_IDS_ENV_VAR) == "1": + return await api_contract_request_json(*args, **kwargs) + try: + with contextlib.redirect_stdout(io.StringIO()): + return await api_contract_request_json(*args, **kwargs) + except Exception as exc: + raise RuntimeError(public_safe_exception_message(exc)) from None + + +async def wait_for_submission_checker_run_public_safe( + client: httpx.AsyncClient, + manager_token: str, + submission_id: str, +) -> dict: + """Wait for a checker run without leaking local ids on failure by default.""" + if os.environ.get(PRINT_RAW_LOCAL_IDS_ENV_VAR) == "1": + return await wait_for_submission_checker_run(client, manager_token, submission_id) + try: + with contextlib.redirect_stdout(io.StringIO()): + return await wait_for_submission_checker_run(client, manager_token, submission_id) + except Exception as exc: + raise RuntimeError(public_safe_exception_message(exc)) from None + + def assert_strict_local_database_url(database_url: str) -> None: """Fail closed unless the drill targets local async Postgres test databases only.""" parsed = urlparse(database_url) @@ -200,10 +261,10 @@ def assert_strict_local_database_url(database_url: str) -> None: def load_fixture(root: Path) -> TerminalBenchmarkFixture: - """Load and validate a Terminal Benchmark reviewer fixture. + """Load and validate a Terminal Benchmark reference fixture. Args: - root: Fixture directory copied from the Termius reviewer workspace. + root: Prepared fixture directory copied from Terminal Benchmark reference material. Returns: Parsed fixture paths and metadata. @@ -224,12 +285,12 @@ def load_fixture(root: Path) -> TerminalBenchmarkFixture: with task_toml.open("rb") as file: task_config = tomllib.load(file) - termius_root = reviewer_root(root) + guide_material_root = guide_root(root) return TerminalBenchmarkFixture( root=root, fixture_id=sanitized_fixture_id(task_toml, submission_zip), - project_guide=termius_root / "PROJECT_GUIDE.md", - reviewer_program=termius_root / "REVIEWER_PROGRAM.md", + project_guide=guide_material_root / "PROJECT_GUIDE.md", + reviewer_program=guide_material_root / "REVIEWER_PROGRAM.md", task_toml=task_toml, submission_zip=submission_zip, static_guard=root / "static_guard.txt", @@ -292,7 +353,7 @@ def evidence_entry(file: FixtureFile, fixture: TerminalBenchmarkFixture) -> dict return { "type": "log", "label": file.label, - "uri": f"local://termius/{fixture.fixture_id}/{file.artifact_name}", + "uri": f"local://terminal-benchmark/{fixture.fixture_id}/{file.artifact_name}", "hash": sha256_token(file.path), "size_bytes": file.path.stat().st_size, "metadata": { @@ -311,7 +372,7 @@ def fixture_files(fixture: TerminalBenchmarkFixture) -> list[FixtureFile]: "submission.zip", fixture.submission_zip, "original submission zip", - "original reviewer-side submission archive", + "original submission archive", ), FixtureFile( "task.toml", @@ -323,13 +384,13 @@ def fixture_files(fixture: TerminalBenchmarkFixture) -> list[FixtureFile]: "static_guard.txt", fixture.static_guard, "platform static guard output", - "static guard output captured by the reviewer", + "static guard output captured with the fixture", ), FixtureFile( "review_packet.md", fixture.review_packet, "automated review packet", - "AutoEval and reviewer packet evidence", + "AutoEval and review packet evidence", ), FixtureFile( "docker_build.log", @@ -361,7 +422,7 @@ def task_payload(fixture: TerminalBenchmarkFixture, run_id: str, suffix: str) -> return { "title": f"Terminal Benchmark {fixture.fixture_id} {suffix}", "description": ( - "Real Terminal Benchmark reviewer fixture with " + "Real Terminal Benchmark reference fixture with " f"{milestone_count} milestones, languages={metadata['languages']}, " f"category={metadata['category']}." ), @@ -394,9 +455,9 @@ def guide_payload(fixture: TerminalBenchmarkFixture, run_id: str) -> dict: "content_markdown": ( f"# Terminal Benchmark Guide {run_id}\n\n" f"Fixture: `{fixture.fixture_id}`\n\n" - "## Termius Project Guide\n\n" + "## Terminal Benchmark Project Guide\n\n" f"{project_guide}\n\n" - "## Termius Reviewer Program\n\n" + "## Terminal Benchmark Reference Program\n\n" f"{reviewer_program}\n\n" "## Selected Terminal Benchmark Task TOML\n\n" "```toml\n" @@ -452,13 +513,13 @@ def submission_payload( files = [file for file in files if file.artifact_name != "static_guard.txt"] summary = ( f"Terminal Benchmark {fixture.fixture_id} packet {suffix} from real " - "reviewer-side fixture evidence." + "reference fixture evidence." ) if low_quality_signal: summary += " Placeholder sample output requires reviewer revision." return { "summary": summary, - "package_uri": f"local://termius/{fixture.fixture_id}/submission.zip", + "package_uri": f"local://terminal-benchmark/{fixture.fixture_id}/submission.zip", "package_hash": sha256_token(fixture.submission_zip), "artifact_hash_manifest": [artifact_entry(file) for file in files], "worker_attestation": STRONG_ATTESTATION, @@ -802,7 +863,11 @@ async def submit_finalize_and_wait( f"/api/v1/submissions/{submission['id']}/finalize", manager_token, ) - run = await wait_for_submission_checker_run(client, manager_token, submission["id"]) + run = await wait_for_submission_checker_run_public_safe( + client, + manager_token, + submission["id"], + ) return submission, locked, run @@ -1236,14 +1301,37 @@ async def exercise_terminal_benchmark_api(base_url: str, env: dict[str, str]) -> print("Terminal Benchmark real API e2e passed") print("scenario_summary:") - print(f"fixture_id={fixture.fixture_id}") - print(f"fixture_label={safe_label(fixture.root.name)}") - print(f"project_id={project['id']}") - print(f"complete_task_id={complete_task['id']}") - print(f"complete_submission_id={complete_submission['id']}") - print(f"revision_task_id={revision_task['id']}") - print(f"revision_v1_submission_id={first_submission['id']}") - print(f"revision_v2_submission_id={second_submission['id']}") + print( + "redaction=" + + public_evidence_value( + "raw_local_ids_enabled", + "public_evidence_safe_by_default", + env, + ) + ) + print( + "fixture_id=" + + public_evidence_value(fixture.fixture_id, "", env) + ) + print( + "fixture_label=" + + public_evidence_value(safe_label(fixture.root.name), "", env) + ) + print("project_id=" + public_evidence_value(project["id"], "", env)) + print("complete_task_id=" + public_evidence_value(complete_task["id"], "", env)) + print( + "complete_submission_id=" + + public_evidence_value(complete_submission["id"], "", env) + ) + print("revision_task_id=" + public_evidence_value(revision_task["id"], "", env)) + print( + "revision_v1_submission_id=" + + public_evidence_value(first_submission["id"], "", env) + ) + print( + "revision_v2_submission_id=" + + public_evidence_value(second_submission["id"], "", env) + ) print("complete_packet=review_pending") print("missing_static_guard=pre_submit_blocked_no_submission") print("low_quality_v1=needs_revision") @@ -1270,9 +1358,15 @@ async def main(env: dict[str, str]) -> None: if __name__ == "__main__": - api_env = api_environment() - require_openai_agent_sdk_environment(api_env) - assert_strict_local_database_url(api_env["WORKSTREAM_DATABASE_URL"]) - os.environ.update(api_env) - command.upgrade(alembic_config(), "head") - asyncio.run(main(api_env)) + try: + api_env = api_environment() + require_openai_agent_sdk_environment(api_env) + assert_strict_local_database_url(api_env["WORKSTREAM_DATABASE_URL"]) + os.environ.update(api_env) + command.upgrade(alembic_config(), "head") + asyncio.run(main(api_env)) + except Exception as exc: + if os.environ.get(PRINT_RAW_LOCAL_IDS_ENV_VAR) == "1": + raise + print(public_safe_script_failure_message(exc), file=sys.stderr) + raise SystemExit(1) from None