Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 13 additions & 10 deletions .agent-loop/LOOP_STATE.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,17 +4,20 @@

- Active initiative: `WS-POL-001` - Submission Artifact Policy Foundation
- Active planning chunk: none
- Active implementation chunk: none
- Branch: `main`
- Status: `WS-POL-001-15` merged through PR #81. The project setup derivation
prompt now explicitly prevents required/forbidden artifact self-conflicts,
keeps derivation project-scoped, and the accepted no-DB Terminal Benchmark
live API drill passes after hardening.
- Active implementation chunk: `WS-POL-001-16` - Terminal Benchmark Live API Drill
- Branch: `codex/ws-pol-001-16-terminal-benchmark-live-api-drill`
- Status: `WS-POL-001-16` completed the final clean Terminal Benchmark live API
drill through real HTTP-visible APIs. The accepted run used sanitized source
material, automatic project setup, live `submission-requirements`-derived
worker packets, blocked pre-submit no-side-effect proof, successful
submission finalization, durable checker-run visibility, and final
`review_pending` task state without database inspection as lifecycle proof.
- Last merged implementation SHA: `b72a5b9`
- Last merge commit: `b1a9851`
- Current gate: post-merge memory update for PR #81, then stop for the user's
next explicit implementation chunk.
- Next chunk: inactive until the user explicitly starts it.
- Current gate: PR creation and human checkpoint for `WS-POL-001-16`; internal
reviewer fanout and evidence gate are complete.
- Next chunk: inactive until this chunk is reviewed, merged, and followed by a
post-merge memory update.

## Operating Rule

Expand Down Expand Up @@ -78,7 +81,7 @@ blockchain, frontend, or agent-runtime behavior.
- `WS-POL-001-06` started on branch `codex/ws-pol-001-06-terminal-benchmark-drill`
after the user's explicit start signal.
- `WS-POL-001-06` real Terminal Benchmark manual HTTP drill passed against a
local Termius reviewer fixture; committed evidence uses placeholder fixture
local Terminal Benchmark reference fixture; committed evidence uses placeholder fixture
paths and local IDs only.
- `WS-POL-001-06` live drill exposed and fixed an OpenAI Agents SDK adapter
strict-schema issue for the policy derivation result's open `policy_body`.
Expand Down
6 changes: 3 additions & 3 deletions .agent-loop/WORK_QUEUE.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@

| Chunk | Title | Risk | Status |
|---|---|---:|---|
| none | none | - | Waiting for user to explicitly start the next chunk |
| `WS-POL-001-16` | Terminal Benchmark Live API Drill | L1 | Active on `codex/ws-pol-001-16-terminal-benchmark-live-api-drill` |

## Completed

Expand All @@ -31,8 +31,8 @@

## Proposed Next

Stop after the PR #81 post-merge memory update. Do not start the next
implementation chunk until the user explicitly starts it.
Stop after `WS-POL-001-16` is implemented, reviewed, and opened for human
review. Do not start another implementation chunk from this branch.

## Blocked

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -10,11 +10,24 @@ valid findings addressed: yes

## Reviewed Revision

Reviewed code SHA: 49101d4ad3fc22ec6e6065b1e593ef04145db953

Reviewed at: 2026-07-09T06:13:59Z

Reviewer run IDs: senior-engineering-final-reviewer-run-id, qa-test-final-reviewer-run-id, security-auth-final-reviewer-run-id, product-ops-final-reviewer-run-id, architecture-final-reviewer-run-id, docs-final-reviewer-run-id, reuse-dedup-final-reviewer-run-id, test-delta-final-reviewer-run-id, ci-integrity-final-reviewer-run-id

Current privacy-scrub chunk: `WS-POL-001-16-terminal-benchmark-live-api-drill`.
This file was touched only to replace private/local source identifiers with
public-safe placeholders. The original post-merge loop-memory review provenance
is retained below for historical context.

Original reviewed revision:

Reviewed code SHA: f4fe5f3c4fbdd626bbc6d3f837aeca1cceb6e9ca

Reviewed at: 2026-06-20T13:15:54Z

Reviewer run IDs: 019ee4bd-d3d5-7830-b042-a46397b2a4f3, 019ee4be-9fd5-78d2-801a-8ccb7541ad19, 019ee4c0-e266-71e3-b65e-3f1afa8af74c, 019ee4c3-8994-7a50-9bb9-49962001a247, 019ee4dd-f49e-72d2-abd4-6391aafe95d3, 019ee4fe-9b01-7741-a130-a4a78f2054b0, 019ee500-050e-7702-99df-a38a87435281, 019ee502-a260-7e01-affe-77867dd21325, 019ee504-e427-76c1-a66f-3fc036207abe
Reviewer run IDs: historical-senior-engineering-review, historical-qa-test-review, historical-security-auth-review, historical-product-ops-review, historical-architecture-review, historical-docs-review, historical-reuse-dedup-review, historical-test-delta-review, historical-ci-integrity-review

After reviewed SHA `f4fe5f3c4fbdd626bbc6d3f837aeca1cceb6e9ca`, the only committed path changed in this PR is this internal review evidence file. No implementation, workflow, test, policy, or loop-memory state file changed after that reviewed SHA.

Expand All @@ -34,7 +47,7 @@ After reviewed SHA `f4fe5f3c4fbdd626bbc6d3f837aeca1cceb6e9ca`, the only committe

## Valid Findings Addressed

- Local Workstream directory confusion: identified `/home/abiorh/flow/workstream` as a separate dirty feature branch, not `main`, and left unrelated checker/test changes untouched.
- Local Workstream directory confusion: identified `<repo-root>` as a separate dirty feature branch, not `main`, and left unrelated checker/test changes untouched.
- Stale merged-loop memory: updated `.agent-loop/LOOP_STATE.md`, initiative `STATUS.md`, `WORK_QUEUE.md`, and `REVIEW_LOG.md` to reflect that PR #23 is merged.
- Missing main enforcement: added the verified workflow path `.github/workflows/loop-memory.yml` so merged loop memory is checked on pushes to `main`.
- Over-broad local-state test risk: changed loop-memory regression tests to use fixture files instead of the live repository state.
Expand All @@ -54,4 +67,4 @@ git diff --check HEAD~1..HEAD

## Remaining Risks

- `/home/abiorh/flow/workstream` remains dirty on `codex/submission-artifact-policy-docs` with unrelated checker/revision testing changes. Those changes were not modified here because they are outside PR #24.
- `<repo-root>` remains dirty on `codex/submission-artifact-policy-docs` with unrelated checker/revision testing changes. Those changes were not modified here because they are outside PR #24.
Original file line number Diff line number Diff line change
Expand Up @@ -483,7 +483,7 @@ Fair worker experience during revision and audit clarity.

Goal:

Use a real Terminal Benchmark reviewer fixture from the local Termius workspace
Use a real Terminal Benchmark reference fixture from the local Terminal Benchmark reference workspace
to prove the current Workstream setup-agent route, project policy bundle, task
locked context, pre-submit feedback, submission versioning, post-submit checker
gate, and fixed revision path over live manual HTTP calls and local Postgres.
Expand Down Expand Up @@ -573,8 +573,8 @@ Acceptance criteria:

Verification:

- Manual live API drill runs against local Postgres and one explicit Termius
reviewer fixture path.
- Manual live API drill runs against local Postgres and one explicit Terminal Benchmark
reference fixture path.
- Targeted adapter regression tests, stale wording scan, ruff, docstring
coverage, markdown link check, and diff whitespace checks pass.

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -3,8 +3,11 @@
## Current Status

`WS-POL-001-01` through `WS-POL-001-15` are merged to `main`.
The post-actor-registry Terminal Benchmark live API drill passed through real
HTTP calls, and task context visibility is now exposed through APIs.
`WS-POL-001-16` completed the final clean Terminal Benchmark live API drill
through real HTTP calls, using sanitized source material and a worker packet
derived from the live `submission-requirements` response. Evidence is recorded
and internal reviewer fanout is complete. The branch is ready for PR/human
checkpoint.
`WS-POL-001-14` replaced public submission lock wording with finalization,
defined system actor audit semantics, and merged PR #79's HTTP-visible Terminal
Benchmark proof evidence. The accepted post-merge no-DB Terminal Benchmark
Expand All @@ -14,7 +17,7 @@ reran that accepted drill successfully before merging through PR #81.

## Active Chunk

None. Waiting for the user's next explicit implementation chunk.
`WS-POL-001-16` - Terminal Benchmark Live API Drill.

## Chunk Status

Expand All @@ -35,6 +38,7 @@ None. Waiting for the user's next explicit implementation chunk.
| `WS-POL-001-13` | Merged | `codex/ws-pol-001-13-task-context-apis` | 77 | Adds task work-context, worker submission-requirements, and operator-only locked-context APIs. |
| `WS-POL-001-14` | Merged | `codex/ws-pol-001-14-submission-finalize` | 79 | Replaces public submission lock with finalize, defines system actor audit semantics, scopes operator visibility, and proves the Terminal Benchmark flow through HTTP-visible lifecycle responses. |
| `WS-POL-001-15` | Merged | `codex/ws-pol-001-15-agent-derivation-hardening` | 81 | Hardens agent-derived submission artifact policy instructions after the no-DB Terminal Benchmark drill exposed a required-artifact/forbidden-pattern self-conflict. |
| `WS-POL-001-16` | Internal review complete | `codex/ws-pol-001-16-terminal-benchmark-live-api-drill` | - | Proved a human-visible Terminal Benchmark drill through real HTTP APIs without DB inspection as lifecycle proof; PR/human checkpoint is pending. |

## Blockers

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ WS-POL-001 - Submission Artifact Policy Foundation

## Goal

Use a real Terminal Benchmark reviewer fixture from the local Termius workspace
Use a real Terminal Benchmark reference fixture from the local Terminal Benchmark reference workspace
to prove the current Workstream project guide, setup-agent, policy bundle, task
locked context, pre-submit feedback, submission versioning, post-submit checker
gate, and revision resubmission path over live HTTP calls and local Postgres.
Expand Down Expand Up @@ -119,7 +119,7 @@ work. Further unrelated runtime bugs still require a separate chunk.
`PreSubmitCheckerPolicy` as the intake contract.
- The drill does not rely on task `required_files` or `required_evidence` as the
source of pre-submit truth.
- The guide source snapshot is built from real Termius material, including the
- The guide source snapshot is built from real Terminal Benchmark reference material, including the
Terminal Benchmark submission guide/program material, reviewer program or
guide material, the selected task TOML, and the selected review packet.
- Persisted fixture identifiers and normal success output do not reveal absolute
Expand Down Expand Up @@ -153,10 +153,10 @@ cd backend && .venv/bin/python -m pytest tests/test_projects.py -k 'openai_agent
cd backend && .venv/bin/python -m pytest tests/test_projects.py
cd backend && .venv/bin/python -m pytest tests/test_tasks.py
cd backend && .venv/bin/python -m pytest tests/test_alembic.py
cd backend && WORKSTREAM_DATABASE_URL=postgresql+asyncpg://workstream:workstream@localhost:5433/workstream_test WORKSTREAM_TERMINAL_BENCH_FIXTURE=/path/to/local/terminal-benchmark-fixture .venv/bin/python ../examples/terminal_benchmark/terminal_benchmark_api_e2e.py
cd backend && WORKSTREAM_DATABASE_URL=<local-test-db-url> WORKSTREAM_TERMINAL_BENCH_FIXTURE=<redacted-local-fixture-path> .venv/bin/python ../examples/terminal_benchmark/terminal_benchmark_api_e2e.py
```

The fixture path may be changed to another local Termius reviewer fixture that
The fixture path may be changed to another local Terminal Benchmark reference fixture that
contains the required files. The command must stay local-only and must never run
against production or shared databases.

Expand Down
Loading
Loading