From a304274412f080817e6da2967e444b31ccd2f864 Mon Sep 17 00:00:00 2001 From: SuPuHe Date: Sun, 6 Sep 2026 14:16:36 +0200 Subject: [PATCH 1/2] chore: establish OneShot agent workflow --- .agent/AGENTS.md | 214 +++++++++--------- .agent/IMPLEMENTATION_LOOP.md | 105 +++++++++ .agent/MILESTONE_IMPLEMENTATION_LOOP.md | 164 -------------- .agent/PROJECT_CONTEXT.md | 73 ++++++ .agent/SECURITY_INVARIANTS.md | 55 +++++ .agent/SPONSOR_REQUIREMENTS.md | 40 ++++ .agent/TEST_MATRIX.md | 33 +++ .../context/20260906T115948Z-agents-setup.md | 107 +++++++++ .agent/context/README.md | 41 ++++ .agent/context/SESSION_TEMPLATE.md | 58 +++++ .agent/context/new-session.sh | 20 ++ .agent/review-prompts/draft-pr-review.md | 67 ------ .agent/review-prompts/freepi-pr-review.md | 69 ++++++ .../review-prompts/freepi-prepush-review.md | 67 ++++++ .../review-prompts/implementation-review.md | 59 ----- .../skills/oneshot-failure-injection/SKILL.md | 28 +++ .agents/skills/oneshot-idempotency/SKILL.md | 31 +++ .agents/skills/sponsor-qualification/SKILL.md | 25 ++ .antigravity/README.md | 17 -- .antigravity/review.md | 14 -- .github/BRANCH_POLICY.md | 52 +++-- .github/PULL_REQUEST_TEMPLATE.md | 74 +++--- .gitignore | 17 ++ AGENTS.md | 31 ++- 24 files changed, 971 insertions(+), 490 deletions(-) create mode 100644 .agent/IMPLEMENTATION_LOOP.md delete mode 100644 .agent/MILESTONE_IMPLEMENTATION_LOOP.md create mode 100644 .agent/PROJECT_CONTEXT.md create mode 100644 .agent/SECURITY_INVARIANTS.md create mode 100644 .agent/SPONSOR_REQUIREMENTS.md create mode 100644 .agent/TEST_MATRIX.md create mode 100644 .agent/context/20260906T115948Z-agents-setup.md create mode 100644 .agent/context/README.md create mode 100644 .agent/context/SESSION_TEMPLATE.md create mode 100755 .agent/context/new-session.sh delete mode 100644 .agent/review-prompts/draft-pr-review.md create mode 100644 .agent/review-prompts/freepi-pr-review.md create mode 100644 .agent/review-prompts/freepi-prepush-review.md delete mode 100644 .agent/review-prompts/implementation-review.md create mode 100644 .agents/skills/oneshot-failure-injection/SKILL.md create mode 100644 .agents/skills/oneshot-idempotency/SKILL.md create mode 100644 .agents/skills/sponsor-qualification/SKILL.md delete mode 100644 .antigravity/README.md delete mode 100644 .antigravity/review.md diff --git a/.agent/AGENTS.md b/.agent/AGENTS.md index 051cdb9..6124f11 100644 --- a/.agent/AGENTS.md +++ b/.agent/AGENTS.md @@ -1,123 +1,121 @@ # OneShot Agent Policy -This file defines mandatory behavior for every coding, documentation, -infrastructure, data, and model agent working in this repository. +This policy applies to code, documentation, infrastructure, data, and agent +work in this repository. ## Instruction order 1. Follow system and user instructions. -2. Follow the root `AGENTS.md` and this file. -3. Follow `.agent/MILESTONE_IMPLEMENTATION_LOOP.md` for milestone work. -4. Follow the narrowest applicable repository documentation and configuration. - -If instructions conflict, stop and surface the conflict. Do not silently choose -the most convenient interpretation. - -## Before making changes - -- Read the task, acceptance criteria, relevant code, and related documentation. -- Inspect `git status`, the current branch, and the diff before editing. -- Preserve unrelated user changes and never include them in a commit. -- Identify the milestone, the smallest reviewable outcome, and explicit non-goals. -- State assumptions that can materially affect behavior, security, privacy, - cost, licensing, or architecture. -- Prefer evidence from the repository and official primary documentation over - memory for version-sensitive technical decisions. - -## Scope and architecture - -- One branch and pull request must represent one milestone or one tightly - related correction. -- Keep scope lean and focused. Do not introduce speculative features, unused - dependencies, or unnecessary abstractions unless the active milestone explicitly - requires them. -- Preserve clean separation of concerns: presentation/interface, orchestration, - domain logic, and external service adapters. -- Validate all untrusted input at boundaries. -- Handle state transitions, loading, empty states, and errors predictably. - -## Implementation rules - -- Write strict, readable code adhering to established style conventions. -- Do not weaken compiler, lint, or type-check configurations to make a change pass. -- Prefer small, typed interfaces at subsystem boundaries. -- Add or update tests for behavior changes and regression fixes. +2. Follow root `AGENTS.md` and this policy. +3. Follow the task-specific documents and repo skills routed by root + `AGENTS.md`. +4. Follow the narrowest applicable repository configuration. + +Surface conflicts. Never silently weaken an invariant, review gate, or security +boundary. + +## Product boundary + +OneShot's core promise is: `One job. Many retries. One settlement.` + +- OneShot owns authoritative business-intent execution state and prevents + duplicate committed settlements. +- Privy provides corporate wallet access and scoped authorization, policy, and + spending permissions. +- Arc is the USDC settlement rail. +- The Graph provides live indexed history and recovery context. It is never the + sole duplicate-payment lock or authority for creating another settlement. + +Read `.agent/PROJECT_CONTEXT.md`, `.agent/SECURITY_INVARIANTS.md`, and +`.agent/SPONSOR_REQUIREMENTS.md` before changing these boundaries. + +## Non-negotiable invariants + +- `1 business intent -> at most 1 committed settlement`. +- Keep one stable `business_intent_id` across retries, restarts, parallel + attempts, and agent instances. +- Treat `UNKNOWN` settlement state as a reconciliation requirement. Never + blindly repay. +- Make state durable and transitions atomic and concurrency-safe. +- Represent money as integer atomic units or `bigint`, never JavaScript + floating point. +- Graph absence or indexing delay is not proof that payment did not happen. +- Normal execution must not bypass Privy policy or OneShot controls. +- Use testnet only unless the user explicitly authorizes another network. +- Never log, expose, persist, commit, or send secrets, private keys, seed + phrases, tokens, wallet credentials, or sensitive runtime configuration. + +## Before changing files + +- Inspect current branch, status, task acceptance criteria, relevant code, and + existing diff. +- Preserve unrelated user work and keep it out of commits. +- Record material assumptions and the active context in `.agent/context/`. +- Use current primary documentation for version-sensitive integrations. +- Do not create the product implementation `plan.md` until required skills and + integration research are ready. Agent infrastructure work is not that plan. + +## Implementation quality + +- Keep one branch and PR focused on one milestone or tightly related change. +- Preserve clear ownership among interface, orchestration, domain state, and + external adapters. +- Validate untrusted input at boundaries. +- Do not weaken compiler, lint, type, test, or security settings to get a pass. +- Add tests for behavior changes and regression fixes. Payment-related changes + must select applicable cases from `.agent/TEST_MATRIX.md`. - Do not leave dead code, unexplained suppressions, placeholder credentials, or - untracked follow-up work hidden in comments. -- Update documentation when behavior or architectural patterns change. + hidden follow-up work. +- Update documentation when behavior, contracts, or architecture change. -## Quality policy +## Skills -Before Review Gate A, run the full local validation suite: +User-level workflow skills expected for Codex: -```bash -# Project quality checks (configure as codebase components are introduced) -# e.g., lint, type check, unit tests, integration checks -``` +- `research`: primary-source integration research before architecture choices. +- `tdd`: one behavior at a time through red-green-refactor. +- `diagnosing-bugs`: evidence-first diagnosis before fixing unclear failures. +- `to-tickets`: split an approved spec or plan into ordered tracer-bullet work. +- `handoff`: compact a session into a durable handoff. +- `resolving-merge-conflicts`: resolve active merge/rebase conflicts safely. +- `prototype`: answer a design question with disposable code. +- `wizard`: guide human-only setup, credentials, or dashboard steps. +- `caveman`: reduce conversational token use; never compress persisted repo + docs, code, review evidence, or security warnings. -Run additional focused tests required by the changed subsystem. A passing build -does not replace behavioral tests, security checks, or manual verification. +Project skills in `.agents/skills/`: -Do not bypass a failed check with `--force`, `--no-verify`, broad ignore rules, -lowered thresholds, or dependency overrides. Fix the cause or document a genuine -blocker for the user. +- `oneshot-idempotency`: mandatory for intent/payment/retry/settlement work. +- `oneshot-failure-injection`: mandatory for external-effect failure boundaries. +- `sponsor-qualification`: mandatory before sponsor/demo/release claims. -## Git and GitHub +Optional generic code-review skills may supplement work. They never satisfy or +replace FreePi Gate A or Gate B. + +## Git and review policy - Never implement directly on `main` or `develop`. -- The default development and integration branch is `develop`. -- All milestone and feature pull requests must target `develop`. -- Use a descriptive branch such as `milestone/-`, `feature/`, or - `fix/`. -- Never force-push, delete a protected branch, rewrite shared history, or use a - destructive reset without explicit user authorization. -- Stage only files belonging to the active milestone and review the staged diff - before committing. -- Draft pull requests may be created only after Review Gate A passes. -- A draft may be marked ready for user review only after required CI and Review - Gate B pass for the exact current head commit. -- Agents must never merge a pull request. The user performs the final review and - explicitly decides whether to merge. - -## Mandatory independent reviews - -Follow `.agent/MILESTONE_IMPLEMENTATION_LOOP.md` exactly. - -- Review Gate A is a fresh, independent review of the complete workspace change - before the draft pull request is created. It evaluates the diff against the target - base branch (`develop`), covering acceptance criteria, correctness, edge cases, - security, and test coverage. -- Review Gate B is a second fresh, independent review after the draft PR exists - and required CI is green. Gate B is bound to the exact PR head commit SHA and verifies - PR readiness, diff integrity, and check results. -- The implementation agent must not act as its own independent reviewer. Reviewers - must be invoked in an independent session using `.agent/review-prompts/implementation-review.md` - for Gate A and `.agent/review-prompts/draft-pr-review.md` for Gate B. -- Do not reuse or resume the Gate A session for Gate B. -- Any content change after Gate A invalidates Gate A. -- Any commit after Gate B invalidates Gate B. -- `WARN`, an incomplete response, unavailable tooling, authentication failure, - or an ambiguous verdict is not a pass. - -## Security, privacy, and secrets - -- Never commit secrets, API keys, tokens, credentials, private certificates, or personal data. -- Maintain a comprehensive `.gitignore` for secrets, local environments, and temporary artifacts. -- Validate input sizes and payloads before running expensive operations. -- Treat dependency and code licensing as core release criteria. - -## Definition of agent-complete - -Work is ready for user review only when: - -- Milestone acceptance criteria are fully met; -- The branch contains only intended changes; -- Local checks and required GitHub checks pass; -- Review Gate A and Review Gate B pass for the current head commit; -- All blocking findings are fixed and re-reviewed; -- The draft PR has been marked ready, but not merged; -- The PR description details scope, risk, validation evidence, and both review verdicts; -- Documentation is updated and accurate. - -When handing off, report the branch, commit SHA, PR URL, checks run, review verdicts, -known limitations, and the exact decision required from the user. +- Branch from current `develop`; target `develop` from short-lived + `feature/*`, `fix/*`, or `milestone/*` branches unless the user explicitly + names a different short-lived branch. +- Never direct-push or force-push protected branches. Never rewrite shared + history without explicit user authorization. +- Follow `.agent/IMPLEMENTATION_LOOP.md` for local checks, both independent + FreePi reviews, CI, PR readiness, and invalidation rules. +- Only explicit `VERDICT: PASS` passes a gate. Missing, ambiguous, truncated, + stale, unauthenticated, or failed review output fails closed. +- Agents never merge. A human must review and explicitly authorize the merge. + +## Context retention + +Follow `.agent/context/README.md`. Update the current record at milestone +boundaries, before handoff or session end, and before deliberate context reset or +compaction when possible. Never store secrets there. + +## Agent-complete + +Handoff only after intended scope is complete, local checks pass, the diff is +cleanly scoped, and the current gate state is recorded. A change is ready for +human review only after Gate A, required CI, and Gate B pass for the exact +applicable content/head SHA. Report branch, commit, PR, checks, gate evidence, +and remaining risks. Never merge. diff --git a/.agent/IMPLEMENTATION_LOOP.md b/.agent/IMPLEMENTATION_LOOP.md new file mode 100644 index 0000000..2f2a8e4 --- /dev/null +++ b/.agent/IMPLEMENTATION_LOOP.md @@ -0,0 +1,105 @@ +# OneShot Implementation Loop + +This is the required, hackathon-friendly path from issue to human review. +`develop` is the base. Gate A binds to complete candidate content; Gate B binds +to the exact draft PR head SHA and its required check state. + +## 1. Scope and branch + +1. Start from current `develop`. +2. Create one short-lived `feature/*`, `fix/*`, or `milestone/*` branch unless + the user explicitly names another short-lived branch. +3. Record goal, acceptance criteria, assumptions, and branch state in + `.agent/context/`. +4. Never implement directly on `develop` or `main`. + +## 2. Implement and validate + +1. Make the smallest coherent change. +2. Use applicable repo skills and `.agent/TEST_MATRIX.md`. +3. Run local format, lint, type, test, build, and focused failure-injection + checks that exist for affected components. +4. Inspect tracked, staged, unstaged, and intended untracked changes against + `develop`. Check scope, generated files, secrets, and unrelated work. +5. Record commands and results in the current context file. + +Do not bypass failures with force flags, skipped checks, broad ignores, lower +thresholds, or disabled hooks. + +## 3. FreePi Gate A: complete pre-push review + +Gate A must run before the first push and before draft PR creation. + +1. From repository root, start a fresh process: + + ```bash + npx free-pi-cli + ``` + +2. In that new FreePi session, provide + `.agent/review-prompts/freepi-prepush-review.md` and the task acceptance + criteria. Do not use invented flags or a `pi --session` command. +3. Let the reviewer inspect the complete intended workspace change against + `develop`, including staged, unstaged, and explicitly intended untracked + files. Never expose ignored or sensitive files. +4. Accept only an explicit `VERDICT: PASS` with reviewed base SHA and sufficient + evidence. + +Each Gate A attempt uses a new `npx free-pi-cli` process. Never resume or reuse a +reviewer context. Any relevant content change after Gate A invalidates it: rerun +local checks and Gate A in another fresh process. + +Missing evidence, an ambiguous or truncated answer, authentication failure, +tool failure, or any verdict other than explicit `VERDICT: PASS` is failure. + +## 4. Commit, push, and draft PR + +Only after Gate A passes for current content: + +1. Stage only reviewed files and inspect the staged diff. +2. Commit with a clear message and push the short-lived branch without force. +3. Create a draft PR targeting `develop`; never target `main` for feature work. +4. Fill `.github/PULL_REQUEST_TEMPLATE.md`, including Gate A evidence. + +## 5. Required CI + +Wait for every required check on the exact draft PR head SHA. Pending, skipped, +missing, or failing required checks are not green. + +If a fix changes content, rerun local validation and a fresh Gate A, commit, +push, and wait for CI again. + +## 6. FreePi Gate B: exact PR review + +After the draft PR exists and required CI is green: + +1. Capture PR URL/number, base, head branch, exact full head SHA, full diff, and + required check results. +2. Start a second fresh process from repository root: + + ```bash + npx free-pi-cli + ``` + +3. Provide `.agent/review-prompts/freepi-pr-review.md` plus the captured PR and + CI evidence. This must not reuse any Gate A process or session. +4. Accept only explicit `VERDICT: PASS` bound to the exact current PR head SHA. + +Any commit or content change after Gate B invalidates Gate B. Return to local +validation, fresh Gate A, commit/push, green CI, then fresh Gate B. + +## 7. Human review and merge + +After Gate A, required CI, and Gate B pass for current content/head: + +1. Record both verdicts and evidence in the PR and context file. +2. Mark the draft ready for human review. +3. Report branch, SHA, PR URL, checks, gates, and residual risks. +4. Stop. Agents never merge. Only a human may authorize and perform merge. + +## Privacy boundary + +FreePi may review code, intended diffs, tests, public docs, and non-sensitive +check output only. Never provide `.env`, `.env.local`, private keys, seed +phrases, access tokens, API secrets, wallet credentials, ignored files, or +sensitive runtime configuration. If safe review is impossible, fail the gate. diff --git a/.agent/MILESTONE_IMPLEMENTATION_LOOP.md b/.agent/MILESTONE_IMPLEMENTATION_LOOP.md deleted file mode 100644 index 0fc7e32..0000000 --- a/.agent/MILESTONE_IMPLEMENTATION_LOOP.md +++ /dev/null @@ -1,164 +0,0 @@ -# Milestone Implementation Loop - -This is the mandatory lifecycle for every milestone and feature implementation in OneShot. -Gate A is bound to reviewed file content against the base branch (`develop`). -Gate B and final readiness are bound to the exact pull-request head commit. - -## Lifecycle - -```mermaid -flowchart LR - A["Phase 1: Scope milestone"] --> B["Phase 2: Implement"] - B --> C["Phase 3: Local validation"] - C --> D["Phase 4: Review Gate A"] - D -->|"FAIL"| B - D -->|"PASS"| E["Phase 5: Commit and create draft PR"] - E --> F["Phase 6: Required CI"] - F -->|"FAIL"| B - F -->|"PASS"| G["Phase 7: Review Gate B"] - G -->|"FAIL"| B - G -->|"PASS"| H["Phase 8: Mark ready for user"] - H --> I["User review"] - I -->|"Changes requested"| B - I -->|"Approved"| J["User-authorized merge"] -``` - -## Phase 1 - Scope the milestone - -Create an implementation plan / proposal record containing: - -- Outcome and user value; -- Acceptance criteria; -- In-scope and out-of-scope behavior; -- Affected components and interfaces; -- Test and measurement plan; -- Security, privacy, operational, and cost risks; -- Rollback or safe-disable strategy. - -Exit gate: The work is scoped small enough for one focused pull request and has -objective, testable acceptance criteria. - -## Phase 2 - Implement - -1. Create a branch from the latest base branch (`develop`): - ```bash - git checkout develop - git pull origin develop - git checkout -b milestone/- - ``` -2. Make the smallest coherent change satisfying the milestone. -3. Add or update tests and documentation alongside code. -4. Inspect the full workspace diff (`git status`, `git diff`, untracked files) - for scope drift, generated files, secrets, and unrelated edits. - -The implementation may remain uncommitted through Gate A. Do not push a branch -or create a PR yet. - -Exit gate: The workspace contains one coherent candidate change and no unrelated work. - -## Phase 3 - Local validation - -Run repository validation commands from the project root: - -```bash -# Execute local quality checks (lint, format check, type check, test suites) -``` - -Record each command and its result. Fix failures and repeat until clean. Do not -classify an expected failure as a pass. - -Exit gate: All applicable local checks pass for the current workspace content. - -## Phase 4 - Review Gate A: workspace implementation review - -Run an independent review session using `.agent/review-prompts/implementation-review.md`. - -The reviewer evaluates: -- Complete workspace diff against `develop`; -- Acceptance criteria coverage; -- Edge cases, error handling, regressions; -- Security, secrets, and licensing; -- Test adequacy and architecture fit. - -Gate decision: -- `PASS`: No blocking correctness, security, data-loss, architecture, test, or - acceptance-criteria findings. -- `FAIL`: At least one blocking finding, missing evidence, incomplete review, or - ambiguous verdict. - -On `FAIL`, resolve every blocking finding, rerun local validation, and repeat Gate A -in a fresh session. - -Exit gate: Gate A returns an explicit `VERDICT: PASS`. - -## Phase 5 - Commit and create draft pull request - -Only after Gate A passes: - -1. Stage only the reviewed milestone files and inspect the staged diff. -2. Commit the reviewed change. -3. Push the branch to origin: - ```bash - git push -u origin milestone/- - ``` -4. Create a **draft** pull request against `develop`. -5. Fill out `.github/PULL_REQUEST_TEMPLATE.md` with: - - Milestone outcome & scope; - - Acceptance criteria checklist; - - Risk assessment; - - Validation evidence; - - Review Gate A verdict and reviewer evidence. -6. Keep the pull request in draft state. - -Exit gate: The draft PR is created against `develop` with complete Gate A evidence. - -## Phase 6 - Required CI - -Wait for all required GitHub Actions checks to finish on the draft PR head commit. - -If CI fails or requires a content change: -1. Fix the issue locally; -2. Rerun local validation (Phase 3); -3. Rerun Review Gate A (Phase 4) for the updated content; -4. Commit, push, and wait for CI. - -Exit gate: All required status checks are green for the exact PR head commit. - -## Phase 7 - Review Gate B: exact draft PR review - -Gate B runs in an independent reviewer session after CI passes, evaluating the draft PR -using `.agent/review-prompts/draft-pr-review.md`. - -The reviewer independently inspects: -- PR title, description, and diff against `develop`; -- Commits and file changes; -- Required CI status and check logs; -- Gate A evidence and resolution of earlier findings; -- Merge readiness and residual risks. - -On `FAIL`, return to Phase 2. Any content change requires rerunning Phases 3 through 7. - -Exit gate: Gate B returns an explicit `VERDICT: PASS` for the exact current PR head SHA. - -## Phase 8 - Ready for user review - -After Gate A, CI, and Gate B all pass: - -1. Add the Gate B verdict and evidence link to the PR description or comment. -2. Mark the draft pull request as ready for review (`gh pr ready `). -3. Notify the user with: - - Branch name & head commit SHA; - - PR URL; - - Checks run & Gate A / B verdicts; - - Known limitations or risks. -4. Stop. **Do not merge.** - -The user performs the final review and explicitly decides whether to merge into `develop`. - -## Fail-closed conditions - -Do not advance a gate when: -- Review output is truncated or lacks an explicit `VERDICT: PASS`; -- Target base branch or PR head SHA is ambiguous; -- Required CI is missing, pending, skipped, or failing; -- Unresolved blocking findings remain. diff --git a/.agent/PROJECT_CONTEXT.md b/.agent/PROJECT_CONTEXT.md new file mode 100644 index 0000000..1189146 --- /dev/null +++ b/.agent/PROJECT_CONTEXT.md @@ -0,0 +1,73 @@ +# OneShot Project Context + +## Product statement + +`One job. Many retries. One settlement.` + +OneShot executes an approved business obligation safely despite retries, +crashes, lost responses, parallel workers, or multiple agent instances. + +The core cardinality is: + +`1 intent / N attempts / <=1 committed settlement` + +This document defines ownership and vocabulary. It is not a product +implementation plan. + +## System ownership + +- OneShot is authoritative for business-intent execution state, attempt state, + settlement state, and permission to create another external settlement. +- Privy is the corporate wallet and scoped authorization boundary. OneShot must + use its policies and spending permissions on the normal execution path. +- Arc is the real USDC settlement rail used by the demo. +- The Graph supplies live indexed recovery/history context and agent decision + support after ambiguous outcomes. It can corroborate or locate activity, but + cannot authorize a duplicate payment. + +When an external submission may have happened but the result is uncertain, +OneShot records `UNKNOWN` and reconciles. Missing Graph data never converts +`UNKNOWN` into permission to submit again. + +## Glossary + +### Business Intent + +The durable identity of one approved business obligation. Its +`business_intent_id` remains stable across retries, process restarts, queue +redelivery, parallel workers, and multiple agents. Payload differences do not +create a second settlement right when the identifier is the same. + +### Attempt + +One execution try for a Business Intent. Attempts are expendable and may fail or +repeat. Any number of Attempts can belong to one Business Intent. + +### Settlement + +The committed external USDC payment for a Business Intent. A Business Intent may +have zero or one committed Settlement, never more than one. + +### Reconciliation + +The process that resolves an ambiguous external effect using durable local +state, provider identifiers and receipts, Arc state, and indexed evidence. +Reconciliation precedes any decision to retry payment when settlement state is +`UNKNOWN`. + +### Recovery View + +A derived, non-authoritative view assembled from durable OneShot records and +live indexed history, including The Graph. It helps operators and agents explain +and recover work but does not grant permission to create a Settlement. + +## Decision test + +Any design affecting retries or payments must answer: + +1. What stable Business Intent does this Attempt belong to? +2. Which durable atomic transition grants the right to submit an external + Settlement? +3. How is an ambiguous submission reconciled without a blind retry? +4. How do parallel workers converge on at most one committed Settlement? +5. Which evidence is authoritative, and which evidence is only a Recovery View? diff --git a/.agent/SECURITY_INVARIANTS.md b/.agent/SECURITY_INVARIANTS.md new file mode 100644 index 0000000..c9c8aa3 --- /dev/null +++ b/.agent/SECURITY_INVARIANTS.md @@ -0,0 +1,55 @@ +# OneShot Security Invariants + +These rules fail closed. A feature, demo, or deadline does not override them. + +## Settlement safety + +- One Business Intent produces at most one committed Settlement. +- Persist a stable `business_intent_id` before any external effect. +- Use durable, atomic, concurrency-safe transitions for settlement ownership. +- A timeout, crash, disconnect, lost response, or provider error after possible + submission creates `UNKNOWN`; reconcile before any payment retry. +- Never infer non-payment from an empty or delayed Graph result. +- Preserve a successful payment result even if a later supplier or API step + fails. +- Do not offer a normal code path that bypasses OneShot state controls or Privy + authorization. + +## Money and authorization + +- Store, compare, calculate, and serialize money as integer atomic units or + `bigint`. Never use JavaScript floating point for monetary values. +- Validate asset, network, recipient, amount, and policy scope before signing or + submitting. +- Privy denial, expired authorization, or an amount above policy produces no + settlement. +- Default to testnet. A non-testnet operation requires explicit user + authorization for that operation. + +## Secrets and privacy + +- Never log, display, persist in context files, commit, or transmit private keys, + seed phrases, access tokens, API secrets, wallet credentials, signing material, + or sensitive runtime configuration. +- Keep secrets in approved runtime secret stores or ignored local environment + files. Commit only safe examples with placeholder values. +- Review staged and untracked files for secret material before every commit. +- FreePi receives only code, intended diffs, tests, public documentation, and + non-sensitive validation evidence. It must not read or receive `.env*`, local + credentials, wallet files, ignored files, or sensitive runtime configuration. +- If a review cannot be completed without sensitive data, the review fails. Do + not send the data. + +## Dependencies and boundaries + +- Prefer official SDKs and primary documentation for Privy, Arc, and The Graph. +- Pin and review dependencies in line with repository conventions. +- Validate all untrusted external data at adapter boundaries. +- Treat provider and indexer output as evidence with explicit freshness and + finality limits, not as implicit authorization. + +## Required response to doubt + +Stop external effects when identity, authorization, amount, network, prior +submission, or settlement state is ambiguous. Persist evidence, enter a safe +state, and reconcile or request human input. diff --git a/.agent/SPONSOR_REQUIREMENTS.md b/.agent/SPONSOR_REQUIREMENTS.md new file mode 100644 index 0000000..6116c2f --- /dev/null +++ b/.agent/SPONSOR_REQUIREMENTS.md @@ -0,0 +1,40 @@ +# Sponsor Requirements + +Use this document before sponsor-facing implementation, demo preparation, +release, or submission claims. + +## Privy + +- Privy must be core corporate wallet authorization, not login-only branding. +- The working path must demonstrate scoped authorization, policies, or spending + permissions that constrain settlement. +- Policy denial or an amount above policy must produce zero settlement. +- The normal agent path must not bypass Privy authorization. + +## Arc + +- The demo must execute a real USDC settlement on the authorized Arc testnet. +- Showing an Arc network label, wallet address, explorer page, or mocked payment + alone does not qualify. +- OneShot must retain the settlement identity and result through retries and + downstream failures. + +## The Graph + +- The integration must use live indexed data for recovery, history, or agent + decision support. +- The demo should show how indexed evidence helps resolve or explain an + ambiguous outcome. +- The Graph must never be the sole duplicate-payment lock, authoritative intent + state, or proof that another settlement may be submitted. +- Empty results and indexing delay must preserve safe behavior. + +## Claim standard + +Do not state or imply sponsor qualification unless the integration exists in +working code and the demo proves the required behavior. Plans, placeholders, +mockups, environment variables, dependency declarations, and network labels are +not implementation evidence. + +Use the `sponsor-qualification` skill to report each sponsor as `QUALIFIED`, +`NOT QUALIFIED`, or `NOT VERIFIED`, with code, test, and demo evidence. diff --git a/.agent/TEST_MATRIX.md b/.agent/TEST_MATRIX.md new file mode 100644 index 0000000..8d60369 --- /dev/null +++ b/.agent/TEST_MATRIX.md @@ -0,0 +1,33 @@ +# OneShot Test Matrix + +Select every applicable case for changes to intents, retries, workers, queues, +payments, settlements, reconciliation, Privy, Arc, or The Graph. Prefer tests at +the public domain boundary plus focused adapter tests. A test must assert durable +state and external settlement count, not only an HTTP response. + +| Case | Fault or concurrency setup | Required result | +| --- | --- | --- | +| Normal job | One valid intent and one worker | Exactly 1 committed settlement | +| Same request twice | Deliver identical request twice | Exactly 1 committed settlement | +| Conflicting payload, same ID | Different request payloads share one `business_intent_id` | At most 1 committed settlement; conflict is explicit | +| Sequential retry storm | Run 10 sequential attempts for one intent | Exactly 1 committed settlement | +| Parallel worker storm | Run 10 workers concurrently for one intent | Exactly 1 committed settlement | +| Crash before submission | Kill process before any external submission | 0 settlements; retry is allowed from durable state | +| Crash after submission | Kill process after possible submission but before local confirmation | Enter `UNKNOWN`; reconcile; no blind retry | +| Lost payment response | Payment succeeds but HTTP response is lost | Exactly 1 committed settlement after reconciliation | +| Graph delay or absence | The Graph temporarily returns nothing or lags | No duplicate settlement; absence is not non-payment proof | +| Privy denial | Policy denies or amount exceeds permission | 0 settlements and explicit authorization failure | +| Service restart | Restart after durable intent creation or in-flight work | Intent and settlement state survive; invariant holds | +| Downstream failure after payment | Supplier/API step fails after settlement | Payment result remains durable; no replacement payment | +| Two agent instances | Same business obligation reaches two agents | Exactly 1 committed settlement | + +## Cross-cutting assertions + +- `business_intent_id` is stable across all attempts. +- Monetary values use integer atomic units or `bigint` end to end. +- State transitions are atomic under real concurrency, not only mocked sequence. +- External submission identifiers and reconciliation evidence survive restart. +- Logs and test fixtures contain no real secrets or wallet material. +- Tests use testnet or isolated fakes; never create unauthorized mainnet effects. + +Record selected cases and results in the session context and pull request. diff --git a/.agent/context/20260906T115948Z-agents-setup.md b/.agent/context/20260906T115948Z-agents-setup.md new file mode 100644 index 0000000..42511ea --- /dev/null +++ b/.agent/context/20260906T115948Z-agents-setup.md @@ -0,0 +1,107 @@ +# Session Context: Agent Infrastructure Setup + +## Date/time + +- UTC: 2026-09-06T11:59:48Z + +## User goal + +Create the `agents-setup` branch from current `develop` and install durable, +OneShot-specific agent policy, skills, context retention, and two independent +FreePi review gates without merging to `develop` or `main`. + +## Original prompt/request + +High-fidelity restatement: work directly in `SuPuHe/OneShot`; preserve the +OneShot invariant `1 business intent -> at most 1 committed settlement`; encode +Privy, Arc, and The Graph ownership; remove superseded review tooling; use two fresh fail-closed +FreePi reviews through actual `npx free-pi-cli` syntax; add project context, +security, sponsor, test, PR, branch, review-prompt, and retention docs; install +the eight selected `mattpocock/skills` plus token-saving Caveman; create three +repo skills; validate, commit, push, and create a draft PR only after Gate A. +Follow-up clarified that the repository is `/home/supuhe/OneShot` in WSL and +Caveman means the token-saving coding-agent skill. + +## Assumptions + +- The user-authorized branch name `agents-setup` is a naming exception only; + all other branch and review rules remain mandatory. +- This task creates agent infrastructure only and no product `plan.md`. +- User-level skills belong in WSL user scope; repo-specific skills belong in + `.agents/skills` per current OpenAI documentation. + +## Plan + +1. Install and verify selected user-level skills. +2. Create `agents-setup` from updated `origin/develop`. +3. Replace stale agent/review policy and add durable context plus repo skills. +4. Validate content and scripts; inspect complete diff. +5. Run fresh FreePi Gate A, fix and repeat until explicit pass. +6. Commit, push, open draft PR to `develop`, wait for required CI, then run fresh + Gate B if CI and available tooling permit. + +## Key decisions + +- Use `.agents/skills` for repo skills because current official Codex docs scan + that path. +- Use `JuliusBrussee/caveman`, not unrelated projects named Caveman, because it + explicitly provides token-saving Codex communication mode. +- Do not install optional `code-review`; FreePi remains the only mandatory Gate + A/B mechanism and generic review could confuse evidence. +- Start each review with bare `npx free-pi-cli`; its 0.2.19 help exposes no review + flags and states every invocation creates a fresh session. + +## Files/components touched + +- Agent policy, context, security, sponsor requirements, test matrix, review + prompts, branch/PR policy, secret ignores, context helper, and repo skills. + +## Commands/checks + +- `git fetch origin develop` - passed; base `6ea00fd0257fb6531184dab7104bacd7a1100b7a`. +- `npx skills add ... --global --agent codex` - installed eight Matt Pocock + skills and Caveman. +- `npx skills list --global --agent codex --json` - all nine discovered. +- `npx free-pi-cli --help` - confirmed interactive command convention. +- `bash -n .agent/context/new-session.sh` - passed. +- Isolated `new-session.sh smoke-test` - created an exact template copy; passed. +- `quick_validate.py` for all three repo skills - passed. +- `git diff --cached --check` - passed. +- Stale review-tool/name search and `plan.md` search - passed; no matches. +- Repository contains policy/docs only, so no application lint, type, test, or + build command exists yet. + +## External-doc findings + +- Official OpenAI Codex docs: repo skills use `.agents/skills//SKILL.md`; + Codex scans from working directory to repository root. +- `skills` CLI 1.5.23 installed selected skills from `mattpocock/skills` HEAD + `3cca18b368ae95cdbdebbff572ccafa662551015` and + `JuliusBrussee/caveman` HEAD + `5184b3d11ac6a1acb7d44b9bfaa31698157cff97` observed at install time. +- `free-pi-cli` 0.2.19 supports `npx free-pi-cli`, `logout`, `--version`, and + `--help`; its README says each run starts a fresh session. + +## Unresolved questions + +- Whether FreePi authentication is already complete and can return a valid Gate + A verdict. +- GitHub CLI is not installed; use an authenticated GitHub web flow after push. + +## Git and PR state + +- Branch: `agents-setup` +- Base: `develop` at `6ea00fd0257fb6531184dab7104bacd7a1100b7a` +- Commit: uncommitted +- PR: not created +- CI: not run + +## Review gates + +- Gate A: NOT RUN +- Gate B: NOT RUN; requires draft PR and green required CI + +## Handoff/next steps + +1. Apply and validate repository changes. +2. Run Gate A in a fresh FreePi process without exposing sensitive files. diff --git a/.agent/context/README.md b/.agent/context/README.md new file mode 100644 index 0000000..4a99e03 --- /dev/null +++ b/.agent/context/README.md @@ -0,0 +1,41 @@ +# Durable Session Context + +Store one Markdown record per meaningful work session so another agent can +continue after a context-window limit, handoff, interruption, or restart. + +## When to update + +- At milestone boundaries or after a material decision. +- Before handoff or end of session. +- Before a deliberate context reset or compaction, when possible. +- After checks, commits, pushes, PR changes, CI results, and Gate A/B results. + +Use `SESSION_TEMPLATE.md`. Keep one active record current rather than creating +many partial notes. Name it `YYYYMMDDTHHMMSSZ-short-topic.md` in UTC. + +To create a record without extra dependencies: + +```bash +.agent/context/new-session.sh short-topic +``` + +The helper copies the template and prints the path. Fill mandatory fields +immediately. Paste the exact original request when safe and practical; otherwise +write a high-fidelity restatement and link the issue or PR. + +## Mandatory fields + +Every record must include date/time, user goal, original prompt/request, +assumptions, plan, key decisions, files/components touched, commands/checks, +external-doc findings, unresolved questions, branch/commit/PR state, Gate A/B +state, and handoff/next steps. + +## Security + +NEVER store secrets or sensitive runtime configuration. Do not include `.env` +contents, private keys, seed phrases, tokens, API secrets, wallet credentials, +authentication responses, or private customer data. Redact sensitive command +output and record only the safe conclusion. + +Context files are operational memory, not authority. Current code, tests, Git +state, provider evidence, and repository policy remain authoritative. diff --git a/.agent/context/SESSION_TEMPLATE.md b/.agent/context/SESSION_TEMPLATE.md new file mode 100644 index 0000000..95d920c --- /dev/null +++ b/.agent/context/SESSION_TEMPLATE.md @@ -0,0 +1,58 @@ +# Session Context: + +## Date/time + +- UTC: + +## User goal + + + +## Original prompt/request + + + +## Assumptions + +- + +## Plan + +1. + +## Key decisions + +- + +## Files/components touched + +- + +## Commands/checks + +- `` - + +## External-doc findings + +- + +## Unresolved questions + +- + +## Git and PR state + +- Branch: +- Base: +- Commit: +- PR: +- CI: + +## Review gates + +- Gate A: +- Gate B: + +## Handoff/next steps + +1. diff --git a/.agent/context/new-session.sh b/.agent/context/new-session.sh new file mode 100755 index 0000000..7880f4f --- /dev/null +++ b/.agent/context/new-session.sh @@ -0,0 +1,20 @@ +#!/usr/bin/env bash +set -euo pipefail + +topic="${1:-session}" +if [[ ! "$topic" =~ ^[a-z0-9][a-z0-9-]*$ ]]; then + printf 'Topic must use lowercase letters, digits, and hyphens.\n' >&2 + exit 2 +fi + +context_dir="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +timestamp="$(date -u +%Y%m%dT%H%M%SZ)" +target="$context_dir/$timestamp-$topic.md" + +if [[ -e "$target" ]]; then + printf 'Context file already exists: %s\n' "$target" >&2 + exit 1 +fi + +cp "$context_dir/SESSION_TEMPLATE.md" "$target" +printf 'Created %s\n' "$target" diff --git a/.agent/review-prompts/draft-pr-review.md b/.agent/review-prompts/draft-pr-review.md deleted file mode 100644 index 806bad8..0000000 --- a/.agent/review-prompts/draft-pr-review.md +++ /dev/null @@ -1,67 +0,0 @@ -# Review Gate B - Draft Pull Request Review - -You are the second independent senior reviewer. Review only; do not edit files, -commit, push, change pull-request state, leave comments, approve, merge, deploy, -or mutate any external system. -Do not attempt to fix a finding yourself. Report findings and return the required verdict. - -This must be a fresh review. Do not rely on memory or a resumed Gate A session. - -## Review target - -- Read the repository `AGENTS.md` and `.agent/AGENTS.md`. -- Verify draft pull request number, base branch (`develop`), head branch, and head SHA. -- Read the PR description, full GitHub PR diff, commits, required status checks, - and Review Gate A evidence. -- Confirm all evidence refers to the exact current PR head SHA. - -## Required analysis - -Independently evaluate: -- Whether the PR delivers the stated milestone without hidden scope or creep; -- Every issue class evaluated in Gate A; -- Whether previous findings from Gate A were fully resolved; -- Whether CI covers changed behavior and all required checks are green; -- Whether documentation, config, and migration paths are complete; -- Whether the PR description provides sufficient detail for human review; -- Whether any commit made after Gate A invalidates its conclusions; -- Whether the PR is safe to mark ready for human review (not whether it should be merged). - -## Verdict standard - -Return `PASS` only if the exact draft head is ready for user review. A stale Gate A -verdict, failing CI, ambiguous evidence, or any blocking finding is `FAIL`. - -Use this exact structure: - -```text -VERDICT: PASS | FAIL -PR: -REVIEWED_HEAD: -REVIEWED_BASE: develop - -BLOCKING_FINDINGS: -- - -- None - -NON_BLOCKING_FINDINGS: -- - -- None - -CI_AND_REVIEW_EVIDENCE: -- - -MILESTONE_READINESS: -- Scope and acceptance criteria - PASS | FAIL -- Tests and required checks - PASS | FAIL -- Security, privacy, and secrets - PASS | FAIL -- Documentation and operations - PASS | FAIL -- Ready for user review - PASS | FAIL - -RESIDUAL_RISKS: -- -- None - -SUMMARY: - -``` diff --git a/.agent/review-prompts/freepi-pr-review.md b/.agent/review-prompts/freepi-pr-review.md new file mode 100644 index 0000000..64bab6b --- /dev/null +++ b/.agent/review-prompts/freepi-pr-review.md @@ -0,0 +1,69 @@ +# FreePi Gate B: Exact Draft PR Review + +You are the second fresh independent senior reviewer. Review only. Do not edit +files, commit, push, change PR state, comment, approve, merge, deploy, or mutate +external state. Do not fix findings. Never reuse or rely on Gate A session +memory; inspect supplied evidence independently. + +## Safety boundary + +Review only repository-tracked code, full PR diff, tests, public docs, and +non-sensitive PR/check evidence. Never read or request ignored files, `.env*`, +private keys, seed phrases, tokens, API secrets, wallet credentials, or +sensitive runtime configuration. If safe evidence is insufficient, fail closed. + +## Target + +- Read `AGENTS.md`, `.agent/AGENTS.md`, applicable `.agent` documents, and PR + acceptance criteria. +- Verify PR URL/number, draft state, base `develop`, head branch, and exact full + head SHA. +- Inspect the complete PR diff and commits against `develop`. +- Inspect required check state and relevant non-sensitive logs for that exact + head SHA. +- Inspect Gate A evidence, but do not treat it as a substitute for this review. + +## Review + +Independently evaluate scope, acceptance coverage, correctness, regressions, +security, privacy, concurrency, retries/idempotency, external effects, tests, +documentation, and readiness for human review. Confirm all required CI is green +and all evidence binds to the exact current head SHA. + +Return `PASS` only when the exact draft head is ready for human review. Missing, +pending, skipped, failing, ambiguous, stale, truncated, or inaccessible evidence +is `FAIL`. + +Use exactly this structure: + +```text +VERDICT: PASS | FAIL +PR: +REVIEWED_HEAD: +REVIEWED_BASE: develop () + +BLOCKING_FINDINGS: +- - +- None + +NON_BLOCKING_FINDINGS: +- - +- None + +CI_AND_REVIEW_EVIDENCE: +- + +READINESS: +- Scope and acceptance criteria - PASS | FAIL +- Tests and required checks - PASS | FAIL +- Security, privacy, and secrets - PASS | FAIL +- Documentation and operations - PASS | FAIL +- Ready for human review - PASS | FAIL + +RESIDUAL_RISKS: +- +- None + +SUMMARY: + +``` diff --git a/.agent/review-prompts/freepi-prepush-review.md b/.agent/review-prompts/freepi-prepush-review.md new file mode 100644 index 0000000..2f3e181 --- /dev/null +++ b/.agent/review-prompts/freepi-prepush-review.md @@ -0,0 +1,67 @@ +# FreePi Gate A: Pre-Push Workspace Review + +You are a fresh independent senior reviewer. Review only. Do not edit files, +commit, push, create a PR, comment, approve, merge, deploy, or mutate external +state. Do not fix findings. + +## Safety boundary + +Review only repository-tracked files, intended diff content, tests, public docs, +and non-sensitive check evidence. Never read or request ignored files, `.env*`, +private keys, seed phrases, tokens, API secrets, wallet credentials, or +sensitive runtime configuration. If required evidence cannot be inspected +safely, fail closed. + +## Target + +- Read `AGENTS.md`, `.agent/AGENTS.md`, applicable `.agent` documents, and task + acceptance criteria. +- Verify base branch is `develop` and report its exact base SHA. +- Inspect complete candidate content against `develop`: committed, staged, + unstaged, and explicitly intended untracked files. +- Confirm no relevant content is omitted and no unrelated content is included. + +## Review + +Evaluate acceptance coverage, correctness, regressions, state transitions, +error handling, security, privacy, secrets, concurrency, retry/idempotency, +external effects, money representation, architecture, dependencies, +documentation, and test adequacy. For settlement-related work, enforce +`.agent/SECURITY_INVARIANTS.md` and `.agent/TEST_MATRIX.md`. + +Return `PASS` only with no blocking finding and sufficient evidence. Incomplete, +ambiguous, stale, or failed inspection is `FAIL`. + +Use exactly this structure: + +```text +VERDICT: PASS | FAIL +REVIEWED_BASE: develop () +REVIEWED_TARGET: + +BLOCKING_FINDINGS: +- - +- None + +NON_BLOCKING_FINDINGS: +- - +- None + +VALIDATION_EVIDENCE: +- + +ACCEPTANCE_CRITERIA: +- - PASS | FAIL | NOT VERIFIED + +SECURITY_AND_INVARIANTS: +- One intent / at most one settlement - PASS | FAIL | NOT APPLICABLE +- UNKNOWN reconciles without blind retry - PASS | FAIL | NOT APPLICABLE +- Secrets and FreePi privacy boundary - PASS | FAIL + +RESIDUAL_RISKS: +- +- None + +SUMMARY: + +``` diff --git a/.agent/review-prompts/implementation-review.md b/.agent/review-prompts/implementation-review.md deleted file mode 100644 index 2a60f9f..0000000 --- a/.agent/review-prompts/implementation-review.md +++ /dev/null @@ -1,59 +0,0 @@ -# Review Gate A - Implementation Review - -You are an independent senior reviewer. Review only; do not edit files, commit, -push, create pull requests, comment on GitHub, or mutate repository state. -Do not attempt to fix a finding yourself. Report findings and return the required verdict. - -## Review target - -- Read the repository `AGENTS.md` and `.agent/AGENTS.md`. -- Target base branch: `develop` (or designated milestone base). -- Review the complete workspace diff against the base, including committed, - staged, unstaged, and untracked files. -- Read the milestone acceptance criteria, requirements, and relevant docs. - -## Required analysis - -Review for: -- Correctness and acceptance-criteria coverage; -- Regressions, edge cases, state transitions, and error handling; -- Security, input validation, secrets, and data safety; -- Concurrency, retry, idempotency, and API boundaries; -- Code clarity, architectural fit, and scope discipline; -- Test adequacy (coverage of new behavior and regressions); -- Configuration, dependencies, and environment assumptions; -- Secrets, credentials, or generated files accidentally tracked. - -## Verdict standard - -Return `PASS` only when there are no blocking findings and evidence is sufficient. -Missing evidence, an ambiguous diff, or an incomplete review is `FAIL`. - -Use this exact structure: - -```text -VERDICT: PASS | FAIL -REVIEWED_TARGET: -REVIEWED_BASE: develop () - -BLOCKING_FINDINGS: -- - -- None - -NON_BLOCKING_FINDINGS: -- - -- None - -VALIDATION_EVIDENCE: -- - -ACCEPTANCE_CRITERIA: -- - PASS | FAIL | NOT VERIFIED - -RESIDUAL_RISKS: -- -- None - -SUMMARY: - -``` diff --git a/.agents/skills/oneshot-failure-injection/SKILL.md b/.agents/skills/oneshot-failure-injection/SKILL.md new file mode 100644 index 0000000..6ef3fc9 --- /dev/null +++ b/.agents/skills/oneshot-failure-injection/SKILL.md @@ -0,0 +1,28 @@ +--- +name: oneshot-failure-injection +description: Design or test OneShot failure boundaries around external effects, including timeouts, lost responses, process kills, duplicate delivery, retries, and parallel execution before, during, or after payment submission. +--- + +# OneShot Failure Injection + +Read `.agent/SECURITY_INVARIANTS.md` and `.agent/TEST_MATRIX.md`. Map each +external effect into three boundaries: definitely before submission, possibly +submitted, and definitely confirmed. + +## Mandatory checks + +- Inject failure before submission and prove zero settlement plus safe retry. +- Inject timeout/process kill/lost response during or after submission and prove + durable `UNKNOWN`, reconciliation, and no blind retry. +- Deliver the same request repeatedly and from 10 parallel workers; prove at + most one committed settlement. +- Restart services between durable transitions and external responses. +- Delay or empty The Graph results; prove no duplicate settlement. +- Deny Privy policy and exceed spending amount; prove zero settlement. +- Fail a supplier/API action after payment; prove settlement result remains. +- Assert external settlement count, durable intent/attempt/settlement state, and + stable identifiers. Do not rely only on returned HTTP status. + +Use testnet or isolated fakes. Never inject failures against unauthorized live +funds or expose secrets in fixtures/logs. Record exact injection points and +results in session context and PR evidence. diff --git a/.agents/skills/oneshot-idempotency/SKILL.md b/.agents/skills/oneshot-idempotency/SKILL.md new file mode 100644 index 0000000..38312b9 --- /dev/null +++ b/.agents/skills/oneshot-idempotency/SKILL.md @@ -0,0 +1,31 @@ +--- +name: oneshot-idempotency +description: Enforce OneShot's at-most-once settlement invariant when work touches business intents, payments, retries, settlements, reconciliation, workers, queues, jobs, invoices, or duplicate delivery. +--- + +# OneShot Idempotency + +Read `.agent/PROJECT_CONTEXT.md`, `.agent/SECURITY_INVARIANTS.md`, and +`.agent/TEST_MATRIX.md` before editing. + +## Mandatory checks + +- Preserve `1 intent / N attempts / <=1 committed settlement`. +- Create and persist one stable `business_intent_id` before external effects; + reuse it across retries, restarts, workers, and agents. +- Keep authoritative intent and settlement state in OneShot durable storage. +- Grant submission rights through an atomic, concurrency-safe transition or + equivalent uniqueness guarantee. +- Use a stable provider idempotency/submission key tied to the Business Intent + where the provider supports it. This supplements, not replaces, OneShot state. +- Treat any possibly submitted but unconfirmed payment as `UNKNOWN`. Reconcile + from durable/provider/Arc evidence before retrying. +- Never use Graph absence or indexing delay as permission to pay. +- Keep money in integer atomic units or `bigint`; validate asset, network, + recipient, amount, and Privy policy before submission. +- Preserve payment results when later supplier/API work fails. + +Select applicable matrix cases, including duplicate input, 10 sequential +retries, 10 parallel workers, process restart, two agents, and ambiguous +submission. Report any invariant that cannot be proven; do not claim safety from +happy-path tests alone. diff --git a/.agents/skills/sponsor-qualification/SKILL.md b/.agents/skills/sponsor-qualification/SKILL.md new file mode 100644 index 0000000..6624883 --- /dev/null +++ b/.agents/skills/sponsor-qualification/SKILL.md @@ -0,0 +1,25 @@ +--- +name: sponsor-qualification +description: Validate Privy, Arc, and The Graph integration evidence before OneShot demos, releases, submissions, sponsor checklists, or qualification claims. +--- + +# Sponsor Qualification + +Read `.agent/SPONSOR_REQUIREMENTS.md`, `.agent/PROJECT_CONTEXT.md`, and relevant +code/tests/demo instructions. Review actual working evidence, not plans. + +## Mandatory checks + +- Privy: prove corporate wallet authorization constrains the normal settlement + path through scoped policy or spending permission. Login-only is insufficient. +- Arc: prove the demo performs a real USDC settlement on the authorized testnet. + A network label, address, explorer link, or mock alone is insufficient. +- The Graph: prove live indexed data supports recovery/history/agent decisions, + while OneShot durable state remains authoritative and empty/indexing-delayed + results cannot unlock another settlement. +- Verify the demo preserves `1 intent / N attempts / <=1 settlement` and never + exposes secrets. + +For each sponsor, report `QUALIFIED`, `NOT QUALIFIED`, or `NOT VERIFIED`, citing +code, tests, live demo evidence, network, and known limitations. Never upgrade +missing or mocked evidence into a qualification claim. diff --git a/.antigravity/README.md b/.antigravity/README.md deleted file mode 100644 index 0e128c3..0000000 --- a/.antigravity/README.md +++ /dev/null @@ -1,17 +0,0 @@ -# Antigravity CLI Review Configuration - -This directory contains configuration, prompts, and documentation for personal code review workflows using Antigravity CLI (gy). - -## Purpose - -- Allows developers to run independent Review Gate A and Review Gate B evaluations using Antigravity CLI without conflicting with other team members' local review tooling. -- Keeps personal Antigravity review logs and configurations decoupled from core repository policies. - -## Review Gates - -- **Gate A (Pre-PR Workspace Review)**: - Run in Antigravity CLI using .agent/review-prompts/implementation-review.md. - Evaluates workspace diff against develop before draft PR creation. -- **Gate B (Post-PR Draft Review)**: - Run in Antigravity CLI using .agent/review-prompts/draft-pr-review.md. - Evaluates draft PR head commit, status checks, and diff against develop. diff --git a/.antigravity/review.md b/.antigravity/review.md deleted file mode 100644 index 5e492e6..0000000 --- a/.antigravity/review.md +++ /dev/null @@ -1,14 +0,0 @@ -# Antigravity CLI Review Guide - -## Workflow - -1. Open Antigravity CLI (gy) in the repository workspace. -2. For **Gate A**: - - Provide the prompt from .agent/review-prompts/implementation-review.md. - - Specify target base branch: develop. - - Verify verdict (VERDICT: PASS). -3. For **Gate B**: - - Provide the prompt from .agent/review-prompts/draft-pr-review.md. - - Supply PR number, head SHA, and CI check status. - - Verify verdict (VERDICT: PASS). -4. Attach review summary and verdict to the pull request. diff --git a/.github/BRANCH_POLICY.md b/.github/BRANCH_POLICY.md index a31497e..d6c4cdd 100644 --- a/.github/BRANCH_POLICY.md +++ b/.github/BRANCH_POLICY.md @@ -2,33 +2,39 @@ ## Branch architecture -- `main`: Production/release branch. Contains only stable, released code. Merges to `main` occur from `develop` through release PRs or tags. -- `develop`: Integration branch. The default base branch for ongoing development, milestones, and features. -- `milestone/-`, `feature/`, `fix/`: Short-lived branches targeting `develop`. +- `main`: stable release branch. Promote reviewed releases from `develop`. +- `develop`: integration branch and base for ongoing work. +- `feature/*`, `fix/*`, `milestone/*`: short-lived branches created from current + `develop` and targeting `develop`. + +No direct pushes, force pushes, history rewrites, or branch deletion on `main` +or `develop`. Agents never merge any PR. A human performs final review and +explicitly authorizes merge. ## Pull request workflow -1. All milestone and feature branches originate from `develop` and create pull requests targeting `develop`. -2. Every pull request begins as a **Draft** PR. -3. Every pull request requires passing: - - Local validation checks (Phase 3); - - Review Gate A (independent pre-PR implementation review); - - All required CI status checks; - - Review Gate B (independent post-PR draft review); - - Human review and approval. -4. Agents must never merge pull requests. Final approval and merging is performed exclusively by the user. +1. Implement on a short-lived branch from `develop`. +2. Pass applicable local lint, type, test, build, and failure-injection checks. +3. Inspect the complete change against `develop` and confirm no secrets or + unrelated files. +4. Run fresh independent FreePi Gate A through `npx free-pi-cli`. Require exact + `VERDICT: PASS` before first push or draft PR creation. +5. Push without force and open a draft PR targeting `develop`. +6. Wait for every required CI check to be green on the exact PR head SHA. +7. Run a second fresh independent FreePi Gate B through `npx free-pi-cli`, bound + to the exact PR head SHA, full PR diff, and check state. +8. After explicit Gate B `VERDICT: PASS`, mark ready for human review. Stop + before merge. + +Gate A and Gate B must use separate new FreePi processes and contexts. Any +relevant content change after Gate A invalidates Gate A. Any commit or content +change after Gate B invalidates Gate B. Ambiguous, incomplete, stale, failed, or +unavailable review output fails closed. ## Required status checks -As CI workflows are established in `.github/workflows/`, branch protection rules for `develop` and `main` must enforce: -- Linting and static analysis; -- Automated test suites; -- Build / compilation checks. - -## Protection rules +As workflows are added, branch protection for `develop` and `main` must require +applicable lint/static analysis, automated tests, and build/compilation checks. +Pending, skipped, missing, or failing required checks are not green. -The `develop` and `main` branches should be protected against: -- Direct pushes (all changes must pass through pull requests); -- Force pushes; -- Branch deletions; -- Merging with unresolved conversations or failing checks. +See `.agent/IMPLEMENTATION_LOOP.md` for the complete workflow and privacy rules. diff --git a/.github/PULL_REQUEST_TEMPLATE.md b/.github/PULL_REQUEST_TEMPLATE.md index 4e6f23e..772a1d0 100644 --- a/.github/PULL_REQUEST_TEMPLATE.md +++ b/.github/PULL_REQUEST_TEMPLATE.md @@ -1,6 +1,6 @@ -## Milestone outcome +## Outcome - + ## Scope @@ -16,47 +16,65 @@ - [ ] +## OneShot invariant impact + +- Business Intent / Attempt / Settlement impact: +- `business_intent_id` stability: +- `UNKNOWN` reconciliation behavior: +- Concurrency and duplicate-settlement protection: +- Money representation: + +## Sponsor impact + +- Privy: +- Arc: +- The Graph: +- Qualification claims made (if any): + ## Risk review -- Security & secrets: -- Data safety & privacy: -- Architecture & performance: -- Rollback or safe disablement: +- Security and secrets: +- Data safety and privacy: +- External effects and rollback/safe disablement: +- Architecture and performance: ## Validation evidence | Command or check | Result | | --- | --- | -| `` | | -| `` | | -| `` | | +| `` | | +| `` | | +| `git diff --check develop...HEAD` | | -## Review Gate A - implementation +## FreePi Gate A: pre-push -- Reviewed head / commit: -- Reviewer / model: -- Verdict: +- Fresh `npx free-pi-cli` process/session: +- Reviewed base SHA: +- Reviewed target/content identity: +- Verdict (must be exact `VERDICT: PASS`): - Blocking findings resolved: -- Evidence / review summary: +- Evidence/summary: ## Required CI -- [ ] Lint and static analysis -- [ ] Automated tests -- [ ] Build validation +- Exact PR head SHA: +- [ ] All required checks are green for this SHA. +- Check names/results: -## Review Gate B - draft PR +## FreePi Gate B: exact draft PR -- Reviewed PR head SHA: -- Reviewer / model: -- Verdict: +- Separate fresh `npx free-pi-cli` process/session: +- PR URL/number: +- Reviewed head SHA: +- Verdict (must be exact `VERDICT: PASS`): - Blocking findings resolved: -- Evidence / review summary: +- Evidence/summary: -## User review +## Human review -- [ ] Review Gate A passed for the current change. -- [ ] Required CI checks are green for the current head. -- [ ] Review Gate B passed for the current head. -- [ ] The PR is marked ready for user review. -- [ ] The user explicitly approved merge. Agents must leave this unchecked. +- [ ] Gate A is valid for current content. +- [ ] Required CI is green for current head. +- [ ] Gate B is valid for current head. +- [ ] Draft is ready for human review. +- [ ] Human explicitly authorized merge. Agents must leave this unchecked and + must never merge. diff --git a/.gitignore b/.gitignore index 07035c0..84de569 100644 --- a/.gitignore +++ b/.gitignore @@ -6,9 +6,24 @@ Thumbs.db .env .env.* !.env.example +.envrc +.direnv/ *.local *.key *.pem +*.p12 +*.pfx +*.jks +*.keystore +*.seed +*.mnemonic +*.token +.npmrc +!.npmrc.example +secrets/ +credentials/ +.free-pi/ +.pi/ # Editor and IDE .vscode/ @@ -27,6 +42,8 @@ dist/ build/ coverage/ node_modules/ +.venv/ +venv/ __pycache__/ .pytest_cache/ .ruff_cache/ diff --git a/AGENTS.md b/AGENTS.md index 9630d3c..41a1643 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -1,14 +1,25 @@ -# Repository Agent Instructions +# OneShot Agent Entry Point -These instructions apply to the entire repository. +Read `.agent/AGENTS.md` before any repository work. -Before planning, editing, reviewing, or publishing any change, read -`.agent/AGENTS.md` and `.agent/MILESTONE_IMPLEMENTATION_LOOP.md` completely and -follow them as mandatory repository policy. +Then load only the documents needed for the task: -CLI agents and assistant tools automatically load this root file. Detailed policies -and milestone loops live under `.agent/` so repository rules remain maintainable without -unnecessarily inflating the top-level prompt. +- Architecture or domain work: `.agent/PROJECT_CONTEXT.md` and + `.agent/SECURITY_INVARIANTS.md`. +- Privy, Arc, The Graph, demo, release, or submission work: + `.agent/SPONSOR_REQUIREMENTS.md`. +- Intent, payment, retry, worker, queue, job, invoice, settlement, or + reconciliation work: the `oneshot-idempotency` repo skill and + `.agent/TEST_MATRIX.md`. +- Failure handling or reliability work: the `oneshot-failure-injection` repo + skill and `.agent/TEST_MATRIX.md`. +- Any implementation, review, commit, push, or pull request: + `.agent/IMPLEMENTATION_LOOP.md`. +- Demo/release sponsor claims: the `sponsor-qualification` repo skill. +- Handoff, milestone boundary, session end, or deliberate context reset: + `.agent/context/README.md` and the current context record. -If either detailed policy file is missing or cannot be read, stop and report the -problem. Do not guess at the review or publishing process. +Repository skills live in `.agents/skills/`. Personal workflow skills may help, +but they never replace OneShot policy or mandatory FreePi Gate A and Gate B. + +If a required document cannot be read, stop and report the missing policy. From b643206f11e40536eb1724d080f541bfd4edf85c Mon Sep 17 00:00:00 2001 From: SuPuHe Date: Sun, 6 Sep 2026 16:26:03 +0200 Subject: [PATCH 2/2] chore: create 1 version of plan.md for the project --- .../context/20260906T134707Z-product-plan.md | 88 +++ .../20260906-integration-decisions.md | 49 ++ plan.md | 521 ++++++++++++++++++ 3 files changed, 658 insertions(+) create mode 100755 .agent/context/20260906T134707Z-product-plan.md create mode 100755 .agent/research/20260906-integration-decisions.md create mode 100755 plan.md diff --git a/.agent/context/20260906T134707Z-product-plan.md b/.agent/context/20260906T134707Z-product-plan.md new file mode 100755 index 0000000..89babea --- /dev/null +++ b/.agent/context/20260906T134707Z-product-plan.md @@ -0,0 +1,88 @@ +# Session Context: Product Implementation Plan + +## Date/time + +- UTC: 2026-09-06T13:47:07Z + +## User goal + +Create a detailed `plan.md` with milestones for exactly three people, minimize cross-person blocking, defer frontend work until the end, run exactly one FreePi review after the plan is complete, and do not create a pull request. + +## Original prompt/request + +The user provided the `SuPuHe/OneShot` `agents-setup` branch URL, noted that the repository may also be found in WSL, requested use of all relevant planning skills, required a single FreePi review after plan creation, prohibited creating a PR, and requested English-only chat responses. + +## Assumptions + +- The application is greenfield because the current repository contains agent/repository policy but no product code. +- The first deliverable is a testnet backend plus minimal late-stage frontend, not a production mainnet system. +- Three implementers work in independent packages/worktrees and converge only after their isolated contract suites pass. +- The current task may add required research and durable context alongside `plan.md`; it does not commit, push, open a PR, or merge. + +## Plan + +1. Load repository policy and applicable repository/user skills. +2. Research current official Privy, Arc, The Graph, PostgreSQL, and worker behavior. +3. Write the cited integration decisions and a detailed three-person milestone plan. +4. Validate scope, links, Markdown, dependencies, invariants, and secrets. +5. Start exactly one fresh `npx free-pi-cli` process for a pre-push-style independent review. +6. Record/report the review result without creating a PR. + +## Key decisions + +- Use three parallel tracks: domain/storage/API; Privy/Arc settlement; reconciliation/The Graph/reliability. +- Freeze port results, state transitions, JSON fixtures, and test seams before implementation so M1–M3 have no cross-person blockers. +- Use PostgreSQL plus Graphile Worker; the queue is at-least-once, while database transitions remain authoritative. +- Use Arc Testnet's six-decimal ERC-20 USDC interface for payments and keep native USDC gas precision separate. +- Treat The Graph as freshness-labeled recovery/history evidence only. +- Begin frontend only after the integrated backend and failure matrix pass. + +## Files/components touched + +- `plan.md` — detailed product architecture, contracts, milestones, ownership, dependencies, tests, risks, and release gates. +- `.agent/research/20260906-integration-decisions.md` — primary-source integration research. +- `.agent/context/20260906T134707Z-product-plan.md` — this durable context record. + +## Commands/checks + +- `git branch --show-current` — `agents-setup`. +- `git status --short` before changes — clean. +- `git rev-parse develop` — `6ea00fd0257fb6531184dab7104bacd7a1100b7a`. +- Repository policy, security, sponsor, test-matrix, implementation-loop, context, and applicable skill documents — read completely. +- Primary-source research — completed; citations saved in the research note. +- Markdown structure, trailing-whitespace, placeholder, obvious-secret-pattern, ownership, dependency, and scope checks — passed. +- One fresh FreePi process/session reviewed `agents-setup` plus the three intended untracked files against `develop` — `VERDICT: PASS`; no blocking findings. +- FreePi non-blocking cautions: keep the cited research note with `plan.md`; external links were not live-revalidated by the reviewer; implementation evidence is intentionally unavailable at planning stage. +- No product lint/type/test/build command exists because the repository still contains no application code. + +## External-doc findings + +- Privy wallet policies constrain authorized wallet actions; request idempotency lasts 24 hours and is only a supplemental duplicate guard. +- Arc Testnet is `eip155:5042002`; ERC-20 USDC is `0x3600000000000000000000000000000000000000` at six decimals, while native USDC gas accounting uses a different precision. +- Arc receipt inclusion is deterministically final, but lost submission responses still require durable UNKNOWN reconciliation. +- The Graph supports `arc-testnet`; `_meta`, indexed block, deployment, health, and lag must accompany recovery evidence. +- PostgreSQL conditional transitions/constraints plus transactional Graphile work delivery fit the at-most-once settlement design. + +## Unresolved questions + +- Exact Node/TypeScript/SDK versions will be pinned after the M1 compatibility tests. +- Privy production webhook availability and Arc-specific rolling-spend policy behavior are optional and must be verified before enablement. + +## Git and PR state + +- Branch: `agents-setup` +- Base: `develop` at `6ea00fd0257fb6531184dab7104bacd7a1100b7a` +- Commit: current HEAD `a304274412f080817e6da2967e444b31ccd2f864`; new planning files uncommitted +- PR: not created by explicit user instruction +- CI: not run; no product code or CI exists yet + +## Review gates + +- Gate A: PASS for reviewed base `develop` (`6ea00fd0257fb6531184dab7104bacd7a1100b7a`) and target `agents-setup` (HEAD `a304274412f080817e6da2967e444b31ccd2f864`) plus `plan.md`, `.agent/research/20260906-integration-decisions.md`, and this context file as it existed before this review-evidence update. Blocking findings: none. This context-only bookkeeping update means the gate must be treated as invalid for any future push; the user limited this task to one review and prohibited a PR, so no rerun is permitted or needed here. +- Gate B: NOT RUN and not applicable because the user prohibited creating a PR. + +## Handoff/next steps + +1. Present `plan.md`, the research note, local validation, and the single FreePi PASS to the user. +2. Do not edit `plan.md`, rerun FreePi, commit, push, or create a PR in this task. +3. Before any future implementation, obtain human plan approval and merge the prerequisite planning/agent-infrastructure work to `develop` under normal repository policy. diff --git a/.agent/research/20260906-integration-decisions.md b/.agent/research/20260906-integration-decisions.md new file mode 100755 index 0000000..f9554dc --- /dev/null +++ b/.agent/research/20260906-integration-decisions.md @@ -0,0 +1,49 @@ +# OneShot Integration Research + +Date: 2026-09-06 +Scope: primary-source facts needed to make the first product implementation plan decision-complete. +Target: Privy-authorized USDC settlement on Arc Testnet with The Graph as a non-authoritative recovery view. + +## Decisions + +### Privy: authorization must constrain the settlement path + +- Use a Privy execution wallet owned by an application authorization key or key quorum. Attach one explicit, fail-closed wallet policy when the wallet is created. Privy owners authorize wallet actions, while wallet policies constrain the actions that an otherwise valid signer may take ([wallet policies and controls](https://docs.privy.io/security/wallet-infrastructure/policy-and-controls), [execution wallets](https://docs.privy.io/recipes/wallets/execution-wallets)). +- Permit only the Arc Testnet ERC-20 USDC `transfer(address,uint256)` path: chain `5042002`, contract `0x3600000000000000000000000000000000000000`, approved recipient, amount at or below the configured cap, and zero native transaction value. Keep key export and all unrelated methods denied. Privy documents default-deny policy behavior and Ethereum transaction conditions ([policy overview](https://docs.privy.io/controls/policies/overview), [Ethereum policy examples](https://docs.privy.io/controls/policies/example-policies/ethereum)). +- Persist the exact Privy request identity and body before submission. Reuse the same `privy-idempotency-key` for the same Business Intent. Privy deduplicates a matching request for only 24 hours, so this is a supplemental guard and never replaces OneShot's durable state and uniqueness constraints ([idempotency keys](https://docs.privy.io/api-reference/idempotency-keys)). +- Attach a stable Privy transaction `reference_id` derived from `business_intent_id` for lookup and reconciliation, not as the authoritative duplicate lock ([transaction reference IDs](https://docs.privy.io/transaction-management/transactions/reference-id)). +- Polling transaction status is the baseline. Webhooks are an optional optimization because availability may depend on the Privy plan; if enabled, verify signatures and process deliveries idempotently ([webhook overview](https://docs.privy.io/api-reference/webhooks/overview)). + +### Arc: use the six-decimal ERC-20 interface for settlement + +- Arc Testnet uses chain ID `5042002`, CAIP-2 `eip155:5042002`, RPC `https://rpc.testnet.arc.network`, WebSocket `wss://rpc.testnet.arc.network`, and explorer `https://testnet.arcscan.app` ([RPC endpoints](https://docs.arc.io/arc/references/rpc-endpoints)). +- Arc's USDC ERC-20 interface is `0x3600000000000000000000000000000000000000`. Application settlement amounts use its six-decimal precision. Arc also exposes the same underlying USDC as an 18-decimal native gas balance, so payment amounts and gas accounting must remain separate and the UI must not double-count the two views ([infrastructure integration](https://docs.arc.io/integrate/infrastructure), [stablecoin-native model](https://docs.arc.io/arc/concepts/stablecoin-native-model)). +- Arc transactions are pending until included, then immediately and deterministically final; there is no accumulating-confirmation state. A receipt with `status: 1` is final success only after validating the expected USDC `Transfer` log. A receipt with `status: 0` is final execution failure and zero settlement ([transaction lifecycle](https://docs.arc.io/integrate/wallets/transaction-lifecycle), [deterministic finality](https://docs.arc.io/arc/concepts/deterministic-finality)). +- Deterministic finality does not eliminate submission ambiguity. A lost Privy/RPC response or process crash after a possible broadcast still becomes `UNKNOWN`; a new-nonce payment is forbidden until reconciliation proves a safe terminal result. + +### The Graph: live recovery evidence, never settlement authority + +- The Graph lists Arc Testnet as `arc-testnet`, protocol Ethereum, CAIP-2 `eip155:5042002` ([Arc Testnet support](https://thegraph.com/docs/en/supported-networks/arc-testnet/)). +- Build a custom Subgraph over the unified USDC `Transfer` event. Identify each event by transaction hash plus log index and retain block number, block timestamp, sender, recipient, and amount as Graph `BigInt` ([Arc event indexing](https://docs.arc.io/integrate/infrastructure/indexing-events), [Subgraph quick start](https://thegraph.com/docs/en/subgraphs/quick-start/)). +- Every query must request `_meta` block data, deployment ID, and `hasIndexingErrors`. Operational health must also compare indexed `latestBlock` with `chainHeadBlock`; `synced` only means the deployment caught up at least once ([GraphQL API](https://thegraph.com/docs/en/subgraphs/querying/graphql-api/), [indexing health](https://thegraph.com/docs/en/subgraphs/developing/deploying-publishing/multiple-networks/)). +- Missing, empty, lagging, or unhealthy indexed results mean only that no matching event was observed through a known indexed block. They never prove non-payment or authorize another submission. Direct Arc receipts plus durable OneShot state remain authoritative. +- Use a deployment-pinned query endpoint for schema-stable demo evidence. Do not claim sponsor qualification until a deployed endpoint returns live Arc data and lag/error behavior is demonstrated ([Subgraph ID versus deployment ID](https://thegraph.com/docs/en/subgraphs/querying/subgraph-id-vs-deployment-id/)). + +### Durable state and work delivery + +- PostgreSQL is the authoritative store. Use primary/unique constraints on `business_intent_id` and one settlement row per intent; use `INSERT ... ON CONFLICT` plus an immutable payload fingerprint to distinguish a replay from a same-ID conflict ([constraints](https://www.postgresql.org/docs/current/ddl-constraints.html), [`INSERT`](https://www.postgresql.org/docs/current/sql-insert.html)). +- Grant submission ownership with a row lock or conditional state transition. Do not keep a database transaction open during Privy or RPC calls. Persist `SUBMITTING`, the request fingerprint, and provider identifiers before crossing the external-effect boundary ([explicit locking](https://www.postgresql.org/docs/current/explicit-locking.html)). +- Use Graphile Worker over the same PostgreSQL database to avoid a second queue datastore. It supports transactional enqueueing and explicitly provides at-least-once delivery ([Graphile Worker](https://worker.graphile.org/docs), [transactional enqueueing](https://worker.graphile.org/docs/sql-add-job)). +- Configure the external-effect `submit_settlement` job for one queue attempt. The task itself classifies the outcome and returns after durably recording `COMMITTED`, `FAILED_SAFE`, or `UNKNOWN`; the queue must never blindly repeat a possibly submitted payment. Read-only reconciliation jobs may retry. `jobKey` is scheduling hygiene, not the settlement lock ([job options](https://worker.graphile.org/docs/library/add-job), [job-key caveats](https://worker.graphile.org/docs/job-key)). + +## Verification gates left for implementation + +1. Pin exact SDK and runtime versions only after a compatibility spike validates Privy request signing, Arc chain support, and policy condition syntax. +2. Assert `eth_chainId == 5042002` and bytecode exists at the configured USDC address during testnet startup checks. +3. Prove the chosen Privy policy denies wrong chain, wrong contract, wrong recipient, wrong method, non-zero native value, and above-cap amount with zero settlement. +4. Prove The Graph deployment health and lag thresholds against live Arc Testnet before sponsor qualification. +5. Keep Privy webhooks outside the critical path until plan availability and signature verification are demonstrated. + +## Planning consequence + +The work can be split into three independent backend tracks after one contract freeze: (A) domain/storage/API, (B) Privy/Arc settlement, and (C) indexing/reconciliation. Each track must ship its own contract simulator and tests so progress does not depend on another track's implementation. Frontend begins only after the integrated backend contract and recovery semantics are stable. diff --git a/plan.md b/plan.md new file mode 100755 index 0000000..d6b0dca --- /dev/null +++ b/plan.md @@ -0,0 +1,521 @@ +# OneShot Product Implementation Plan + +Status: implementation-ready proposal +Team: exactly three engineers +Planning horizon: 20 working days, recalibrated after Milestone 1 +Base for implementation: current `develop` after the agent-infrastructure work is human-reviewed and merged +Research basis: `.agent/research/20260906-integration-decisions.md` + +## 1. Outcome + +Deliver a testnet application that accepts one approved Business Intent, safely survives retries, crashes, duplicate delivery, parallel workers, and ambiguous provider responses, and produces at most one committed USDC settlement on Arc Testnet through a Privy-controlled corporate wallet. The Graph supplies live indexed history and recovery evidence without becoming an authorization source. + +The release claim is: + +`1 Business Intent / N Attempts / <= 1 committed Settlement` + +The milestone plan is deliberately backend-first. No production frontend work begins until the backend integration and failure suite pass in Milestone 4. + +## 2. Success criteria + +- A caller creates a Business Intent with a stable `business_intent_id`; identical replays return the same durable result and conflicting payloads under the same ID fail explicitly. +- Privy authorization and wallet policy constrain every normal settlement path. Wrong network, asset, recipient, method, or above-cap amount results in zero settlement. +- A valid intent produces a real ERC-20 USDC transfer on Arc Testnet and stores the final receipt and transfer identity. +- A timeout, lost response, or crash after possible submission produces durable `UNKNOWN`; no new settlement submission is allowed until reconciliation resolves it. +- Ten sequential retries, ten parallel workers, a restart, and two agent instances cannot produce more than one committed settlement. +- Live The Graph data explains settlement history and supports recovery. Empty, delayed, or unhealthy indexed data never unlocks another payment. +- Money remains an integer string/`bigint` in six-decimal ERC-20 USDC atomic units from API through policy evaluation, storage, calldata, indexing, and UI. +- The final demo proves Privy, Arc, and The Graph requirements with testnet evidence and exposes no secrets. + +## 3. Scope + +### In scope + +- TypeScript backend, worker, shared contracts, PostgreSQL state, and migrations. +- Privy execution-wallet authorization and a fail-closed wallet policy. +- Arc Testnet ERC-20 USDC submission, receipt verification, and explorer evidence. +- Durable reconciliation using OneShot state, Privy identifiers/status, Arc RPC receipts, and The Graph evidence. +- A custom Subgraph plus freshness and indexing-health classification. +- Failure injection, concurrency tests, service restart tests, audit-safe structured logs, metrics, and a demo runbook. +- A minimal operator/user frontend only after backend acceptance. + +### Explicit non-goals + +- Mainnet, multi-chain, multi-asset, swaps, bridging, fiat on/off ramps, or custody beyond the configured Privy testnet wallet. +- Treating The Graph as authoritative settlement state or as permission to retry. +- Automatic same-nonce transaction replacement in the first release. +- General workflow automation, arbitrary supplier integrations, accounting/ERP integrations, or production compliance certification. +- Production-scale multi-region deployment, high availability, or a native mobile client. + +## 4. Fixed technical decisions + +These decisions are frozen for the first implementation. Changing one requires a short ADR, updated contract fixtures, and approval from all affected owners. + +| Area | Decision | Reason | +| --- | --- | --- | +| Runtime | Node.js LTS + strict TypeScript; exact versions pinned in Milestone 1 | One language across API, worker, Privy, Arc, Subgraph tooling, and frontend | +| Repository | `pnpm` workspace with independently testable packages | Each owner can build and test without waiting for root integration | +| API | HTTP JSON described by OpenAPI; generated schemas are checked for drift | Stable seam for simulators and the late frontend | +| Durable state | PostgreSQL | Atomic conditional transitions, constraints, transactional enqueueing | +| Work delivery | Graphile Worker in the same PostgreSQL database | At-least-once work without adding Redis; transactional enqueueing | +| Money | Decimal-free integer strings at boundaries and `bigint` internally | Prevents floating-point loss and JSON `bigint` ambiguity | +| Settlement asset | Arc Testnet ERC-20 USDC at `0x3600000000000000000000000000000000000000`, six decimals | One canonical payment representation; native USDC is gas accounting only | +| Privy | Execution wallet with authorization owner/key quorum and one fail-closed policy | Privy remains a real authorization boundary, not branding | +| Chain access | Arc RPC with startup checks for chain ID `5042002` and USDC bytecode | Fails closed on misconfiguration | +| Indexed view | Custom Subgraph on `arc-testnet`, queried with `_meta` and explicit freshness | Live recovery/history with visible limitations | +| External-effect queueing | `submit_settlement` gets one queue attempt; reconciliation reads may retry | Prevents the queue from blindly repeating an ambiguous payment | + +## 5. Architecture and ownership boundaries + +```text +Caller / late frontend + | + v +HTTP API ---------> PostgreSQL authoritative ledger <------ Worker claims + | | | + | +---- durable outbox/jobs ---+ + | + +--> AuthorizationPort --> Privy policy + wallet + +--> SettlementPort ----> Arc ERC-20 USDC + +--> EvidencePort ------> Privy status + Arc RPC + +--> IndexViewPort -----> The Graph (non-authoritative) +``` + +### Person A — Domain, storage, API, and work delivery + +Owns `packages/contracts`, `packages/domain`, `packages/storage-postgres`, `apps/api`, `apps/worker`, migrations, OpenAPI, and the domain adapter simulator. Person A does not implement Privy, Arc, or The Graph clients. + +### Person B — Privy authorization and Arc settlement + +Owns `packages/privy-adapter`, `packages/arc-adapter`, policy fixtures, Arc chain configuration, transaction construction, receipt verification, provider error classification, and the settlement-adapter simulator. Person B does not change domain states or database tables directly. + +### Person C — Reconciliation, The Graph, and reliability evidence + +Owns `packages/reconciliation`, `packages/graph-client`, `subgraph`, recovery-view contracts, freshness/health classification, and the failure-injection harness. Person C may propose state transitions only through the frozen reconciliation command port. + +### Shared files and conflict rule + +- Only Person A edits root workspace/build configuration after Milestone 1. +- Every package must have a package-local test command so Persons B and C can run independently before root composition exists. +- Contract changes are additive during a milestone. Breaking changes require an ADR and all three owners' approval; consumers retain the old form until migration is complete. +- No owner imports another owner's implementation package. Integration happens only through ports and JSON fixtures defined below. + +## 6. Contract freeze — the mechanism that removes day-to-day blockers + +The following semantics are the Milestone 0 contract. Implementation details may vary, but no track may reinterpret them. + +### 6.1 Create-intent command + +Required input: + +- `business_intent_id`: caller-supplied UUID/opaque stable ID. +- `recipient`: checksummed or normalized EVM address. +- `amount_atomic`: canonical base-10, non-negative integer string; no signs, decimals, exponent, or leading whitespace. +- `asset`: exactly `USDC`. +- `network`: exactly `eip155:5042002`. +- `purpose`: non-secret, length-bounded human description used only for display/audit. + +The server computes an immutable payload fingerprint from normalized recipient, amount, asset, network, and purpose. Reusing the ID with the same fingerprint is a replay; reusing it with a different fingerprint is a conflict and never creates another settlement right. + +### 6.2 Public HTTP seam + +| Operation | Required behavior | +| --- | --- | +| `POST /v1/intents` | Create or replay an intent; return `202` for accepted, `200` for identical replay, `409` for same-ID conflict, and no external effect in the request transaction | +| `GET /v1/intents/{id}` | Return intent, attempts, settlement state, sanitized evidence, and stable version | +| `POST /v1/intents/{id}/reconcile` | Enqueue/read-trigger reconciliation only; never directly submit settlement | +| `GET /v1/intents/{id}/recovery-view` | Return authoritative local state plus clearly labeled indexed/provider evidence and freshness | +| `GET /health/live` | Process liveness without external dependency claims | +| `GET /health/ready` | Database plus configuration readiness; fail on wrong Arc chain ID or invalid required configuration | + +Every mutation uses service authentication, request-size limits, schema validation, a correlation ID, and rate limiting. API errors use stable machine codes and never expose provider secrets or raw authorization material. + +### 6.3 Port result contracts + +| Port | Terminal result families | Required meaning | +| --- | --- | --- | +| `AuthorizationPort.evaluate` | `AUTHORIZED`, `DENIED`, `UNAVAILABLE` | `DENIED` and invalid scope produce zero submission; `UNAVAILABLE` is retryable only before submission | +| `SettlementPort.submit` | `CONFIRMED`, `DEFINITELY_NOT_SUBMITTED`, `POSSIBLY_SUBMITTED` | The adapter must never collapse an ambiguous response into a safe retry | +| `EvidencePort.lookup` | `FINAL_SUCCESS`, `FINAL_REVERT`, `PENDING`, `NOT_FOUND`, `UNAVAILABLE` | `NOT_FOUND` alone cannot authorize a new submission | +| `IndexViewPort.lookup` | evidence plus indexed block, timestamp, deployment, lag, health | Data is explanatory; missing/unhealthy data cannot transition `UNKNOWN` to retryable | + +All port requests contain the stable Business Intent ID, immutable payload fingerprint, Arc/USDC identifiers, persisted provider idempotency key, correlation ID, and attempt ID. Simulators must read and emit the same checked JSON fixtures as production adapters. + +### 6.4 Durable state model + +Keep separate records for Business Intent, Attempt, and Settlement. A compact settlement state machine is: + +| Current | Trigger | Next | External submission allowed? | +| --- | --- | --- | --- | +| `NONE` | validated intent accepted | `AUTHORIZING` | No | +| `AUTHORIZING` | Privy policy authorizes | `READY` | No | +| `AUTHORIZING` | policy denies | `REJECTED` | No, terminal | +| `READY` | atomic owner grant persists request identity | `SUBMITTING` | Exactly one owner may cross the boundary | +| `SUBMITTING` | verified final receipt and expected Transfer log | `COMMITTED` | No, terminal | +| `SUBMITTING` | narrow proof of no broadcast | `FAILED_SAFE` | A new attempt may be scheduled by policy | +| `SUBMITTING` | timeout, disconnect, lost response, crash, or doubt | `UNKNOWN` | No | +| `UNKNOWN` | reconciliation finds verified success | `COMMITTED` | No, terminal | +| `UNKNOWN` | reconciliation proves final revert/no settlement with authoritative evidence | `FAILED_SAFE` | Only then may policy schedule a new attempt | +| `UNKNOWN` | pending, not found, lagging, unhealthy, or contradictory evidence | `UNKNOWN` | No; operator attention if deadline exceeded | + +Mandatory storage constraints and records: + +- Primary/unique Business Intent ID plus immutable fingerprint. +- At most one Settlement row per Business Intent; provider transaction hash unique when present. +- N append-only Attempt rows with stage, timestamps, sanitized error class, and correlation ID. +- Persisted Privy idempotency key, reference ID, request fingerprint, wallet ID, recipient, amount, chain, token contract, transaction ID/hash/nonce when learned, receipt block/hash/status, and verified Transfer log identity. +- Compare-and-set state transitions with a monotonically increasing version. No database transaction spans an external network call. +- Transactional outbox/job insertion. Queue delivery and API retries are assumed duplicate and out of order. +- On worker startup, any orphaned `SUBMITTING` record is conservatively moved/treated as `UNKNOWN` for reconciliation; lease expiry never grants a blind resubmission. + +### 6.5 Agreed test seams + +Tests observe behavior through these public seams only: + +1. HTTP API plus returned durable state. +2. Worker task input/output plus durable state and external-submission counter. +3. Adapter ports with official-response fixtures. +4. Reconciliation command plus durable transition and evidence record. +5. Subgraph mappings/GraphQL query plus indexed entity and `_meta` classification. +6. Browser UI through the public API contract in Milestone 5. + +Each implementation ticket uses one red-green vertical slice at a time. Tests must assert both durable state and settlement count; HTTP status alone is insufficient. + +## 7. Milestone overview and dependency graph + +| Milestone | Days | Exit outcome | +| --- | ---: | --- | +| M0 — Contract and safety freeze | 0.5 | This plan, research, ports, fixtures, states, ownership, and test seams accepted | +| M1 — Three independent walking skeletons | 1–4 | Each track runs locally with its own simulator and no cross-track implementation import | +| M2 — Safety-critical vertical slices | 5–8 | Domain concurrency, real policy/transaction adapter, and reconciliation logic pass independently | +| M3 — Failure and operational hardening | 9–12 | Each track passes its assigned fault, restart, and observability evidence | +| M4 — Integrated backend and live testnet proof | 13–15 | All adapters compose; full matrix passes; one real authorized settlement is recorded and indexed | +| M5 — Frontend, last | 16–18 | Minimal intent, status, and recovery UI works against the stable backend | +| M6 — Demo qualification and release candidate | 19–20 | Scripted demo, sponsor evidence, runbooks, and release checks pass | + +```text +A1 -> A2 -> A3 --\ +B1 -> B2 -> B3 ----> M4 integrated backend -> M5 frontend -> M6 demo/release +C1 -> C2 -> C3 --/ +``` + +There are no cross-person blockers through M3. Each task depends only on the same owner's prior task. M4 is the first convergence dependency; its build work can continue against simulators, but its exit test requires all three artifacts. This is intentional and cannot be removed without pretending integration is optional. + +## 8. Detailed milestones + +### M0 — Contract and safety freeze (all three, half day) + +Deliverables: + +- Accept Sections 4–6 as the initial ADR-equivalent contract. +- Create versioned JSON fixtures for identical replay, conflicting replay, authorization denial, confirmed transfer, final revert, pending transaction, lost response, empty Graph result, lagging Graph result, and indexing error. +- Confirm package/file ownership and the no-cross-implementation-import rule. +- Record required environment-variable names in `.env.example` with placeholders only; classify each as secret or public. +- Confirm the six public test seams before any test is written. + +Exit criteria: + +- Each person can run their package tests with local fakes and no credentials. +- Every contract field has one owner, type, normalization rule, and redaction rule. +- No unresolved decision can change settlement cardinality, money representation, or the classification of `UNKNOWN`. + +### M1 — Three independent walking skeletons (days 1–4) + +#### A1 — Durable intent skeleton (Person A; blockers: M0 only) + +What it delivers: an intent can be accepted, replayed, queried, queued, and observed end to end using fake authorization/settlement ports. + +Work: + +- Create the workspace, strict compiler/lint/test/build commands, API/worker entry points, OpenAPI validation, and package-local commands. +- Add PostgreSQL migrations for intents, attempts, settlements, outbox/jobs, evidence, and schema versioning. +- Implement normalized fingerprinting, create/replay/conflict behavior, GET status, atomic state transitions, and a fake adapter with an external-settlement counter. +- Add containerized PostgreSQL test support and deterministic clock/ID seams. + +Acceptance: + +- Identical request twice returns the same Business Intent and one queued execution. +- Conflicting payload under the same ID returns `409`, records the conflict safely, and creates zero extra settlement rights. +- State survives API and worker restarts. +- Package tests prove constraints using a real PostgreSQL transaction, not only mocks. + +#### B1 — Privy/Arc adapter skeleton (Person B; blockers: M0 only) + +What it delivers: a standalone adapter can validate configuration, build exactly one canonical ERC-20 transfer request, classify official-response fixtures, and verify receipts without a real domain service. + +Work: + +- Pin and validate the Privy Node SDK plus Arc client library in a package-local compatibility test. +- Define Arc Testnet configuration and readiness checks for chain ID, USDC contract code, wallet address, and amount precision. +- Build six-decimal ERC-20 transfer calldata and the Privy request with stable idempotency/reference identifiers. +- Implement receipt verification: chain, sender, token contract, status, recipient, amount, transaction hash, block, and unique Transfer log. +- Draft the fail-closed Privy wallet policy fixture; do not store credentials. + +Acceptance: + +- Golden fixtures produce byte-for-byte stable request fingerprints and calldata. +- Wrong chain, token, recipient, amount format, or native value is rejected before signing. +- `status: 1` without the expected Transfer log is not `CONFIRMED`; `status: 0` is final revert. +- Timeout/lost-response fixtures return `POSSIBLY_SUBMITTED`, never safe retry. + +#### C1 — Indexed recovery skeleton (Person C; blockers: M0 only) + +What it delivers: a standalone Subgraph and recovery package map Arc USDC Transfer fixtures and expose a freshness-labeled recovery view against a fake domain/evidence host. + +Work: + +- Create Subgraph schema, manifest, mapping, and Matchstick/unit fixtures for the Arc USDC contract. +- Create the Graph client query including `_meta`, deployment, block, timestamp, and indexing errors. +- Implement freshness states: `FRESH`, `LAGGING`, `UNHEALTHY`, `UNAVAILABLE`, `UNKNOWN_FRESHNESS`. +- Create the reconciliation decision table and a simulator for local state, Privy evidence, Arc receipts, and indexed evidence. + +Acceptance: + +- Mapping identity is transaction hash plus log index; amount remains Graph `BigInt`/decimal string. +- Empty or lagging Graph fixtures never return permission to resubmit. +- Recovery output labels which facts are authoritative and which are indexed observations. +- Package tests run with no network or credentials. + +M1 exit: all three package suites pass independently. Re-estimate M2–M6 from actual SDK, chain, and Subgraph friction; do not reduce safety acceptance to preserve the date. + +### M2 — Safety-critical vertical slices (days 5–8) + +#### A2 — Atomic at-most-once engine (Person A; blockers: A1 only) + +What it delivers: duplicate deliveries and concurrent workers converge on one submission owner and at most one committed settlement in the fake-adapter system. + +Work and acceptance: + +- Implement transactional authorization-to-ready and ready-to-submitting compare-and-set transitions. +- Configure `submit_settlement` with one queue attempt and catch/classify all adapter results into durable states before returning. +- Prove one normal job, 10 sequential retries, 10 parallel workers, and two worker/agent instances produce exactly one external submission/commit. +- Kill before external call: zero settlement and safe retry. Kill after the boundary: durable `UNKNOWN` and no new submission. +- Preserve a committed Settlement when a downstream/supplier simulation fails. + +#### B2 — Authorized Arc Testnet settlement (Person B; blockers: B1 only) + +What it delivers: the adapter executes one policy-constrained testnet USDC settlement from the standalone harness and returns verified normalized evidence. + +Work and acceptance: + +- Provision the execution wallet, owner/key quorum, and one attached policy through a human-run setup procedure; write secrets only to ignored runtime storage/approved CI secrets. +- Test policy allow and deny cases against Arc Testnet: wrong chain, wrong contract/method, wrong recipient, above cap, non-zero native value, expired authorization. +- Submit via Privy with `eip155:5042002`, persisted idempotency key, and stable reference ID; capture Privy transaction ID/hash and Arc receipt. +- Prove one allowed transfer commits once and every denial produces zero settlement. +- Produce sanitized fixtures from real response shapes for Person A and C without exposing secrets. + +#### C2 — UNKNOWN reconciliation engine (Person C; blockers: C1 only) + +What it delivers: a deterministic read-only reconciliation decision engine resolves authoritative evidence or holds safely without ever submitting a payment. + +Work and acceptance: + +- Implement evidence precedence: durable committed record and verified Arc receipt are authoritative; Privy status locates provider activity; Graph corroborates/history only. +- Resolve verified receipt success to `COMMITTED` and final revert with matching identity to `FAILED_SAFE`. +- Keep `UNKNOWN` for pending, provider unavailable, RPC unavailable, Graph empty/lagging/unhealthy, identity mismatch, or contradictory evidence. +- Persist every observation with source, retrieval time, block height, health, and sanitized reason. +- Prove repeated reconciliation and duplicate webhook/provider events are idempotent and create zero submissions. + +### M3 — Failure and operational hardening (days 9–12) + +#### A3 — Restart-safe orchestration and auditability (Person A; blockers: A2 only) + +What it delivers: the API/worker system recovers after process/database interruptions, exposes useful safe telemetry, and has a deterministic safe-disable path. + +Acceptance: + +- Restart after intent creation, job claim, `SUBMITTING` persistence, and adapter return; invariant holds at every point. +- Stale/orphaned work becomes reconciliation work, not a new submission lease. +- Structured logs carry Business Intent/Attempt IDs and state transitions but redact payload purpose as configured and never include credentials, authorization signatures, raw provider bodies, or private wallet material. +- Metrics cover state counts, transition failures, queue lag, UNKNOWN age, reconciliation outcomes, policy denials, and duplicate/conflict counts. +- A kill switch stops new submissions while status and reconciliation reads remain available. + +#### B3 — Provider ambiguity and policy hardening (Person B; blockers: B2 only) + +What it delivers: provider/RPC outcomes are conservatively classified across realistic failures and the policy remains effective after restart/config changes. + +Acceptance: + +- Inject DNS failure, connection refusal, timeout before response, truncated response, 429/5xx, malformed payload, lost success response, pending/evicted transaction, final revert, and mismatched receipt. +- Only documented, proven pre-broadcast failures become `DEFINITELY_NOT_SUBMITTED`; every doubtful result becomes `POSSIBLY_SUBMITTED`. +- Reusing the persisted Privy idempotency key and identical body is tested; its 24-hour limit is documented and never treated as permanent protection. +- Policy fingerprint/ID and expected restrictions are checked at readiness; mismatch fails closed. +- If webhooks are available, signature verification and duplicate/out-of-order delivery tests pass; otherwise polling remains complete and webhooks stay disabled. + +#### C3 — Indexer lag, contradiction, and chaos evidence (Person C; blockers: C2 only) + +What it delivers: recovery remains safe when The Graph or other evidence sources are delayed, empty, unhealthy, inconsistent, or unavailable. + +Acceptance: + +- Delay and empty The Graph results, set `hasIndexingErrors`, trail chain head, remove `_meta`, fail the query, and return duplicate/out-of-order events; none unlock a payment. +- Inject crash/lost response after possible submission and show the record remains `UNKNOWN` until authoritative evidence resolves it. +- Verify recovery evidence survives service restart and can be replayed for audit without provider secrets. +- Define UNKNOWN-age alerts and a human escalation runbook; the runbook never tells an operator to “just retry.” +- Produce a single command that runs the cross-source fixture matrix against the reconciliation package. + +M3 exit: each owner passes their package suite and provides a versioned artifact plus fixtures. No cross-track package implementation is required to reach this exit. + +### M4 — Integrated backend and live testnet proof (days 13–15) + +This is the first cross-track convergence. Each person prepares against simulators immediately; only the final acceptance run waits for all three M3 artifacts. + +#### A4 — Composition and migration integration (Person A) + +- Wire production ports without importing provider details into the domain package. +- Run migrations from an empty database and from the previous schema; verify rollback/safe-disable behavior. +- Validate OpenAPI, generated contract fixtures, root lint/type/test/build, and service readiness. +- Own conflict resolution only in shared/root files; provider owners resolve their packages. + +#### B4 — Live authorization/settlement evidence (Person B) + +- Run one allowed Arc Testnet transfer through the integrated worker. +- Run policy-denied and above-cap intents and prove zero settlement. +- Capture sanitized transaction ID/hash, receipt, expected Transfer log, chain, policy identity, and explorer link for demo evidence. +- Trace an intentionally lost local response into `UNKNOWN` without permitting a second transaction. + +#### C4 — Integrated reconciliation and matrix (Person C) + +- Reconcile the lost-response scenario to the original final transaction using durable/Privy/Arc evidence. +- Demonstrate live Subgraph history with `_meta`, then simulate lag/empty/error and show safe behavior. +- Run the full `.agent/TEST_MATRIX.md` suite and publish a sanitized results table with durable state and external settlement count. +- Verify alerts and recovery-view output distinguish authoritative and indexed evidence. + +M4 exit criteria: + +- All lint, static analysis, type, unit, integration, contract, build, migration, and focused failure-injection checks pass from the repository root. +- Every required test-matrix row records stable ID, final durable state, and external settlement count. +- Real testnet happy path has exactly one committed settlement; denial paths have zero; ambiguous path has no duplicate. +- Backend API/OpenAPI and recovery semantics are frozen for the frontend. Breaking changes after this point use expand-migrate-contract. + +### M5 — Frontend, last (days 16–18) + +Frontend work starts only after M4 passes. All three slices use the frozen OpenAPI and mock server, so component work remains parallel. + +#### F-A — Intent shell and status (Person A) + +- App shell, service-auth handoff suitable for the demo environment, create-intent form, exact atomic-amount parsing/formatting, and status polling. +- Show replay and same-ID conflict clearly; never generate a new Business Intent ID on a retry unless the user starts a genuinely new obligation. +- Display only USDC, Arc Testnet, and six payment decimals; keep native gas details separate. + +#### F-B — Authorization and settlement details (Person B) + +- Policy scope summary, authorization denied state, submission/pending/final state, sanitized transaction details, and Arc explorer link. +- No bypass button and no “force pay” action. `UNKNOWN` disables new settlement submission. +- Do not show a confirmation counter: Arc is pending or final. + +#### F-C — Recovery timeline and indexed history (Person C) + +- Attempt/reconciliation timeline, authoritative local state, Privy/Arc evidence, Graph observations, indexed-through block/time, lag, and health. +- Empty Graph data is labeled “not observed through block N,” never “not paid.” +- UNKNOWN state provides safe explanation/escalation, not a retry shortcut. + +M5 exit criteria: + +- Browser tests cover create, identical replay, conflict, denial, committed, UNKNOWN, reconciliation, Graph lag/error, and service-unavailable paths. +- Accessibility smoke tests, responsive layout, lint/type/build, and no-secret/source-map checks pass. +- The UI cannot invoke an unguarded settlement path. + +### M6 — Demo qualification and release candidate (days 19–20) + +#### Person A — Invariant and operational demo + +- Script duplicate requests, 10 parallel workers, restart, lost response, and downstream failure; show durable states and settlement count. +- Verify clean database bootstrap, safe-disable switch, logs/metrics, README, architecture diagram, and operator runbook. + +#### Person B — Privy and Arc evidence + +- Demonstrate policy-constrained corporate wallet execution, one real authorized USDC transfer on Arc Testnet, and zero-settlement denials. +- Record sanitized policy scope, transaction/receipt/Transfer proof, network, explorer URL, and limitations. + +#### Person C — The Graph and recovery evidence + +- Demonstrate live indexed Arc data in the recovery view and the same flow under delayed/empty/unhealthy indexed data. +- Run sponsor qualification against working code/tests/demo evidence and report each sponsor `QUALIFIED`, `NOT QUALIFIED`, or `NOT VERIFIED`; never promote missing evidence. + +Release-candidate exit: + +- Full test matrix and root checks pass on the exact candidate content. +- Secret scan and intended-file review pass; `.env*`, credentials, wallet material, and authorization responses are absent from review inputs. +- Each implementation change follows `.agent/IMPLEMENTATION_LOOP.md`: local checks, a fresh FreePi Gate A, draft PR to `develop`, green required CI, a separate fresh Gate B, then human review. Agents never merge. +- Demo can be reset and repeated using testnet-only funds without manual database surgery. + +## 9. Test ownership matrix + +| Required case | Primary owner | Independent harness | Integrated verifier | +| --- | --- | --- | --- | +| Normal job | A | Fake settlement counter | B | +| Same request twice / conflicting payload | A | HTTP + PostgreSQL | C observes recovery output | +| 10 sequential retries | A | Worker + fake port | B validates one adapter call | +| 10 parallel workers | A | Real PostgreSQL concurrency | C captures evidence timeline | +| Crash before submission | A | Worker kill point | B proves zero call | +| Crash after possible submission | B | Adapter fault point | C reconciles; A verifies state | +| Lost payment response | B | Proxy/fixture fault | C resolves original transaction | +| Graph delay or absence | C | Graph simulator | A verifies no submission grant | +| Privy denial / above policy | B | Policy testnet harness | A verifies zero settlement | +| Service restart | A | Process orchestration | C verifies evidence durability | +| Downstream failure after payment | A | Supplier fake | B verifies original receipt retained | +| Two agent instances | A | Two workers/processes | C verifies one settlement history | + +The primary owner builds the failure fixture and focused proof. Integrated verification is a Milestone 4 responsibility, not a prerequisite for the owner to finish M1–M3. + +## 10. Branching, review, and merge train + +- Do not implement product code on `agents-setup`, `develop`, or `main`. Once this planning/infrastructure change is human-merged to `develop`, create short-lived `milestone/a-*`, `milestone/b-*`, and `milestone/c-*` branches from the same `develop` SHA. +- One branch/PR delivers one task above. Within M1–M3, each branch is blocked only by the prior branch in the same lettered track. +- Merge independent package PRs before root composition. If two changes touch a shared contract, use expand-migrate-contract: add the new form, migrate all consumers in independent PRs, then remove the old form. +- Only Person A edits root composition files during M4. Persons B/C supply reviewed package commits and fixtures, preventing three-way conflicts. +- Every PR lists exact blockers, acceptance evidence, selected test-matrix cases, invariant impact, safe-disable strategy, and both required review gates. +- A human controls merge order and performs every merge. + +## 11. Human-only configuration plan + +Person B owns a repeatable interactive setup wizard after B1 fixes the variable contract. It must guide a human through Privy application/wallet/key-quorum/policy creation, Arc testnet funding, The Graph Studio deployment credentials, and CI secret entry. It must: + +- Open current official URLs before each instruction. +- Capture secrets with hidden input and write them only to ignored `.env` or approved CI secrets. +- Keep public chain/contract/deployment identifiers in non-secret variables. +- Confirm before policy replacement, wallet ownership change, funding, deployment, or any irreversible action. +- Be statically validated but never run end to end by an agent without the human. + +The backend must still boot in fake/local mode with no third-party credentials, so configuration work never blocks Persons A or C. + +## 12. Observability and safe operation + +- Correlation keys: Business Intent ID, Attempt ID, settlement version, Privy reference/transaction ID, Arc transaction hash, and Graph deployment/indexed block. Never log authorization signatures, credentials, private keys, or raw sensitive payloads. +- Alerts: oldest UNKNOWN age, count of UNKNOWN intents, repeated reconciliation failures, Graph block lag/health, policy denials, provider/RPC failure rate, queue lag, and state-transition conflicts. +- Safe disable: stop accepting/claiming new settlement submissions while keeping GET status, evidence ingestion, and reconciliation reads operational. +- Manual escalation: operators inspect durable request identity and evidence; there is no generic retry button. Any future override requires a separate audited design and is outside this plan. + +## 13. Risks and mitigations + +| Risk | Mitigation / fail-closed response | Owner | +| --- | --- | --- | +| Privy idempotency expires after 24 hours | Durable OneShot constraint remains authoritative; reuse stored key/body only as supplemental protection | A/B | +| SDK or policy syntax changes | Pin after B1 compatibility test; readiness verifies policy ID/fingerprint and network; deny on mismatch | B | +| Arc native/ERC-20 precision confusion or double counting | Settlement uses six-decimal ERC-20 only; native balance is gas; verify one canonical Transfer identity | B/C | +| Lost response or process crash after broadcast | Persist request identity before call, enter UNKNOWN, reconcile, forbid new nonce/payment | All | +| Arc transaction pending/evicted with no receipt | Hold UNKNOWN; no automatic replacement in v1; escalate after threshold | B/C | +| Graph lag, error, endpoint version drift, or empty result | Query `_meta`, compare chain head, pin deployment for demo, label stale/unhealthy, never authorize from absence | C | +| Queue redelivery | One queue attempt for submission, domain CAS/unique constraints, idempotent reconciliation | A | +| Shared-file merge conflicts | File ownership plus independent package commands; root composition owned by A | A | +| Credential/setup delays | Fakes unblock all tracks; human wizard and live setup occur in B2, before M4 | B | +| Schedule pressure | Preserve safety acceptance; cut optional webhooks, rolling policies, visual polish, and nonessential telemetry first | All | + +## 14. Definition of done for every implementation task + +- Outcome and non-goals match the task above; no hidden follow-up is required for claimed behavior. +- Public-seam test is written red first, then the smallest vertical behavior is implemented; tests avoid private implementation coupling. +- Relevant unit, contract, integration, concurrency, failure-injection, migration, lint, type, and build checks pass. +- Durable state and external settlement count are asserted where money or retries are involved. +- Security boundaries, input validation, integer money, logging redaction, testnet restriction, and safe-disable behavior are reviewed. +- Documentation, OpenAPI/fixtures, runbooks, and `.env.example` are updated without secrets. +- The branch/diff is focused and passes the repository's FreePi/CI/human-review policy. No agent merges. + +## 15. First implementation actions after plan approval + +1. Human merges the planning/agent-infrastructure change to `develop`; record the exact base SHA. +2. All three people complete M0 together and create their independent branches/worktrees. +3. Person A starts A1; Person B starts B1; Person C starts C1 simultaneously. +4. Hold one 15-minute daily contract check limited to proposed breaking changes, UNKNOWN classification, and risks. Status reporting must not become an approval dependency. +5. At M1 exit, re-estimate the calendar from evidence while preserving M2–M6 acceptance criteria and the rule that frontend remains last.