Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
29 changes: 25 additions & 4 deletions AUDIT_OPEN.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,23 @@
# Open audit items

## 2026-10-04 — wave 5 (live: S02 batch, production upgrade, shadows on)

- Production ship web+worker upgraded `dd5b13e` -> `80c9187` (coordinated
stop; pre-deploy backup `pre-80c9187-shadow-deploy-2026-10-04` verified;
the five waiting runs survived intact; gateway stayed at `c0b07dd`).
- Shadow flags set on production: `SHIP_BUDGET_RESERVATION`,
`SHIP_MODEL_ROUTING`, `SHIP_PLACEMENT`, `SHIP_TOOL_MANIFEST`,
`SHIP_KNOWLEDGE_PROVENANCE` = shadow, `SHIP_POLICY_SHADOW=on`.
`SHIP_MODEL_ROUTING` shadow logs "no usable policy" until
`SHIP_MODEL_ROUTING_POLICY` is set — setting that policy is now the
unblock for that shadow's logs. Shadow-log review before any `on` flip
is still the rule (see NEXT_SESSION.md).
- S02 executed (A only, $0 real spend): see
`evals/product-journeys/BATCH_2026-10-04.md`. Upstream reports filed
(neutron#6, teploy-cli#18/#19/#20 — links above/below).
- The integration-check test-quoting defect (spaced checkout paths) was
fixed in PR #56.

## 2026-10-04 — wave 4 (S07, S13, S15/S17, S19, S21, S22, S24, S26, S27)

All off by default or observe-only; flags are in the programme's wave-4 table.
Expand All @@ -8,9 +26,12 @@ All off by default or observe-only; flags are in the programme's wave-4 table.
legacy path (unreadable status, including a brand-new destination, holds).
A manual `teploy deploy` landing inside Ship's deploy is overwritten and still
reads back confirmed; the lease serialises Ship's own releases on one host
only. Upstream feature requests (not filed): deploy dry-run; deploy generation
and compare-and-set in `status --json`; documented `logs` flags and a
no-deployment-versus-unreachable signal.
only. Upstream feature requests (filed 2026-10-04): deploy dry-run
([teploy-cli#18](http://100.108.123.49:49152/Tyler/teploy-cli/issues/18)); deploy generation
and compare-and-set in `status --json`
([teploy-cli#19](http://100.108.123.49:49152/Tyler/teploy-cli/issues/19)); documented `logs` flags and a
no-deployment-versus-unreachable signal
([teploy-cli#20](http://100.108.123.49:49152/Tyler/teploy-cli/issues/20)).
- **S21 (PR #51):** provenance is a sidecar, not fields on the note. Project
scope is repo-only. Pre-flag notes read as unknown. Run `shadow` and read the
logs before `on`.
Expand Down Expand Up @@ -114,7 +135,7 @@ detail is in the programme's wave-2 status table. Items that matter for audit:
`teploy deploy` can race a tracked delivery; a destination lease is needed
before a Teploy adapter may declare fencing. Upstream feature requests (not
defects): a Teploy CLI dry-run, and a status field exposing a deploy
generation. Not yet filed.
generation. Filed 2026-10-04: teploy-cli#18 and #19.
- **S13 (PR #28):** the steer-route refusal has no automated test.
- **S21, S22, S23, S24, S06:** modules exist unwired (PRs #22, #19, #21, #23,
#29); see each PR's wiring list. The tool-manifest schema is strict, so a
Expand Down
38 changes: 18 additions & 20 deletions NEXT_SESSION.md
Original file line number Diff line number Diff line change
@@ -1,35 +1,33 @@
# Next session

The single forward plan is [docs/SHIP_RELEASE_PROGRAMME_2026-09-21.md](docs/SHIP_RELEASE_PROGRAMME_2026-09-21.md); its **Execution status** section (the wave-2, wave-3 and wave-4 tables) is the only status list. Open audit items and deferred findings live in [AUDIT_OPEN.md](AUDIT_OPEN.md). Update both rather than creating another plan.
The single forward plan is [docs/SHIP_RELEASE_PROGRAMME_2026-09-21.md](docs/SHIP_RELEASE_PROGRAMME_2026-09-21.md); its **Execution status** section (the wave-2, wave-3 and wave-4 tables) plus the wave-5 notes at the top of [AUDIT_OPEN.md](AUDIT_OPEN.md) are the only status lists. Update both rather than creating another plan.

## Where things stand (2026-10-04)
## Where things stand (2026-10-04, after wave 5)

Everything below is **implemented and checked by automated tests only**. No row has a verified outcome: no live Ship, Nucleus, sandbox, forge, preview target or model gateway was reachable, and nothing was spent. Most wave-2 modules are tested but not called by any route or worker; waves 3 and 4 wired many of them behind flags that are **off by default** (or observe-only), so default behaviour is unchanged.
Wave 4's state ("implemented and checked by automated tests only") is now partly superseded by live receipts:

## Needs a human decision first
- **S02 batch executed** (A only, glm-5.3 via the gateway's coding-plan key, $0 real spend): canary + 33 runs, 24/33 first-attempt passes, $1.81 priced. Read [evals/product-journeys/BATCH_2026-10-04.md](evals/product-journeys/BATCH_2026-10-04.md) first — including the `pj-s-question` 0/3 variance signal.
- **Production upgraded** `dd5b13e` -> `80c9187`: coordinated stop, pre-deploy backup verified, five waiting runs intact, gateway untouched at `c0b07dd`.
- **Shadows are ON in production** (budget-reservation, model-routing, policy, tool-manifest, knowledge-provenance = shadow; policy-shadow = on). They need soak time and then a log review before any `on` flip.
- **Founder decisions recorded 2026-10-04**: S28 n>=20 confirmed; S25 service-account stays id-only; S08 findings WILL reach PR body + webhook (approved, **not yet built** — new work); the 79 sub-24px targets stay compact; upstream reports filed (neutron#6 404 page, teploy-cli#18/#19/#20).
- **Known live variance**: `pj-s-question` canary pass + 0/3 batch — compare the four transcripts before touching prompts.

1. **Paid evaluation batch** (S02, propose-only): 66 runs, hard cap $25, details in the programme's S02 slice. It must run from somewhere that can reach a Ship instance. The graders were tightened in PR #34, so the retained results were scored under looser graders and were not re-graded.
2. S28 percentile threshold (n>=20) was an agent's choice.
3. S25 service-account role (shadow treats service accounts as id-only matches).
4. Whether S08 test-integrity findings should reach the PR body and webhook (would change the worker path).
5. The 79 sub-24px compact targets from the UI audit (keep or enlarge).
6. File the upstream reports: Neutron plain-text 404 page (`docs/UI_AUDIT_2026-10-03.md`), and the Teploy CLI feature requests in `AUDIT_OPEN.md`. Nothing has been filed.
7. When each wired flag should be turned on (shadow first, read the logs).
## Next work, in order

## Needs live systems or the owner's side

S01 credential proofs on a real sandbox with a private repo; preview isolation (teploy-cli); end-to-end journeys; real-Nucleus checks; the S03 storage migration rehearsal on a restored copy; real deployment adapters; real placement targets; human observation of dashboard users; re-grading retained runs where their trees exist; shadow-log review before enabling any flag.

## Safe next work (no live system needed)

- Turn shadow findings into decisions once logs exist; otherwise wire the remaining pure modules (S15/S17 into delivery and incidents, S07 into the plan-park point without adding a durable step, S19 snapshot producer and restore-check command, S18 worker wiring).
- The untouched or barely started packages: S04, S12, S10 follow-ups (offline and error states, other roles), remaining S07 parts.
- Add `./plan-grounding`, `./deployment-adapter`, `./policy-inheritance`, `./tool-manifest` and `./teploy-adapter` subpath exports only when something imports them.
1. **Shadow-log review** after soak (placement/policy/budget/tool-manifest/knowledge-provenance JSONL + reports); set `SHIP_MODEL_ROUTING_POLICY` so the routing shadow has something to record. Only then discuss any `on` flips.
2. **S08 wiring** (approved): findings into PR body + webhook behind a default-off flag; worker-path change, so shadow-first per the standing rule.
3. **Live proofs still open**: S01 credential proofs on a real sandbox with a private repo; S03 storage migration (write the additive migration, rehearse on a restored copy — backup `pre-80c9187-shadow-deploy-2026-10-04` is available and verified); S27 real teploy-adapter run against a scratch target; S19 doctor probes against the real store/clock; human observation of dashboard users.
4. **Code still unfinished**: S04 and S12 barely started; most of S07 beyond the grounding check; S10 follow-ups (offline/error states, other roles, screen readers); wiring the inert modules (S15/S17 into delivery+incidents, S19 snapshot producer + restore-check, S18 worker tree provisioning, S07 into the plan-park point).
5. **Re-grade retained runs where possible**: only transcripts were preserved for pre-batch runs, so regrades are limited to transcript+fixture-verifiable scenarios; write supplementary `regrade-*.json` beside the records, never replacing originals.
6. **Deferred**: configuration B of the S02 comparison (needs real API spend, owner-gated).

## Working rules that paid off

- One branch and PR per slice; merge only on green CI, pinned to the checked head; resolve conflicts by merging main in, never rewriting history.
- Agents in parallel worktrees: give each its own scratch directory (the shared one caused a backup collision) and tell them to skip `scripts/grader-sensitivity.test.mjs` locally (fixed port 8901); CI runs it.
- Every wiring change: default-off equivalence test, a real-path test with the flag on, a negative control per rule.
- Quote interpolated paths in any command template a test executes (PR #56 was the spaced-checkout lesson).
- Never redeploy a teploy app from a hand-copied or stripped teploy.yml — deploy from the repo that owns the full one, or you fight the original deployment's shape (the 2026-10-04 gateway incident: ad-hoc deploy from a stripped yml removed the running container; restored by redeploying the exact revision from the real repo with no data loss, but it did not need to happen).
- Production changes: coordinated stop + verified backup first (scripts/ship-backup.sh), waiting runs must survive, gateway is a separate app and stays up.

Before finishing any slice: `pnpm run lint`, `pnpm test`, and for `web/` changes `cd web && pnpm test && pnpm run build`. Install `web` dependencies first (`cd web && pnpm install --frozen-lockfile`) or the deployment-pin script test fails for an environmental reason.
19 changes: 15 additions & 4 deletions docs/SHIP_RELEASE_PROGRAMME_2026-09-21.md
Original file line number Diff line number Diff line change
Expand Up @@ -512,14 +512,14 @@ This section supersedes the earlier table where rows overlap (S13, S28, S05, S06
| S21 | #22 | `src/knowledge-record.ts` provenance, verification, freshness, correction, redaction closure, visibility | 11 tests + 10 negative controls | Not wired; needs provenance fields and a project concept; decision acceptance, cross-project sharing, retrieval baseline not built |
| S24 | #23 | `src/tool-manifest.ts` validation, intersection-only permissions, conformance audit, signed and deduped event envelope | 25 tests + 8 negative controls | Not wired; enforcement stays with the credential and executor layer; manifest schema is strict (newer-minor manifests are rejected) |
| S16 | #24 | `src/schedule-time.ts` timezone and DST-aware daily/weekly slots, missed-run policy, overlap, debounce | 34 tests, naive-implementation control, 7 mutations | Overlap is inert in production (the worker does not pass `isRunning`); no UI or API to create `at` schedules; no event triggers |
| S25 | #25 | `src/policy-inheritance.ts` org/project/user/service-account resolution, intersection only, fail-closed | 21 tests + union-merge control + 3 mutations | Not wired; no layer store, model/tool/retention enforcement or lifecycle; **service-account role needs product confirmation** |
| S25 | #25 | `src/policy-inheritance.ts` org/project/user/service-account resolution, intersection only, fail-closed | 21 tests + union-merge control + 3 mutations | Not wired; no layer store, model/tool/retention enforcement or lifecycle; **service-account role: owner confirmed 2026-10-04 — keep id-only matches** |
| S26 | #26 | `src/execution-target.ts` placement with reasons, host-loss recovery, conformance | 23 tests, 10 mutations, two fake backends | Not wired into `SandboxPool` or `durable.ts`; no real Windows/macOS/mobile/GPU adapter |
| S27 | #27 | `src/deployment-adapter.ts` adapter contract, delivery and recovery and provisioning journeys, 27-check conformance suite | 48 tests, 2 faithful + 9 faulty fakes, 8 mutations | No real Teploy/CI/Kubernetes adapter, so S27 acceptance is not met; Teploy has no destination-level fence |
| S13 | #28 | Full capability declaration for native, claude-code and opencode; shared plan-review refusal; pure `conformanceCheck` | 8 new tests + negative control | Per-harness journeys on live runs; `conformanceCheck` not wired to adapter probes; steer-route refusal has no automated test |
| S06 | #29 | `src/environment-recipe.ts` recipe, planning, teardown, lifecycle validation | 18 tests, 6 mutation-verified controls, fake runner only | Not wired; no real runner or cache store; large/binary transport, browser tooling, snapshot secret scan open |
| S02 graders | #34 | `scripts/grader-sensitivity.mjs` plus 73 hand-written variants (60 wrong, 12 correct references, 1 known limit) across all 12 scenarios, committed matrix; **eight grader tightenings after 11 wrong variants passed against the previous graders** (deploy-recovery health gate satisfied by a comment and by a lying server, permissions accepted 403 where the manifest says 401, API catch-all 409, migration edited in place, "30-minute" wording, scheduled-job store bypass, applied diff in a plan) | 79 tests; 60/60 wrong variants rejected for their targeted reason, 12/12 references accepted; negative controls (stub grader passes everything, sabotaged reference, reverted fix) | Variants are finite; the review grader stays lexical; the tightened graders have not run on a live baseline; **retained results were not re-graded, so earlier pass counts were scored under the looser graders**; `pj-s-feature` accepts a marked-placeholder contact page by design (known limit); deploy-recovery grading uses fixed port 8901 and cannot run in parallel |
| S10 | #33 | `scripts/ui-audit.mjs` (47 routes x 5 viewports, optional axe-core, JSON and markdown report) and five defect-class fixes (input labels, empty table headers, underlined in-text links, keyboard-reachable scrolling regions, narrow-viewport overflow); report in `docs/UI_AUDIT_2026-10-03.md` | 14 helper tests with negative controls, 3 web source guards, one full audit run (230 page loads; no error-level finding or axe violation on audited pages apart from the framework 404) | No human observation of target users; 79 targets under 24px kept as compact controls (design decision); framework plain-text 404 (upstream-report candidate, not filed); keyboard activation, offline and error states, other roles, screen readers, `/connect*`, `/oidc/*`, `/hooks/*`, `/api/*` not audited |
| S28 | #30 | `timing` in `audit --format json` (not CSV), `timingSummary` with an n>=20 gate | 8 tests, CSV byte-pinned, mutation control; CLI wiring type-checked only | No real-store run; **the n>=20 threshold was an agent's choice and needs confirmation**; capacity table not recomputed |
| S10 | #33 | `scripts/ui-audit.mjs` (47 routes x 5 viewports, optional axe-core, JSON and markdown report) and five defect-class fixes (input labels, empty table headers, underlined in-text links, keyboard-reachable scrolling regions, narrow-viewport overflow); report in `docs/UI_AUDIT_2026-10-03.md` | 14 helper tests with negative controls, 3 web source guards, one full audit run (230 page loads; no error-level finding or axe violation on audited pages apart from the framework 404) | No human observation of target users; 79 targets under 24px kept as compact controls (**owner-confirmed 2026-10-04**); framework plain-text 404 (filed upstream as [Tyler/neutron#6](http://100.108.123.49:49152/Tyler/neutron/issues/6)); keyboard activation, offline and error states, other roles, screen readers, `/connect*`, `/oidc/*`, `/hooks/*`, `/api/*` not audited |
| S28 | #30 | `timing` in `audit --format json` (not CSV), `timingSummary` with an n>=20 gate | 8 tests, CSV byte-pinned, mutation control; CLI wiring type-checked only | No real-store run; **the n>=20 threshold was confirmed by the owner 2026-10-04**; capacity table not recomputed |

#### Wave 3 (2026-10-04): wiring behind flags

Expand Down Expand Up @@ -583,7 +583,18 @@ Small specifications per the slice discipline above, recorded against their exis

Spec, evidence and the open prerequisites are in `AUDIT_OPEN.md` (2026-10-02/03 entry). Rollout order: confirm the Sandbox daemon forwards per-exec `env` and the image's git is 2.31+; prove a private-repo clone and push on Forgejo and GitHub through the real sandbox with `SHIP_GIT_CREDENTIAL=env` on a canary worker; observe; only then change the default. Revert is unsetting the variable.

#### S02 — proposed current-versus-candidate batch (**awaiting your approval; nothing run, nothing spent**)
#### S02 — batch EXECUTED 2026-10-04, configuration A only (owner decision)

The owner approved **A only**: the gateway's `ZAI_API_KEY` was rotated to a
z.ai coding-plan (subscription) credential, so configuration A
(`zai/glm-5.3`) ran with **$0 real spend**; configuration B
(`anthropic/claude-sonnet-5`) was not run and no API dollars were spent.
Canary `eval-20261003-2` (pass) then the 11-scenario × 3 batch
(`eval-20261003-3/-4`, `eval-20261004-1/-2/-3/-4`): **24/33 first-attempt
passes, $1.81 priced total** (accounting rate table, not dollars).
Summary and honest reads: [../evals/product-journeys/BATCH_2026-10-04.md](../evals/product-journeys/BATCH_2026-10-04.md).
Not done: the B leg of the comparison (needs real API spend, owner-gated);
re-grading retained pre-tightened-grader results.

- **Purpose:** a regression comparison on the existing 12-scenario product-journey set. It is **not** a held-out result: these scenarios have been seen, graders amended, and the review grader tuned, so they are a development/regression set. Do not quote it as general quality or market rank.
- **Configurations (exact):** A, current default: native harness, `zai/glm-5.3`, existing project settings. B, candidate: native harness, `anthropic/claude-sonnet-5`, same project settings and prompts. Everything else identical; harness revision and fixture hashes are recorded by the runner. The prompts were tuned against GLM only (docs/MODELS.md), so B is disadvantaged by construction; a loss for B is not evidence about the model.
Expand Down
Loading
Loading