Skip to content

feat: Architect and Orchestrator autonomy (direct work, tool discovery, Room amendments, durable waits) - #624

Draft
monobyte wants to merge 54 commits into
mainfrom
feat/improve-architect-orchestrator-autonomy
Draft

monobyte wants to merge 54 commits into
mainfrom
feat/improve-architect-orchestrator-autonomy

Conversation

@monobyte

@monobyte monobyte commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Implements OpenSpec change improve-architect-orchestrator-autonomy (#620) as one PR. 32 of 44 tasks are ticked. The 12 left are listed under "Not done".

What you can do after this PR

  • The Architect can do a milestone itself. The owner begins, continues and reports work without a Workflow or Room, in Workspace mode and in Worktree mode. In Worktree mode each milestone gets its own checkout and branch inside the project folder, and Sero commits the work to that branch at the report. The work has its own identity, saved before any file changes, so a restart resumes it. A report is a claim: evidence and acceptance still apply.
  • Agents find approved tools as they need them. An Architect owner, a Room member and a Workflow worker start with a small loaded set and load other approved tools with tool_search. A tool outside the approval is never registered. New Architect owners ask for the workspace's enabled skills at the start approval.
  • Code Mode in managed sessions. A new Architect owner asks for Code Mode at the start approval, and a Room can give it to a member. A script can call only the tools that session was approved for.
  • A running Room can change. Change a member's model, thinking, tools or skills, add a member, or replace one, on the same approval with history kept. A change that adds access is held until you approve it in the new Team view.
  • Waits resume by themselves. An owner or a Goal can wait on linked child work with a deadline and is woken once when it finishes. An expired wait is a hold, never a completion.
  • No fixed 10 minute turn limit. A turn that keeps working runs on. Ten silent minutes get a nudge to save and stop, five more an interruption, and a second stall in a row holds the project. Your cost cap now stops a running turn.
  • You can see who set each limit, and an agent can no longer raise or remove a Goal limit you set through the goal tool.

Two faults the paid pilot found, both fixed here

  1. The Architect could lose sero-cli. A run from source cached sero-cli as a plugin's tool, and the approval step then dropped it for managed sessions. The owner changed the files correctly and had no way to record anything. Fix: be0d682. This is on main today.
  2. Running sero from the shell in a managed session answered "Unauthorized session". The model spent its turn probing it. The refusal now says to use the sero-cli tool. Fix: 7edaf65.

Paid runs

deepseek/deepseek-flash at low thinking, one small-fix scenario, $1 and 8 minutes per run. Records and a fuller table are in openspec/changes/improve-architect-orchestrator-autonomy/evidence/results.md.

Run Hidden test Outcome recorded Elapsed Cost
Architect, before this change passed none 8 min (bound) $0.022
Single agent, before this change passed accepted 23 s $0.006
Architect, this branch, Workspace mode passed milestone accepted in about 90 s, project not closed at the bound 8 min (bound) $0.036
Architect, this branch, Worktree mode passed accepted, project closed 90 s $0.016

No improvement is claimed from these. One scenario with one repeat is too few. The pilot was planned as 20 runs and cut to keep paid runs short.

In the Worktree run the owner got a checkout on its own branch, edited inside it, and Sero committed the work there. The hidden test passed on that branch. A new owner's approval on this branch listed Code Mode (read from a run I stopped early); no owner ran a script. An earlier Worktree run has no test result because my runner read the checkout after Sero had released it. The runner now reads the branch.

In the Workspace run the owner used direct work end to end with a real model. Its first milestone picked up a preview check the project could not pass, so it opened a second milestone and asked how to close the first. That preview behaviour is not from this change.

Decisions I made that are yours to overturn

  • In Worktree mode a delivered milestone's result is its branch. Nothing merges it into the default branch for you. The receipt is the branch name.
  • Only new owners get Code Mode. An owner that already has an approval keeps exactly what it had.
  • Rooms got a new Team view with a button in the Room top bar. The drawn table needs four aligned columns and the 230px roster rail cannot hold them.
  • A tool the host does not grant is dropped, and the change still applies with the reason naming the tool. A missing model or skill holds the change with the member paused.
  • The pilot ran on deepseek-flash, low thinking, at reduced size, after your steer on test length.

Needs your eyes

  • Prototypes, approved on 2026-10-06 after they were built: apps/styleguide/public/prototypes/room-amendment-status.html and architect-wait-status.html, with captures under prototypes/screenshots/.
  • The production Team view and the Architect wait card were checked as components, not inside the running app.
  • In production the wait card sits at the top of the Live tab and the Limits list at the top of the Inspector. Limit names come from the code and are longer than the drawing. A wait names its source as "Child work ", not a PR title.
  • "Resume work" on an expired wait calls the existing resume, which may refuse if the project is neither paused nor blocked.

Not done

  • Process and CI waits (5.3). Both are refused. No existing seam reports a process exit or a specific check result by identity.
  • An Architect action that asks to widen an existing owner's access. The host supports the amendment. Nothing in the Architect requests it.
  • The full pilot and the replays (1.5, 2.9, 3.7, 4.8, 5.8, 6.6). No paid run exercised a Room amendment, a durable wait, tool_search, a Code Mode script, a preview check in a worktree, or stall recovery. Their evidence is tests only.
  • Reporting comparison and prompt removals (6.1, 6.2). They need counted protocol failures, which the runner does not record. No prompt procedure was removed.
  • Rendered UI check and final consistency pass (6.3 to 6.5). The docs for each phase are written. The final pass is not done.

Review

Routed rounds on gpt-6-astra at high effort, posted below.

  • Round 1 found 17 defects. 16 are fixed and one is tracked as Plugin UI tools are callable by the model, so the Goals tool can raise a user-set limit #623 (a plugin's UI tool can be called by a model, so the Goals tool can still raise a limit you set).

  • Round 2 checked those fixes and raised 9 points. 7 are fixed.

  • Round 3 checked the round 2 fixes and raised 4 points. All 4 are fixed. Those last fixes have not had a further review.

  • Round 4 reviewed direct work in Worktree mode and Code Mode. Code Mode: no finding. Worktree: 6 points, 5 fixed.

  • Round 5 checked those fixes and raised 4 follow-on points. All 4 are fixed.

Two round 2 points were not acted on:

  • Amendment entries saved by a "pre-fix build": none exist, because amendments have never shipped.
  • A model's provider removed between a hold and its approval leaves that member paused until the change is retried successfully.

Checks

  • pnpm typecheck --force: 29 of 29 tasks pass.
  • Architect plugin: 820 tests pass. Orchestrator plugin: 1846 pass, 1 skipped. Desktop: 3203 pass, 5 skipped.
  • No changed source file is over 500 lines.
  • @sero-ai/common is bumped to 0.24.0 (grant amendment API, toolsAreLoadout). It is not published.

How to test

  1. Build and run the desktop app from this branch.
  2. Start an Architect project in Workspace mode on a small repo and ask for a small fix. Watch for a milestone the Architect works on itself, then evidence.
  3. Do the same in Worktree mode. The work lands on a branch named after the milestone, in .sero/worktrees/ until the milestone is delivered.
  4. On a new project, read the start approval: it lists Code Mode. Ask the owner to use codemode for a multi-step read.
  5. In a running Room, ask the Conductor to change a member's model, then to add a tool the member was not approved for. Open the Team view and approve or decline.
  6. Open the Architect Inspector and read the Limits list.

Found on the way, filed separately

Phase 0 of improve-architect-orchestrator-autonomy (issue #620).

- Baseline records carry run identity, strategy, matched inputs, checks,
  interventions, recoveries and owner tokens per turn.
- An outcome comes only from the task's independent checks. A record is
  never compared with itself, and mismatched inputs are named.
- The e2e runner covers five scenarios for the current Architect and a
  persistent single agent on one global tier, with a spend bound.
- A contract spec proves the manifests and checks before any paid run.
- Restriction map verified against source at ba4c3fb.
Approved prototype for task 2.1 of improve-architect-orchestrator-autonomy.
Task 2.2 of improve-architect-orchestrator-autonomy. A milestone can link
to work the owner does itself, with begin, continue, report and interrupt
rules. Older records load unchanged.
Tasks 2.3 to 2.5 of improve-architect-orchestrator-autonomy, in part.

- A work action begins, continues and reports the owner's own work
  against a saved execution identity.
- Continue is a wake outcome. The scheduler wakes the owner again only
  while the project may start work, after directives and work events.
- Evidence runs on a direct completion claim. Acceptance and delivery
  keep their own steps, and workspace-files is the only direct receipt.
- Direct and delegated writers hold each other out of the folder.
- A stopped turn or a restart marks the work interrupted, never done.

A Worktree project is refused for now: the owner's access covers the
project folder only.
A paused, blocked or capped project kept its direct work marked running
after a restart, with no turn behind it. The interrupt now runs for every
project. The outcome rule also names work continue everywhere the owner
reads it, and runtime tests cover continuation, the cap and restart.
Tasks 2.6 and 2.7 of improve-architect-orchestrator-autonomy.

- The milestone rail lines up kind, status and link in three columns. A
  direct milestone reads architect and links to Watch work or Evidence.
- The Live tab and the overview name the milestone the Architect is doing.
- A turn spent on direct work is charged once, to the run the work
  started under, and its wake names the execution.
- The run inspector counts continuations in a row that changed no file.
Task 2.8 of improve-architect-orchestrator-autonomy.
Task 5.1 of improve-architect-orchestrator-autonomy.
…kspace

The Architect refused a folder that already existed. Both candidates now
get the same committed starting files, and the Architect starts on a
workspace registered for that folder.
…h for managed sessions and workers

A session now registers every tool its approval allows and declares only
the initial loadout. Pi's tool_search loads the rest in the same session.
A tool outside the approval is never registered, so it cannot be found or
called. Tools loaded by search come back when a session reopens.
…runs

Adds amendGrant to the persistent sessions API with a typed request and
result, and toolsAreLoadout on subagent run params. Workflow steps now
pass the planner's tool picks as a starting loadout. Bumps
@sero-ai/common to 0.24.0.
An Architect owner or a Goal can register a wait on linked child work
with an optional deadline. The existing driver reserves and consumes one
wake when the condition is met, after its usual authority, budget and
ownership checks. Old reason-only waits stay manual. Process and CI
sources are refused for now: no existing seam reports them by identity.
amendGrant changes subject policies or retires a subject under the grant
store's lock, keeping the grant id, session directory and bindings. A
change inside the current approval applies with no dialog. Added
authority is held, or approved by the user when asked. Results are stored
by amendment id, so a repeat changes nothing.
Approve matched the charter, plan and cap buttons of whichever project
was on screen, so a run could press them on an older project.
A member's model, thinking, tools or skills, an added member and a
replacement now amend the Room's existing grant. The intent is saved
first, the member reaches the end of its turn, the host applies the
change, and the same session reopens before the revision is marked
applied. Added authority holds for approval. A restart finishes or holds
the same revision. Workflow steps pass their tool picks as a loadout.
…ser set

Each Goal limit records who set it. An agent may tighten a user limit or
set one the user left unset. A Goal saved before this is read as
user-set.
The approved set is recorded apart from what a turn loads. An owner that
already holds a grant asks again for exactly what it had.
A turn that keeps producing events may run on. Ten silent minutes steer
the owner to save its work and end the turn. Five more abort it as an
interruption. A second stall in a row holds the project. Adds the list
of effective limits with who set each.
The Live tab shows what the project waits for, its deadline, and active
and waiting time as separate numbers. An expired wait reads as on hold,
never complete. The Inspector lists each effective limit with who set it.
…ed member changes

A held change shows what it adds with Approve and Decline. A retired
member keeps a History link and its replacement names the handover.
… runs sero from the shell

The shell path is closed to Architect owners and Room members. The
refusal read Unauthorized session, and a paid pilot run spent its whole
turn probing it and never recorded an outcome.
@github-actions

github-actions Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

React Doctor found 5 new issues in 4 files · 5 warnings · score 91 / 100 (Great) · 0 fixed · vs main

5 warnings

shared/__tests__/direct-execution.test.ts

  • ⚠️ L65 JSON parse/stringify deep clone no-json-parse-stringify-clone

ui/components/InspectorLimits.tsx

  • ⚠️ L10 Role used instead of HTML tag prefer-tag-over-role

ui/components/RoomDetail.tsx

  • ⚠️ L60 Large component is hard to read and change no-giant-component

ui/lib/room-team.ts

  • ⚠️ L52 Array lookup inside a loop js-set-map-lookups
  • ⚠️ L53 Array lookup inside a loop js-set-map-lookups

Reviewed by React Doctor for commit f05a940. See inline comments for fixes.

@monobyte

monobyte commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

Routed review, round 1 (gpt-6-astra, high effort, three seams)

Only defects that fail now on a reachable path were asked for. 17 findings. Fixes follow in this PR.

Host: grant amendments and tool loadout

  1. Major — Disabled tools can be restored by a worker’s initial loadout.
    File: apps/desktop/electron/features/subagent/runtime/session-policy.ts:192–205
    Trigger: Run an Orchestrator background step with toolsAreLoadout: true, a planner-selected tool such as web_search in tools, and that same tool in the user’s disabledTools.
    What goes wrong: Only policyTools is filtered against disabled tools; ...allowlist adds the disabled tool straight back into the authorized set. Plugin extensions still provide its implementation, so the worker can execute it. The subsequent initial-tool list also retains it. Separately, tool_search is added whenever tools are deferred, even if the user disabled it.
    Fix: Remove disabled names from the final authorized set and initial loadout, including automatically added tools, and pass them through the SDK’s excludeTools boundary.

  2. Major — An in-flight session can commit after its subject is retired or its authority is narrowed.
    File: apps/desktop/electron/features/apps/runtime/capabilities/persistent-sessions/grant-store.ts:248–252, 343
    Trigger: Start create or open and let validation and reservation succeed. While asynchronous session construction is running, apply an amendment that retires the subject or removes its write access. Then allow construction to finish.
    What goes wrong: Retirement is checked only during validation/reservation. Neither commit path checks retirement or whether the validated policy changed. Construction retains the old policy object, so a newly returned session can still have the removed tools. For creation, the store also durably binds the retired subject’s session and increments its lifetime count. Subsequent prompting checks only whether the grant is active, so this session remains usable.
    Fix: Record the validated grant revision with each reservation and, under the commit lock, reject/dispose construction if retirement or a policy change invalidated it.

  3. Minor — Reusing an amendment ID for a different change falsely acknowledges the old change.
    File: apps/desktop/electron/features/apps/runtime/capabilities/persistent-sessions/amendment-runner.ts:20–22; host.ts:198–200
    Trigger: Apply amendment ID a to narrow subject A. Submit another amendment with ID a that instead retires subject B, even with the correct current revision. The same problem occurs when conflicting requests arrive during an approval dialog.
    What goes wrong: Both the stored-result cache and the in-flight cache use only the grant/amendment IDs. They never compare the requested subjects or retirement list. The second request receives applied for the first change although B was not retired; its requested change is silently ignored.
    Fix: Bind each amendment ID to a fingerprint of its subject changes and retirement list, and refuse conflicting reuse while allowing genuine retries.

Orchestrator: Room amendments, Goal waits and limits

  1. Major — An agent can raise a user-set Goal limit through the management tool.
    File: plugins/sero-orchestrator-plugin/extension/goal-app.ts:72
    Trigger: The user sets a Goal’s turn limit to 10. A chat agent invokes the bridged goals tool with action: "set_limits", that Goal’s ID, and maxTurns: 100.
    What goes wrong: goals is exposed to chat sessions through the CLI bridge, but this handler unconditionally passes 'user' to setLimits. The origin check is therefore bypassed, and the increased limit is saved without user approval. Unlike the ordinary goal tool, this route also accepts another session’s Goal ID.
    Fix: Require authenticated UI authority for user-origin limit changes; treat agent-accessible bridge calls as agent-origin requests.

  2. Major — A wait wake can overwrite a completed pause or stop.
    File: plugins/sero-orchestrator-plugin/runtime/goals/goal-wait-watcher.ts:97–103
    Trigger: A Workflow finishes while its Goal is waiting. The watcher reads the waiting Goal and awaits claim(). During that await, the user pauses or stops the Goal and that operation finishes saving. The watcher then completes.
    What goes wrong: The watcher saves an activated copy of its earlier snapshot. This overwrites the pause—or removes the stop’s closedAt, incremented control revision, and cancelled waits—and signals another turn. The store serializes writes, not these read/modify/write operations, so the control-revision checks on the stale snapshot do not prevent resurrection.
    Fix: Serialize wake decisions with Goal controls and commit only after rechecking the current status, control revision, and limits.

  3. Major — A clamped model change leaves the host changed while the Room resumes the old configuration.
    File: plugins/sero-orchestrator-plugin/runtime/rooms/room-amendment.ts:289–292
    Trigger: For an existing member with an open session, propose change-configuration with a model name unavailable to the host. The parser accepts it; the host clamps the requested model list to empty and commits the amendment as a narrowing.
    What goes wrong: clampProblem notices the missing model only after the host commit. The Room leaves the amendment in the intent phase, marks it held with workPaused: false, and keeps its old configuration. Its cached session can continue using a model no longer in its grant; after eviction or restart, that configuration cannot reopen. Decline is also allowed here, despite the host change already being committed.
    Fix: Reject unusable clamped policies before host commit, and never resume or locally decline an old configuration after an applied host response.

  4. Major — Pending roster additions can exceed the approved member limit.
    File: plugins/sero-orchestrator-plugin/runtime/rooms/room-amendment.ts:169–174
    Trigger: A running Room has one member slot left and delegation authority permitting two proposed newcomers. Submit both add-member revisions before the first finishes applying.
    What goes wrong: Both proposals pass the roster limit check because pending additions are not counted. The amendment queue subsequently applies both: it rebases the second host request but never revalidates the roster constraints. The Room durably exceeds maxMembers without an approved limit increase. The same missing reservation also allows two pending proposals to claim the same new member key.
    Fix: Reserve member keys and roster budgets when saving intent, and revalidate unattempted amendments before committing them to the host.

  5. Major — A Room reports a requested tool change as applied even when the host removed that tool.
    File: plugins/sero-orchestrator-plugin/runtime/rooms/room-amendment.ts:153–157
    Trigger: Propose adding an unavailable tool, or adding a write tool to a read-only member. The host removes that tool during clamping and returns the remaining policy as applied.
    What goes wrong: clampProblem checks models, thinking levels, and skills, but not tools. The Room saves the requested tool list into configuration.tools, opens the session using the smaller grantedTools list, and reports the revision as applied. The user and Conductor are told the member received a tool it cannot use.
    Fix: Compare requested tools with the granted set and report the actual accepted configuration instead of marking the full request applied.

  6. Major — One wait completion can queue two Goal continuations.
    File: plugins/sero-orchestrator-plugin/extension/goal-loop.ts:214–218
    Trigger: A wait completes while agent_settled is processing the turn that registered it. That handler has already cleared boundaryOpen, but is still awaiting Goal reads or accounting writes. The watcher saves the Goal as active and sends its wake notification during this interval.
    What goes wrong: The wake listener sees boundaryOpen === false and calls startTurn. The still-running settled handler also observes the active Goal and reaches its own startTurn call. There is no shared continuation reservation or queued-turn check, so one wait completion produces two continuation messages and can run an extra turn without a limit check between them.
    Fix: Keep the boundary claimed until settled processing finishes and make all kickoff paths share one atomic continuation reservation.

Architect: direct work, waits, stall recovery

  1. Major — plugins/sero-architect-plugin/runtime/owner-session.ts:413–423: A busy owner can spend indefinitely past the user’s cost cap.
    Start direct work with a $1 cap, exceed it during the turn, and keep making tool calls less than ten minutes apart. Usage accounting changes the project to limited, but never aborts the running session. Every event resets the new watchdog, so removing the fixed timeout leaves this over-budget turn unbounded; the cap only prevents subsequent wakes.
    Fix: Enforce the cost cap against the running owner session independently of the silence watchdog.

  2. Major — plugins/sero-architect-plugin/runtime/execution-location.ts:26–29: Direct and delegated work can acquire the same milestone concurrently.
    A dispatch passes its initial availability check, then waits in resolveDispatchProject. Before it reserves pendingDispatch, the owner successfully begins direct work on that milestone. When dispatch resumes, its queued reservation checks neither the target’s current status nor its active direct execution; projectWriter excludes that target. Both executions are then recorded and permitted to edit the same files. This can occur while the user approves a previously proposed dispatch and the owner begins work.
    Fix: Recheck the target’s status and active direct execution inside the queued dispatch reservation.

  3. Major — plugins/sero-architect-plugin/shared/direct-execution.ts:156–159: A direct report bypasses an unanswered decision that parked the milestone.
    Begin direct work on m1, then raise a decision that parks m1 pending the user’s answer. Calling work report with its execution ID succeeds and changes parked to verifying. Evidence and acceptance can then complete the milestone without that answer. Repeating begin also succeeds because the active-execution shortcut precedes the parked-status check.
    Fix: Reject begin, continue and report operations on decision-parked milestones until their blocking decisions are answered.

  4. Major — plugins/sero-architect-plugin/shared/direct-execution.ts:90: Changing requirements permanently traps an active direct execution.
    Begin a milestone at working revision 1, then use working to add or change a criterion, producing revision 2. Reporting correctly refuses the old revision and instructs the owner to begin again. But every subsequent begin returns the same revision-1 execution. There is no action that supersedes or cancels it, and its active writer reservation also prevents other milestones and evidence from using the folder.
    Fix: Allow an explicit restart against changed requirements to supersede the old execution and create a new identity.

  5. Major — plugins/sero-architect-plugin/runtime/index.ts:38–45: Restarting after a direct completion report can leave the project stuck permanently.
    In an already-started project with no other pending work, successfully report the direct milestone, then close or crash Sero before requesting evidence. On restart, the execution is reported, so it is not considered active; the milestone is verifying without passed evidence, so plannedWorkRemains also returns false. There is no delegated completion event or pending evidence operation to wake the owner. Verification never starts without manual intervention.
    Fix: Recover reported direct milestones that still need evidence as actionable owner work.

  6. Major — plugins/sero-architect-plugin/runtime/index.ts:240–250: A wait wake can be permanently lost before its turn starts.
    Let a deadline expire while its Workflow is still active. The runtime durably marks the wake consumed, then crashes while opening the owner session, before sending its prompt. On restart, the consumed wait is not retried, the still-active Workflow produces no completion transition, and its running milestone suppresses the quiet-work wake. The promised deadline notification is lost.
    Fix: Associate consumption with a durable turn identity and recover deliveries whose owner turn never started.

  7. Major — plugins/sero-architect-plugin/shared/waits.ts:177–181: A reserved wait can start a paid turn after the project is paused or capped.
    Let delivery pass its initial eligibility check, then pause the project or exhaust its budget before consume runs—for example, while delivery awaits maintenance setup. Consumption checks only the stop revision, not waitMayWake. It consumes the reservation and starts runTurn using the now-paused or over-budget record; that method does not recheck eligibility.
    Fix: Recheck waitMayWake against the fresh record when consuming the reservation, leaving it pending when work is prohibited.

  8. Minor — plugins/sero-architect-plugin/runtime/wait-reconciler.ts:125–126: One child completion produces two owner wakes.
    Register a wait on a running milestone’s Workflow or Room, then let it complete normally. The dispatch watcher invokes this reconciler, which queues a wait wake, and then independently queues its ordinary dispatch-complete wake for the same completion. The scheduler merges only wakes of the same kind, so both become paid owner turns.
    Fix: Combine the matched wait and dispatch notification into one wake, or suppress the ordinary completion wake when the wait already covers it.

A run from source reports sero-cli with a path inside the desktop app's
own package. The catalogue then treated it as a plugin tool and dropped
it from every managed session's approval, so an Architect owner had no
way to record an action. Found by the paid pilot.
@monobyte

monobyte commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

Routed review, round 3: delta check of the round 2 fixes (gpt-6-astra, high effort)

Round 2 fixes: bba8550 (Orchestrator), 8af3e24 (Architect). The host seam had nothing left to recheck.

Orchestrator

Point 2 — Fixed, but the fix introduced two defects. The watcher now preserves the session claim when a user resume has already made the Goal active.

  • Major — Deleting the Goal during the pending claim leaves the session permanently claimed.
    File: plugins/sero-orchestrator-plugin/runtime/goals/goal-wait-watcher.ts:139–142
    Trigger: The watcher awaits its session lookup. The user stops and deletes the Goal. The lookup then finishes and the watcher claims the session.
    What goes wrong: Both subsequent store.update calls return null because the Goal was deleted. The cleanup callback never runs, so the claim is never released. New Goals and Workflows cannot acquire that session, and the deleted Goal can no longer be stopped to release it.
    Fix: Explicitly release the acquired claim when the serialized cleanup finds no Goal record.

  • Major — The new cleanup wait can finish after Stop, then dispatch the cancelled wake.
    File: plugins/sero-orchestrator-plugin/runtime/goals/goal-wait-watcher.ts:139–146
    Trigger: The watcher activates the Goal. While that write is finishing, the user’s Stop operation enters the store queue ahead of the new cleanup update. Stop finishes, and cleanup reads the stopped Goal.
    What goes wrong: final correctly contains the stopped record, but notification still uses the earlier active woke.goal. With the session idle, the Goal-loop listener starts another automatic turn despite the completed Stop.
    Fix: Dispatch the notification only within a serialized recheck that confirms the Goal is still active and the wake is still valid.

Architect

Two points remain not fixed. The other six are fixed.

1. Cost-cap enforcement — Not fixed (major)

Location: plugins/sero-architect-plugin/runtime/owner-usage.ts:45–57; runtime/owner-session.ts:289–291; runtime/index.ts:263.

Trigger: An ordinary work wake is admitted while the project is under its cap. While model selection or session opening is awaiting completion, the user lowers the cap below current spend. Session preparation then reads that already-limited record into turnRecord.

What goes wrong: startedUnderCap becomes false, permanently disabling the turn’s cap-abort check. Only wait wakes receive the new pre-prompt eligibility check, so this ordinary work wake still starts and can keep spending without being aborted at subsequent usage reads. It incorrectly receives the exemption intended for explicit directive/decision wakes.

The previously reported sequence where the cap changes after preparation is fixed; the preparation race remains.

One-line fix: Recheck eligibility before prompting every ordinary work wake, and grant the already-over-cap exception explicitly by wake type rather than inferring it from turnRecord.

2. Duplicate Room completion/receipt wakes — Not fixed (minor)

Location: plugins/sero-architect-plugin/runtime/dispatch-watch.ts:318–352.

Trigger: Register a completion wait on a Room. The Room publishes its delivery receipt while its status is still completing, and Architect processes that update before the subsequent completed update.

This is the actual completion ordering: plugins/sero-orchestrator-plugin/runtime/rooms/room-completion.ts:48–60 delivers first, releases authority, then records completed. The Room index exposes both the intermediate status and receipt.

What goes wrong: At the receipt update, the wait is still open, so waitCoversCompletion returns false and the receipt generates a dispatch-complete wake. The later completed update satisfies the wait and generates a separate wait wake. Different wake kinds do not coalesce, so one Room completion still causes two paid turns.

The fix works when Architect observes the receipt and completed status together, but not when it observes the real intermediate state.

One-line fix: When a Room has an outstanding completion wait, retain its receipt without issuing a separate receipt wake, and carry it in the eventual wait outcome.

Fixed points

  1. Concurrent direct and delegated writers — Fixed.
    The queued dispatch reservation still rejects a target owned by an active direct execution. The second-round changes do not remove that guard.

  2. Bypassing decision parking — Fixed.
    Direct begin, continue and report retain the decision-parking checks.

  3. Requirement changes trapping direct work on a previously delegated milestone — Fixed.
    Beginning direct repair marks the retained dispatch as finished. Replacing the direct execution after a requirements change no longer mistakes that historical dispatch for a live delegate.

  4. Restart after replacement direct completion — Fixed.
    Beginning replacement work marks previous evidence stale, so reporting and restarting before new evidence now qualifies for recovery.

  5. Wait acknowledgement delayed until the turn finishes — Fixed.
    Acknowledgement now starts on the host’s actual turn_start event, rather than waiting for the whole api.prompt call to finish.

  6. Wait wake starting after pause or cap during preparation — Fixed.
    Wait delivery now rereads eligibility immediately before prompting. A refused start leaves the wake unacknowledged and returns its reservation to pending.

The three focused test files passed: 64 tests. The two remaining failures above were traced through the current source. No source files were changed.

Triage

All four remaining points are being fixed in this PR.

A wait watcher releases its claim when the Goal was deleted and sends no
wake after a Stop. Only directive and decision wakes are exempt from the
cap, and every work wake is rechecked before the prompt. A Room receipt
that arrives before completion no longer adds a second turn.
An Architect owner or a Room member whose approval names codemode now
loads it, switched on from the first turn because tool search cannot
find it. A script can call only the tools the session registered, so a
read-only member's script cannot reach a write tool.
…m can give it to a member

A new Architect owner names codemode at the start approval when the
catalogue offers it. An owner that already holds a grant asks again for
exactly what it had. The session request loads only the five owner
tools. Docs say what is now true.
The owner gets a managed checkout per milestone, saved with the
execution before any file changes. Report commits the checkout to its
branch, and evidence runs there. A worktree execution is not a
project-folder writer. An accepted or parked milestone releases its
checkout without force. The pilot runner can start a Worktree project.
@monobyte

monobyte commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator Author

Review round 4: direct work in Worktree mode, and Code Mode for managed sessions

Routed review on gpt-6-astra at high effort, of commits 93d3f033c, a094567b7 and 2fce1ab2b.

Code Mode (93d3f033c, a094567b7)

No qualifying defects found.

Direct work in Worktree mode (2fce1ab2b)

  1. P1 — Preview evidence runs against the project folder, not the worktree.
    plugins/sero-architect-plugin/runtime/services.ts:157
    Request preview evidence for a direct Worktree execution. Replacing record.folder makes runPreviewCapture pass the worktree as both workspacePath and cwdPath. The managed-server API translates that pair to /workspace, but the unchanged workspace ID maps /workspace to the original project folder. It therefore starts—or reuses—the project-folder server. A greenfield preview fails; an existing app can supply passing screenshots of the wrong code.
    Fix: Keep the registered project folder as workspacePath and pass the worktree separately as the server’s cwdPath.

  2. P1 — Starting work again during evidence can accept changes that were never checked or reported.
    plugins/sero-architect-plugin/runtime/owner-direct.ts:58–62
    Report a milestone, start background evidence, then call begin again for that milestone while evidence is running. The Worktree guard ignores pendingEvidence, and the new execution reuses the same checkout. If the owner edits after the tests finish but before the evidence runner fingerprints the files—for example, during a final waiting command—the old passing test results get paired with the new files’ fingerprint. The evidence completion overwrites the new execution’s milestone status with verifying; acceptance then succeeds even though that execution never reported and its changes were not tested.
    Fix: Refuse begin while evidence for that milestone is pending, including in the queued freshness check.

  3. P1 — A repeated begin can durably record the project folder as a worktree.
    plugins/sero-architect-plugin/runtime/owner-direct.ts:97–115
    With an active direct execution, issue a repeated begin concurrently with a working action that changes the requirement revision. The initial snapshot says this is the same execution, so checkout opening is skipped and placement defaults to record.folder. If the revision update lands while readState runs, beginDirectExecution sees the changed revision and creates a replacement execution using that fallback placement: mode worktree, directory equal to the project folder, and no branch. Subsequent instructions and checkpoints operate on the project folder rather than the managed checkout.
    Fix: Use the saved execution’s placement for repeated begins and retry checkout/state resolution if the revision changes before the queued write.

  4. P2 — A report can record a branch that does not contain the reported work.
    plugins/sero-architect-plugin/runtime/direct-worktree.ts:56–66
    Begin on branch A, then use the owner’s shell tools to switch the managed checkout to branch B and make changes there. ensure only checks that the directory exists; checkpoint commits whichever branch is currently checked out. Reporting with --destination workspace-files succeeds and records A as the receipt, while evidence checks B’s files. Acceptance can therefore mark A delivered although the implementation exists only on B.
    Fix: Verify the checkout’s repository and symbolic HEAD against the saved placement before checkpointing or running evidence; refuse a different branch or detached HEAD.

  5. P2 — Parking an unchanged milestone deletes the branch needed to resume it.
    plugins/sero-architect-plugin/runtime/direct-worktree.ts:122
    Begin a milestone in a freshly bootstrapped project, then park it with a decision before editing files. The checkpoint has nothing to commit, so the branch still equals the default branch. Cleanup removes the checkout and successfully deletes its “merged” branch. After the user answers, continue and report try to restore that now-nonexistent branch and fail. A repeated begin merely returns the existing execution and missing directory.
    Fix: Retain branches for parked, resumable executions instead of requesting merged-branch deletion.

  6. P2 — Delivery cleanup deletes the saved preview evidence.
    plugins/sero-architect-plugin/runtime/services.ts:157
    Complete a successful preview evidence run, report a delivery destination, accept the milestone, and end the wake. Passing the worktree as record.folder also places the screenshot under <worktree>/.sero/apps/architect/evidence/. That directory is ignored by Git, so the release checkpoint does not preserve it, and normal worktree removal deletes it. The durable evidence record still points to a screenshot that no longer exists.
    Fix: Save evidence artifacts under the permanent project/state directory, independently of the checkout being tested.

Triage

  • 1, 2, 3, 5 and 6 are being fixed.
  • 4 is not acted on. The owner session has no mutating git command: the shell guard refuses git switch and git checkout, so the owner cannot move the managed checkout to another branch.

… and a repeated begin

Review round 4 on PR 624. The preview server now runs in the checkout while
the workspace stays the project folder, and its screenshot is kept in the
project folder. Begin is refused while evidence for the milestone is running.
A repeated begin keeps the saved checkout. A parked milestone keeps its branch.
The pilot runner reads a released checkout from its branch.
… a fresh preview server

Review round 5 on PR 624. A preview server started for a checkout is stopped
after the capture. Evidence is refused while the owner's own work on the
milestone is running. A repeated begin restores a released checkout. A checkout
on any branch but its saved one is refused.
@monobyte

monobyte commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

Review round 5: re-check of the round 4 fixes (bd13f6d)

Same reviewer session, gpt-6-astra at high effort.

  1. Point 1 — Fixed, but the fix exposes stale preview-server reuse after checkout removal. P1.
    plugins/sero-architect-plugin/runtime/preview-capture.ts:76
    The workspace-root translation is now correct. However, run preview evidence using a Node server that keeps its application code in memory, then park the milestone. Cleanup removes the checkout without stopping its server. After the user answers, begin again, change the restored checkout, report, and request preview evidence. HostDevServerManager.start (host-dev-server-manager.ts:87–90) reuses the old running process because its command and directory string match. The screenshot can therefore verify the previous implementation rather than the restored checkout’s current code.
    Fix: Stop the checkout’s managed preview server before releasing the checkout and start a fresh server after restoration.

  2. Point 2 — Not fixed for concurrent begin/evidence calls. P1.
    plugins/sero-architect-plugin/runtime/services.ts:418–424
    The original sequential case is blocked, but the opposite reservation order remains possible. Issue begin and evidence concurrently for a reported milestone. The evidence action can read the old reported record while begin’s queued write is underway, then reserve evidence after begin saves the new running execution. That reservation checks only project-folder writers, not whether this worktree execution still has the report the request validated. It changes the running milestone to verifying and starts evidence. Edits made after its tests but before fingerprinting can still acquire passing evidence and be accepted without the replacement execution reporting.
    Fix: Inside the evidence reservation’s queued write, revalidate the reported execution’s identity and state; refuse if begin replaced it.

  3. Point 3 — Fixed, but repeated begin now fails after parking releases the checkout. P2.
    plugins/sero-architect-plugin/runtime/owner-direct.ts:101–107
    The revision race no longer saves the project folder as a worktree. However, begin a milestone, park it before reporting, let cleanup release its checkout, and have the user answer without changing the requirement revision. A repeated begin takes the same path, skips checkout restoration, and now fingerprints the saved—but nonexistent—directory. The action throws a Git fingerprint error instead of returning the existing execution. continue can recover, but repeated begin itself is broken on this normal lifecycle path.
    Fix: Ensure the saved worktree exists before reading its state on the same path.

  4. Point 4 — Not fixed; the shell guard has a reachable absolute-path bypass. P2.
    apps/desktop/electron/platform/security/git-command-filter.ts:110–111
    The guard does reject bare git switch, but its matcher requires git immediately after whitespace or the start of the command. An owner shell command such as cd <checkout> && /usr/bin/git switch -c branch-b does not match and is not classified as mutating Git. After editing there, report checkpoints B while the saved placement and delivery receipt still name A; evidence checks B. This preserves the original wrong-branch delivery path.
    Fix: Verify that the checkout’s symbolic HEAD matches the recorded branch before checkpointing and reporting.

  5. Point 5 — Fixed.
    plugins/sero-architect-plugin/runtime/direct-worktree.ts:123
    Parked milestones now pass deleteMergedBranch: false. An unchanged branch survives checkout removal, so continue can restore it after the decision is answered.

  6. Point 6 — Fixed.
    plugins/sero-architect-plugin/runtime/preview-capture.ts:95–96
    Screenshots now live under the permanent project folder, while their filename uses the checkout’s commit. Releasing the worktree no longer deletes the recorded screenshot.

Triage

Points 1 to 4 are fixed in the next commit. Point 4 now has a check: a checkout that is not on its saved branch is refused at continue, report and evidence. The absolute-path gap in the shell guard itself is added to #625.

@monobyte

monobyte commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

Review round 6: re-check of the round 5 fixes (cec47f6)

Same reviewer session, gpt-6-astra at high effort. All four points are fixed, and no new defect was found in the fixes or what they touch.

sero-cli returns a failed command as text with a non-zero exit code, and the
agent loop reported it as a success. A Code Mode script only stops on a call
the loop reports as failed, so it ran every later line. A live script ran past
a blocked click and then past the 50 commands per turn limit, getting the same
refusal for each remaining call.
@monobyte

monobyte commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator Author

Found in manual testing: a Code Mode script did not stop on a failed sero command

A script that drives the browser through sero-cli kept running after a click was refused, and then after the 50 commands per turn limit, so every remaining call returned the same refusal.

Cause: sero-cli returns a failed command as text with a non-zero exit code. Pi stops a script only on a call the agent loop reports as failed, and the loop reported these as successes. Fixed by marking a non-zero sero-cli exit as a failed call at the same hook that already does this for bash. It applies to chat sessions, managed sessions and workers.

Not changed: the limit of 50 sero commands per turn still counts each call a script makes.

… a script issues

How many commands a Code Mode script runs is the agent's decision. A call
another tool made carries Pi's nested call id, and the limit skips it. A
command the model issues itself is still counted.
…over several tool calls in a row

The rule is added to the system prompt of chat sessions, managed sessions and
workers, and only when codemode is switched on for that session. It also says
when separate calls are right: when a result must be read before the next step
can be decided.
…a Code Mode script stops on it

The earlier fix marked the failure at agent.afterToolCall, which Pi does not
run for a call a script makes. The tool now returns isError itself, as bash
does. A real-session test runs a script through Pi's codemode and the real
sero-cli tool: the line after a failed command never runs, and 60 commands in
one script run with no rate limit.
…answering

A page stuck in its own code, such as an endless loop, blocks the browser
daemon: every later command waits out its limit, and so do close and launch.
When a command times out and the session cannot answer a cheap question, the
daemon and its browser are stopped and the agent is told the page is stuck.
A slow command on a session that still answers is left alone.

A test runs the real tool against a real looping page when a browser pack is
present: red without the reset, green with it.
…ser is reset on Windows too

sero-cli split its input at every line break, so a script passed to
--expression "..." on several lines became several broken commands
("Unterminated quoted string"). A line break inside quotes now stays in the
argument.

The stuck-browser reset now runs on a Windows host through taskkill. That
path is written from the Git Bash rules and has not been run on Windows.
A host workspace may only touch files inside its own roots. The screenshot
went to the browser pack's temp folder, which is not one, so every screenshot
failed with "Host path must be inside a workspace root". A recording was
given /workspace, which does not exist on a host, so it could not be saved.
Both now use the workspace's real path.

The tests for this tool faked the runtime, so the host rule never ran. The
real-browser test now runs screenshot and recording through the real host
backend.
…seconds

The browser's own "wait for load" waits for an event that has already
passed once the page is open, and gives up after its 25s limit. Every launch
or navigate that asked for load paid that: measured at 26s per call, which
ran a seven-viewport script past its time limit. The page's ready state is
asked for instead.

The session helpers move to their own file to keep the tool under 500 lines.

it('reads a milestone saved before direct execution unchanged', () => {
const saved = milestone({ status: 'verifying', verification: 'reported' });
const reloaded = JSON.parse(JSON.stringify(project([saved]))) as ProjectRecord;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

React Doctor · react-doctor/no-json-parse-stringify-clone (warning)

JSON.parse(JSON.stringify(x)) deep-clones by re-serializing: it is slow on large objects and silently drops undefined, functions, Date/Map/Set, and cyclic references. Use structuredClone(x).

Fix → Replace JSON.parse(JSON.stringify(value)) with structuredClone(value). It is faster and preserves Dates, Maps, Sets, and cyclic references.

Docs

return (
<section aria-label="Limits">
<div className="ar-lim-title"><span>Limits</span><span>{rows.length}</span></div>
<div className="ar-lim-card" role="table" aria-label="Limits">

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

React Doctor · react-doctor/prefer-tag-over-role (warning)

Screen reader users get more reliable semantics from <table> than role="table", so use <table> instead.

Fix → Use the matching HTML element when one exists so browsers and assistive tech get native semantics.

Docs

/** What a tool change does to the list, as the drawing words it: `add shell`. */
function toolsMove(current: string[], next: string | string[] | undefined): string | null {
if (!Array.isArray(next)) return null;
const added = next.filter((tool) => !current.includes(tool)).map((tool) => `add ${tool}`);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

React Doctor · react-doctor/js-set-map-lookups (warning)

This scales poorly because array.includes() inside a loop scans the whole list every time. Use a Set for constant-time lookups.

Fix → Use a Set or Map when you check for the same items over and over. Array.includes/find scans the whole list each time

Docs

function toolsMove(current: string[], next: string | string[] | undefined): string | null {
if (!Array.isArray(next)) return null;
const added = next.filter((tool) => !current.includes(tool)).map((tool) => `add ${tool}`);
const removed = current.filter((tool) => !next.includes(tool)).map((tool) => `remove ${tool}`);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

React Doctor · react-doctor/js-set-map-lookups (warning)

This scales poorly because array.includes() inside a loop scans the whole list every time. Use a Set for constant-time lookups.

Fix → Use a Set or Map when you check for the same items over and over. Array.includes/find scans the whole list each time

Docs

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant