Repository navigation
Conversation
… orchestrator automation improvements
Phase 0 of improve-architect-orchestrator-autonomy (issue #620). - Baseline records carry run identity, strategy, matched inputs, checks, interventions, recoveries and owner tokens per turn. - An outcome comes only from the task's independent checks. A record is never compared with itself, and mismatched inputs are named. - The e2e runner covers five scenarios for the current Architect and a persistent single agent on one global tier, with a spend bound. - A contract spec proves the manifests and checks before any paid run. - Restriction map verified against source at ba4c3fb.
Approved prototype for task 2.1 of improve-architect-orchestrator-autonomy.
Task 2.2 of improve-architect-orchestrator-autonomy. A milestone can link to work the owner does itself, with begin, continue, report and interrupt rules. Older records load unchanged.
Tasks 2.3 to 2.5 of improve-architect-orchestrator-autonomy, in part. - A work action begins, continues and reports the owner's own work against a saved execution identity. - Continue is a wake outcome. The scheduler wakes the owner again only while the project may start work, after directives and work events. - Evidence runs on a direct completion claim. Acceptance and delivery keep their own steps, and workspace-files is the only direct receipt. - Direct and delegated writers hold each other out of the folder. - A stopped turn or a restart marks the work interrupted, never done. A Worktree project is refused for now: the owner's access covers the project folder only.
A paused, blocked or capped project kept its direct work marked running after a restart, with no turn behind it. The interrupt now runs for every project. The outcome rule also names work continue everywhere the owner reads it, and runtime tests cover continuation, the cap and restart.
Tasks 2.6 and 2.7 of improve-architect-orchestrator-autonomy. - The milestone rail lines up kind, status and link in three columns. A direct milestone reads architect and links to Watch work or Evidence. - The Live tab and the overview name the milestone the Architect is doing. - A turn spent on direct work is charged once, to the run the work started under, and its wake names the execution. - The run inspector counts continuations in a row that changed no file.
Task 2.8 of improve-architect-orchestrator-autonomy.
Task 5.1 of improve-architect-orchestrator-autonomy.
…kspace The Architect refused a folder that already existed. Both candidates now get the same committed starting files, and the Architect starts on a workspace registered for that folder.
…h for managed sessions and workers A session now registers every tool its approval allows and declares only the initial loadout. Pi's tool_search loads the rest in the same session. A tool outside the approval is never registered, so it cannot be found or called. Tools loaded by search come back when a session reopens.
…runs Adds amendGrant to the persistent sessions API with a typed request and result, and toolsAreLoadout on subagent run params. Workflow steps now pass the planner's tool picks as a starting loadout. Bumps @sero-ai/common to 0.24.0.
An Architect owner or a Goal can register a wait on linked child work with an optional deadline. The existing driver reserves and consumes one wake when the condition is met, after its usual authority, budget and ownership checks. Old reason-only waits stay manual. Process and CI sources are refused for now: no existing seam reports them by identity.
amendGrant changes subject policies or retires a subject under the grant store's lock, keeping the grant id, session directory and bindings. A change inside the current approval applies with no dialog. Added authority is held, or approved by the user when asked. Results are stored by amendment id, so a repeat changes nothing.
Approve matched the charter, plan and cap buttons of whichever project was on screen, so a run could press them on an older project.
A member's model, thinking, tools or skills, an added member and a replacement now amend the Room's existing grant. The intent is saved first, the member reaches the end of its turn, the host applies the change, and the same session reopens before the revision is marked applied. Added authority holds for approval. A restart finishes or holds the same revision. Workflow steps pass their tool picks as a loadout.
…ser set Each Goal limit records who set it. An agent may tighten a user limit or set one the user left unset. A Goal saved before this is read as user-set.
The approved set is recorded apart from what a turn loads. An owner that already holds a grant asks again for exactly what it had.
A turn that keeps producing events may run on. Ten silent minutes steer the owner to save its work and end the turn. Five more abort it as an interruption. A second stall in a row holds the project. Adds the list of effective limits with who set each.
The Live tab shows what the project waits for, its deadline, and active and waiting time as separate numbers. An expired wait reads as on hold, never complete. The Inspector lists each effective limit with who set it.
…ed member changes A held change shows what it adds with Approve and Decline. A retired member keeps a History link and its replacement names the handover.
… runs sero from the shell The shell path is closed to Architect owners and Room members. The refusal read Unauthorized session, and a paid pilot run spent its whole turn probing it and never recorded an outcome.
|
React Doctor found 5 new issues in 4 files · 5 warnings · score 91 / 100 (Great) · 0 fixed · vs 5 warnings
Reviewed by React Doctor for commit |
Routed review, round 1 (gpt-6-astra, high effort, three seams)Only defects that fail now on a reachable path were asked for. 17 findings. Fixes follow in this PR. Host: grant amendments and tool loadout
Orchestrator: Room amendments, Goal waits and limits
Architect: direct work, waits, stall recovery
|
A run from source reports sero-cli with a path inside the desktop app's own package. The catalogue then treated it as a plugin tool and dropped it from every managed session's approval, so an Architect owner had no way to record an action. Found by the paid pilot.
Routed review, round 3: delta check of the round 2 fixes (gpt-6-astra, high effort)Round 2 fixes: bba8550 (Orchestrator), 8af3e24 (Architect). The host seam had nothing left to recheck. OrchestratorPoint 2 — Fixed, but the fix introduced two defects. The watcher now preserves the session claim when a user resume has already made the Goal active.
ArchitectTwo points remain not fixed. The other six are fixed. 1. Cost-cap enforcement — Not fixed (major)Location: Trigger: An ordinary work wake is admitted while the project is under its cap. While model selection or session opening is awaiting completion, the user lowers the cap below current spend. Session preparation then reads that already-limited record into What goes wrong: The previously reported sequence where the cap changes after preparation is fixed; the preparation race remains. One-line fix: Recheck eligibility before prompting every ordinary work wake, and grant the already-over-cap exception explicitly by wake type rather than inferring it from 2. Duplicate Room completion/receipt wakes — Not fixed (minor)Location: Trigger: Register a completion wait on a Room. The Room publishes its delivery receipt while its status is still This is the actual completion ordering: What goes wrong: At the receipt update, the wait is still open, so The fix works when Architect observes the receipt and completed status together, but not when it observes the real intermediate state. One-line fix: When a Room has an outstanding completion wait, retain its receipt without issuing a separate receipt wake, and carry it in the eventual wait outcome. Fixed points
The three focused test files passed: 64 tests. The two remaining failures above were traced through the current source. No source files were changed. TriageAll four remaining points are being fixed in this PR. |
A wait watcher releases its claim when the Goal was deleted and sends no wake after a Stop. Only directive and decision wakes are exempt from the cap, and every work wake is rechecked before the prompt. A Room receipt that arrives before completion no longer adds a second turn.
An Architect owner or a Room member whose approval names codemode now loads it, switched on from the first turn because tool search cannot find it. A script can call only the tools the session registered, so a read-only member's script cannot reach a write tool.
…m can give it to a member A new Architect owner names codemode at the start approval when the catalogue offers it. An owner that already holds a grant asks again for exactly what it had. The session request loads only the five owner tools. Docs say what is now true.
The owner gets a managed checkout per milestone, saved with the execution before any file changes. Report commits the checkout to its branch, and evidence runs there. A worktree execution is not a project-folder writer. An accepted or parked milestone releases its checkout without force. The pilot runner can start a Worktree project.
…e owner's own milestones as work
Review round 4: direct work in Worktree mode, and Code Mode for managed sessionsRouted review on gpt-6-astra at high effort, of commits 93d3f033c, a094567b7 and 2fce1ab2b. Code Mode (93d3f033c, a094567b7)No qualifying defects found. Direct work in Worktree mode (2fce1ab2b)
Triage
|
… and a repeated begin Review round 4 on PR 624. The preview server now runs in the checkout while the workspace stays the project folder, and its screenshot is kept in the project folder. Begin is refused while evidence for the milestone is running. A repeated begin keeps the saved checkout. A parked milestone keeps its branch. The pilot runner reads a released checkout from its branch.
… a fresh preview server Review round 5 on PR 624. A preview server started for a checkout is stopped after the capture. Evidence is refused while the owner's own work on the milestone is running. A repeated begin restores a released checkout. A checkout on any branch but its saved one is refused.
Review round 5: re-check of the round 4 fixes (bd13f6d)Same reviewer session, gpt-6-astra at high effort.
TriagePoints 1 to 4 are fixed in the next commit. Point 4 now has a check: a checkout that is not on its saved branch is refused at continue, report and evidence. The absolute-path gap in the shell guard itself is added to #625. |
Review round 6: re-check of the round 5 fixes (cec47f6)Same reviewer session, gpt-6-astra at high effort. All four points are fixed, and no new defect was found in the fixes or what they touch. |
… states which approval was read
sero-cli returns a failed command as text with a non-zero exit code, and the agent loop reported it as a success. A Code Mode script only stops on a call the loop reports as failed, so it ran every later line. A live script ran past a blocked click and then past the 50 commands per turn limit, getting the same refusal for each remaining call.
Found in manual testing: a Code Mode script did not stop on a failed
|
… a script issues How many commands a Code Mode script runs is the agent's decision. A call another tool made carries Pi's nested call id, and the limit skips it. A command the model issues itself is still counted.
…over several tool calls in a row The rule is added to the system prompt of chat sessions, managed sessions and workers, and only when codemode is switched on for that session. It also says when separate calls are right: when a result must be read before the next step can be decided.
…a Code Mode script stops on it The earlier fix marked the failure at agent.afterToolCall, which Pi does not run for a call a script makes. The tool now returns isError itself, as bash does. A real-session test runs a script through Pi's codemode and the real sero-cli tool: the line after a failed command never runs, and 60 commands in one script run with no rate limit.
…answering A page stuck in its own code, such as an endless loop, blocks the browser daemon: every later command waits out its limit, and so do close and launch. When a command times out and the session cannot answer a cheap question, the daemon and its browser are stopped and the agent is told the page is stuck. A slow command on a session that still answers is left alone. A test runs the real tool against a real looping page when a browser pack is present: red without the reset, green with it.
…ser is reset on Windows too
sero-cli split its input at every line break, so a script passed to
--expression "..." on several lines became several broken commands
("Unterminated quoted string"). A line break inside quotes now stays in the
argument.
The stuck-browser reset now runs on a Windows host through taskkill. That
path is written from the Git Bash rules and has not been run on Windows.
A host workspace may only touch files inside its own roots. The screenshot went to the browser pack's temp folder, which is not one, so every screenshot failed with "Host path must be inside a workspace root". A recording was given /workspace, which does not exist on a host, so it could not be saved. Both now use the workspace's real path. The tests for this tool faked the runtime, so the host rule never ran. The real-browser test now runs screenshot and recording through the real host backend.
…seconds The browser's own "wait for load" waits for an event that has already passed once the page is open, and gives up after its 25s limit. Every launch or navigate that asked for load paid that: measured at 26s per call, which ran a seven-viewport script past its time limit. The page's ready state is asked for instead. The session helpers move to their own file to keep the tool under 500 lines.
|
|
||
| it('reads a milestone saved before direct execution unchanged', () => { | ||
| const saved = milestone({ status: 'verifying', verification: 'reported' }); | ||
| const reloaded = JSON.parse(JSON.stringify(project([saved]))) as ProjectRecord; |
There was a problem hiding this comment.
React Doctor · react-doctor/no-json-parse-stringify-clone (warning)
JSON.parse(JSON.stringify(x)) deep-clones by re-serializing: it is slow on large objects and silently drops undefined, functions, Date/Map/Set, and cyclic references. Use structuredClone(x).
Fix → Replace JSON.parse(JSON.stringify(value)) with structuredClone(value). It is faster and preserves Dates, Maps, Sets, and cyclic references.
| return ( | ||
| <section aria-label="Limits"> | ||
| <div className="ar-lim-title"><span>Limits</span><span>{rows.length}</span></div> | ||
| <div className="ar-lim-card" role="table" aria-label="Limits"> |
There was a problem hiding this comment.
React Doctor · react-doctor/prefer-tag-over-role (warning)
Screen reader users get more reliable semantics from <table> than role="table", so use <table> instead.
Fix → Use the matching HTML element when one exists so browsers and assistive tech get native semantics.
| /** What a tool change does to the list, as the drawing words it: `add shell`. */ | ||
| function toolsMove(current: string[], next: string | string[] | undefined): string | null { | ||
| if (!Array.isArray(next)) return null; | ||
| const added = next.filter((tool) => !current.includes(tool)).map((tool) => `add ${tool}`); |
There was a problem hiding this comment.
React Doctor · react-doctor/js-set-map-lookups (warning)
This scales poorly because array.includes() inside a loop scans the whole list every time. Use a Set for constant-time lookups.
Fix → Use a Set or Map when you check for the same items over and over. Array.includes/find scans the whole list each time
| function toolsMove(current: string[], next: string | string[] | undefined): string | null { | ||
| if (!Array.isArray(next)) return null; | ||
| const added = next.filter((tool) => !current.includes(tool)).map((tool) => `add ${tool}`); | ||
| const removed = current.filter((tool) => !next.includes(tool)).map((tool) => `remove ${tool}`); |
There was a problem hiding this comment.
React Doctor · react-doctor/js-set-map-lookups (warning)
This scales poorly because array.includes() inside a loop scans the whole list every time. Use a Set for constant-time lookups.
Fix → Use a Set or Map when you check for the same items over and over. Array.includes/find scans the whole list each time
Implements OpenSpec change
improve-architect-orchestrator-autonomy(#620) as one PR. 32 of 44 tasks are ticked. The 12 left are listed under "Not done".What you can do after this PR
tool_search. A tool outside the approval is never registered. New Architect owners ask for the workspace's enabled skills at the start approval.goaltool.Two faults the paid pilot found, both fixed here
sero-cli. A run from source cachedsero-clias a plugin's tool, and the approval step then dropped it for managed sessions. The owner changed the files correctly and had no way to record anything. Fix: be0d682. This is on main today.serofrom the shell in a managed session answered "Unauthorized session". The model spent its turn probing it. The refusal now says to use thesero-clitool. Fix: 7edaf65.Paid runs
deepseek/deepseek-flashat low thinking, one small-fix scenario, $1 and 8 minutes per run. Records and a fuller table are inopenspec/changes/improve-architect-orchestrator-autonomy/evidence/results.md.No improvement is claimed from these. One scenario with one repeat is too few. The pilot was planned as 20 runs and cut to keep paid runs short.
In the Worktree run the owner got a checkout on its own branch, edited inside it, and Sero committed the work there. The hidden test passed on that branch. A new owner's approval on this branch listed Code Mode (read from a run I stopped early); no owner ran a script. An earlier Worktree run has no test result because my runner read the checkout after Sero had released it. The runner now reads the branch.
In the Workspace run the owner used direct work end to end with a real model. Its first milestone picked up a preview check the project could not pass, so it opened a second milestone and asked how to close the first. That preview behaviour is not from this change.
Decisions I made that are yours to overturn
Needs your eyes
apps/styleguide/public/prototypes/room-amendment-status.htmlandarchitect-wait-status.html, with captures underprototypes/screenshots/.Not done
tool_search, a Code Mode script, a preview check in a worktree, or stall recovery. Their evidence is tests only.Review
Routed rounds on gpt-6-astra at high effort, posted below.
Round 1 found 17 defects. 16 are fixed and one is tracked as Plugin UI tools are callable by the model, so the Goals tool can raise a user-set limit #623 (a plugin's UI tool can be called by a model, so the Goals tool can still raise a limit you set).
Round 2 checked those fixes and raised 9 points. 7 are fixed.
Round 3 checked the round 2 fixes and raised 4 points. All 4 are fixed. Those last fixes have not had a further review.
Round 4 reviewed direct work in Worktree mode and Code Mode. Code Mode: no finding. Worktree: 6 points, 5 fixed.
Round 5 checked those fixes and raised 4 follow-on points. All 4 are fixed.
Two round 2 points were not acted on:
Checks
pnpm typecheck --force: 29 of 29 tasks pass.@sero-ai/commonis bumped to 0.24.0 (grant amendment API,toolsAreLoadout). It is not published.How to test
.sero/worktrees/until the milestone is delivered.codemodefor a multi-step read.Found on the way, filed separately
git --no-pager diff, and missesgitcalled by absolute path.