diff --git a/agentcore/adr/ADR-035-unified-model-profiles-and-role-bindings.md b/agentcore/adr/ADR-035-unified-model-profiles-and-role-bindings.md index 0a6b8cd60..adb735c3c 100644 --- a/agentcore/adr/ADR-035-unified-model-profiles-and-role-bindings.md +++ b/agentcore/adr/ADR-035-unified-model-profiles-and-role-bindings.md @@ -194,3 +194,7 @@ Branch 9W4D adds OpenAI Responses as a Commander conformance-policy-v3 selection. The model-configuration schema and Executor mapping stay v1, and the compatibility matrix remains descriptive evidence rather than profile, conformance, mapping, or readiness authority. + +Branch 9W4E constructs this unchanged vocabulary from a fixed six-recipe +catalog and persists only credential-free recipe identities and semantic +hashes. It does not introduce a second selection schema. diff --git a/agentcore/adr/ADR-036-runtime-model-profile-registry-and-role-readiness.md b/agentcore/adr/ADR-036-runtime-model-profile-registry-and-role-readiness.md index 1fb4cb911..c1b67e221 100644 --- a/agentcore/adr/ADR-036-runtime-model-profile-registry-and-role-readiness.md +++ b/agentcore/adr/ADR-036-runtime-model-profile-registry-and-role-readiness.md @@ -83,3 +83,8 @@ External research MCP remains post-v1. `resume_supported=false`, `provider_tool_loop_enabled=false`, and `external_read_execution_enabled=false` remain unchanged. + +Branch 9W4E activates persisted setup only during the next RuntimeServer +construction. Executor readiness uses the exact packaged OpenCode executable +selected for launch and remains separate non-authoritative evidence; the +registry and projection hashes remain immutable for the process lifetime. diff --git a/agentcore/adr/ADR-040-first-run-model-setup-and-role-selection.md b/agentcore/adr/ADR-040-first-run-model-setup-and-role-selection.md new file mode 100644 index 000000000..32bc6e288 --- /dev/null +++ b/agentcore/adr/ADR-040-first-run-model-setup-and-role-selection.md @@ -0,0 +1,94 @@ +# ADR-040 - First-run model setup and role selection + +## Status + +Accepted for Branch 9W4E. + +## Context + +ADR-035 defines credential-free model connections, profiles, and independent +Commander/Executor bindings. ADR-036 activates those snapshots in one immutable +RuntimeServer registry. ADR-037 and ADR-038 add native Gemini and OpenAI +Responses. None provides durable user selection or a first-run product flow. + +Model setup must not turn compatibility evidence, OpenCode discovery, or TUI +state into provider authority. It also cannot mutate the active registry or +store connector credentials. + +## Decision + +1. Add the code-owned `nexusloop_model_setup_catalog_v1`. It offers exactly + three Commander recipes backed by static conformance and three primary + Executor recipes backed by the static provider mapping: Anthropic Claude + Sonnet 4.5, Google Gemini 2.5 Flash, and OpenAI GPT-4.1 mini. Either role may + remain explicitly unconfigured. +2. RuntimeServer reconstructs a complete ADR-035 candidate from exact recipe + IDs. The compatibility matrix, OpenCode catalog/authentication, and caller + provider/model assertions cannot add recipes or conformance. +3. Preview is read-only and returns the current revision plus candidate, + configuration, and role-projection hashes. Confirmation requires the exact + revision, candidate hash, bounded human identity, and + `CONFIRM_MODEL_SETUP`. +4. The sole durable transition is + `runtime_model_setup_committed` under + `nexusloop_model_setup_event_v1`. Its allowlist contains recipe IDs, + contiguous revision linkage, semantic hashes, bounded human identity, and + time. It contains no complete configuration, connector URL, header, + credential reference name/value, environment name, provider object, or + OpenCode authentication state. +5. `EventStore.appendIfLatestKind` provides atomic expected-kind comparison + without allowing unrelated journal traffic to starve setup confirmation. + Exact duplicate confirmation is idempotent; stale, different, malformed, + truncated, duplicate-revision, unknown-version, or hash-invalid authority + fails closed. +6. Startup projects the append-only event stream before RuntimeServer + construction and builds the existing immutable ADR-036 registry. Persisted + setup, explicit registry/provider authority, and legacy Commander + environment authority are mutually exclusive. +7. A commit never replaces the active registry. It is pending for the next + process start. Setup writes are RuntimeServer-owned, acquire or reuse the + run lock, and drain before `runtime_shutdown`. +8. Commander and Executor readiness remain independent evidence. A safe + selection may persist while connector or credential readiness is blocked or + unknown. Neither role falls back to the other. +9. Production Executor readiness invokes the 9W4E0 command `opencode + nexusloop executor-readiness-v1`. Runtime sends only the exact immutable + Executor projection. The packaged OpenCode command checks pinned + catalog/config/plugin/auth semantics, discards unrelated identities and raw + state, and returns only matching tri-state evidence. Runtime owns + timeout, output bounds, concurrency, identity validation, cancellation, and + shutdown drain. The observation cannot select, map, normalize, recommend, + fall back, or authorize Commander. + Production uses the exact validated OpenCode executable configured for + Executor launch and fixes readiness arguments internally. No separate + executable, source module, preload, dependency path, or environment + assertion may replace it; process injection is package-internal test + machinery only. +10. OpenTUI extends the existing initialization/onboarding surface with + keyboard selection, preview, separate explicit confirmation, current and + pending hashes, blocked readiness, cancellation, and restart-required + rendering. Cached UI state is never mutation authority. +11. Executor selection still reaches only the primary tactical OpenCode run as + one exact `--model provider/model` argument. Auxiliary models and OpenCode + global/user configuration are unchanged. +12. The unset TUI runtime-client mode is `auto`: legacy fake behavior is + limited to pre-spec onboarding, while an approved project constructs the + real RuntimeServer client. Explicit fake mode remains non-production + fixture authority and cannot prove a durable setup commit. + +## Consequences + +First-run and later model selection now have one credential-free, append-only, +restart-only authority path. Connector construction and both role-owned +credential resolvers remain separate. Custom endpoints, arbitrary model +discovery, credentials in TUI state, OpenCode `auth.json` mutation, hot reload, +fallback, retry, streaming, auxiliary model selection, MCP, proposals, +governance, and mutation remain out of scope. + +OpenCode's public provider-list response is not Runtime authority and is not +crossed into RuntimeServer. The packaged command performs no provider execution +request, retry, mutation, or catalog refresh. Observation failure, +partial state, timeout, truncation, cancellation, or shutdown remains unknown. + +`resume_supported=false`, `provider_tool_loop_enabled=false`, and +`external_read_execution_enabled=false` remain unchanged. diff --git a/agentcore/runtime/src/authority/command-authority-registry.ts b/agentcore/runtime/src/authority/command-authority-registry.ts index a7935d122..be128f6ea 100644 --- a/agentcore/runtime/src/authority/command-authority-registry.ts +++ b/agentcore/runtime/src/authority/command-authority-registry.ts @@ -83,6 +83,7 @@ const profiles = { apply: profile(["tests/e2e_user/scenarios/test_commander_cycle_tui.py"]), externalApi: profile(["tests/e2e_user/scenarios/test_reasoning_provider_tui.py"]), commanderRecovery: profile(["tests/e2e_user/scenarios/test_commander_recovery_operator_controls_tui.py"], ["tests/e2e_user/scenarios/test_command_authority_inventory_tui.py", "tests/e2e_user/scenarios/test_commander_continuity_packet_tui.py"]), + modelSetup: profile(["tests/e2e_user/scenarios/test_model_setup_executor_readiness_tui.py"], ["tests/e2e_user/scenarios/test_command_authority_inventory_tui.py"]), } type BaseRecord = Omit @@ -178,6 +179,7 @@ function write(args: { } export const COMMAND_AUTHORITY_REGISTRY: CommandAuthorityRecord[] = [ + record({ slash_command: "/model-setup", runtime_command: "runtime.confirm_model_setup", risk: "medium_risk_write", gate: "model_setup_runtime", owner: "model_setup", mutates_events: true, creates_external_process: false, calls_provider: false, requires_active_runtime: false, requires_run_lock: true, requires_approval: true, approval_surface: "/model-setup", expected_event_kinds: ["runtime_model_setup_committed"], blocked_by_default: true, current_phase_status: "implemented", validation_profile: profiles.modelSetup, notes: ["Explicitly confirms one credential-free model setup candidate and appends one setup record for restart-only activation. Confirmation itself starts no external process; subsequent status and launch-readiness reads may invoke the bounded packaged OpenCode observer."], out_of_scope: ["credential storage", "provider calls", "OpenCode launch", "hot reload", "automatic restart", "provider discovery", "role fallback"] }), read("/authority", "runtime.command_authority_summary", "runtime_status", "none", profiles.authority, ["/command-authority", "/command-map"]), read("/authority-summary", "runtime.command_authority_summary", "runtime_status", "none", profiles.authority), read("/authority-list", "runtime.command_authority_list", "runtime_status", "none", profiles.authority), diff --git a/agentcore/runtime/src/authority/command-authority-types.ts b/agentcore/runtime/src/authority/command-authority-types.ts index 1c4f70981..bc39e8dee 100644 --- a/agentcore/runtime/src/authority/command-authority-types.ts +++ b/agentcore/runtime/src/authority/command-authority-types.ts @@ -22,6 +22,7 @@ export type CommandAuthorityGate = | "opencode_runtime" | "external_api_runtime" | "commander_recovery_runtime" + | "model_setup_runtime" | "unknown" export type CommandAuthorityOwner = @@ -52,6 +53,7 @@ export type CommandAuthorityOwner = | "playbook" | "commander_apply" | "commander_recovery" + | "model_setup" | "unknown" export type CommandPhaseStatus = diff --git a/agentcore/runtime/src/events/event-store.ts b/agentcore/runtime/src/events/event-store.ts index b93efef18..73a9ec68c 100644 --- a/agentcore/runtime/src/events/event-store.ts +++ b/agentcore/runtime/src/events/event-store.ts @@ -72,6 +72,46 @@ export class EventStore { return operation.finally(() => { this.pendingAppends -= 1 }) } + async appendIfLatestKind( + event: JsonlEvent, + kind: string, + expectedLatestEventId: string | null, + operational: { before_write?: () => void } = {}, + ): Promise { + if (event.kind !== kind) throw new Error("event kind does not match append authority") + this.appendGeneration += 1 + this.pendingAppends += 1 + const operation = this.appendQueue.then(async () => { + await mkdir(dirname(this.eventsPath), { recursive: true }) + const events = await this.readAllSnapshot() + let latest: JsonlEvent | undefined + for (let index = events.length - 1; index >= 0; index -= 1) { + if (events[index]?.kind === kind) { + latest = events[index] + break + } + } + const latestEventId = latest?.event_id ? String(latest.event_id) : null + if (latestEventId !== expectedLatestEventId) throw new Error("event kind changed before append") + const safeEvent = redactValue({ + ...event, + event_id: event.event_id ?? makeEventId(), + timestamp: event.timestamp ?? new Date().toISOString(), + }) + const handle = await open(this.eventsPath, "a") + try { + operational.before_write?.() + await handle.write(JSON.stringify(safeEvent) + "\n") + await handle.sync() + } finally { + await handle.close() + } + return String(safeEvent.event_id) + }) + this.appendQueue = operation.catch(() => undefined) + return operation.finally(() => { this.pendingAppends -= 1 }) + } + async readAll(): Promise { while (true) { const generationBefore = this.appendGeneration diff --git a/agentcore/runtime/src/index.ts b/agentcore/runtime/src/index.ts index ceb811a80..a79ad994d 100644 --- a/agentcore/runtime/src/index.ts +++ b/agentcore/runtime/src/index.ts @@ -1,5 +1,6 @@ export { RuntimeServer } from "./server" export type { RuntimeResearchDbProjection, RuntimeResearchDbReader, RuntimeServerOptions } from "./server" +export { locateProjectRoot } from "./project/project-root" export { createRuntimeServerFromLaunchConfig, readRuntimeServerLaunchOptionsFromEnv, readWakeSchedulerBootstrapConfigFromEnv } from "./launch-config" export type { RuntimeServerLaunchConfig } from "./launch-config" export { EventStore } from "./events/event-store" @@ -13,6 +14,8 @@ export { adaptLegacyCommanderModelAuthority } from "./model-configuration/model- export * from "./model-configuration/model-profile-runtime-registry-types" export { providerCompatibilityMatrix, PROVIDER_COMPATIBILITY_MATRIX_POLICY_VERSION, PROVIDER_COMPATIBILITY_MATRIX_SCHEMA_VERSION } from "./model-configuration/provider-compatibility-matrix" export type { CommanderProviderCompatibilityEvidence, CommanderProviderCompatibilityMatrix } from "./model-configuration/provider-compatibility-matrix" +export { ModelSetupService, buildModelSetupCandidate, modelSetupCatalog, projectModelSetupEvents, readPersistedModelSetupAuthority } from "./model-configuration/model-setup" +export type { ModelSetupCandidate, ModelSetupCatalog, ModelSetupChoices, ModelSetupCommitInput, ModelSetupCommitResult, ModelSetupPreview, ModelSetupProjection, ModelSetupRecipe, PersistedModelSetupAuthority } from "./model-configuration/model-setup" export type { WakeSchedulerBootstrapConfig, WakeSchedulerBootstrapStatus, WakeSchedulerStaleRunInfo } from "./schedules/wake-scheduler-bootstrap-types" export type { WakeSchedulerRecovery, WakeSchedulerRecoveryAcknowledgeInput, WakeSchedulerRecoveryCommand, WakeSchedulerRecoveryPreview, WakeSchedulerRecoveryRecord, WakeSchedulerRecoveryStatus } from "./schedules/wake-scheduler-recovery-types" export type { WakeSchedulerRecoveryWorkflow, WakeSchedulerRecoveryWorkflowCancelInput, WakeSchedulerRecoveryWorkflowInput, WakeSchedulerRecoveryWorkflowPreview, WakeSchedulerRecoveryWorkflowRecord, WakeSchedulerRecoveryWorkflowStep, WakeSchedulerRecoveryWorkflowStepRecordInput, WakeSchedulerRecoveryWorkflowVerification } from "./schedules/wake-scheduler-recovery-workflow-types" diff --git a/agentcore/runtime/src/launch-config.ts b/agentcore/runtime/src/launch-config.ts index b63c493a1..ac993a555 100644 --- a/agentcore/runtime/src/launch-config.ts +++ b/agentcore/runtime/src/launch-config.ts @@ -4,8 +4,11 @@ import { readExternalApiConnectorsFromEnv } from "./external-api/api-connector-r import { readReasoningProviderConfigFromEnv } from "./reasoning/reasoning-provider-config" import { readCommanderInvestigationProviderConfigFromEnv } from "./commander-agent" import type { WakeSchedulerBootstrapConfig } from "./schedules/wake-scheduler-bootstrap-types" +import { locateProjectRoot } from "./project/project-root" +import { readPersistedModelSetupAuthority } from "./model-configuration/model-setup" +import { createPackagedOpenCodeExecutorReadinessResolver, snapshotPackagedOpenCodeAdapterConfig } from "./model-configuration/opencode-executor-readiness-resolver" -export interface RuntimeServerLaunchConfig extends RuntimeServerOptions { +export interface RuntimeServerLaunchConfig extends Omit { env?: Record } @@ -13,6 +16,7 @@ export function readRuntimeServerLaunchOptionsFromEnv( env: Record, baseOptions: RuntimeServerOptions = {}, ): RuntimeServerOptions { + rejectExecutorObserverEnvironment(env) const options: RuntimeServerOptions = { ...baseOptions } if (options.modelProfileRuntimeRegistry && hasLegacyCommanderEnvironmentAuthority(env)) { throw new Error("explicit model-profile registry cannot be combined with legacy Commander environment authority") @@ -56,8 +60,65 @@ function hasLegacyCommanderEnvironmentAuthority(env: Record { + try { + return readPersistedModelSetupAuthority(projectDir) + } catch (error) { + if (error instanceof Error && error.message === "model setup journal is malformed") return undefined + throw error + } +} + +function rejectExecutorObserverEnvironment(env: Record): void { + if (env.NXL_OPENCODE_EXECUTOR_READINESS_COMMAND !== undefined + || env.NXL_OPENCODE_EXECUTOR_READINESS_ARGS_JSON !== undefined) { + throw new Error("custom Executor readiness observer environment configuration is not supported") + } } export function readWakeSchedulerBootstrapConfigFromEnv(env: Record): WakeSchedulerBootstrapConfig | undefined { diff --git a/agentcore/runtime/src/model-configuration/model-profile-runtime-registry-types.ts b/agentcore/runtime/src/model-configuration/model-profile-runtime-registry-types.ts index 4a5d1bc0a..22faaa34c 100644 --- a/agentcore/runtime/src/model-configuration/model-profile-runtime-registry-types.ts +++ b/agentcore/runtime/src/model-configuration/model-profile-runtime-registry-types.ts @@ -51,7 +51,9 @@ export type ExecutorModelReadinessObservation = Readonly<{ }> export interface ExecutorModelReadinessResolver { + start?(): Promise | void observe(selection: ExecutorModelSelectionProjection): Promise | unknown + shutdown?(): Promise | void } export type CommanderModelReadinessInput = CommanderInvestigationProviderReadiness diff --git a/agentcore/runtime/src/model-configuration/model-setup.test.ts b/agentcore/runtime/src/model-configuration/model-setup.test.ts new file mode 100644 index 000000000..da7f5610f --- /dev/null +++ b/agentcore/runtime/src/model-configuration/model-setup.test.ts @@ -0,0 +1,237 @@ +import { describe, expect, test } from "bun:test" +import { createHash } from "node:crypto" +import { mkdtemp, readFile, writeFile } from "node:fs/promises" +import { join } from "node:path" +import { tmpdir } from "node:os" +import { EventStore } from "../events/event-store" +import { + MODEL_SETUP_CATALOG_POLICY_VERSION, + MODEL_SETUP_EVENT_KIND, + ModelSetupService, + buildModelSetupCandidate, + modelSetupCatalog, + projectModelSetupEvents, + readPersistedModelSetupAuthority, +} from "./model-setup" + +describe("9W4E model setup authority", () => { + test("catalog is immutable, credential-free, and limits Commander to exact recipes", () => { + const catalog = modelSetupCatalog() + expect(catalog.policy_version).toBe(MODEL_SETUP_CATALOG_POLICY_VERSION) + expect(catalog.commander_recipes.map((item) => item.recipe_id)).toEqual([ + "commander-anthropic-claude-sonnet-4-5", + "commander-google-gemini-2-5-flash", + "commander-openai-gpt-4-1-mini-responses", + ]) + expect(Object.isFrozen(catalog)).toBe(true) + expect(Object.isFrozen(catalog.commander_recipes)).toBe(true) + const serialized = JSON.stringify(catalog) + for (const forbidden of ["base_url", "env_name", "api_key", "authorization", "header", "package", "plugin"]) { + expect(serialized.toLowerCase()).not.toContain(forbidden) + } + }) + + test("builds same, different, and explicitly unconfigured role choices without fallback", () => { + const same = buildModelSetupCandidate({ + commander_recipe_id: "commander-anthropic-claude-sonnet-4-5", + executor_recipe_id: "executor-anthropic-claude-sonnet-4-5", + }) + expect(same.commander_selection?.model_id).toBe("claude-sonnet-4-5-20250929") + expect(same.executor_selection?.model_id).toBe("claude-sonnet-4-5-20250929") + expect(same.commander_selection?.projection_hash).not.toBe(same.executor_selection?.projection_hash) + + const different = buildModelSetupCandidate({ + commander_recipe_id: "commander-openai-gpt-4-1-mini-responses", + executor_recipe_id: "executor-google-gemini-2-5-flash", + }) + expect(different.commander_selection?.provider_kind).toBe("openai") + expect(different.executor_selection?.provider_kind).toBe("google") + + const commanderOnly = buildModelSetupCandidate({ commander_recipe_id: "commander-google-gemini-2-5-flash", executor_recipe_id: null }) + expect(commanderOnly.commander_selection).toBeDefined() + expect(commanderOnly.executor_selection).toBeUndefined() + const executorOnly = buildModelSetupCandidate({ commander_recipe_id: null, executor_recipe_id: "executor-openai-gpt-4-1-mini" }) + expect(executorOnly.commander_selection).toBeUndefined() + expect(executorOnly.executor_selection).toBeDefined() + }) + + test("rejects unknown recipes, unknown fields, accessors, symbols, and caller mutation", () => { + expect(() => buildModelSetupCandidate({ commander_recipe_id: "unknown", executor_recipe_id: null })).toThrow("unknown Commander setup recipe") + expect(() => buildModelSetupCandidate({ commander_recipe_id: null, executor_recipe_id: null, extra: true } as never)).toThrow("unknown") + let getterCalls = 0 + const accessor = Object.defineProperty({ executor_recipe_id: null }, "commander_recipe_id", { + enumerable: true, + get() { getterCalls += 1; return null }, + }) + expect(() => buildModelSetupCandidate(accessor as never)).toThrow() + expect(getterCalls).toBe(0) + const symbolic = { commander_recipe_id: null, executor_recipe_id: null } as Record + symbolic[Symbol("hidden")] = "secret" + expect(() => buildModelSetupCandidate(symbolic as never)).toThrow() + + const input = { commander_recipe_id: "commander-google-gemini-2-5-flash", executor_recipe_id: null } + const candidate = buildModelSetupCandidate(input) + input.commander_recipe_id = "commander-openai-gpt-4-1-mini-responses" + expect(candidate.commander_selection?.model_id).toBe("gemini-2.5-flash") + expect(Object.isFrozen(candidate.configuration)).toBe(true) + }) + + test("rejects inherited fields, arrays, live proxies, and revoked proxies without executing caller code", () => { + let inheritedCalls = 0 + const prototype = Object.create(null) + Object.defineProperty(prototype, "commander_recipe_id", { get() { inheritedCalls += 1; return null } }) + const inherited = Object.create(prototype) + Object.defineProperty(inherited, "executor_recipe_id", { value: null, enumerable: true }) + expect(() => buildModelSetupCandidate(inherited)).toThrow() + expect(inheritedCalls).toBe(0) + expect(() => buildModelSetupCandidate(new Array(3))).toThrow("plain object") + + let traps = 0 + const proxied = new Proxy({ commander_recipe_id: null, executor_recipe_id: null }, { + ownKeys() { traps += 1; return [] }, + getOwnPropertyDescriptor() { traps += 1; return undefined }, + getPrototypeOf() { traps += 1; return Object.prototype }, + }) + expect(() => buildModelSetupCandidate(proxied)).toThrow("Proxy") + expect(traps).toBe(0) + const revoked = Proxy.revocable({ commander_recipe_id: null, executor_recipe_id: null }, {}) + revoked.revoke() + expect(() => buildModelSetupCandidate(revoked.proxy)).toThrow("Proxy") + }) + + test("rejects authority-shaped operator identity without persisting it", async () => { + const dir = await mkdtemp(join(tmpdir(), "nxl-model-setup-identity-")) + const store = new EventStore(join(dir, ".nxl", "events.jsonl")) + const service = new ModelSetupService({ eventStore: store }) + const choices = { commander_recipe_id: null, executor_recipe_id: null } as const + const preview = await service.preview(choices) + for (const confirmedBy of ["AWS_PROFILE", "AUTH_JSON", "NXL_REGION", "https://host.example", "Authorization"]) { + await expect(service.confirm({ ...choices, expected_revision: 0, candidate_hash: preview.candidate_hash, confirmed_by: confirmedBy, confirmation: "CONFIRM_MODEL_SETUP" })).rejects.toThrow() + } + expect(await store.readAll()).toHaveLength(0) + }) + + test("previews and atomically commits one revision with idempotent exact replay", async () => { + const dir = await mkdtemp(join(tmpdir(), "nxl-model-setup-")) + const store = new EventStore(join(dir, ".nxl", "events.jsonl")) + const service = new ModelSetupService({ eventStore: store, now: () => new Date("2026-08-22T00:00:00.000Z") }) + const choices = { commander_recipe_id: "commander-openai-gpt-4-1-mini-responses", executor_recipe_id: "executor-anthropic-claude-sonnet-4-5" } as const + const preview = await service.preview(choices) + expect(preview.expected_revision).toBe(0) + expect(preview.restart_required).toBe(true) + const input = { ...choices, expected_revision: 0, candidate_hash: preview.candidate_hash, confirmed_by: "human-operator", confirmation: "CONFIRM_MODEL_SETUP" } as const + const [left, right] = await Promise.all([service.confirm(input), service.confirm(input)]) + expect([left.status, right.status].sort()).toEqual(["committed", "idempotent"]) + const events = await store.readAll() + expect(events.filter((event) => event.kind === MODEL_SETUP_EVENT_KIND)).toHaveLength(1) + const setup = events.find((event) => event.kind === MODEL_SETUP_EVENT_KIND)! + expect(setup.revision).toBe(1) + expect(JSON.stringify(setup)).not.toMatch(/secret|env_name|base_url|authorization|header/i) + + const currentPreview = await service.preview(choices) + const unchanged = await service.confirm({ ...input, expected_revision: currentPreview.expected_revision }) + expect(unchanged).toMatchObject({ status: "idempotent", revision: 1, setup_hash: setup.event_payload_hash }) + expect((await store.readAll()).filter((event) => event.kind === MODEL_SETUP_EVENT_KIND)).toHaveLength(1) + await expect(service.confirm({ ...input, confirmed_by: "different-operator" })).rejects.toThrow("stale") + expect((await store.readAll()).filter((event) => event.kind === MODEL_SETUP_EVENT_KIND)).toHaveLength(1) + }) + + test("stale revisions and candidate hashes fail without appending", async () => { + const dir = await mkdtemp(join(tmpdir(), "nxl-model-setup-stale-")) + const store = new EventStore(join(dir, ".nxl", "events.jsonl")) + const service = new ModelSetupService({ eventStore: store }) + const choices = { commander_recipe_id: null, executor_recipe_id: "executor-google-gemini-2-5-flash" } as const + const preview = await service.preview(choices) + await expect(service.confirm({ ...choices, expected_revision: 1, candidate_hash: preview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" })).rejects.toThrow("stale") + await expect(service.confirm({ ...choices, expected_revision: 0, candidate_hash: "0".repeat(64), confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" })).rejects.toThrow("candidate hash") + expect((await store.readAll()).filter((event) => event.kind === MODEL_SETUP_EVENT_KIND)).toHaveLength(0) + }) + + test("concurrent different confirmations produce one durable winner", async () => { + const dir = await mkdtemp(join(tmpdir(), "nxl-model-setup-concurrent-")) + const store = new EventStore(join(dir, ".nxl", "events.jsonl")) + const left = new ModelSetupService({ eventStore: store }) + const right = new ModelSetupService({ eventStore: store }) + const leftChoices = { commander_recipe_id: null, executor_recipe_id: "executor-google-gemini-2-5-flash" } as const + const rightChoices = { commander_recipe_id: null, executor_recipe_id: "executor-openai-gpt-4-1-mini" } as const + const [leftPreview, rightPreview] = await Promise.all([left.preview(leftChoices), right.preview(rightChoices)]) + const settled = await Promise.allSettled([ + left.confirm({ ...leftChoices, expected_revision: 0, candidate_hash: leftPreview.candidate_hash, confirmed_by: "operator-left", confirmation: "CONFIRM_MODEL_SETUP" }), + right.confirm({ ...rightChoices, expected_revision: 0, candidate_hash: rightPreview.candidate_hash, confirmed_by: "operator-right", confirmation: "CONFIRM_MODEL_SETUP" }), + ]) + expect(settled.filter((item) => item.status === "fulfilled")).toHaveLength(1) + expect(settled.filter((item) => item.status === "rejected")).toHaveLength(1) + expect((await store.readAll()).filter((event) => event.kind === MODEL_SETUP_EVENT_KIND)).toHaveLength(1) + }) + + test("concurrent identical confirmations reconcile one semantic winner across different timestamps", async () => { + const dir = await mkdtemp(join(tmpdir(), "nxl-model-setup-identical-concurrent-")) + const store = new EventStore(join(dir, ".nxl", "events.jsonl")) + const left = new ModelSetupService({ eventStore: store, now: () => new Date("2026-08-22T00:00:00.000Z") }) + const right = new ModelSetupService({ eventStore: store, now: () => new Date("2026-08-22T00:00:01.000Z") }) + const choices = { commander_recipe_id: null, executor_recipe_id: "executor-openai-gpt-4-1-mini" } as const + const preview = await left.preview(choices) + const input = { ...choices, expected_revision: 0, candidate_hash: preview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" } as const + const settled = await Promise.all([left.confirm(input), right.confirm(input)]) + expect(settled.map((item) => item.status).sort()).toEqual(["committed", "idempotent"]) + expect(new Set(settled.map((item) => item.setup_hash)).size).toBe(1) + expect((await store.readAll()).filter((event) => event.kind === MODEL_SETUP_EVENT_KIND)).toHaveLength(1) + }) + + test("confirmation ignores sustained unrelated tail appends without weakening setup revision authority", async () => { + const dir = await mkdtemp(join(tmpdir(), "nxl-model-setup-unrelated-tail-")) + const store = new EventStore(join(dir, ".nxl", "events.jsonl")) + const service = new ModelSetupService({ eventStore: store }) + const choices = { commander_recipe_id: null, executor_recipe_id: "executor-openai-gpt-4-1-mini" } as const + const preview = await service.preview(choices) + const original = store.appendIfLatestKind.bind(store) + let attempts = 0 + store.appendIfLatestKind = (async (...args: Parameters) => { + attempts += 1 + for (let index = 0; index < 5; index += 1) { + await store.append({ kind: "runtime_status_observed", observation_sequence: index }) + } + return original(...args) + }) as typeof store.appendIfLatestKind + + await expect(service.confirm({ ...choices, expected_revision: 0, candidate_hash: preview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" })).resolves.toMatchObject({ status: "committed", revision: 1 }) + expect(attempts).toBe(1) + expect((await store.readAll()).filter((event) => event.kind === MODEL_SETUP_EVENT_KIND)).toHaveLength(1) + }) + + test("projection and startup reader fail closed on corrupt, duplicate, and truncated setup records", async () => { + const dir = await mkdtemp(join(tmpdir(), "nxl-model-setup-corrupt-")) + const path = join(dir, ".nxl", "events.jsonl") + const store = new EventStore(path) + const service = new ModelSetupService({ eventStore: store }) + const choices = { commander_recipe_id: "commander-anthropic-claude-sonnet-4-5", executor_recipe_id: null } as const + const preview = await service.preview(choices) + await service.confirm({ ...choices, expected_revision: 0, candidate_hash: preview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" }) + const events = await store.readAll() + const authority = readPersistedModelSetupAuthority(dir) + expect(authority?.candidate.commander_selection?.model_id).toBe("claude-sonnet-4-5-20250929") + const setupEvent = events.find((event) => event.kind === MODEL_SETUP_EVENT_KIND)! + const { event_id: _eventId, timestamp: _timestamp, ...payloadOnly } = setupEvent + expect(() => projectModelSetupEvents([payloadOnly as typeof setupEvent])).toThrow("EventStore envelope") + expect(() => projectModelSetupEvents([{ ...setupEvent, event_id: "" }])).toThrow("event_id") + expect(() => projectModelSetupEvents([{ ...setupEvent, timestamp: "not-a-timestamp" }])).toThrow("timestamp") + const noncanonicalCommit: Record = { ...setupEvent, committed_at: "2026-08-22" } + noncanonicalCommit.event_payload_hash = setupPayloadHash(noncanonicalCommit) + expect(() => projectModelSetupEvents([noncanonicalCommit])).toThrow("committed_at") + expect(() => projectModelSetupEvents([...events, events.find((event) => event.kind === MODEL_SETUP_EVENT_KIND)!])).toThrow("revision") + await writeFile(path, `${await readFile(path, "utf8")}{\"kind\":\"${MODEL_SETUP_EVENT_KIND}\"`, "utf8") + expect(() => readPersistedModelSetupAuthority(dir)).toThrow("model setup journal is malformed") + }) +}) + +function setupPayloadHash(event: Record): string { + const { kind: _kind, event_id: _eventId, timestamp: _timestamp, event_payload_hash: _hash, ...payload } = event + return createHash("sha256").update(canonicalJson(payload)).digest("hex") +} + +function canonicalJson(value: unknown): string { + if (value === null || typeof value === "string" || typeof value === "boolean" || typeof value === "number") return JSON.stringify(value) + if (Array.isArray(value)) return `[${value.map(canonicalJson).join(",")}]` + const record = value as Record + return `{${Object.keys(record).sort().map((key) => `${JSON.stringify(key)}:${canonicalJson(record[key])}`).join(",")}}` +} diff --git a/agentcore/runtime/src/model-configuration/model-setup.ts b/agentcore/runtime/src/model-configuration/model-setup.ts new file mode 100644 index 000000000..182f009b7 --- /dev/null +++ b/agentcore/runtime/src/model-configuration/model-setup.ts @@ -0,0 +1,543 @@ +import { createHash } from "node:crypto" +import { readFileSync } from "node:fs" +import { join } from "node:path" +import { types as nodeUtilTypes } from "node:util" +import type { CommanderInvestigationProviderConfig } from "../commander-agent/commander-investigation-provider-types" +import { validateCommanderInvestigationProviderConfig } from "../commander-agent/commander-investigation-provider-config" +import type { JsonlEvent } from "../events/event-types" +import type { EventStore } from "../events/event-store" +import { + COMMANDER_MODEL_CONFORMANCE_POLICY_VERSION_V3, + EXECUTOR_PROVIDER_MAPPING_POLICY_VERSION, + MODEL_CONFIGURATION_POLICY_VERSION, + projectCommanderModelSelection, + projectExecutorModelSelection, + validateAuthoritySafeIdentifier, + validateCommanderModelConformanceRegistry, + validateExecutorProviderMappingRegistry, + validateModelConfiguration, +} from "./model-configuration-kernel" +import type { + CommanderModelSelectionProjection, + ExecutorModelSelectionProjection, + ModelConfiguration, +} from "./model-configuration-types" +import { ModelProfileRuntimeRegistry } from "./model-profile-runtime-registry" + +export const MODEL_SETUP_CATALOG_SCHEMA_VERSION = 1 as const +export const MODEL_SETUP_CATALOG_POLICY_VERSION = "nexusloop_model_setup_catalog_v1" as const +export const MODEL_SETUP_EVENT_SCHEMA_VERSION = 1 as const +export const MODEL_SETUP_EVENT_POLICY_VERSION = "nexusloop_model_setup_event_v1" as const +export const MODEL_SETUP_EVENT_KIND = "runtime_model_setup_committed" as const + +type SetupRole = "commander" | "executor" + +export type ModelSetupRecipe = Readonly<{ + recipe_version: 1 + recipe_id: string + role: SetupRole + display_name: string + provider_kind: "anthropic" | "google" | "openai" + provider_id: string + model_id: string + credential_binding_id: string + recipe_hash: string +}> + +type CommanderRecipeInternal = ModelSetupRecipe & Readonly<{ + connector_id: string + conformance_id: string + transport_kind: CommanderInvestigationProviderConfig["transport_kind"] +}> + +type ExecutorRecipeInternal = ModelSetupRecipe & Readonly<{ mapping_id: string }> + +export type ModelSetupCatalog = Readonly<{ + schema_version: 1 + policy_version: typeof MODEL_SETUP_CATALOG_POLICY_VERSION + commander_recipes: readonly ModelSetupRecipe[] + executor_recipes: readonly ModelSetupRecipe[] + catalog_hash: string +}> + +export type ModelSetupChoices = Readonly<{ + commander_recipe_id: string | null + executor_recipe_id: string | null +}> + +export type ModelSetupCandidate = Readonly<{ + candidate_version: 1 + choices: ModelSetupChoices + catalog_hash: string + configuration: ModelConfiguration + commander_selection?: CommanderModelSelectionProjection + executor_selection?: ExecutorModelSelectionProjection + candidate_hash: string +}> + +export type ModelSetupPreview = Readonly<{ + preview_version: 1 + expected_revision: number + current_setup_hash?: string + candidate_hash: string + catalog_hash: string + configuration_hash: string + commander_selection?: CommanderModelSelectionProjection + executor_selection?: ExecutorModelSelectionProjection + restart_required: true +}> + +export type ModelSetupCommitInput = ModelSetupChoices & Readonly<{ + expected_revision: number + candidate_hash: string + confirmed_by: string + confirmation: "CONFIRM_MODEL_SETUP" +}> + +export type ModelSetupCommitResult = Readonly<{ + status: "committed" | "idempotent" + revision: number + setup_hash: string + candidate_hash: string + restart_required: true +}> + +export type ModelSetupProjection = Readonly<{ + status: "missing" | "ready" + revision: number + latest_event_id: string | null + setup_hash?: string + candidate?: ModelSetupCandidate + confirmed_by?: string + committed_at?: string +}> + +export type PersistedModelSetupAuthority = Readonly<{ + revision: number + setup_hash: string + candidate: ModelSetupCandidate + registry: ModelProfileRuntimeRegistry + commander_provider_config?: CommanderInvestigationProviderConfig +}> + +const COMMANDER_RECIPES: readonly CommanderRecipeInternal[] = freezeRecipes([ + commanderRecipe("commander-anthropic-claude-sonnet-4-5", "Anthropic Claude Sonnet 4.5", "anthropic", "anthropic-native", "claude-sonnet-4-5-20250929", "credential-commander-anthropic-primary", "anthropic-main", "anthropic-claude-sonnet-4-5-native-v1", "anthropic_messages_connector"), + commanderRecipe("commander-google-gemini-2-5-flash", "Google Gemini 2.5 Flash", "google", "google-native", "gemini-2.5-flash", "credential-commander-google-primary", "google-main", "google-gemini-2-5-flash-native-v1", "google_generative_ai_connector"), + commanderRecipe("commander-openai-gpt-4-1-mini-responses", "OpenAI GPT-4.1 mini Responses", "openai", "openai-native", "gpt-4.1-mini", "credential-commander-openai-primary", "openai-main", "openai-gpt-4-1-mini-responses-v1", "openai_responses_connector"), +]) + +const EXECUTOR_RECIPES: readonly ExecutorRecipeInternal[] = freezeRecipes([ + executorRecipe("executor-anthropic-claude-sonnet-4-5", "Anthropic Claude Sonnet 4.5", "anthropic", "anthropic", "claude-sonnet-4-5-20250929", "credential-executor-anthropic-primary", "executor-anthropic-v1"), + executorRecipe("executor-google-gemini-2-5-flash", "Google Gemini 2.5 Flash", "google", "google", "gemini-2.5-flash", "credential-executor-google-primary", "executor-google-v1"), + executorRecipe("executor-openai-gpt-4-1-mini", "OpenAI GPT-4.1 mini", "openai", "openai", "gpt-4.1-mini", "credential-executor-openai-primary", "executor-openai-v1"), +]) + +const CATALOG = (() => { + const publicCommander = COMMANDER_RECIPES.map(publicRecipe) + const publicExecutor = EXECUTOR_RECIPES.map(publicRecipe) + const stable = { + schema_version: MODEL_SETUP_CATALOG_SCHEMA_VERSION, + policy_version: MODEL_SETUP_CATALOG_POLICY_VERSION, + commander_recipes: publicCommander, + executor_recipes: publicExecutor, + } + return deepFreeze({ ...stable, catalog_hash: hash(stable) }) +})() + +export function modelSetupCatalog(): ModelSetupCatalog { return CATALOG } + +export function buildModelSetupCandidate(value: unknown): ModelSetupCandidate { + const choices = parseChoices(value) + const commander = choices.commander_recipe_id === null ? undefined : COMMANDER_RECIPES.find((item) => item.recipe_id === choices.commander_recipe_id) + const executor = choices.executor_recipe_id === null ? undefined : EXECUTOR_RECIPES.find((item) => item.recipe_id === choices.executor_recipe_id) + if (choices.commander_recipe_id !== null && !commander) throw new Error("unknown Commander setup recipe") + if (choices.executor_recipe_id !== null && !executor) throw new Error("unknown Executor setup recipe") + + const connections: Record[] = [] + const profiles: Record[] = [] + const roleBindings: Record[] = [] + if (commander) { + connections.push({ + connection_id: `setup-${commander.recipe_id}`, + provider_kind: commander.provider_kind, + credential_binding_id: commander.credential_binding_id, + commander: { connector_id: commander.connector_id, conformance_id: commander.conformance_id }, + }) + profiles.push({ profile_id: `profile-${commander.recipe_id}`, connection_id: `setup-${commander.recipe_id}`, model_id: commander.model_id, display_name: commander.display_name }) + roleBindings.push({ role: "commander", profile_id: `profile-${commander.recipe_id}` }) + } + if (executor) { + connections.push({ + connection_id: `setup-${executor.recipe_id}`, + provider_kind: executor.provider_kind, + credential_binding_id: executor.credential_binding_id, + executor: { provider_id: executor.provider_id }, + }) + profiles.push({ profile_id: `profile-${executor.recipe_id}`, connection_id: `setup-${executor.recipe_id}`, model_id: executor.model_id, display_name: executor.display_name }) + roleBindings.push({ role: "executor", profile_id: `profile-${executor.recipe_id}` }) + } + const configuration = validateModelConfiguration({ + schema_version: 1, + policy_version: MODEL_CONFIGURATION_POLICY_VERSION, + connections, + profiles, + role_bindings: roleBindings, + }) + const conformance = validateCommanderModelConformanceRegistry({ + registry_version: 1, + policy_version: COMMANDER_MODEL_CONFORMANCE_POLICY_VERSION_V3, + entries: COMMANDER_RECIPES.map((item) => ({ + conformance_version: 1, + conformance_id: item.conformance_id, + provider_kind: item.provider_kind, + transport_kind: item.transport_kind, + provider_id: item.provider_id, + model_id: item.model_id, + })), + }) + const executorMapping = validateExecutorProviderMappingRegistry({ + registry_version: 1, + policy_version: EXECUTOR_PROVIDER_MAPPING_POLICY_VERSION, + entries: EXECUTOR_RECIPES.map((item) => ({ + mapping_version: 1, + mapping_id: item.mapping_id, + provider_kind: item.provider_kind, + provider_ids: [item.provider_id], + })), + }) + const commanderSelection = commander ? projectCommanderModelSelection(configuration, conformance) : undefined + const executorSelection = executor ? projectExecutorModelSelection(configuration, executorMapping) : undefined + const stable = { + candidate_version: 1 as const, + choices, + catalog_hash: CATALOG.catalog_hash, + configuration, + ...(commanderSelection ? { commander_selection: commanderSelection } : {}), + ...(executorSelection ? { executor_selection: executorSelection } : {}), + } + return deepFreeze({ + ...stable, + candidate_hash: hash({ + candidate_version: 1, + choices, + catalog_hash: CATALOG.catalog_hash, + configuration_hash: configuration.configuration_hash, + commander_projection_hash: commanderSelection?.projection_hash ?? null, + executor_projection_hash: executorSelection?.projection_hash ?? null, + }), + }) +} + +export function buildPersistedModelSetupAuthority(candidate: ModelSetupCandidate, revision: number, setupHash: string): PersistedModelSetupAuthority { + const conformance = validateCommanderModelConformanceRegistry({ + registry_version: 1, + policy_version: COMMANDER_MODEL_CONFORMANCE_POLICY_VERSION_V3, + entries: COMMANDER_RECIPES.map((item) => ({ conformance_version: 1, conformance_id: item.conformance_id, provider_kind: item.provider_kind, transport_kind: item.transport_kind, provider_id: item.provider_id, model_id: item.model_id })), + }) + const executorMapping = validateExecutorProviderMappingRegistry({ + registry_version: 1, + policy_version: EXECUTOR_PROVIDER_MAPPING_POLICY_VERSION, + entries: EXECUTOR_RECIPES.map((item) => ({ mapping_version: 1, mapping_id: item.mapping_id, provider_kind: item.provider_kind, provider_ids: [item.provider_id] })), + }) + const registry = new ModelProfileRuntimeRegistry({ authority_source: "explicit", configuration: candidate.configuration, commander_conformance: conformance, executor_provider_mapping: executorMapping }) + const commander = candidate.choices.commander_recipe_id === null ? undefined : COMMANDER_RECIPES.find((item) => item.recipe_id === candidate.choices.commander_recipe_id) + return deepFreeze({ + revision, + setup_hash: setupHash, + candidate, + registry, + ...(commander ? { commander_provider_config: commanderProviderConfig(commander) } : {}), + }) +} + +export class ModelSetupService { + private readonly eventStore: EventStore + private readonly now: () => Date + + constructor(options: { eventStore: EventStore; now?: () => Date }) { + this.eventStore = options.eventStore + this.now = options.now ?? (() => new Date()) + } + + catalog(): ModelSetupCatalog { return modelSetupCatalog() } + + async status(activeSetupHash?: string): Promise { + const projection = projectModelSetupEvents(await this.eventStore.readAll()) + return deepFreeze({ ...projection, ...(activeSetupHash ? { active_setup_hash: activeSetupHash } : {}), pending_restart: projection.setup_hash !== activeSetupHash }) + } + + async preview(value: unknown): Promise { + const candidate = buildModelSetupCandidate(value) + const projection = projectModelSetupEvents(await this.eventStore.readAll()) + return deepFreeze({ + preview_version: 1, + expected_revision: projection.revision, + ...(projection.setup_hash ? { current_setup_hash: projection.setup_hash } : {}), + candidate_hash: candidate.candidate_hash, + catalog_hash: candidate.catalog_hash, + configuration_hash: candidate.configuration.configuration_hash, + ...(candidate.commander_selection ? { commander_selection: candidate.commander_selection } : {}), + ...(candidate.executor_selection ? { executor_selection: candidate.executor_selection } : {}), + restart_required: true, + }) + } + + async confirm(value: unknown): Promise { + const input = parseCommitInput(value) + const candidate = buildModelSetupCandidate({ + commander_recipe_id: input.commander_recipe_id, + executor_recipe_id: input.executor_recipe_id, + }) + if (candidate.candidate_hash !== input.candidate_hash) throw new Error("model setup candidate hash does not match current authority") + const events = await this.eventStore.readAll() + const projection = projectModelSetupEvents(events) + if (projection.candidate?.candidate_hash === candidate.candidate_hash + && projection.confirmed_by === input.confirmed_by + && (input.expected_revision === projection.revision || input.expected_revision === projection.revision - 1)) { + return result("idempotent", projection.revision, projection.setup_hash!, candidate.candidate_hash) + } + if (input.expected_revision !== projection.revision) throw new Error("model setup revision is stale") + const revision = projection.revision + 1 + const committedAt = this.now().toISOString() + const payload = { + schema_version: MODEL_SETUP_EVENT_SCHEMA_VERSION, + policy_version: MODEL_SETUP_EVENT_POLICY_VERSION, + revision, + previous_setup_hash: projection.setup_hash ?? null, + commander_recipe_id: candidate.choices.commander_recipe_id, + executor_recipe_id: candidate.choices.executor_recipe_id, + catalog_hash: candidate.catalog_hash, + candidate_hash: candidate.candidate_hash, + configuration_hash: candidate.configuration.configuration_hash, + commander_projection_hash: candidate.commander_selection?.projection_hash ?? null, + executor_projection_hash: candidate.executor_selection?.projection_hash ?? null, + confirmed_by: input.confirmed_by, + committed_at: committedAt, + } + const eventPayloadHash = hash(payload) + const event = { kind: MODEL_SETUP_EVENT_KIND, ...payload, event_payload_hash: eventPayloadHash } + try { + await this.eventStore.appendIfLatestKind(event, MODEL_SETUP_EVENT_KIND, projection.latest_event_id) + return result("committed", revision, eventPayloadHash, candidate.candidate_hash) + } catch (error) { + const reconciled = projectModelSetupEvents(await this.eventStore.readAll()) + if (reconciled.revision === input.expected_revision + 1 + && reconciled.setup_hash + && reconciled.candidate?.candidate_hash === candidate.candidate_hash + && reconciled.confirmed_by === input.confirmed_by) { + return result("idempotent", reconciled.revision, reconciled.setup_hash, candidate.candidate_hash) + } + throw error + } + } +} + +export function projectModelSetupEvents(events: readonly JsonlEvent[]): ModelSetupProjection { + let revision = 0 + let setupHash: string | undefined + let candidate: ModelSetupCandidate | undefined + let latestEventId: string | null = null + let confirmedBy: string | undefined + let committedAt: string | undefined + for (let index = 0; index < events.length; index += 1) { + const event = events[index]! + if (event.kind !== MODEL_SETUP_EVENT_KIND) continue + const input = strictRecord(event, "model setup event", [ + "kind", "schema_version", "policy_version", "revision", "previous_setup_hash", + "commander_recipe_id", "executor_recipe_id", "catalog_hash", "candidate_hash", "configuration_hash", + "commander_projection_hash", "executor_projection_hash", "confirmed_by", "committed_at", "event_payload_hash", + ], ["event_id", "timestamp"]) + if (typeof input.event_id !== "string" || !/^rt_[a-z0-9]{10}_[a-z0-9]{8}$/.test(input.event_id)) { + throw new Error("model setup EventStore envelope event_id is invalid") + } + canonicalIsoTimestamp(input.timestamp, "model setup EventStore envelope timestamp") + if (input.schema_version !== MODEL_SETUP_EVENT_SCHEMA_VERSION || input.policy_version !== MODEL_SETUP_EVENT_POLICY_VERSION) throw new Error("model setup event version is unsupported") + if (!Number.isInteger(input.revision) || input.revision !== revision + 1) throw new Error("model setup revision is not contiguous") + if (input.previous_setup_hash !== (setupHash ?? null)) throw new Error("model setup previous hash does not match") + const choices = parseChoices({ commander_recipe_id: input.commander_recipe_id, executor_recipe_id: input.executor_recipe_id }) + const rebuilt = buildModelSetupCandidate(choices) + if (input.catalog_hash !== rebuilt.catalog_hash || input.candidate_hash !== rebuilt.candidate_hash || input.configuration_hash !== rebuilt.configuration.configuration_hash) throw new Error("model setup event authority hash is invalid") + if (input.commander_projection_hash !== (rebuilt.commander_selection?.projection_hash ?? null) || input.executor_projection_hash !== (rebuilt.executor_selection?.projection_hash ?? null)) throw new Error("model setup role projection hash is invalid") + const canonicalConfirmedBy = setupOperatorIdentifier(input.confirmed_by) + const canonicalCommittedAt = canonicalIsoTimestamp(input.committed_at, "model setup committed_at") + const payload = { + schema_version: input.schema_version, + policy_version: input.policy_version, + revision: input.revision, + previous_setup_hash: input.previous_setup_hash, + commander_recipe_id: input.commander_recipe_id, + executor_recipe_id: input.executor_recipe_id, + catalog_hash: input.catalog_hash, + candidate_hash: input.candidate_hash, + configuration_hash: input.configuration_hash, + commander_projection_hash: input.commander_projection_hash, + executor_projection_hash: input.executor_projection_hash, + confirmed_by: input.confirmed_by, + committed_at: canonicalCommittedAt, + } + if (input.event_payload_hash !== hash(payload)) throw new Error("model setup event payload hash is invalid") + revision = input.revision + setupHash = input.event_payload_hash as string + candidate = rebuilt + latestEventId = input.event_id + confirmedBy = canonicalConfirmedBy + committedAt = canonicalCommittedAt + } + return deepFreeze(revision === 0 + ? { status: "missing", revision: 0, latest_event_id: null } + : { status: "ready", revision, latest_event_id: latestEventId, setup_hash: setupHash, candidate, confirmed_by: confirmedBy, committed_at: committedAt }) +} + +export function readPersistedModelSetupAuthority(projectDir: string): PersistedModelSetupAuthority | undefined { + let text: string + try { + text = readFileSync(join(projectDir, ".nxl", "events.jsonl"), "utf8") + } catch (error) { + if ((error as NodeJS.ErrnoException).code === "ENOENT") return undefined + throw error + } + let events: JsonlEvent[] + try { + events = text.split(/\r?\n/).filter(Boolean).map((line) => JSON.parse(line) as JsonlEvent) + } catch { + throw new Error("model setup journal is malformed") + } + const projection = projectModelSetupEvents(events) + if (!projection.candidate || !projection.setup_hash) return undefined + return buildPersistedModelSetupAuthority(projection.candidate, projection.revision, projection.setup_hash) +} + +function commanderProviderConfig(recipe: CommanderRecipeInternal): CommanderInvestigationProviderConfig { + return validateCommanderInvestigationProviderConfig({ + transport_kind: recipe.transport_kind, + provider_id: recipe.provider_id, + provider_kind: recipe.provider_kind, + connector_id: recipe.connector_id, + model_id: recipe.model_id, + enabled_phases: ["proposal_investigation"], + timeout_ms: 30_000, + max_request_bytes: 65_536, + max_response_bytes: 65_536, + max_context_bytes: 65_536, + max_context_tokens: 32_768, + max_output_tokens: 4_096, + supports_tools: true, + supports_json_schema: false, + supports_long_context: "unknown", + supports_local_execution: false, + }) +} + +function commanderRecipe(recipeId: string, displayName: string, providerKind: "anthropic" | "google" | "openai", providerId: string, modelId: string, credentialBindingId: string, connectorId: string, conformanceId: string, transportKind: CommanderRecipeInternal["transport_kind"]): CommanderRecipeInternal { + const stable = { recipe_version: 1 as const, recipe_id: recipeId, role: "commander" as const, display_name: displayName, provider_kind: providerKind, provider_id: providerId, model_id: modelId, credential_binding_id: credentialBindingId, connector_id: connectorId, conformance_id: conformanceId, transport_kind: transportKind } + return deepFreeze({ ...stable, recipe_hash: hash(stable) }) +} + +function executorRecipe(recipeId: string, displayName: string, providerKind: "anthropic" | "google" | "openai", providerId: string, modelId: string, credentialBindingId: string, mappingId: string): ExecutorRecipeInternal { + const stable = { recipe_version: 1 as const, recipe_id: recipeId, role: "executor" as const, display_name: displayName, provider_kind: providerKind, provider_id: providerId, model_id: modelId, credential_binding_id: credentialBindingId, mapping_id: mappingId } + return deepFreeze({ ...stable, recipe_hash: hash(stable) }) +} + +function publicRecipe(item: CommanderRecipeInternal | ExecutorRecipeInternal): ModelSetupRecipe { + return deepFreeze({ recipe_version: 1, recipe_id: item.recipe_id, role: item.role, display_name: item.display_name, provider_kind: item.provider_kind, provider_id: item.provider_id, model_id: item.model_id, credential_binding_id: item.credential_binding_id, recipe_hash: item.recipe_hash }) +} + +function freezeRecipes(items: T[]): readonly T[] { return deepFreeze(items) } + +function parseChoices(value: unknown): ModelSetupChoices { + const input = strictRecord(value, "model setup choices", ["commander_recipe_id", "executor_recipe_id"]) + return deepFreeze({ + commander_recipe_id: nullableIdentifier(input.commander_recipe_id, "commander_recipe_id"), + executor_recipe_id: nullableIdentifier(input.executor_recipe_id, "executor_recipe_id"), + }) +} + +function parseCommitInput(value: unknown): ModelSetupCommitInput { + const input = strictRecord(value, "model setup confirmation", ["commander_recipe_id", "executor_recipe_id", "expected_revision", "candidate_hash", "confirmed_by", "confirmation"]) + if (!Number.isInteger(input.expected_revision) || Number(input.expected_revision) < 0 || Number(input.expected_revision) > 1_000_000) throw new Error("model setup expected revision is invalid") + if (typeof input.candidate_hash !== "string" || !/^[a-f0-9]{64}$/.test(input.candidate_hash)) throw new Error("model setup candidate hash is invalid") + if (input.confirmation !== "CONFIRM_MODEL_SETUP") throw new Error("explicit model setup confirmation is required") + return deepFreeze({ + ...parseChoices({ commander_recipe_id: input.commander_recipe_id, executor_recipe_id: input.executor_recipe_id }), + expected_revision: Number(input.expected_revision), + candidate_hash: input.candidate_hash, + confirmed_by: setupOperatorIdentifier(input.confirmed_by), + confirmation: "CONFIRM_MODEL_SETUP", + }) +} + +function setupOperatorIdentifier(value: unknown): string { + const identifier = validateAuthoritySafeIdentifier(value, "model setup confirmed_by", 160) + if (/^(?:authorization|cookie|host|x-api-key|x-goog-api-key|anthropic-version|content-type)$/i.test(identifier)) { + throw new Error("model setup confirmed_by must not be header-shaped") + } + return identifier +} + +function strictRecord(value: unknown, label: string, requiredKeys: readonly string[], optionalKeys: readonly string[] = []): Record { + rejectProxy(value, label) + if (typeof value !== "object" || value === null || Array.isArray(value)) throw new Error(`${label} must be a plain object`) + const prototype = Object.getPrototypeOf(value) + if (prototype !== Object.prototype && prototype !== null) throw new Error(`${label} must be a plain object`) + const allowed = new Set([...requiredKeys, ...optionalKeys]) + const keys = Reflect.ownKeys(value) + if (keys.some((key) => typeof key === "symbol")) throw new Error(`${label} contains symbol fields`) + for (const key of keys) if (typeof key !== "string" || !allowed.has(key)) throw new Error(`${label} contains unknown fields`) + const output = Object.create(null) as Record + for (const key of requiredKeys) { + const descriptor = Object.getOwnPropertyDescriptor(value, key) + if (!descriptor || !descriptor.enumerable || !("value" in descriptor)) throw new Error(`${label} requires own enumerable data fields`) + output[key] = descriptor.value + } + for (const key of optionalKeys) { + const descriptor = Object.getOwnPropertyDescriptor(value, key) + if (!descriptor) continue + if (!descriptor.enumerable || !("value" in descriptor)) throw new Error(`${label} requires own enumerable data fields`) + output[key] = descriptor.value + } + return output +} + +function rejectProxy(value: unknown, label: string): void { + try { + if (typeof value === "object" && value !== null && nodeUtilTypes.isProxy(value)) throw new Error(`${label} must not be a Proxy`) + } catch { + throw new Error(`${label} must not be a Proxy`) + } +} + +function nullableIdentifier(value: unknown, label: string): string | null { + return value === null ? null : validateAuthoritySafeIdentifier(value, label, 160) +} + +function canonicalIsoTimestamp(value: unknown, label: string): string { + if (typeof value !== "string" || value.length > 40) throw new Error(`${label} is invalid`) + const parsed = new Date(value) + if (Number.isNaN(parsed.getTime()) || parsed.toISOString() !== value) throw new Error(`${label} is invalid`) + return value +} + +function result(status: "committed" | "idempotent", revision: number, setupHash: string, candidateHash: string): ModelSetupCommitResult { + return deepFreeze({ status, revision, setup_hash: setupHash, candidate_hash: candidateHash, restart_required: true }) +} + +function hash(value: unknown): string { + return createHash("sha256").update(canonicalJson(value)).digest("hex") +} + +function canonicalJson(value: unknown): string { + if (value === null || typeof value === "string" || typeof value === "boolean") return JSON.stringify(value) + if (typeof value === "number") { + if (!Number.isFinite(value)) throw new Error("model setup hash input must be finite") + return JSON.stringify(value) + } + if (Array.isArray(value)) return `[${value.map(canonicalJson).join(",")}]` + if (typeof value !== "object" || value === null) throw new Error("model setup hash input must be JSON") + return `{${Object.keys(value).sort().map((key) => `${JSON.stringify(key)}:${canonicalJson((value as Record)[key])}`).join(",")}}` +} + +function deepFreeze(value: T): T { + if (typeof value !== "object" || value === null || Object.isFrozen(value)) return value + for (const key of Object.keys(value)) deepFreeze((value as Record)[key]) + return Object.freeze(value) +} diff --git a/agentcore/runtime/src/model-configuration/opencode-executor-readiness-resolver.test.ts b/agentcore/runtime/src/model-configuration/opencode-executor-readiness-resolver.test.ts new file mode 100644 index 000000000..308b746e8 --- /dev/null +++ b/agentcore/runtime/src/model-configuration/opencode-executor-readiness-resolver.test.ts @@ -0,0 +1,337 @@ +import { describe, expect, test } from "bun:test" +import { EventEmitter } from "node:events" +import { createHash } from "node:crypto" +import { + OpenCodeExecutorModelReadinessResolver, + createPackagedOpenCodeExecutorReadinessResolver, +} from "./opencode-executor-readiness-resolver" +import type { ExecutorModelSelectionProjection } from "./model-configuration-types" +import type { OpenCodeSpawn, OpenCodeSpawnedProcess } from "../opencode/process-adapter" + +const selection = Object.freeze({ + projection_version: 1, + role: "executor", + selection_status: "selected", + availability_status: "role_readiness_unknown", + connection_status: "role_readiness_unknown", + profile_id: "profile-executor-anthropic", + connection_id: "connection-executor-anthropic", + provider_kind: "anthropic", + provider_id: "anthropic", + model_id: "claude-sonnet-4-5-20250929", + credential_binding_id: "credential-executor-anthropic-primary", + provider_mapping_id: "executor-anthropic-v1", + connection_authority_hash: "1".repeat(64), + profile_hash: "2".repeat(64), + binding_hash: "5".repeat(64), + provider_mapping_hash: "6".repeat(64), + provider_mapping_policy_hash: "7".repeat(64), + projection_hash: "3".repeat(64), +}) satisfies ExecutorModelSelectionProjection + +function observation(overrides: Record = {}) { + const value = { + observation_version: 1 as const, + selection_projection_hash: selection.projection_hash, + provider_id: selection.provider_id, + model_id: selection.model_id, + credential_binding_id: selection.credential_binding_id, + provider_availability_status: "available" as const, + credential_connection_status: "connected" as const, + ...overrides, + } + const semantic = { + policy_version: "nexusloop_opencode_executor_readiness_policy_v1", + selection_projection_hash: value.selection_projection_hash, + provider_id: value.provider_id, + model_id: value.model_id, + credential_binding_id: value.credential_binding_id, + provider_availability_status: value.provider_availability_status, + credential_connection_status: value.credential_connection_status, + } + return { + ...value, + evidence_id: "opencode-readiness-v1-" + createHash("sha256").update(JSON.stringify(semantic)).digest("hex"), + } +} + +function spawnFixture(response: string, options: { code?: number | null; delayMs?: number; signal?: NodeJS.Signals | null } = {}) { + const calls: Array<{ command: string; args: string[]; input: string; cwd: string }> = [] + const spawn: OpenCodeSpawn = (command, args, spawnOptions) => { + const child = new EventEmitter() as EventEmitter & OpenCodeSpawnedProcess + const stdout = new EventEmitter() + const stderr = new EventEmitter() + let input = "" + child.stdout = stdout + child.stderr = stderr + child.stdin = { + write(data: string, callback?: (error?: Error | null) => void) { input += data; callback?.(null); return true }, + end() { + calls.push({ command, args, input, cwd: spawnOptions.cwd }) + setTimeout(() => { + stdout.emit("data", Buffer.from(response)) + child.emit("close", options.code === undefined ? 0 : options.code, options.signal ?? null) + }, options.delayMs ?? 0) + }, + on() { return child.stdin }, + } + child.kill = () => { child.emit("close", null, "SIGTERM"); return true } + queueMicrotask(() => child.emit("spawn")) + return child + } + return { spawn, calls } +} + +describe("packaged OpenCode Executor readiness resolver", () => { + test("production construction uses the exact launch executable and fixed packaged subcommand", async () => { + const fixture = spawnFixture(JSON.stringify(observation()) + "\n") + const resolver = createPackagedOpenCodeExecutorReadinessResolver({ + projectDir: "/tmp/project", + openCodeAdapterConfig: { + kind: "process", + command: "/opt/nexusloop/opencode", + args: ["run", "--format", "json"], + cwd: "/tmp/project", + }, + spawn: fixture.spawn, + }) + expect(await resolver.observe(selection)).toEqual(observation()) + expect(fixture.calls).toHaveLength(1) + expect(fixture.calls[0]).toMatchObject({ + command: "/opt/nexusloop/opencode", + args: ["nexusloop", "executor-readiness-v1"], + cwd: "/tmp/project", + }) + expect(JSON.parse(fixture.calls[0]!.input)).toEqual({ + request_version: "nexusloop_opencode_executor_readiness_request_v1", + selection_projection_hash: selection.projection_hash, + provider_id: selection.provider_id, + model_id: selection.model_id, + credential_binding_id: selection.credential_binding_id, + }) + }) + + test("caller mutation cannot redirect the snapshotted executable authority", async () => { + const fixture = spawnFixture(JSON.stringify(observation()) + "\n") + const config = { + kind: "process" as const, + command: "/opt/nexusloop/opencode", + args: ["run"], + env: { SAFE_MODE: "1" }, + } + const resolver = createPackagedOpenCodeExecutorReadinessResolver({ + projectDir: "/tmp/project", + openCodeAdapterConfig: config, + spawn: fixture.spawn, + }) + config.command = "/tmp/forged" + config.args[0] = "forged" + config.env.SAFE_MODE = "0" + await resolver.observe(selection) + expect(fixture.calls[0]).toMatchObject({ command: "/opt/nexusloop/opencode", args: ["nexusloop", "executor-readiness-v1"] }) + }) + + test("fake adapters and alternate readiness authority fail closed", () => { + expect(() => createPackagedOpenCodeExecutorReadinessResolver({ + projectDir: "/tmp/project", + openCodeAdapterConfig: { kind: "fake" }, + })).toThrow("process") + expect(() => createPackagedOpenCodeExecutorReadinessResolver({ + projectDir: "/tmp/project", + openCodeAdapterConfig: { + kind: "process", + command: "/opt/opencode", + args: [], + readinessCommand: "/tmp/forged", + } as never, + })).toThrow() + }) + + test("identity, protocol, extra fields, multiple records, and trailing output fail closed", async () => { + for (const response of [ + JSON.stringify(observation({ provider_id: "google" })) + "\n", + JSON.stringify(observation({ model_id: "gemini-2.5-flash" })) + "\n", + JSON.stringify(observation({ selection_projection_hash: "4".repeat(64) })) + "\n", + JSON.stringify(observation({ credential_binding_id: "credential-executor-other" })) + "\n", + JSON.stringify(observation({ observation_version: 2 })) + "\n", + JSON.stringify(observation({ extra: true })) + "\n", + "{\"observation_version\":1,\"observation_version\":1}\n", + "not-json\n", + JSON.stringify(observation()) + "\n" + JSON.stringify(observation()) + "\n", + JSON.stringify(observation()) + " trailing", + ]) { + const fixture = spawnFixture(response) + const resolver = new OpenCodeExecutorModelReadinessResolver({ + command: "/opt/opencode", cwd: "/tmp/project", spawn: fixture.spawn, + }) + await expect(resolver.observe(selection)).rejects.toThrow() + } + }) + + test("availability and credential observations remain independent tri-state evidence", async () => { + for (const [providerStatus, credentialStatus] of [ + ["available", "disconnected"], + ["unavailable", "connected"], + ["unknown", "unknown"], + ] as const) { + const expected = observation({ + provider_availability_status: providerStatus, + credential_connection_status: credentialStatus, + }) + const fixture = spawnFixture(JSON.stringify(expected) + "\n") + const resolver = new OpenCodeExecutorModelReadinessResolver({ + command: "/opt/opencode", cwd: "/tmp/project", spawn: fixture.spawn, + }) + await expect(resolver.observe(selection)).resolves.toMatchObject({ + provider_availability_status: providerStatus, + credential_connection_status: credentialStatus, + }) + expect(fixture.calls).toHaveLength(1) + } + }) + + test("oversized output, timeout, nonzero exit, and shutdown settle without retry", async () => { + const oversized = spawnFixture("x".repeat(4097)) + const oversizedResolver = new OpenCodeExecutorModelReadinessResolver({ + command: "/opt/opencode", cwd: "/tmp/project", spawn: oversized.spawn, + }) + await expect(oversizedResolver.observe(selection)).rejects.toThrow() + expect(oversized.calls).toHaveLength(1) + + const failed = spawnFixture("", { code: 2 }) + const failedResolver = new OpenCodeExecutorModelReadinessResolver({ + command: "/opt/opencode", cwd: "/tmp/project", spawn: failed.spawn, + }) + await expect(failedResolver.observe(selection)).rejects.toThrow() + expect(failed.calls).toHaveLength(1) + + const delayed = spawnFixture(JSON.stringify(observation()) + "\n", { delayMs: 50 }) + const resolver = new OpenCodeExecutorModelReadinessResolver({ + command: "/opt/opencode", cwd: "/tmp/project", spawn: delayed.spawn, timeoutMs: 10, + }) + await expect(resolver.observe(selection)).rejects.toThrow() + await resolver.shutdown() + expect(resolver.activeCount()).toBe(0) + }) + + test("synchronous spawn failure and signal termination are bounded and do not retry", async () => { + let spawnCalls = 0 + const throwingSpawn: OpenCodeSpawn = () => { + spawnCalls += 1 + throw new Error("https://secret.example Authorization NXL_PRIVATE_KEY") + } + const throwing = new OpenCodeExecutorModelReadinessResolver({ + command: "/opt/opencode", cwd: "/tmp/project", spawn: throwingSpawn, + }) + await expect(throwing.observe(selection)).rejects.toThrow("Executor readiness observation failed") + await expect(throwing.observe(selection)).rejects.not.toThrow("secret.example") + expect(spawnCalls).toBe(2) + expect(throwing.activeCount()).toBe(0) + + const signalled = spawnFixture("", { code: null, signal: "SIGTERM" }) + const signalledResolver = new OpenCodeExecutorModelReadinessResolver({ + command: "/opt/opencode", cwd: "/tmp/project", spawn: signalled.spawn, + }) + await expect(signalledResolver.observe(selection)).rejects.toThrow("did not complete") + expect(signalled.calls).toHaveLength(1) + }) + + test("concurrency is capped at two and shutdown owns both active children", async () => { + const fixture = spawnFixture(JSON.stringify(observation()) + "\n", { delayMs: 100 }) + const resolver = new OpenCodeExecutorModelReadinessResolver({ + command: "/opt/opencode", cwd: "/tmp/project", spawn: fixture.spawn, + }) + const first = resolver.observe(selection) + const second = resolver.observe(selection) + await expect(resolver.observe(selection)).rejects.toThrow("capacity") + expect(resolver.activeCount()).toBe(2) + await resolver.shutdown() + await expect(first).rejects.toThrow("shutdown") + await expect(second).rejects.toThrow("shutdown") + expect(fixture.calls).toHaveLength(2) + expect(resolver.activeCount()).toBe(0) + }) + + test("startup waits for an owned pre-start observation without cancelling it", async () => { + const fixture = spawnFixture(JSON.stringify(observation()) + "\n", { delayMs: 50 }) + const resolver = new OpenCodeExecutorModelReadinessResolver({ + command: "/opt/opencode", cwd: "/tmp/project", spawn: fixture.spawn, + }) + const pending = resolver.observe(selection) + const starting = resolver.start() + await expect(Promise.race([ + starting.then(() => "started" as const), + new Promise<"waiting">((resolve) => setTimeout(() => resolve("waiting"), 10)), + ])).resolves.toBe("waiting") + await expect(resolver.observe(selection)).rejects.toThrow("startup is in progress") + await expect(pending).resolves.toEqual(observation()) + await expect(starting).resolves.toBeUndefined() + await expect(resolver.observe(selection)).resolves.toEqual(observation()) + expect(fixture.calls).toHaveLength(2) + await resolver.shutdown() + }) + + test("startup remains bounded when a timed-out observer never closes", async () => { + const signals: string[] = [] + const spawn: OpenCodeSpawn = () => { + const child = new EventEmitter() as EventEmitter & OpenCodeSpawnedProcess + child.stdout = new EventEmitter() + child.stderr = new EventEmitter() + child.stdin = { + write(_data: string, callback?: (error?: Error | null) => void) { callback?.(null); return true }, + end() {}, + on() { return child.stdin }, + } + child.kill = (signal) => { + signals.push(String(signal)) + if (signal === "SIGTERM") queueMicrotask(() => child.emit("error", new Error("termination failed"))) + return true + } + return child + } + const resolver = new OpenCodeExecutorModelReadinessResolver({ + command: "/opt/opencode", cwd: "/tmp/project", spawn, timeoutMs: 5, + }) + const pending = resolver.observe(selection) + const starting = resolver.start() + await expect(pending).rejects.toThrow("timed out") + await expect(starting).resolves.toBeUndefined() + expect(signals).toEqual(["SIGTERM", "SIGKILL"]) + expect(resolver.activeCount()).toBe(0) + await resolver.shutdown() + + signals.length = 0 + const shutdownResolver = new OpenCodeExecutorModelReadinessResolver({ + command: "/opt/opencode", cwd: "/tmp/project", spawn, timeoutMs: 50, + }) + const shutdownPending = shutdownResolver.observe(selection) + await expect(shutdownResolver.shutdown()).resolves.toBeUndefined() + await expect(shutdownPending).rejects.toThrow("shutdown") + expect(signals).toEqual(["SIGTERM", "SIGKILL"]) + expect(shutdownResolver.activeCount()).toBe(0) + }) + + test("production factory rejects proxy and accessor authority without executing caller code", () => { + let traps = 0 + const proxied = new Proxy({ + projectDir: "/tmp/project", + openCodeAdapterConfig: { kind: "process", command: "/opt/opencode" }, + }, { + get() { traps += 1; return undefined }, + ownKeys() { traps += 1; return [] }, + getOwnPropertyDescriptor() { traps += 1; return undefined }, + }) + expect(() => createPackagedOpenCodeExecutorReadinessResolver(proxied as never)).toThrow("Proxy") + expect(traps).toBe(0) + + let getterCalls = 0 + const accessor = Object.defineProperty({ + projectDir: "/tmp/project", + }, "openCodeAdapterConfig", { + enumerable: true, + get() { getterCalls += 1; return { kind: "process", command: "/opt/opencode" } }, + }) + expect(() => createPackagedOpenCodeExecutorReadinessResolver(accessor as never)).toThrow("own data") + expect(getterCalls).toBe(0) + }) +}) diff --git a/agentcore/runtime/src/model-configuration/opencode-executor-readiness-resolver.ts b/agentcore/runtime/src/model-configuration/opencode-executor-readiness-resolver.ts new file mode 100644 index 000000000..1ea6c7d4e --- /dev/null +++ b/agentcore/runtime/src/model-configuration/opencode-executor-readiness-resolver.ts @@ -0,0 +1,442 @@ +import { createHash } from "node:crypto" +import { spawn as nodeSpawn } from "node:child_process" +import { types as nodeUtilTypes } from "node:util" +import type { ExecutorModelSelectionProjection } from "./model-configuration-types" +import { + validateOpenCodeAdapterConfig, + type OpenCodeAdapterConfig, +} from "../opencode/adapter-config" +import type { + OpenCodeSpawn, + OpenCodeSpawnedProcess, +} from "../opencode/process-adapter" +import type { + ExecutorModelReadinessObservation, + ExecutorModelReadinessResolver, +} from "./model-profile-runtime-registry-types" + +export const OPENCODE_EXECUTOR_READINESS_PROTOCOL_VERSION = 1 as const +export const OPENCODE_EXECUTOR_READINESS_PROTOCOL_POLICY = "nexusloop_opencode_executor_readiness_observation_v1" as const +export const OPENCODE_EXECUTOR_READINESS_REQUEST_VERSION = "nexusloop_opencode_executor_readiness_request_v1" as const + +const HASH = /^[a-f0-9]{64}$/ +const SAFE_ID = /^[A-Za-z0-9_.:-]+$/ +const DEFAULT_TIMEOUT_MS = 5_000 +const DEFAULT_MAX_OUTPUT_BYTES = 4_096 +const DEFAULT_MAX_CONCURRENCY = 2 +const MAX_INPUT_BYTES = 2_048 +const TERMINATION_GRACE_MS = 100 + +type ResolverOptions = Readonly<{ + command: string + args?: readonly string[] + cwd: string + env?: Readonly> + timeoutMs?: number + maxOutputBytes?: number + maxConcurrency?: number + spawn?: OpenCodeSpawn +}> + +type ActiveObservation = { + child: OpenCodeSpawnedProcess + promise: Promise + rejectForShutdown: () => void +} + +export class OpenCodeExecutorModelReadinessResolver implements ExecutorModelReadinessResolver { + readonly #command: string + readonly #args: readonly string[] + readonly #cwd: string + readonly #env?: Readonly> + readonly #timeoutMs: number + readonly #maxOutputBytes: number + readonly #maxConcurrency: number + readonly #spawn: OpenCodeSpawn + readonly #active = new Set() + #lifecycleState: "available" | "starting" | "stopping" = "available" + #startTask: Promise | null = null + + constructor(options: ResolverOptions) { + const input = resolverOptions(options) + this.#command = boundedText(input.command, "observer command", 1_024) + this.#args = Object.freeze(copyDenseStrings(input.args ?? [], "observer arguments", 1_024)) + this.#cwd = boundedText(input.cwd, "observer working directory", 4_096) + this.#env = input.env ? Object.freeze(copyEnvironment(input.env)) : undefined + this.#timeoutMs = positiveBoundedInteger(input.timeoutMs ?? DEFAULT_TIMEOUT_MS, "observer timeout", 60_000) + this.#maxOutputBytes = positiveBoundedInteger(input.maxOutputBytes ?? DEFAULT_MAX_OUTPUT_BYTES, "observer output limit", 65_536) + this.#maxConcurrency = positiveBoundedInteger(input.maxConcurrency ?? DEFAULT_MAX_CONCURRENCY, "observer concurrency", 8) + this.#spawn = (input.spawn as OpenCodeSpawn | undefined) ?? defaultSpawn + } + + observe(selection: ExecutorModelSelectionProjection): Promise { + if (this.#lifecycleState === "stopping") return Promise.reject(new Error("Executor readiness observation failed: Runtime shutdown is in progress")) + if (this.#lifecycleState === "starting") return Promise.reject(new Error("Executor readiness observation failed: Runtime startup is in progress")) + if (this.#active.size >= this.#maxConcurrency) return Promise.reject(new Error("Executor readiness observation failed: observer capacity is exhausted")) + const input = requestFor(selection) + const requestText = `${JSON.stringify(input)}\n` + if (Buffer.byteLength(requestText, "utf8") > MAX_INPUT_BYTES) { + return Promise.reject(new Error("Executor readiness observation failed: request exceeded its bound")) + } + let child: OpenCodeSpawnedProcess + try { + child = this.#spawn(this.#command, [...this.#args], { + cwd: this.#cwd, + env: childEnvironment(this.#env), + }) + } catch { + return Promise.reject(new Error("Executor readiness observation failed: process start failed")) + } + let rejectForShutdown = () => {} + const promise = new Promise((resolve, reject) => { + let stdout = Buffer.alloc(0) + let settled = false + let terminationError: Error | null = null + let timer: ReturnType | undefined + let terminationTimer: ReturnType | undefined + const finish = (error?: Error, value?: ExecutorModelReadinessObservation) => { + if (settled) return + settled = true + if (timer) clearTimeout(timer) + if (terminationTimer) clearTimeout(terminationTimer) + if (error) reject(error) + else resolve(value!) + } + let terminating = false + const terminate = (error: Error) => { + if (terminating) return + terminating = true + terminationError = error + if (timer) clearTimeout(timer) + try { child.kill?.("SIGTERM") } catch {} + terminationTimer = setTimeout(() => { + try { child.kill?.("SIGKILL") } catch {} + finish(terminationError ?? new Error("Executor readiness observation failed: observer process did not complete")) + }, TERMINATION_GRACE_MS) + } + rejectForShutdown = () => { + terminate(new Error("Executor readiness observation failed: Runtime shutdown cancelled the observation")) + } + child.stdout?.on("data", (chunk: unknown) => { + if (settled) return + const bytes = Buffer.isBuffer(chunk) ? chunk : Buffer.from(String(chunk)) + if (stdout.length + bytes.length > this.#maxOutputBytes) { + terminate(new Error("Executor readiness observation failed: output exceeded its bound")) + return + } + stdout = Buffer.concat([stdout, bytes]) + }) + child.stderr?.on("data", () => {}) + child.on("error", () => { + if (!terminating) finish(new Error("Executor readiness observation failed: process start failed")) + }) + child.on("close", (code, signal) => { + if (settled) return + if (terminationError) return finish(terminationError) + if (code !== 0 || signal !== null) return finish(new Error("Executor readiness observation failed: observer process did not complete")) + try { + finish(undefined, parseResponse(stdout.toString("utf8"), input)) + } catch (error) { + finish(new Error(error instanceof ObservationIdentityError + ? "Executor readiness observation failed: observation identity mismatch" + : "Executor readiness observation failed: malformed observer output")) + } + }) + timer = setTimeout(() => { + terminate(new Error(`Executor readiness observation failed: timed out after ${this.#timeoutMs}ms`)) + }, this.#timeoutMs) + timer.unref() + child.stdin?.on?.("error", () => { + terminate(new Error("Executor readiness observation failed: request delivery failed")) + }) + child.stdin?.write?.(requestText) + child.stdin?.end?.() + }) + const active: ActiveObservation = { child, promise, rejectForShutdown: () => rejectForShutdown() } + this.#active.add(active) + void promise.finally(() => this.#active.delete(active)).catch(() => {}) + return promise + } + + start(): Promise { + if (this.#startTask) return this.#startTask + this.#lifecycleState = "starting" + const active = [...this.#active] + const task = (async () => { + await Promise.allSettled(active.map((item) => item.promise)) + if (this.#lifecycleState === "stopping") { + throw new Error("Executor readiness observation failed: Runtime shutdown is in progress") + } + this.#lifecycleState = "available" + })() + this.#startTask = task + void task.finally(() => { + if (this.#startTask === task) this.#startTask = null + }).catch(() => {}) + return task + } + + async shutdown(): Promise { + this.#lifecycleState = "stopping" + const active = [...this.#active] + for (const item of active) item.rejectForShutdown() + await Promise.allSettled(active.map((item) => item.promise)) + } + + activeCount(): number { return this.#active.size } +} + +export function createProductionOpenCodeExecutorReadinessResolver(options: { + projectDir: string + openCodeAdapterConfig: OpenCodeAdapterConfig + spawn?: OpenCodeSpawn +}): OpenCodeExecutorModelReadinessResolver { + return createPackagedOpenCodeExecutorReadinessResolver(options) +} + +export function createPackagedOpenCodeExecutorReadinessResolver(options: { + projectDir: string + openCodeAdapterConfig: OpenCodeAdapterConfig + spawn?: OpenCodeSpawn +}): OpenCodeExecutorModelReadinessResolver { + const input = snapshotRecord(options, "packaged Executor readiness options", ["projectDir", "openCodeAdapterConfig"], ["spawn"]) + const projectDir = boundedText(input.projectDir, "project directory", 4_096) + const config = snapshotPackagedOpenCodeAdapterConfig(input.openCodeAdapterConfig) + if (config.kind !== "process") throw new Error("packaged Executor readiness requires the process OpenCode adapter") + return new OpenCodeExecutorModelReadinessResolver({ + command: config.command!, + args: ["nexusloop", "executor-readiness-v1"], + cwd: config.cwd ?? projectDir, + env: config.env, + ...(input.spawn === undefined ? {} : { spawn: input.spawn as OpenCodeSpawn }), + }) +} + +export function snapshotPackagedOpenCodeAdapterConfig(value: unknown): Readonly { + const validated = validateOpenCodeAdapterConfig(snapshotOpenCodeAdapterConfig(value)) + if (validated.args) Object.freeze(validated.args) + if (validated.env) Object.freeze(validated.env) + return Object.freeze(validated) +} + +function snapshotOpenCodeAdapterConfig(value: unknown): OpenCodeAdapterConfig { + const input = snapshotRecord(value, "OpenCode adapter config", ["kind"], [ + "command", "args", "cwd", "env", "spawnTimeoutMs", "writeTimeoutMs", "shutdownTimeoutMs", + ]) + const output: Record = Object.create(null) + output.kind = input.kind + if (input.command !== undefined) output.command = input.command + if (input.args !== undefined) output.args = copyDenseStrings(input.args, "OpenCode adapter arguments", 4_096) + if (input.cwd !== undefined) output.cwd = input.cwd + if (input.env !== undefined) output.env = copyEnvironment(input.env) + for (const key of ["spawnTimeoutMs", "writeTimeoutMs", "shutdownTimeoutMs"] as const) { + if (input[key] !== undefined) output[key] = input[key] + } + return output as unknown as OpenCodeAdapterConfig +} + +function snapshotRecord(value: unknown, label: string, required: readonly string[], optional: readonly string[]): Record { + rejectProxy(value, label) + if (typeof value !== "object" || value === null || Array.isArray(value) + || (Object.getPrototypeOf(value) !== Object.prototype && Object.getPrototypeOf(value) !== null)) { + throw new Error(`${label} must be a plain object`) + } + const allowed = new Set([...required, ...optional]) + const keys = Reflect.ownKeys(value) + if (keys.some((key) => typeof key !== "string" || !allowed.has(key))) throw new Error(`${label} contains unknown fields`) + const output = Object.create(null) as Record + for (const key of required) { + const descriptor = Object.getOwnPropertyDescriptor(value, key) + if (!descriptor || !descriptor.enumerable || !("value" in descriptor)) throw new Error(`${label} requires own data fields`) + output[key] = descriptor.value + } + for (const key of optional) { + const descriptor = Object.getOwnPropertyDescriptor(value, key) + if (!descriptor) continue + if (!descriptor.enumerable || !("value" in descriptor)) throw new Error(`${label} requires own data fields`) + output[key] = descriptor.value + } + return output +} + +function requestFor(selection: ExecutorModelSelectionProjection) { + rejectProxy(selection, "Executor selection") + if (typeof selection !== "object" || selection === null || (Object.getPrototypeOf(selection) !== Object.prototype && Object.getPrototypeOf(selection) !== null)) throw new Error("Executor selection must be a plain object") + const projectionHash = ownData(selection, "projection_hash", "Executor selection") + const providerId = ownData(selection, "provider_id", "Executor selection") + const modelId = ownData(selection, "model_id", "Executor selection") + const credentialBindingId = ownData(selection, "credential_binding_id", "Executor selection") + return Object.freeze({ + request_version: OPENCODE_EXECUTOR_READINESS_REQUEST_VERSION, + selection_projection_hash: hash(projectionHash, "selection projection hash"), + provider_id: identifier(providerId, "provider ID", 160), + model_id: inertText(modelId, "model ID", 240), + credential_binding_id: identifier(credentialBindingId, "credential binding ID", 160), + }) +} + +function parseResponse(text: string, expected: ReturnType): ExecutorModelReadinessObservation { + if (!text.endsWith("\n") || text.includes("\r") || text.slice(0, -1).includes("\n")) { + throw new Error("observer output must contain one canonical JSON record") + } + const line = text.slice(0, -1) + const rawKeys = [...line.matchAll(/"((?:\\.|[^"\\])*)"\s*:/g)].map((match) => JSON.parse(`"${match[1]}"`) as string) + if (rawKeys.length !== 8 || new Set(rawKeys).size !== rawKeys.length) throw new Error("observer output contains duplicate or nested fields") + const value = JSON.parse(line) + const input = strictRecord(value, [ + "observation_version", "selection_projection_hash", "provider_id", "model_id", "credential_binding_id", + "provider_availability_status", "credential_connection_status", "evidence_id", + ]) + if (input.observation_version !== OPENCODE_EXECUTOR_READINESS_PROTOCOL_VERSION) throw new Error("protocol mismatch") + if (input.selection_projection_hash !== expected.selection_projection_hash || input.provider_id !== expected.provider_id + || input.model_id !== expected.model_id || input.credential_binding_id !== expected.credential_binding_id) { + throw new ObservationIdentityError() + } + const selectionProjectionHash = hash(input.selection_projection_hash, "selection projection hash") + const providerId = identifier(input.provider_id, "provider ID", 160) + const modelId = inertText(input.model_id, "model ID", 240) + const credentialBindingId = identifier(input.credential_binding_id, "credential binding ID", 160) + const providerAvailability = enumValue(input.provider_availability_status, ["available", "unavailable", "unknown"] as const) + const credentialConnection = enumValue(input.credential_connection_status, ["connected", "disconnected", "unknown"] as const) + const evidence = { + policy_version: "nexusloop_opencode_executor_readiness_policy_v1", + selection_projection_hash: selectionProjectionHash, + provider_id: providerId, + model_id: modelId, + credential_binding_id: credentialBindingId, + provider_availability_status: providerAvailability, + credential_connection_status: credentialConnection, + } + const expectedEvidenceId = `opencode-readiness-v1-${createHash("sha256").update(JSON.stringify(evidence)).digest("hex")}` + if (input.evidence_id !== expectedEvidenceId) throw new ObservationIdentityError() + return Object.freeze({ + observation_version: 1, + selection_projection_hash: selectionProjectionHash, + provider_id: providerId, + model_id: modelId, + credential_binding_id: credentialBindingId, + provider_availability_status: providerAvailability, + credential_connection_status: credentialConnection, + evidence_id: expectedEvidenceId, + }) +} + +function childEnvironment(overrides?: Readonly>): Record { + const output: Record = Object.create(null) + for (const [key, value] of Object.entries(process.env)) if (value !== undefined) output[key] = value + for (const [key, value] of Object.entries(overrides ?? {})) { + if (value === undefined) delete output[key] + else output[key] = value + } + return output +} + +function resolverOptions(value: unknown): Record { + rejectProxy(value, "observer options") + if (typeof value !== "object" || value === null || Array.isArray(value) || (Object.getPrototypeOf(value) !== Object.prototype && Object.getPrototypeOf(value) !== null)) throw new Error("observer options must be a plain object") + const allowed = new Set(["command", "args", "cwd", "env", "timeoutMs", "maxOutputBytes", "maxConcurrency", "spawn"]) + const keys = Reflect.ownKeys(value) + if (keys.some((key) => typeof key !== "string" || !allowed.has(key))) throw new Error("observer options contain unknown fields") + const output = Object.create(null) as Record + for (const key of keys as string[]) { + const descriptor = Object.getOwnPropertyDescriptor(value, key) + if (!descriptor || !descriptor.enumerable || !("value" in descriptor)) throw new Error("observer options must contain own data fields") + output[key] = descriptor.value + } + if (!Object.hasOwn(output, "command") || !Object.hasOwn(output, "cwd")) throw new Error("observer options are incomplete") + return output +} + +function copyEnvironment(value: unknown): Record { + rejectProxy(value, "observer environment") + if (typeof value !== "object" || value === null || Array.isArray(value) || (Object.getPrototypeOf(value) !== Object.prototype && Object.getPrototypeOf(value) !== null)) throw new Error("observer environment must be a plain object") + const output = Object.create(null) as Record + if (Object.getOwnPropertySymbols(value).length > 0) throw new Error("observer environment must not contain symbols") + for (const key of Object.keys(value)) { + const descriptor = Object.getOwnPropertyDescriptor(value, key) + if (!descriptor || !descriptor.enumerable || !("value" in descriptor)) throw new Error("observer environment must contain own data properties") + if (descriptor.value !== undefined && typeof descriptor.value !== "string") throw new Error("observer environment values must be strings") + output[key] = descriptor.value + } + return output +} + +function ownData(value: object, key: string, label: string): unknown { + const descriptor = Object.getOwnPropertyDescriptor(value, key) + if (!descriptor || !descriptor.enumerable || !("value" in descriptor)) throw new Error(`${label} must contain own data fields`) + return descriptor.value +} + +function copyDenseStrings(value: unknown, label: string, max: number): string[] { + rejectProxy(value, label) + if (!Array.isArray(value)) throw new Error(`${label} must be an array`) + const keys = Reflect.ownKeys(value) + const allowed = new Set(["length", ...Array.from({ length: value.length }, (_, index) => String(index))]) + if (keys.some((key) => typeof key !== "string" || !allowed.has(key))) throw new Error(`${label} must be a dense plain array`) + const output: string[] = [] + for (let index = 0; index < value.length; index += 1) { + const descriptor = Object.getOwnPropertyDescriptor(value, String(index)) + if (!descriptor || !descriptor.enumerable || !("value" in descriptor)) throw new Error(`${label} must be a dense own-data array`) + output.push(boundedText(descriptor.value, "observer argument", max)) + } + return output +} + +class ObservationIdentityError extends Error {} + +function strictRecord(value: unknown, keys: readonly string[]): Record { + rejectProxy(value, "Executor readiness response") + if (typeof value !== "object" || value === null || Array.isArray(value) || Object.getPrototypeOf(value) !== Object.prototype) throw new Error("response must be a plain object") + const own = Reflect.ownKeys(value) + const expected = new Set(keys) + if (own.length !== expected.size || own.some((key) => typeof key !== "string" || !expected.has(key))) throw new Error("response contains unknown or missing fields") + const output = Object.create(null) as Record + for (const key of keys) { + const descriptor = Object.getOwnPropertyDescriptor(value, key) + if (!descriptor || !descriptor.enumerable || !("value" in descriptor)) throw new Error("response fields must be own data properties") + output[key] = descriptor.value + } + return output +} + +function identifier(value: unknown, label: string, max: number): string { + if (typeof value !== "string" || value.length < 1 || value.length > max || value !== value.trim() || !SAFE_ID.test(value)) throw new Error(`${label} is invalid`) + return value +} +function inertText(value: unknown, label: string, max: number): string { + if (typeof value !== "string" || value.length < 1 || value.length > max || value !== value.trim() || /[\u0000-\u001f\u007f]/.test(value)) throw new Error(`${label} is invalid`) + return value +} +function boundedText(value: unknown, label: string, max: number): string { + if (typeof value !== "string" || value.length < 1 || value.length > max || value !== value.trim() || /[\u0000\r\n]/.test(value)) throw new Error(`${label} is invalid`) + return value +} +function positiveBoundedInteger(value: unknown, label: string, max: number): number { + if (!Number.isSafeInteger(value) || (value as number) < 1 || (value as number) > max) throw new Error(`${label} is invalid`) + return value as number +} +function hash(value: unknown, label: string): string { + if (typeof value !== "string" || !HASH.test(value)) throw new Error(`${label} is invalid`) + return value +} +function enumValue(value: unknown, allowed: T): T[number] { + if (typeof value !== "string" || !allowed.includes(value)) throw new Error("observation status is invalid") + return value +} +function rejectProxy(value: unknown, label: string): void { + if ((typeof value === "object" && value !== null || typeof value === "function") && nodeUtilTypes.isProxy(value)) throw new Error(`${label} must not be a Proxy`) +} +function canonicalJson(value: unknown): string { + if (value === null || typeof value === "string" || typeof value === "boolean") return JSON.stringify(value) + if (typeof value === "number" && Number.isFinite(value)) return JSON.stringify(value) + if (Array.isArray(value)) return `[${Array.from(value, canonicalJson).join(",")}]` + if (typeof value !== "object") throw new Error("non-canonical evidence") + const record = value as Record + return `{${Object.keys(record).sort().map((key) => `${JSON.stringify(key)}:${canonicalJson(record[key])}`).join(",")}}` +} + +const defaultSpawn: OpenCodeSpawn = (command, args, options) => nodeSpawn(command, args, { + cwd: options.cwd, + env: options.env, + stdio: ["pipe", "pipe", "pipe"], +}) as unknown as OpenCodeSpawnedProcess diff --git a/agentcore/runtime/src/runtime.test.ts b/agentcore/runtime/src/runtime.test.ts index 464016e4e..2137c1344 100644 --- a/agentcore/runtime/src/runtime.test.ts +++ b/agentcore/runtime/src/runtime.test.ts @@ -1,5 +1,6 @@ import { afterEach, describe, expect, test } from "bun:test" import { createHash } from "node:crypto" +import { EventEmitter } from "node:events" import { existsSync } from "node:fs" import { chmod, mkdir, mkdtemp, readFile, rename, rm, symlink, writeFile } from "node:fs/promises" import { delimiter, join } from "node:path" @@ -8,6 +9,7 @@ import { Database } from "bun:sqlite" import { RuntimeServer } from "./server" import type { RuntimeResearchDbProjection } from "./server" import { createRuntimeServerFromLaunchConfig, readRuntimeServerLaunchOptionsFromEnv, readWakeSchedulerBootstrapConfigFromEnv } from "./launch-config" +import type { RuntimeServerLaunchConfig } from "./launch-config" import { RuntimeServerClient } from "./tui/runtime-server-client" import { EventStore } from "./events/event-store" import { RuntimeEventBus } from "./events/event-bus" @@ -34,7 +36,7 @@ import type { CommanderCycleProvider, CommanderCycleProviderInput, CommanderCycl import type { CommanderExecutorReviewProvider, CommanderExecutorReviewProviderInput, CommanderExecutorReviewProviderResult } from "./commander-executor-review/commander-executor-review-provider" import { SpecService } from "./spec/spec-service" import { FakeOpenCodeAdapter } from "./opencode/fake-adapter" -import { ProcessOpenCodeAdapter, type OpenCodeSpawnedProcess, type OpenCodeProcessEventSource } from "./opencode/process-adapter" +import { ProcessOpenCodeAdapter, type OpenCodeSpawn, type OpenCodeSpawnedProcess, type OpenCodeProcessEventSource } from "./opencode/process-adapter" import { createOpenCodeAdapter, readOpenCodeAdapterConfigFromEnv, redactOpenCodeAdapterConfig, validateOpenCodeAdapterConfig } from "./opencode/adapter-config" import { buildOpenCodeSessionContract } from "./opencode/session-contract" import { OpenCodeSessionInstructionPackService } from "./opencode-session/opencode-session-instruction-pack-service" @@ -84,6 +86,8 @@ import { CommanderToolService } from "./commander-tools/commander-tool-service" import { COMMANDER_TOOL_REGISTRY } from "./commander-tools/commander-tool-registry" import { readBoundedDirectoryEntries } from "./commander-tools/commander-repo-read-service" import { RestrictedGitReadAdapter, restrictedGitLogArgs, restrictedGitReadEnv } from "./commander-tools/restricted-git-read-adapter" +import { MODEL_SETUP_EVENT_KIND, ModelSetupService } from "./model-configuration/model-setup" +import { OpenCodeExecutorModelReadinessResolver } from "./model-configuration/opencode-executor-readiness-resolver" const cleanup: string[] = [] const NON_BLOCKING_START_TIMEOUT_MS = 1000 @@ -94,6 +98,92 @@ async function tempProject(): Promise { return dir } +async function openCodeObservationEnvironment( + dir: string, + providerId: "anthropic" | "google" | "openai", + modelId: string, + options: { connected?: boolean; launchCapture?: string; runtimeSession?: boolean } = {}, +): Promise> { + const commandPath = join(dir, "opencode") + await writeFile(commandPath, `#!/usr/bin/env bun +import { createHash } from "node:crypto" +import { writeFileSync } from "node:fs" +if (process.argv[2] !== "nexusloop" || process.argv[3] !== "executor-readiness-v1") { + ${options.launchCapture ? `if (process.argv.includes("run") && process.argv.includes("--model")) { writeFileSync(${JSON.stringify(options.launchCapture)}, JSON.stringify(process.argv.slice(2))); process.exit(0) }` : ""} + ${options.runtimeSession ? `process.stdin.resume() + process.on("SIGTERM", () => process.exit(0)) + process.on("SIGINT", () => process.exit(0)) + await new Promise(() => {})` : ""} + process.exit(3) +} +const input = JSON.parse(await new Response(Bun.stdin.stream()).text()) +const connected = Boolean(process.env.OPENCODE_AUTH_CONTENT) +const semantic = { + policy_version: "nexusloop_opencode_executor_readiness_policy_v1", + selection_projection_hash: input.selection_projection_hash, + provider_id: input.provider_id, + model_id: input.model_id, + credential_binding_id: input.credential_binding_id, + provider_availability_status: "available", + credential_connection_status: connected ? "connected" : "disconnected", +} +process.stdout.write(JSON.stringify({ + observation_version: 1, + selection_projection_hash: input.selection_projection_hash, + provider_id: input.provider_id, + model_id: input.model_id, + credential_binding_id: input.credential_binding_id, + provider_availability_status: "available", + credential_connection_status: connected ? "connected" : "disconnected", + evidence_id: "opencode-readiness-v1-" + createHash("sha256").update(JSON.stringify(semantic)).digest("hex"), +}) + "\\n") +`, "utf8") + await chmod(commandPath, 0o755) + const modelsPath = join(dir, `opencode-models-${providerId}.json`) + await writeFile(modelsPath, JSON.stringify({ + [providerId]: { + id: providerId, + name: providerId, + env: [`${providerId.toUpperCase()}_API_KEY`], + models: { + [modelId]: { + id: modelId, + name: modelId, + release_date: "2025-01-01", + attachment: false, + reasoning: false, + temperature: true, + tool_call: true, + limit: { context: 128_000, output: 8_192 }, + }, + }, + }, + }), "utf8") + return { + PATH: `${dir}${delimiter}${process.env.PATH ?? ""}`, + HOME: join(dir, "opencode-home"), + XDG_CONFIG_HOME: join(dir, "opencode-config"), + XDG_DATA_HOME: join(dir, "opencode-data"), + OPENCODE_MODELS_PATH: modelsPath, + ...(options.connected === false ? {} : { + OPENCODE_AUTH_CONTENT: JSON.stringify({ [providerId]: { type: "api", key: "fixture-secret-never-published" } }), + }), + } +} + +async function commitUnconfiguredModelSetup(dir: string): Promise { + const service = new ModelSetupService({ eventStore: new EventStore(join(dir, ".nxl", "events.jsonl")) }) + const choices = { commander_recipe_id: null, executor_recipe_id: null } as const + const preview = await service.preview(choices) + await service.confirm({ + ...choices, + expected_revision: preview.expected_revision, + candidate_hash: preview.candidate_hash, + confirmed_by: "runtime-test-operator", + confirmation: "CONFIRM_MODEL_SETUP", + }) +} + function executorOnlyRuntimeRegistry(): ModelProfileRuntimeRegistry { return new ModelProfileRuntimeRegistry({ authority_source: "explicit", @@ -126,6 +216,699 @@ afterEach(async () => { while (cleanup.length) await rm(cleanup.pop()!, { recursive: true, force: true }) }) +describe("9W4E runtime model setup", () => { + test("production auto-start rejects missing model authority before Runtime lifecycle ownership", async () => { + const dir = await tempProject() + await makeProject(dir, { approvedSpec: true }) + const adapter = new LongLivedAdapter() + const server = createRuntimeServerFromLaunchConfig({ + projectDir: dir, + env: {}, + adapter, + wakeSchedulerBootstrapConfig: { + autostart_enabled: true, + interval_ms: 1_000, + max_due_items: 1, + dry_run: true, + stop_on_error: true, + }, + }) + const client = new RuntimeServerClient({ server, autoStart: true, ownsServer: false }) + + await expect(client.command("runtime.status")).rejects.toThrow("model setup is required before Runtime startup") + expect(adapter.startCalls).toBe(0) + expect(await server.status()).toMatchObject({ runtimeStatus: "created", lockHeld: false }) + expect(await readEventKinds(dir)).toEqual([]) + }) + + test("startup waits for an active model-setup status readiness inspection", async () => { + const dir = await tempProject() + await makeProject(dir, { approvedSpec: true }) + let observationStarted!: () => void + let releaseObservation!: () => void + const entered = new Promise((resolve) => { observationStarted = resolve }) + const gate = new Promise((resolve) => { releaseObservation = resolve }) + const spawn: OpenCodeSpawn = (_command, _args, _options) => { + const child = new EventEmitter() as EventEmitter & OpenCodeSpawnedProcess + const stdout = new EventEmitter() + const stderr = new EventEmitter() + let requestText = "" + child.stdout = stdout + child.stderr = stderr + child.stdin = { + write(data: string, callback?: (error?: Error | null) => void) { + requestText += data + callback?.(null) + return true + }, + end() { + observationStarted() + void gate.then(() => { + const request = JSON.parse(requestText) as Record + const semantic = { + policy_version: "nexusloop_opencode_executor_readiness_policy_v1", + selection_projection_hash: request.selection_projection_hash, + provider_id: request.provider_id, + model_id: request.model_id, + credential_binding_id: request.credential_binding_id, + provider_availability_status: "available", + credential_connection_status: "connected", + } + stdout.emit("data", Buffer.from(JSON.stringify({ + observation_version: 1, + selection_projection_hash: request.selection_projection_hash, + provider_id: request.provider_id, + model_id: request.model_id, + credential_binding_id: request.credential_binding_id, + provider_availability_status: "available", + credential_connection_status: "connected", + evidence_id: `opencode-readiness-v1-${createHash("sha256").update(JSON.stringify(semantic)).digest("hex")}`, + }) + "\n")) + child.emit("close", 0, null) + }) + }, + on() { return child.stdin }, + } + child.kill = () => { child.emit("close", null, "SIGTERM"); return true } + return child + } + const resolver = new OpenCodeExecutorModelReadinessResolver({ + command: "/opt/opencode", + cwd: dir, + spawn, + }) + const adapter = new LongLivedAdapter() + const server = new RuntimeServer({ + projectDir: dir, + adapter, + modelProfileRuntimeRegistry: executorOnlyRuntimeRegistry(), + executorModelReadinessResolver: resolver, + researchProjectionMode: "disabled", + }) + + const status = server.command("runtime.model_setup_status") as Promise> + await entered + const starting = server.start() + await expect(Promise.race([ + starting.then(() => "started" as const), + Bun.sleep(20).then(() => "waiting" as const), + ])).resolves.toBe("waiting") + expect(adapter.startCalls).toBe(0) + + releaseObservation() + await expect(status).resolves.toMatchObject({ executor_role_readiness: { ready: true } }) + await expect(starting).resolves.toBeUndefined() + expect(adapter.startCalls).toBe(1) + expect((await server.eventStore.readAll()).some((event) => event.kind === "runtime_started")).toBe(true) + await server.shutdown() + }) + + test("setup confirmation cannot append across RuntimeServer startup", async () => { + const dir = await tempProject() + await makeProject(dir, { approvedSpec: true }) + let entered!: () => void + let release!: () => void + const started = new Promise((resolve) => { entered = resolve }) + const gate = new Promise((resolve) => { release = resolve }) + const server = new RuntimeServer({ + projectDir: dir, + adapter: new LongLivedAdapter(), + executorModelReadinessResolver: { + start: async () => { + entered() + await gate + }, + observe: async () => { throw new Error("readiness is not selected") }, + }, + }) + const starting = server.start() + await started + const choices = { commander_recipe_id: null, executor_recipe_id: "executor-openai-gpt-4-1-mini" } as const + const preview = await server.command("runtime.preview_model_setup", choices) as { expected_revision: number; candidate_hash: string } + await expect(server.command("runtime.confirm_model_setup", { + ...choices, + expected_revision: preview.expected_revision, + candidate_hash: preview.candidate_hash, + confirmed_by: "operator", + confirmation: "CONFIRM_MODEL_SETUP", + })).rejects.toThrow("runtime lifecycle is starting") + expect((await server.eventStore.readAll()).filter((event) => event.kind === MODEL_SETUP_EVENT_KIND)).toHaveLength(0) + release() + await starting + expect((await server.eventStore.readAll()).filter((event) => event.kind === MODEL_SETUP_EVENT_KIND)).toHaveLength(0) + await server.shutdown() + }) + + test("startup owns lifecycle when setup confirmation is already waiting for the run lock", async () => { + const dir = await tempProject() + await makeProject(dir, { approvedSpec: true }) + const server = new RuntimeServer({ projectDir: dir, adapter: new LongLivedAdapter() }) + const choices = { commander_recipe_id: null, executor_recipe_id: "executor-openai-gpt-4-1-mini" } as const + const preview = await server.command("runtime.preview_model_setup", choices) as { expected_revision: number; candidate_hash: string } + const internal = server as unknown as { runLock: { acquire(): Promise } } + const originalAcquire = internal.runLock.acquire.bind(internal.runLock) + let entered!: () => void + let release!: () => void + const waiting = new Promise((resolve) => { entered = resolve }) + const gate = new Promise((resolve) => { release = resolve }) + let acquisitions = 0 + internal.runLock.acquire = async () => { + acquisitions += 1 + if (acquisitions === 1) { + entered() + await gate + } + await originalAcquire() + } + const confirmation = server.command("runtime.confirm_model_setup", { + ...choices, + expected_revision: preview.expected_revision, + candidate_hash: preview.candidate_hash, + confirmed_by: "operator", + confirmation: "CONFIRM_MODEL_SETUP", + }) + await waiting + const starting = server.start() + release() + await expect(confirmation).rejects.toThrow("runtime lifecycle is starting") + await starting + expect((await server.eventStore.readAll()).filter((event) => event.kind === MODEL_SETUP_EVENT_KIND)).toHaveLength(0) + await server.shutdown() + }) + + test("pre-start shutdown drains setup lock acquisition and prevents a later append", async () => { + const dir = await tempProject() + const server = createRuntimeServerFromLaunchConfig({ projectDir: dir, env: {} }) + const choices = { commander_recipe_id: null, executor_recipe_id: "executor-openai-gpt-4-1-mini" } + const preview = await server.command("runtime.preview_model_setup", choices) as { expected_revision: number; candidate_hash: string } + const internal = server as unknown as { runLock: { acquire(): Promise } } + const originalAcquire = internal.runLock.acquire.bind(internal.runLock) + let release!: () => void + const gate = new Promise((resolve) => { release = resolve }) + internal.runLock.acquire = async () => { + await gate + await originalAcquire() + } + const confirmation = server.command("runtime.confirm_model_setup", { ...choices, expected_revision: preview.expected_revision, candidate_hash: preview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" }) + await Bun.sleep(10) + const shutdown = server.shutdown() + expect(await Promise.race([shutdown.then(() => "settled"), Bun.sleep(20).then(() => "waiting")])).toBe("waiting") + release() + await expect(confirmation).rejects.toThrow("stopping") + await shutdown + expect((await server.eventStore.readAll()).some((event) => event.kind === MODEL_SETUP_EVENT_KIND)).toBe(false) + await Bun.sleep(10) + expect((await server.eventStore.readAll()).some((event) => event.kind === MODEL_SETUP_EVENT_KIND)).toBe(false) + }) + + test("shutdown drains an owned setup append before runtime_shutdown", async () => { + const dir = await tempProject() + await makeProject(dir, { approvedSpec: true }) + await commitUnconfiguredModelSetup(dir) + const server = createRuntimeServerFromLaunchConfig({ projectDir: dir, env: {} }) + await server.start() + const choices = { commander_recipe_id: null, executor_recipe_id: "executor-openai-gpt-4-1-mini" } + const preview = await server.command("runtime.preview_model_setup", choices) as { expected_revision: number; candidate_hash: string } + const store = server.eventStore + const original = store.appendIfLatestKind.bind(store) + let release!: () => void + const gate = new Promise((resolve) => { release = resolve }) + store.appendIfLatestKind = (async (...args: Parameters) => { + await gate + return await original(...args) + }) as typeof store.appendIfLatestKind + const confirmation = server.command("runtime.confirm_model_setup", { ...choices, expected_revision: preview.expected_revision, candidate_hash: preview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" }) + await Bun.sleep(10) + const shutdown = server.shutdown() + expect(await Promise.race([shutdown.then(() => "settled"), Bun.sleep(20).then(() => "waiting")])).toBe("waiting") + release() + await confirmation + await shutdown + expect((await store.readAll()).map((event) => event.kind).slice(-2)).toEqual([MODEL_SETUP_EVENT_KIND, "runtime_shutdown"]) + }) + test("routes catalog status preview and pre-start confirmation through RuntimeServer", async () => { + const dir = await tempProject() + const server = createRuntimeServerFromLaunchConfig({ projectDir: dir, env: {} }) + const catalog = await server.command("runtime.model_setup_catalog") as { commander_recipes: Array<{ recipe_id: string }> } + expect(catalog.commander_recipes).toHaveLength(3) + expect(await server.command("runtime.model_setup_status")).toMatchObject({ status: "missing", revision: 0, pending_restart: false }) + const choices = { commander_recipe_id: "commander-openai-gpt-4-1-mini-responses", executor_recipe_id: "executor-anthropic-claude-sonnet-4-5" } + const preview = await server.command("runtime.preview_model_setup", choices) as { expected_revision: number; candidate_hash: string } + const committed = await server.command("runtime.confirm_model_setup", { + ...choices, + expected_revision: preview.expected_revision, + candidate_hash: preview.candidate_hash, + confirmed_by: "runtime-test-operator", + confirmation: "CONFIRM_MODEL_SETUP", + }) + expect(committed).toMatchObject({ status: "committed", revision: 1, restart_required: true }) + expect(existsSync(join(dir, ".nxl", "run.lock"))).toBe(false) + expect(await server.command("runtime.model_setup_status")).toMatchObject({ status: "ready", revision: 1, pending_restart: true }) + }) + + test("setup status distinguishes active registry authority from missing durable setup", async () => { + const dir = await tempProject() + const server = createRuntimeServerFromLaunchConfig({ + projectDir: dir, + env: {}, + modelProfileRuntimeRegistry: executorOnlyRuntimeRegistry(), + openCodeAdapterConfig: { kind: "process", command: "/opt/opencode" }, + }) + expect(await server.command("runtime.model_setup_status")).toMatchObject({ + status: "missing", + revision: 0, + pending_restart: false, + active_authority_source: "explicit", + }) + const choices = { commander_recipe_id: null, executor_recipe_id: "executor-openai-gpt-4-1-mini" } as const + const preview = await server.command("runtime.preview_model_setup", choices) as { expected_revision: number; candidate_hash: string } + await expect(server.command("runtime.confirm_model_setup", { + ...choices, + expected_revision: preview.expected_revision, + candidate_hash: preview.candidate_hash, + confirmed_by: "operator", + confirmation: "CONFIRM_MODEL_SETUP", + })).rejects.toThrow("non-setup model authority is active") + expect((await server.eventStore.readAll()).filter((event) => event.kind === MODEL_SETUP_EVENT_KIND)).toHaveLength(0) + await server.shutdown() + }) + + test("legacy Commander authority rejects model setup confirmation without appending", async () => { + const dir = await tempProject() + const server = createRuntimeServerFromLaunchConfig({ + projectDir: dir, + env: { + NXL_COMMANDER_INVESTIGATION_PROVIDER_ENABLED: "1", + NXL_COMMANDER_INVESTIGATION_TRANSPORT_KIND: "anthropic_messages_connector", + NXL_COMMANDER_INVESTIGATION_PROVIDER_ID: "anthropic-primary", + NXL_COMMANDER_INVESTIGATION_PROVIDER_KIND: "anthropic", + NXL_COMMANDER_INVESTIGATION_CONNECTOR_ID: "anthropic-main", + NXL_COMMANDER_INVESTIGATION_MODEL_ID: "claude-sonnet-4-5-20250929", + NXL_COMMANDER_INVESTIGATION_ENABLED_PHASES: "proposal_investigation", + NXL_COMMANDER_INVESTIGATION_TIMEOUT_MS: "30000", + NXL_COMMANDER_INVESTIGATION_MAX_REQUEST_BYTES: "65536", + NXL_COMMANDER_INVESTIGATION_MAX_RESPONSE_BYTES: "65536", + NXL_COMMANDER_INVESTIGATION_MAX_CONTEXT_BYTES: "48000", + NXL_COMMANDER_INVESTIGATION_MAX_CONTEXT_TOKENS: "16000", + NXL_COMMANDER_INVESTIGATION_MAX_OUTPUT_TOKENS: "4096", + NXL_COMMANDER_INVESTIGATION_SUPPORTS_TOOLS: "1", + NXL_COMMANDER_INVESTIGATION_SUPPORTS_JSON_SCHEMA: "0", + NXL_COMMANDER_INVESTIGATION_SUPPORTS_LONG_CONTEXT: "1", + NXL_COMMANDER_INVESTIGATION_SUPPORTS_LOCAL_EXECUTION: "0", + }, + }) + expect(await server.command("runtime.model_setup_status")).toMatchObject({ + status: "missing", + active_authority_source: "legacy_commander_environment", + }) + const choices = { commander_recipe_id: null, executor_recipe_id: null } as const + const preview = await server.command("runtime.preview_model_setup", choices) as { expected_revision: number; candidate_hash: string } + await expect(server.command("runtime.confirm_model_setup", { + ...choices, + expected_revision: preview.expected_revision, + candidate_hash: preview.candidate_hash, + confirmed_by: "operator", + confirmation: "CONFIRM_MODEL_SETUP", + })).rejects.toThrow("non-setup model authority is active") + expect((await server.eventStore.readAll()).filter((event) => event.kind === MODEL_SETUP_EVENT_KIND)).toHaveLength(0) + await server.shutdown() + }) + + test("RuntimeServerClient keeps setup preview and confirmation pre-start", async () => { + const dir = await tempProject() + const server = createRuntimeServerFromLaunchConfig({ projectDir: dir, env: {} }) + const client = new RuntimeServerClient({ server, autoStart: true, ownsServer: true }) + const choices = { commander_recipe_id: null, executor_recipe_id: "executor-google-gemini-2-5-flash" } as const + const preview = await client.command("runtime.preview_model_setup", choices) + await client.command("runtime.confirm_model_setup", { ...choices, expected_revision: preview.expected_revision, candidate_hash: preview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" }) + expect(await server.status()).toMatchObject({ runtimeStatus: "created", lockHeld: false, specApproved: false }) + await client.shutdown() + }) + + test("activates persisted setup only on the next RuntimeServer construction", async () => { + const dir = await tempProject() + const service = new ModelSetupService({ eventStore: new EventStore(join(dir, ".nxl", "events.jsonl")) }) + const choices = { commander_recipe_id: "commander-google-gemini-2-5-flash", executor_recipe_id: "executor-openai-gpt-4-1-mini" } as const + const preview = await service.preview(choices) + await service.confirm({ ...choices, expected_revision: 0, candidate_hash: preview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" }) + + const server = createRuntimeServerFromLaunchConfig({ + projectDir: dir, + env: {}, + openCodeAdapterConfig: { kind: "process", command: "/opt/opencode" }, + }) + expect(await server.command("runtime.model_setup_status")).toMatchObject({ + status: "ready", + revision: 1, + pending_restart: false, + active_candidate: { choices }, + }) + const unchangedPreview = await server.command("runtime.preview_model_setup", choices) as { expected_revision: number; candidate_hash: string } + const eventCount = (await server.eventStore.readAll()).filter((event) => event.kind === MODEL_SETUP_EVENT_KIND).length + await expect(server.command("runtime.confirm_model_setup", { + ...choices, + expected_revision: unchangedPreview.expected_revision, + candidate_hash: unchangedPreview.candidate_hash, + confirmed_by: "operator", + confirmation: "CONFIRM_MODEL_SETUP", + })).resolves.toMatchObject({ status: "idempotent", revision: 1, restart_required: false }) + expect((await server.eventStore.readAll()).filter((event) => event.kind === MODEL_SETUP_EVENT_KIND)).toHaveLength(eventCount) + const internal = server as unknown as { modelProfileRuntimeRegistry?: ModelProfileRuntimeRegistry; commanderInvestigationProviderConfig?: { transport_kind: string; model_id: string } } + expect(internal.modelProfileRuntimeRegistry?.commanderSelection()).toMatchObject({ provider_kind: "google", model_id: "gemini-2.5-flash" }) + expect(internal.modelProfileRuntimeRegistry?.executorSelection()).toMatchObject({ provider_kind: "openai", model_id: "gpt-4.1-mini" }) + expect(internal.commanderInvestigationProviderConfig).toMatchObject({ transport_kind: "google_generative_ai_connector", model_id: "gemini-2.5-flash" }) + + const nextChoices = { commander_recipe_id: null, executor_recipe_id: "executor-anthropic-claude-sonnet-4-5" } as const + const nextPreview = await server.command("runtime.preview_model_setup", nextChoices) as { expected_revision: number; candidate_hash: string } + await server.command("runtime.confirm_model_setup", { ...nextChoices, expected_revision: nextPreview.expected_revision, candidate_hash: nextPreview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" }) + expect(internal.modelProfileRuntimeRegistry?.commanderSelection()?.model_id).toBe("gemini-2.5-flash") + expect(await server.command("runtime.model_setup_status")).toMatchObject({ + revision: 2, + pending_restart: true, + active_candidate: { choices }, + candidate: { choices: nextChoices }, + }) + }) + + test("startup revalidates persisted setup after acquiring the run lock", async () => { + const dir = await tempProject() + await makeProject(dir, { approvedSpec: true }) + const staleServer = createRuntimeServerFromLaunchConfig({ projectDir: dir, env: {}, adapter: new LongLivedAdapter() }) + const writer = createRuntimeServerFromLaunchConfig({ projectDir: dir, env: {} }) + const choices = { commander_recipe_id: null, executor_recipe_id: "executor-openai-gpt-4-1-mini" } as const + const preview = await writer.command("runtime.preview_model_setup", choices) as { expected_revision: number; candidate_hash: string } + await writer.command("runtime.confirm_model_setup", { + ...choices, + expected_revision: preview.expected_revision, + candidate_hash: preview.candidate_hash, + confirmed_by: "operator", + confirmation: "CONFIRM_MODEL_SETUP", + }) + + await expect(staleServer.start()).rejects.toThrow("persisted model setup changed before runtime start") + expect((await staleServer.eventStore.readAll()).some((event) => event.kind === "runtime_started")).toBe(false) + expect(existsSync(join(dir, ".nxl", "run.lock"))).toBe(false) + }) + + test("reports Commander and Executor readiness independently from persisted selection", async () => { + const dir = await tempProject() + const service = new ModelSetupService({ eventStore: new EventStore(join(dir, ".nxl", "events.jsonl")) }) + const choices = { commander_recipe_id: "commander-google-gemini-2-5-flash", executor_recipe_id: "executor-google-gemini-2-5-flash" } as const + const preview = await service.preview(choices) + await service.confirm({ ...choices, expected_revision: 0, candidate_hash: preview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" }) + const observationEnv = await openCodeObservationEnvironment(dir, "google", "gemini-2.5-flash") + const server = createRuntimeServerFromLaunchConfig({ + projectDir: dir, + env: {}, + openCodeAdapterConfig: { kind: "process", command: "opencode", env: observationEnv }, + }) + const status = await server.command("runtime.model_setup_status") as Record + expect(status.commander_role_readiness).toMatchObject({ selection_status: "selected", credential_connection_status: "unknown", ready: false }) + expect(status.executor_role_readiness).toMatchObject({ selection_status: "selected", credential_connection_status: "connected", ready: true }) + expect(status.commander_role_readiness.readiness_hash).not.toBe(status.executor_role_readiness.readiness_hash) + }) + + test("shared Commander and Executor setup keeps exact context capability role-isolated and usable", async () => { + const dir = await tempProject() + const setup = new ModelSetupService({ eventStore: new EventStore(join(dir, ".nxl", "events.jsonl")) }) + const choices = { + commander_recipe_id: "commander-anthropic-claude-sonnet-4-5", + executor_recipe_id: "executor-anthropic-claude-sonnet-4-5", + } as const + const preview = await setup.preview(choices) + await setup.confirm({ ...choices, expected_revision: 0, candidate_hash: preview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" }) + const observationEnv = await openCodeObservationEnvironment(dir, "anthropic", "claude-sonnet-4-5-20250929", { runtimeSession: true }) + const server = createRuntimeServerFromLaunchConfig({ + projectDir: dir, + env: {}, + openCodeAdapterConfig: { kind: "process", command: "opencode", env: observationEnv }, + researchProjectionMode: "disabled", + }) + const commander = await server.command("runtime.preview_context_budget", { + purpose: "commander_investigation", + role: "commander", + provider_kind: "anthropic", + model_id: "claude-sonnet-4-5-20250929", + }) as { blockers: string[] } + const executor = await server.command("runtime.preview_context_budget", { + purpose: "opencode_executor_session", + role: "executor", + provider_kind: "anthropic", + model_id: "claude-sonnet-4-5-20250929", + }) as { blockers: string[] } + expect(commander.blockers).not.toContain("selected model capability does not support requested role") + expect(executor.blockers).not.toContain("selected model capability does not support requested role") + const commanderCapability = await server.command("runtime.get_model_capability", { + providerKind: "anthropic", + modelId: "claude-sonnet-4-5-20250929", + role: "commander", + }) as Record + const executorCapability = await server.command("runtime.get_model_capability", { + providerKind: "anthropic", + modelId: "claude-sonnet-4-5-20250929", + role: "executor", + }) as Record + expect(commanderCapability).toMatchObject({ role_support: ["commander"], supports_tools: true }) + expect(executorCapability).toMatchObject({ role_support: ["executor"], supports_tools: "unknown", supports_long_context: "unknown" }) + expect(executorCapability.capability_id).not.toBe(commanderCapability.capability_id) + await server.shutdown() + }) + + test("launch configuration constructs production Executor readiness without package injection", async () => { + const dir = await tempProject() + const service = new ModelSetupService({ eventStore: new EventStore(join(dir, ".nxl", "events.jsonl")) }) + const choices = { commander_recipe_id: null, executor_recipe_id: "executor-google-gemini-2-5-flash" } as const + const preview = await service.preview(choices) + await service.confirm({ ...choices, expected_revision: 0, candidate_hash: preview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" }) + const observationEnv = await openCodeObservationEnvironment(dir, "google", "gemini-2.5-flash", { runtimeSession: true }) + const server = createRuntimeServerFromLaunchConfig({ + projectDir: dir, + env: {}, + openCodeAdapterConfig: { kind: "process", command: "opencode", env: observationEnv }, + }) + const internal = server as unknown as { executorModelReadinessResolver?: { constructor: { name: string } } } + expect(internal.executorModelReadinessResolver?.constructor.name).toBe("OpenCodeExecutorModelReadinessResolver") + await server.shutdown() + }) + + test("production Executor readiness observes process-adapter environment and working-directory authority", async () => { + const dir = await tempProject() + const service = new ModelSetupService({ eventStore: new EventStore(join(dir, ".nxl", "events.jsonl")) }) + const choices = { commander_recipe_id: null, executor_recipe_id: "executor-google-gemini-2-5-flash" } as const + const preview = await service.preview(choices) + await service.confirm({ ...choices, expected_revision: 0, candidate_hash: preview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" }) + const launchCwd = join(dir, "opencode-launch-cwd") + await mkdir(launchCwd, { recursive: true }) + const observationEnv = await openCodeObservationEnvironment(launchCwd, "google", "gemini-2.5-flash", { connected: false }) + const server = createRuntimeServerFromLaunchConfig({ + projectDir: dir, + env: { + GOOGLE_GENERATIVE_AI_API_KEY: "parent-credential-must-be-overridden", + NXL_DETACHED_EXECUTOR_CREDENTIAL: "detached-config-must-not-reach-child", + }, + openCodeAdapterConfig: { + kind: "process", + command: "opencode", + cwd: launchCwd, + env: { ...observationEnv, GOOGLE_GENERATIVE_AI_API_KEY: "" }, + }, + }) + + await expect(server.previewExecutorModelRoleReadiness()).resolves.toMatchObject({ + ready: false, + provider_availability_status: "available", + credential_connection_status: "disconnected", + }) + await server.shutdown() + }) + + test("startup defers a torn setup snapshot until strict projection under the run lock", async () => { + const dir = await tempProject() + await makeProject(dir, { approvedSpec: true }) + const service = new ModelSetupService({ eventStore: new EventStore(join(dir, ".nxl", "events.jsonl")) }) + const choices = { commander_recipe_id: null, executor_recipe_id: "executor-openai-gpt-4-1-mini" } as const + const preview = await service.preview(choices) + const committed = await service.confirm({ ...choices, expected_revision: 0, candidate_hash: preview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" }) + const eventsPath = join(dir, ".nxl", "events.jsonl") + const complete = await readFile(eventsPath, "utf8") + await writeFile(eventsPath, complete.slice(0, -2), "utf8") + + const server = createRuntimeServerFromLaunchConfig({ projectDir: dir, env: {}, adapter: new LongLivedAdapter() }) + await writeFile(eventsPath, complete, "utf8") + + await expect(server.start()).rejects.toThrow("persisted model setup changed before runtime start") + expect(committed).toMatchObject({ status: "committed" }) + expect((await server.eventStore.readAll()).some((event) => event.kind === "runtime_started")).toBe(false) + expect(existsSync(join(dir, ".nxl", "run.lock"))).toBe(false) + }) + + test("startup fails closed when a corrupt setup journal remains corrupt", async () => { + const dir = await tempProject() + await makeProject(dir, { approvedSpec: true }) + const service = new ModelSetupService({ eventStore: new EventStore(join(dir, ".nxl", "events.jsonl")) }) + const choices = { commander_recipe_id: null, executor_recipe_id: "executor-openai-gpt-4-1-mini" } as const + const preview = await service.preview(choices) + await service.confirm({ ...choices, expected_revision: 0, candidate_hash: preview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" }) + const eventsPath = join(dir, ".nxl", "events.jsonl") + const complete = await readFile(eventsPath, "utf8") + await writeFile(eventsPath, complete.slice(0, -2), "utf8") + + const server = createRuntimeServerFromLaunchConfig({ projectDir: dir, env: {}, adapter: new LongLivedAdapter() }) + await expect(server.start()).rejects.toThrow("model setup journal is malformed") + expect((await server.eventStore.readAll().catch(() => [])).some((event) => event.kind === "runtime_started")).toBe(false) + expect(existsSync(join(dir, ".nxl", "run.lock"))).toBe(false) + }) + + test("environment input cannot replace the code-owned production Executor readiness observer", async () => { + const dir = await tempProject() + const service = new ModelSetupService({ eventStore: new EventStore(join(dir, ".nxl", "events.jsonl")) }) + const choices = { commander_recipe_id: null, executor_recipe_id: "executor-openai-gpt-4-1-mini" } as const + const preview = await service.preview(choices) + await service.confirm({ ...choices, expected_revision: 0, candidate_hash: preview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" }) + const marker = join(dir, "forged-observer-ran") + const forged = join(dir, "forged-observer.ts") + await writeFile(forged, `await Bun.write(${JSON.stringify(marker)}, "forged");`, "utf8") + for (const env of [ + { NXL_OPENCODE_EXECUTOR_READINESS_COMMAND: process.execPath }, + { NXL_OPENCODE_EXECUTOR_READINESS_ARGS_JSON: JSON.stringify([forged]) }, + { + NXL_OPENCODE_EXECUTOR_READINESS_COMMAND: process.execPath, + NXL_OPENCODE_EXECUTOR_READINESS_ARGS_JSON: JSON.stringify([forged]), + }, + ]) { + expect(() => createRuntimeServerFromLaunchConfig({ projectDir: dir, env })).toThrow("not supported") + } + expect(() => createRuntimeServerFromLaunchConfig({ + projectDir: dir, + env: { NXL_OPENCODE_EXECUTOR_READINESS_COMMAND: process.execPath }, + executorModelReadinessResolver: { observe: () => { throw new Error("must not run") } }, + } as unknown as RuntimeServerLaunchConfig)).toThrow("not supported") + expect(() => createRuntimeServerFromLaunchConfig({ + projectDir: dir, + env: {}, + executorModelReadinessResolver: { observe: () => { throw new Error("must not run") } }, + } as unknown as RuntimeServerLaunchConfig)).toThrow("not supported") + expect(existsSync(marker)).toBe(false) + }) + + test("persisted Executor setup rejects alternate production launch process seams", async () => { + const dir = await tempProject() + const service = new ModelSetupService({ eventStore: new EventStore(join(dir, ".nxl", "events.jsonl")) }) + const choices = { commander_recipe_id: null, executor_recipe_id: "executor-openai-gpt-4-1-mini" } as const + const preview = await service.preview(choices) + await service.confirm({ ...choices, expected_revision: 0, candidate_hash: preview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" }) + const base = { projectDir: dir, env: {}, openCodeAdapterConfig: { kind: "process" as const, command: "/opt/opencode" } } + const directAdapter = new LongLivedAdapter() + expect(() => createRuntimeServerFromLaunchConfig({ + ...base, + adapter: directAdapter, + })).toThrow("same packaged OpenCode execution target") + expect(directAdapter.startCalls).toBe(0) + let injectedReadinessCalls = 0 + expect(() => createRuntimeServerFromLaunchConfig({ + ...base, + executorModelReadinessResolver: { + observe: () => { + injectedReadinessCalls += 1 + throw new Error("must not run") + }, + }, + } as unknown as RuntimeServerLaunchConfig)).toThrow("custom Executor readiness resolver is not supported") + expect(injectedReadinessCalls).toBe(0) + expect(() => createRuntimeServerFromLaunchConfig({ + projectDir: dir, + env: {}, + })).toThrow("requires a packaged process OpenCode execution target") + expect(() => createRuntimeServerFromLaunchConfig({ + projectDir: dir, + env: {}, + openCodeAdapterConfig: { kind: "fake" }, + })).toThrow("requires a packaged process OpenCode execution target") + expect(() => createRuntimeServerFromLaunchConfig({ + ...base, + opencodeLaunchAdapter: {} as never, + })).toThrow("same packaged OpenCode execution target") + expect(() => createRuntimeServerFromLaunchConfig({ + ...base, + opencodeLaunchSpawn: (() => { throw new Error("must not run") }) as never, + })).toThrow("same packaged OpenCode execution target") + let aliasedSpawnCalls = 0 + expect(() => createRuntimeServerFromLaunchConfig({ + ...base, + openCodeAdapterFactoryOptions: { + spawn: (() => { + aliasedSpawnCalls += 1 + throw new Error("must not run") + }) as never, + }, + })).toThrow("same packaged OpenCode execution target") + expect(aliasedSpawnCalls).toBe(0) + expect(await readEventKinds(dir)).not.toContain("runtime_started") + }) + + test("shutdown terminates and drains production Executor observation before runtime_shutdown", async () => { + const dir = await tempProject() + await makeProject(dir, { approvedSpec: true }) + const service = new ModelSetupService({ eventStore: new EventStore(join(dir, ".nxl", "events.jsonl")) }) + const choices = { commander_recipe_id: null, executor_recipe_id: "executor-google-gemini-2-5-flash" } as const + const preview = await service.preview(choices) + await service.confirm({ ...choices, expected_revision: 0, candidate_hash: preview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" }) + const observationEnv = await openCodeObservationEnvironment(dir, "google", "gemini-2.5-flash") + const server = createRuntimeServerFromLaunchConfig({ + projectDir: dir, + env: {}, + openCodeAdapterConfig: { kind: "process", command: "opencode", env: observationEnv }, + researchProjectionMode: "disabled", + }) + await server.start() + const pending = server.previewExecutorModelRoleReadiness() + await server.shutdown() + await expect(pending).resolves.toMatchObject({ ready: false, provider_availability_status: "unknown" }) + const events = await server.eventStore.readAll() + expect(events.at(-1)?.kind).toBe("runtime_shutdown") + const internal = server as unknown as { executorModelReadinessResolver?: { activeCount(): number } } + expect(internal.executorModelReadinessResolver?.activeCount()).toBe(0) + }) + + test("RuntimeServer restart reactivates the production Executor readiness observer", async () => { + const dir = await tempProject() + await makeProject(dir, { approvedSpec: true }) + const service = new ModelSetupService({ eventStore: new EventStore(join(dir, ".nxl", "events.jsonl")) }) + const choices = { commander_recipe_id: null, executor_recipe_id: "executor-google-gemini-2-5-flash" } as const + const preview = await service.preview(choices) + await service.confirm({ ...choices, expected_revision: 0, candidate_hash: preview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" }) + const observationEnv = await openCodeObservationEnvironment(dir, "google", "gemini-2.5-flash", { runtimeSession: true }) + const server = createRuntimeServerFromLaunchConfig({ + projectDir: dir, + env: {}, + openCodeAdapterConfig: { kind: "process", command: "opencode", env: observationEnv }, + researchProjectionMode: "disabled", + }) + await server.start() + await expect(server.previewExecutorModelRoleReadiness()).resolves.toMatchObject({ ready: true }) + await server.shutdown() + await server.start() + await expect(server.previewExecutorModelRoleReadiness()).resolves.toMatchObject({ ready: true }) + await server.shutdown() + }) + + test("persisted setup conflicts with explicit registry and legacy Commander environment authority", async () => { + const dir = await tempProject() + const service = new ModelSetupService({ eventStore: new EventStore(join(dir, ".nxl", "events.jsonl")) }) + const choices = { commander_recipe_id: "commander-anthropic-claude-sonnet-4-5", executor_recipe_id: null } as const + const preview = await service.preview(choices) + await service.confirm({ ...choices, expected_revision: 0, candidate_hash: preview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" }) + expect(() => createRuntimeServerFromLaunchConfig({ projectDir: dir, env: {}, modelProfileRuntimeRegistry: executorOnlyRuntimeRegistry() })).toThrow("persisted model setup") + expect(() => createRuntimeServerFromLaunchConfig({ + projectDir: dir, + env: { + NXL_COMMANDER_INVESTIGATION_PROVIDER_ENABLED: "1", + NXL_COMMANDER_INVESTIGATION_TRANSPORT_KIND: "anthropic_messages_connector", + }, + })).toThrow("persisted model setup") + }) +}) + function timeout(ms: number): Promise<"timeout"> { return new Promise((resolve) => setTimeout(() => resolve("timeout"), ms)) } @@ -1043,6 +1826,7 @@ describe("CommandAuthorityService", () => { test("registry classifies critical authority and risk boundaries", () => { const service = new CommandAuthorityService(() => "2026-06-19T00:00:00.000Z") + expect(service.get("/model-setup")).toMatchObject({ runtime_command: "runtime.confirm_model_setup", risk: "medium_risk_write", gate: "model_setup_runtime", owner: "model_setup", mutates_events: true, creates_external_process: false, calls_provider: false, requires_active_runtime: false, requires_run_lock: true, requires_approval: true, blocked_by_default: true, expected_event_kinds: ["runtime_model_setup_committed"] }) expect(service.get("/commander-recoveries")).toMatchObject({ risk: "safe_read", runtime_command: "runtime.list_commander_investigation_recoveries", owner: "commander_recovery", mutates_events: false, calls_provider: false }) expect(service.get("/commander-recovery-show")).toMatchObject({ risk: "safe_read", runtime_command: "runtime.get_commander_investigation_recovery", owner: "commander_recovery", mutates_events: false }) expect(service.get("/commander-recovery-preview")).toMatchObject({ risk: "safe_read", runtime_command: "runtime.preview_commander_investigation_recovery", mutates_events: false, calls_provider: false }) @@ -7952,9 +8736,13 @@ describe("RuntimeServer core", () => { const first = service.executeStep({ plan_id: "plan_continue_concurrent_1", index: 0, requested_by: "operator" }) await readStarted const second = service.executeStep({ plan_id: "plan_continue_concurrent_1", index: 0, requested_by: "operator" }) + const secondResult = second.then( + () => ({ error: null }), + (error: unknown) => ({ error }), + ) releaseRead([]) await expect(first).resolves.toMatchObject({ status: "succeeded", command: "/missions" }) - await expect(second).rejects.toThrow("continuation plan is completed") + expect((await secondResult).error).toEqual(new Error("continuation plan is completed")) expect(readCalls).toBe(1) const events = await eventStore.readAll() @@ -11882,6 +12670,7 @@ describe("RuntimeServer core", () => { timers.shift()?.() await waitForCondition(() => service.status().tick_count === 1, "first failed scheduler tick did not settle") + await waitForCondition(() => timers.length === 1, "failed scheduler tick did not schedule its replacement timer") expect(service.status()).toMatchObject({ status: "running", tick_count: 1 }) expect(timers).toHaveLength(1) @@ -14324,6 +15113,7 @@ describe("RuntimeServer launch OpenCode env wiring", () => { test("no env config preserves fake default launch behavior", async () => { const dir = await tempProject() await makeProject(dir, { approvedSpec: true }) + await commitUnconfiguredModelSetup(dir) const server = createRuntimeServerFromLaunchConfig({ projectDir: dir, env: {} }) expect(server.adapter).toBeInstanceOf(FakeOpenCodeAdapter) @@ -14335,6 +15125,7 @@ describe("RuntimeServer launch OpenCode env wiring", () => { test("env fake explicitly selects fake adapter", async () => { const dir = await tempProject() await makeProject(dir, { approvedSpec: true }) + await commitUnconfiguredModelSetup(dir) const options = readRuntimeServerLaunchOptionsFromEnv({ NXL_OPENCODE_ADAPTER: "fake" }, { projectDir: dir }) const server = createRuntimeServerFromLaunchConfig({ ...options }) @@ -14370,6 +15161,7 @@ describe("RuntimeServer launch OpenCode env wiring", () => { test("env external API config and credentials are wired through launch boundary", async () => { const dir = await tempProject() await makeProject(dir, { approvedSpec: true }) + await commitUnconfiguredModelSetup(dir) const transport = new FakeExternalApiTransport([{ status_code: 200, body: "{\"ok\":true}" }]) const server = createRuntimeServerFromLaunchConfig({ projectDir: dir, @@ -14416,6 +15208,7 @@ describe("RuntimeServer launch OpenCode env wiring", () => { test("env process config starts runtime with fake spawn and writes session start", async () => { const dir = await tempProject() await makeProject(dir, { approvedSpec: true }) + await commitUnconfiguredModelSetup(dir) const process = new FakeSpawnedProcess() const server = createRuntimeServerFromLaunchConfig({ projectDir: dir, @@ -14444,6 +15237,7 @@ describe("RuntimeServer launch OpenCode env wiring", () => { test("env process config allows submitUserMessage and writes mission packet", async () => { const dir = await tempProject() await makeProject(dir, { approvedSpec: true }) + await commitUnconfiguredModelSetup(dir) const process = new FakeSpawnedProcess() const server = createRuntimeServerFromLaunchConfig({ projectDir: dir, @@ -14500,6 +15294,7 @@ describe("RuntimeServer launch OpenCode env wiring", () => { test("secret-looking env launch values do not leak into runtime status events or errors", async () => { const dir = await tempProject() await makeProject(dir, { approvedSpec: true }) + await commitUnconfiguredModelSetup(dir) const process = new FakeSpawnedProcess() const server = createRuntimeServerFromLaunchConfig({ projectDir: dir, @@ -14526,6 +15321,7 @@ describe("RuntimeServer launch OpenCode env wiring", () => { test("direct adapter injection still takes precedence over env launch config", async () => { const dir = await tempProject() await makeProject(dir, { approvedSpec: true }) + await commitUnconfiguredModelSetup(dir) const adapter = new LongLivedAdapter() const server = createRuntimeServerFromLaunchConfig({ projectDir: dir, @@ -21079,6 +21875,7 @@ describe("OpenCode launch readiness", () => { test("launch config env opt-in reaches real launch gate", async () => { const dir = await tempProject() await makeProject(dir, { approvedSpec: true }) + await commitUnconfiguredModelSetup(dir) const db = ResearchDb.open(dir, { appendEvents: true, idFactory: () => "unused" }) db.createTopic({ id: "topic_launch_config", title: "Launch config" }) db.proposeResearchResult({ @@ -21241,6 +22038,37 @@ describe("OpenCode launch readiness", () => { await server.shutdown() }) + test("persisted 9W4E Executor selection reaches production launch as one exact primary model argument", async () => { + const dir = await tempProject() + await makeProject(dir, { approvedSpec: true }) + const setup = new ModelSetupService({ eventStore: new EventStore(join(dir, ".nxl", "events.jsonl")) }) + const choices = { commander_recipe_id: null, executor_recipe_id: "executor-google-gemini-2-5-flash" } as const + const setupPreview = await setup.preview(choices) + await setup.confirm({ ...choices, expected_revision: 0, candidate_hash: setupPreview.candidate_hash, confirmed_by: "operator", confirmation: "CONFIRM_MODEL_SETUP" }) + const launchCapture = join(dir, "opencode-launch-args.json") + const observationEnv = await openCodeObservationEnvironment(dir, "google", "gemini-2.5-flash", { launchCapture, runtimeSession: true }) + const server = createRuntimeServerFromLaunchConfig({ + projectDir: dir, + env: {}, + researchProjectionMode: "disabled", + openCodeAdapterConfig: { kind: "process", command: "opencode", args: ["--format", "json"], env: observationEnv }, + opencodeLaunchEnv: { NXL_REAL_OPENCODE_LAUNCH: "1" }, + opencodeLaunchId: () => "launch_persisted_setup", + }) + await server.start() + const session = await server.command("runtime.create_opencode_session_plan", { objective: "persisted setup launch" }) as { session_id: string } + const pack = await server.command("runtime.write_opencode_session_instruction_pack", { sessionId: session.session_id }) as { pack_id: string } + await expect(server.command("runtime.launch_opencode_session", { sessionId: session.session_id, packId: pack.pack_id })).resolves.toMatchObject({ status: "launch_started" }) + for (let attempt = 0; attempt < 50 && !existsSync(launchCapture); attempt += 1) await Bun.sleep(5) + const spawnedArgs = JSON.parse(await readFile(launchCapture, "utf8")) as string[] + const modelIndex = spawnedArgs.indexOf("--model") + expect(modelIndex).toBeGreaterThan(-1) + expect(spawnedArgs.filter((argument) => argument === "--model" || argument === "-m")).toHaveLength(1) + expect(spawnedArgs[modelIndex + 1]).toBe("google/gemini-2.5-flash") + expect(spawnedArgs.join(" ")).not.toMatch(/small_model|title|summary|compaction|agent|subagent/) + await server.shutdown() + }) + test("RuntimeServerClient no-start covers launch gate preview list get and dry-run", async () => { const dir = await tempProject() await makeProject(dir, { approvedSpec: true }) diff --git a/agentcore/runtime/src/server.ts b/agentcore/runtime/src/server.ts index 332d947bb..7bd680c5f 100644 --- a/agentcore/runtime/src/server.ts +++ b/agentcore/runtime/src/server.ts @@ -7,6 +7,7 @@ import type { RuntimeEvent, RuntimeMode, RuntimeResearchProjectionHealth, Runtim import { modeRequiresApprovedSpec } from "./project/project-status" import { locateProjectRoot, projectName } from "./project/project-root" import { RunLock } from "./project/run-lock" +import { buildModelSetupCandidate, ModelSetupService, readPersistedModelSetupAuthority, type ModelSetupCandidate } from "./model-configuration/model-setup" import { FakeOpenCodeAdapter } from "./opencode/fake-adapter" import type { ExecutorToolHandlerAdapter, OpenCodeRuntimeAdapter } from "./opencode/adapter" import { createOpenCodeAdapter, type OpenCodeAdapterConfig, type OpenCodeAdapterFactoryOptions } from "./opencode/adapter-config" @@ -259,6 +260,7 @@ import { redactText, redactValue } from "./security/redaction" import { adaptLegacyCommanderModelAuthority } from "./model-configuration/model-profile-legacy-commander-adapter" import { evaluateCommanderModelRoleReadiness, evaluateExecutorModelRoleReadiness, ModelProfileRuntimeRegistry } from "./model-configuration/model-profile-runtime-registry" import type { ExecutorModelReadinessResolver, ModelRoleReadinessEvidence } from "./model-configuration/model-profile-runtime-registry-types" +import type { ExecutorModelSelectionProjection } from "./model-configuration/model-configuration-types" import { ResearchDb, type ListResearchEventsOptions, @@ -283,6 +285,38 @@ const READ_ONLY_RESEARCH_INGESTION_DB: ResearchIngestionDbWriter = { }, } +function modelProfileRuntimeCapabilities( + commanderConfig: CommanderInvestigationProviderConfig | undefined, + executorSelection: ExecutorModelSelectionProjection | undefined, +): ModelCapability[] { + const commander = commanderConfig ? commanderInvestigationModelCapability(commanderConfig) : undefined + if (!executorSelection) return commander ? [commander] : [] + const executor: ModelCapability = { + capability_id: `runtime-executor-${stableHash({ + provider_kind: executorSelection.provider_kind, + provider_id: executorSelection.provider_id, + model_id: executorSelection.model_id, + projection_hash: executorSelection.projection_hash, + }).slice(0, 16)}`, + provider_kind: executorSelection.provider_kind, + provider_id: executorSelection.provider_id, + model_id: executorSelection.model_id, + display_name: `${executorSelection.provider_kind} primary Executor model`, + role_support: ["executor"], + supports_tools: "unknown", + supports_json_schema: "unknown", + supports_mcp: "unknown", + supports_long_context: "unknown", + supports_streaming: "unknown", + supports_local_execution: "unknown", + safety_margin_ratio: 0.25, + source: "runtime_config", + warnings: ["exact primary Executor selection; context limits remain unknown and conservative"], + created_at: "1970-01-01T00:00:00.000Z", + } + return commander ? [commander, executor] : [executor] +} + export interface RuntimeServerOptions { projectDir?: string mode?: RuntimeMode @@ -350,6 +384,9 @@ export interface RuntimeServerOptions { commanderModelStepAdapter?: CommanderModelStepAdapter commanderInvestigationProviderConfig?: CommanderInvestigationProviderConfig modelProfileRuntimeRegistry?: ModelProfileRuntimeRegistry + modelSetupActiveHash?: string + modelSetupActiveCandidate?: ModelSetupCandidate + revalidatePersistedModelSetupOnStart?: boolean executorModelReadinessResolver?: ExecutorModelReadinessResolver commanderInvestigationControlGate?: CommanderInvestigationControlGate commanderGithubGatewayConfig?: CommanderGithubGatewayConfig @@ -447,6 +484,11 @@ export class RuntimeServer { private readonly commanderModelStepAdapter?: CommanderModelStepAdapter private readonly commanderInvestigationProviderConfig?: CommanderInvestigationProviderConfig private readonly modelProfileRuntimeRegistry?: ModelProfileRuntimeRegistry + private readonly modelSetupActiveHash?: string + private readonly modelSetupActiveCandidate?: ModelSetupCandidate + private readonly revalidatePersistedModelSetupOnStart: boolean + private readonly activeModelSetupWrites = new Set>() + private modelSetupStartupMutex = Promise.resolve() private readonly executorModelReadinessResolver?: ExecutorModelReadinessResolver private readonly commanderInvestigationControlGate?: CommanderInvestigationControlGate private readonly commanderGithubGatewayConfig?: CommanderGithubGatewayConfig @@ -489,6 +531,7 @@ export class RuntimeServer { private commanderInvestigationRecoveryExecutionServiceInstance: CommanderInvestigationRecoveryExecutionService | null = null private commanderInvestigationRecoveryTransactionServiceInstance: CommanderInvestigationRecoveryTransactionService | null = null private commanderInvestigationRecoveryOperatorServiceInstance: CommanderInvestigationRecoveryOperatorService | null = null + private modelSetupServiceInstance: ModelSetupService | null = null private opencodeSessionContinuityServiceInstance: OpenCodeSessionContinuityService | null = null private opencodeContextRefreshServiceInstance: OpenCodeContextRefreshService | null = null private contextBudgetServiceInstance: ContextBudgetService | null = null @@ -586,6 +629,18 @@ export class RuntimeServer { this.commanderInvestigationProviderConfig = options.commanderInvestigationProviderConfig ? validateCommanderInvestigationProviderConfig(options.commanderInvestigationProviderConfig) : undefined this.modelProfileRuntimeRegistry = options.modelProfileRuntimeRegistry ?? (this.commanderInvestigationProviderConfig ? adaptLegacyCommanderModelAuthority(this.commanderInvestigationProviderConfig).registry : undefined) + if ((options.modelSetupActiveHash === undefined) !== (options.modelSetupActiveCandidate === undefined)) { + throw new Error("active model setup hash and candidate must be supplied together") + } + this.modelSetupActiveHash = options.modelSetupActiveHash + this.revalidatePersistedModelSetupOnStart = options.revalidatePersistedModelSetupOnStart === true + if (options.modelSetupActiveCandidate) { + const rebuilt = buildModelSetupCandidate(options.modelSetupActiveCandidate.choices) + if (rebuilt.candidate_hash !== options.modelSetupActiveCandidate.candidate_hash) { + throw new Error("active model setup candidate does not match current setup authority") + } + this.modelSetupActiveCandidate = rebuilt + } this.executorModelReadinessResolver = options.executorModelReadinessResolver if (options.modelProfileRuntimeRegistry && this.commanderInvestigationProviderConfig) { requireCommanderRegistryAssertion(options.modelProfileRuntimeRegistry, this.commanderInvestigationProviderConfig) @@ -596,7 +651,10 @@ export class RuntimeServer { this.reasoningProviderConfig = validateReasoningProviderConfig(options.reasoningProviderConfig ?? defaultReasoningProviderConfig()) this.modelCapabilityRegistry = new ModelCapabilityRegistry({ reasoningProviderConfig: this.reasoningProviderConfig, - runtimeCapabilities: this.commanderInvestigationProviderConfig ? [commanderInvestigationModelCapability(this.commanderInvestigationProviderConfig)] : [], + runtimeCapabilities: modelProfileRuntimeCapabilities( + this.commanderInvestigationProviderConfig, + this.modelProfileRuntimeRegistry?.executorSelection(), + ), }) const minimaxProvider = this.reasoningProviderConfig.kind === "minimax" ? this.createMiniMaxReasoningProvider() : null this.researchSynthesisProvider = options.researchSynthesisProvider ?? (minimaxProvider ?? new FakeResearchSynthesisProvider()) @@ -662,7 +720,7 @@ export class RuntimeServer { async start(): Promise { if (this.lifecycleShutdownRequested || this.lifecycleState === "stopping") throw new Error("runtime lifecycle is stopping") if (this.lifecycleStartTask) return this.lifecycleStartTask - const task = this.startUnserialized() + const task = this.withModelSetupStartupMutex(() => this.startUnserialized()) this.lifecycleStartTask = task try { await task @@ -684,6 +742,8 @@ export class RuntimeServer { this.commanderInvestigationLifecycleAbort = new AbortController() } try { + this.requireCurrentPersistedModelSetupAuthority() + await this.executorModelReadinessResolver?.start?.() this.ensureResearchProjectionUsable("startup") this.started = true if (this.mode === "active") { @@ -704,6 +764,20 @@ export class RuntimeServer { } } + private requireCurrentPersistedModelSetupAuthority(): void { + if (!this.revalidatePersistedModelSetupOnStart) return + const current = readPersistedModelSetupAuthority(this.projectDir) + if (!current && !this.modelSetupActiveHash && !this.modelProfileRuntimeRegistry) { + throw new Error("model setup is required before Runtime startup") + } + const currentHash = current?.setup_hash + const currentCandidateHash = current?.candidate.candidate_hash + if (currentHash !== this.modelSetupActiveHash + || currentCandidateHash !== this.modelSetupActiveCandidate?.candidate_hash) { + throw new Error("persisted model setup changed before runtime start; reconstruct RuntimeServer") + } + } + private emitStartupReadyEvents(recordsCount: number, lastRunId: string): void { this.eventBus.emit({ type: "RuntimeReady", @@ -763,6 +837,15 @@ export class RuntimeServer { this.executorStreamAbort = true this.lifecycleState = "stopping" this.commanderInvestigationLifecycleAbort.abort(new Error("RuntimeServer startup failed before Commander investigations became ready")) + try { + await this.executorModelReadinessResolver?.shutdown?.() + } catch (error) { + this.eventBus.emit({ + type: "ExecutorLifecycle", + phase: "runtime-executor-readiness-observer-startup-cleanup-error", + message: "Executor readiness observer startup cleanup failed", + }) + } await this.drainConfiguredCommanderInvestigations() this.started = false try { @@ -808,6 +891,23 @@ export class RuntimeServer { switch (name) { case "runtime.status": return this.status() + case "runtime.model_setup_catalog": + return this.modelSetupService().catalog() + case "runtime.model_setup_status": + return this.modelSetupStatus() + case "runtime.preview_model_setup": + return this.modelSetupService().preview(payload) + case "runtime.confirm_model_setup": + return this.withModelSetupWriteLock(async () => { + if (this.modelProfileRuntimeRegistry && !this.modelSetupActiveHash) { + throw new Error("model setup cannot be committed while non-setup model authority is active") + } + const confirmation = await this.modelSetupService().confirm(payload) + return Object.freeze({ + ...confirmation, + restart_required: confirmation.setup_hash !== this.modelSetupActiveHash, + }) + }) case "runtime.reasoning_provider_status": return this.reasoningProviderStatus() case "runtime.command_authority_summary": @@ -3446,6 +3546,26 @@ export class RuntimeServer { ) } + private async modelSetupStatus(): Promise> & { + active_authority_source?: "explicit" | "legacy_commander_environment" + active_candidate?: ModelSetupCandidate + commander_role_readiness?: ModelRoleReadinessEvidence + executor_role_readiness?: ModelRoleReadinessEvidence + }> { + const status = await this.modelSetupService().status(this.modelSetupActiveHash) + const commander = this.previewCommanderModelRoleReadiness() + const executor = await this.previewExecutorModelRoleReadiness() + return { + ...status, + ...(this.modelProfileRuntimeRegistry + ? { active_authority_source: this.modelProfileRuntimeRegistry.snapshot().authority_source } + : {}), + ...(this.modelSetupActiveCandidate ? { active_candidate: this.modelSetupActiveCandidate } : {}), + ...(commander ? { commander_role_readiness: commander } : {}), + ...(executor ? { executor_role_readiness: executor } : {}), + } + } + async previewExecutorModelRoleReadiness(): Promise { if (!this.modelProfileRuntimeRegistry) return undefined return evaluateExecutorModelRoleReadiness( @@ -4158,6 +4278,17 @@ export class RuntimeServer { firstError ??= error } } + try { + await this.executorModelReadinessResolver?.shutdown?.() + } catch (error) { + firstError ??= error + this.eventBus.emit({ + type: "ExecutorLifecycle", + phase: "runtime-executor-readiness-observer-shutdown-error", + message: "Executor readiness observer shutdown failed", + }) + } + await this.drainModelSetupWrites() if (this.started || this.runLock.isHeld()) { this.lifecycleState = "stopping" this.commanderInvestigationLifecycleAbort.abort(new Error("RuntimeServer shutdown cancelled Commander investigation")) @@ -4313,6 +4444,66 @@ export class RuntimeServer { } } + private async withModelSetupWriteLock(operation: () => Promise): Promise { + if (this.lifecycleStartTask) throw new Error("runtime lifecycle is starting") + const task = this.withModelSetupStartupMutex(() => this.runModelSetupWrite(operation)) + this.activeModelSetupWrites.add(task) + try { + return await task + } finally { + this.activeModelSetupWrites.delete(task) + } + } + + private async withModelSetupStartupMutex(operation: () => Promise): Promise { + const predecessor = this.modelSetupStartupMutex + let release!: () => void + const owned = new Promise((resolve) => { release = resolve }) + this.modelSetupStartupMutex = predecessor.then(() => owned) + await predecessor + try { + return await operation() + } finally { + release() + } + } + + private async runModelSetupWrite(operation: () => Promise): Promise { + if (this.lifecycleStartTask) { + throw new Error("runtime lifecycle is starting") + } + if (this.modelSetupWritesBlocked()) { + throw new Error("runtime lifecycle is stopping") + } + if (this.runLock.isHeld()) return operation() + await this.runLock.acquire() + try { + if (this.lifecycleStartTask) { + throw new Error("runtime lifecycle is starting") + } + if (this.modelSetupWritesBlocked()) { + throw new Error("runtime lifecycle is stopping") + } + return await operation() + } finally { + await this.runLock.release() + } + } + + private modelSetupWritesBlocked(): boolean { + return this.lifecycleShutdownRequested || this.lifecycleState === "stopping" || this.lifecycleState === "stopped" + } + + private async drainModelSetupWrites(): Promise { + while (this.activeModelSetupWrites.size > 0) { + await Promise.allSettled([...this.activeModelSetupWrites]) + } + } + + private modelSetupService(): ModelSetupService { + return this.modelSetupServiceInstance ??= new ModelSetupService({ eventStore: this.eventStore }) + } + private updateResearchProjectionHealth(integrity: ResearchProjectionIntegrity): void { let status: ResearchProjectionStatus | null = null try { diff --git a/agentcore/runtime/src/tui/runtime-client.ts b/agentcore/runtime/src/tui/runtime-client.ts index 8111739f6..b3b0205c4 100644 --- a/agentcore/runtime/src/tui/runtime-client.ts +++ b/agentcore/runtime/src/tui/runtime-client.ts @@ -1,4 +1,6 @@ import type { RuntimeEvent, RuntimeResearchProjectionHealth, RuntimeStatus } from "../events/event-types" +import type { ModelSetupCatalog, ModelSetupChoices, ModelSetupCommitInput, ModelSetupCommitResult, ModelSetupPreview, ModelSetupProjection } from "../model-configuration/model-setup" +import type { ModelRoleReadinessEvidence } from "../model-configuration/model-profile-runtime-registry-types" import type { ExecutorClaim, MissionProgress, MissionRecord, MissionResult } from "../missions/mission-types" import type { ReviewRequest, ReviewRequestInput, ReviewStatusSummary } from "../missions/review-types" import type { CommanderProposal, CommanderProposalInput, ProposalStatusSummary } from "../missions/proposal-types" @@ -85,6 +87,10 @@ export interface SubmitUserMessageResult { } export interface RuntimeClient { + command(name: "runtime.model_setup_catalog"): Promise + command(name: "runtime.model_setup_status"): Promise + command(name: "runtime.preview_model_setup", payload: ModelSetupChoices): Promise + command(name: "runtime.confirm_model_setup", payload: ModelSetupCommitInput): Promise command(name: "runtime.status"): Promise command(name: "runtime.reasoning_provider_status"): Promise command(name: "runtime.command_authority_summary"): Promise diff --git a/agentcore/runtime/src/tui/runtime-server-client.ts b/agentcore/runtime/src/tui/runtime-server-client.ts index de94603ad..c53b85124 100644 --- a/agentcore/runtime/src/tui/runtime-server-client.ts +++ b/agentcore/runtime/src/tui/runtime-server-client.ts @@ -5,6 +5,10 @@ import type { RuntimeClient, SubmitUserMessageResult } from "./runtime-client" const serverStartTasks = new WeakMap>() const noStartCommands = new Set([ + "runtime.model_setup_catalog", + "runtime.model_setup_status", + "runtime.preview_model_setup", + "runtime.confirm_model_setup", "runtime.command_authority_summary", "runtime.command_authority_list", "runtime.command_authority_get", diff --git a/agentcore/tui/src/app.tsx b/agentcore/tui/src/app.tsx index 8aec6a41f..eece1c9da 100644 --- a/agentcore/tui/src/app.tsx +++ b/agentcore/tui/src/app.tsx @@ -2,8 +2,13 @@ import { createCliRenderer, type CliRendererConfig, type TextareaRenderable } fr import { render, useKeyboard, useRenderer, useTerminalDimensions } from "@opentui/solid" import { createEffect, For, onMount, Show } from "solid-js" import { createStore } from "solid-js/store" -import { applyKeyCommandWithEffects, type KeyCommand } from "./keyboard" -import { reduceRuntimeEvent } from "./reducer" +import { applyKeyCommandWithModelSetupStartupGate, type KeyCommand } from "./keyboard" +import { + modelSetupStartupGateAllowsInput, + modelSetupStartupGateAllowsCommand, + reduceRuntimeEventDuringModelSetupGate, + type ModelSetupStartupGate, +} from "./reducer" import { applyRuntimeUiEffect, refreshRuntimeRecords } from "./runtime-effects" import { mergeRuntimeEffectState } from "./runtime-state-merge" import { commanderRecoveryApprovalDisplay, commanderRecoveryAuthorityValues, commanderRecoveryPreviewDiagnostics } from "./commander-recovery-view" @@ -11,6 +16,7 @@ import { snapshotUiState } from "./state-snapshot" import { initialState, type FocusTarget, type StreamLine, type UiState } from "./state" import type { RuntimeClient } from "./runtime" import { redactText } from "./redaction" +import { modelSetupCompletionCopy } from "./model-setup-view" const color = { bg: "#0b0f14", @@ -153,6 +159,59 @@ function ChoiceScreen(props: { state: UiState; kind: "init" | "resume" }) { ) } +function ModelSetupScreen(props: { state: UiState }) { + const setup = props.state.modelSetup + const completionCopy = () => modelSetupCompletionCopy(setup.pendingRestart) + const choices = setup.stage === "commander" ? setup.commanderChoices : setup.executorChoices + const selection = setup.stage === "commander" ? setup.commanderSelection : setup.executorSelection + const selectedCommander = setup.commanderChoices[setup.commanderSelection]?.label ?? "Unconfigured" + const selectedExecutor = setup.executorChoices[setup.executorSelection]?.label ?? "Unconfigured" + return ( + + + Model setup + Candidate Commander: {selectedCommander} + Candidate Executor: {selectedExecutor} + Active Commander: {setup.activeCommanderLabel} + Active Executor: {setup.activeExecutorLabel} + Active Commander readiness: {setup.commanderReadiness} + Active Executor readiness: {setup.executorReadiness} + Pending Commander: {setup.pendingCommanderLabel}; Executor: {setup.pendingExecutorLabel} + + {setup.stage === "commander" ? "Select Commander model" : "Select primary Executor model"} + {(choice, index) => ( + + {index() === selection ? "> " : " "}{choice.label} + + )} + + Loading setup authority... + + Preview ready. Candidate {setup.candidateHash?.slice(0, 12)} configuration {setup.configurationHash?.slice(0, 12)} + Enter opens explicit confirmation. + + + Confirm credential-free selection for the next RuntimeServer start. + + + Recording setup authority. Wait for the durable result. + + + {completionCopy().headline} + + {(value) => setup error: {value()}} + {setup.stage === "committed" + ? completionCopy().instructions + : setup.stage === "confirming" + ? "Confirmation is in progress and cannot be cancelled locally." + : setup.stage === "loading" && setup.commandError + ? "Enter retries the no-start setup authority check." + : "Enter selects. Up/Down changes selection. Esc returns."} + + + ) +} + function StreamPanel(props: { title: string; focus: FocusTarget; state: UiState; items: StreamLine[]; empty: string }) { return ( @@ -459,6 +518,13 @@ function OnboardingPanel(props: { state: UiState }) { model: {provider.model} credential: {provider.credentialSource} connection: {provider.connectionStatus} + Active Commander model: {props.state.modelSetup.activeCommanderLabel} + Active Executor model: {props.state.modelSetup.activeExecutorLabel} + Active Commander readiness: {props.state.modelSetup.commanderReadiness} + Active Executor readiness: {props.state.modelSetup.executorReadiness} + Pending Commander model: {props.state.modelSetup.pendingCommanderLabel} + Pending Executor model: {props.state.modelSetup.pendingExecutorLabel} + Model selection pending next start gpu quota: {project.gpuQuota} wake hooks: {project.wakeHooks} max parallel runs: {project.maxParallelRuns} @@ -564,14 +630,25 @@ function toCommand(evt: { name: string; shift: boolean; ctrl: boolean; raw?: str export function NexusLoopTui(props: { runtime: RuntimeClient; initial: UiState }) { const renderer = useRenderer() const [state, setState] = createStore(props.initial) + let modelSetupStartupGate: ModelSetupStartupGate = "pending" + + function updateModelSetupStartupGate(next: UiState) { + modelSetupStartupGate = next.modelSetup.startupCheckStatus === "required" + ? "required" + : next.modelSetup.startupCheckStatus === "clear" + ? "clear" + : "blocked" + } function apply(command: KeyCommand) { - const result = applyKeyCommandWithEffects(state, command) + const result = applyKeyCommandWithModelSetupStartupGate(state, command, modelSetupStartupGate) + if (result.state === state && result.effects.length === 0) return setState(result.state) for (const effect of result.effects) { const baseline = snapshotUiState(result.state) void (async () => { const next = await applyRuntimeUiEffect(baseline, props.runtime, effect) + if (effect.type === "load-model-setup") updateModelSetupStartupGate(next) setState((current) => mergeRuntimeEffectState(current, next, baseline.systemActions.length, baseline)) renderer.requestRender() })() @@ -582,14 +659,20 @@ export function NexusLoopTui(props: { runtime: RuntimeClient; initial: UiState } onMount(() => { void (async () => { for await (const event of props.runtime.stream()) { - setState((current) => reduceRuntimeEvent(current, event)) + setState((current) => reduceRuntimeEventDuringModelSetupGate(current, event, modelSetupStartupGate)) renderer.requestRender() } })() const baseline = snapshotUiState(state) void (async () => { const next = await refreshRuntimeRecords(baseline, props.runtime) + updateModelSetupStartupGate(next) setState((current) => mergeRuntimeEffectState(current, next, 0, baseline)) + if (modelSetupStartupGate === "clear") { + setState((current) => current.screen === "boot" + ? { ...current, screen: "resume", focus: "resume-choice" } + : current) + } renderer.requestRender() })() }) @@ -597,6 +680,7 @@ export function NexusLoopTui(props: { runtime: RuntimeClient; initial: UiState } useKeyboard((evt) => { const command = toCommand(evt) if (!command) return + if (!modelSetupStartupGateAllowsCommand(state, command, modelSetupStartupGate)) return evt.preventDefault() apply(command) }) @@ -608,9 +692,15 @@ export function NexusLoopTui(props: { runtime: RuntimeClient; initial: UiState } return ( } + fallback={state.screen === "model-setup" ? : } > - setState("messageDraft", value)} onSubmit={() => apply({ type: "submit" })} /> + { + if (modelSetupStartupGateAllowsInput(modelSetupStartupGate)) setState("messageDraft", value) + }} + onSubmit={() => apply({ type: "submit" })} + /> ) } diff --git a/agentcore/tui/src/keyboard.ts b/agentcore/tui/src/keyboard.ts index e1ccde748..5d7764c77 100644 --- a/agentcore/tui/src/keyboard.ts +++ b/agentcore/tui/src/keyboard.ts @@ -1,5 +1,6 @@ -import type { FocusTarget, UiState } from "./state" +import type { FocusTarget, ModelSetupState, UiState } from "./state" import { redactText } from "./redaction" +import { modelSetupStartupGateAllowsCommand, type ModelSetupStartupGate } from "./reducer" export type KeyCommand = | { type: "focus-next" } @@ -14,12 +15,25 @@ export type KeyCommand = export type KeySideEffect = | { type: "send-command"; command: string; args?: string[] } | { type: "send-user-message"; message: string } + | { type: "load-model-setup"; continueInitializationIfActive?: boolean; enterIfMissing?: boolean } + | { type: "preview-model-setup"; commanderRecipeId: string | null; executorRecipeId: string | null } + | { type: "confirm-model-setup"; commanderRecipeId: string | null; executorRecipeId: string | null; expectedRevision: number; candidateHash: string } export type KeyCommandResult = { state: UiState effects: KeySideEffect[] } +export function applyKeyCommandWithModelSetupStartupGate( + state: UiState, + command: KeyCommand, + gate: ModelSetupStartupGate, +): KeyCommandResult { + return modelSetupStartupGateAllowsCommand(state, command, gate) + ? applyKeyCommandWithEffects(state, command) + : { state, effects: [] } +} + const mainFocusOrder: FocusTarget[] = [ "executor", "commander", @@ -46,6 +60,7 @@ export function applyKeyCommandWithEffects(state: UiState, command: KeyCommand): case "focus-prev": return { state: state.screen === "main" ? { ...state, focus: moveFocus(state.focus, -1) } : state, effects: [] } case "select-next": + if (state.screen === "model-setup") return selectModelSetup(state, 1) if (state.screen === "init") { return { state: { ...state, initSelection: (state.initSelection + 1) % state.initChoices.length }, effects: [] } } @@ -54,6 +69,7 @@ export function applyKeyCommandWithEffects(state: UiState, command: KeyCommand): } return { state, effects: [] } case "select-prev": + if (state.screen === "model-setup") return selectModelSetup(state, -1) if (state.screen === "init") { return { state: { ...state, initSelection: (state.initSelection - 1 + state.initChoices.length) % state.initChoices.length }, @@ -81,13 +97,14 @@ export function applyKeyCommandWithEffects(state: UiState, command: KeyCommand): focus: "message-box", lastCommand: "initialize", commander: { ...state.commander, workIntent: "TUI onboarding shell" }, - systemActions: [...state.systemActions, { title: "Initialize selected", detail: "Entered onboarding shell" }], + systemActions: [...state.systemActions, { title: "Initialize selected", detail: "Entered spec onboarding shell" }], }, effects: [{ type: "send-command", command: "initialize" }], } } return { state: { ...state, lastCommand: "cancel" }, effects: [{ type: "send-command", command: "cancel" }] } } + if (state.screen === "model-setup") return submitModelSetup(state) if (state.screen === "resume") { const choice = state.resumeChoices[state.resumeSelection] const selected = choice?.id ?? "resume" @@ -131,6 +148,28 @@ export function applyKeyCommandWithEffects(state: UiState, command: KeyCommand): effects: [{ type: "send-user-message", message: state.messageDraft }], } case "cancel": + if (state.screen === "model-setup") { + if (state.modelSetup.stage === "confirming" + || (state.modelSetup.stage === "committed" && state.modelSetup.pendingRestart)) { + return { state, effects: [] } + } + if (state.modelSetup.startupCheckStatus === "required" + && (state.modelSetup.stage === "loading" || state.modelSetup.stage === "commander")) { + return { state, effects: [] } + } + if (state.modelSetup.stage === "executor") return { state: { ...state, modelSetup: { ...clearModelSetupPreview(state.modelSetup), stage: "commander" } }, effects: [] } + if (state.modelSetup.stage === "preview" || state.modelSetup.stage === "confirmation") return { state: { ...state, modelSetup: { ...clearModelSetupPreview(state.modelSetup), stage: "executor" } }, effects: [] } + const screen = state.modelSetup.origin + return { + state: { + ...state, + screen, + focus: screen === "main" ? "message-box" : "init-choice", + modelSetup: { ...state.modelSetup, stage: "loading" }, + }, + effects: [], + } + } return { state: state.screen === "init" || state.screen === "resume" ? { ...state, lastCommand: "cancel" } : { ...state, messageDraft: "" }, effects: [], @@ -148,6 +187,73 @@ export function applyKeyCommandWithEffects(state: UiState, command: KeyCommand): } } +function selectModelSetup(state: UiState, direction: 1 | -1): KeyCommandResult { + const setup = state.modelSetup + const key = setup.stage === "commander" ? "commanderSelection" : setup.stage === "executor" ? "executorSelection" : undefined + if (!key) return { state, effects: [] } + const choices = setup.stage === "commander" ? setup.commanderChoices : setup.executorChoices + const next = (setup[key] + direction + choices.length) % choices.length + return { state: { ...state, modelSetup: { ...clearModelSetupPreview(setup), [key]: next } }, effects: [] } +} + +function submitModelSetup(state: UiState): KeyCommandResult { + const setup = state.modelSetup + if (setup.stage === "loading" && setup.commandError !== undefined) { + return { + state: { ...state, modelSetup: { ...setup, commandError: undefined } }, + effects: [{ type: "load-model-setup", enterIfMissing: true, continueInitializationIfActive: true }], + } + } + if (setup.stage === "commander") return { state: { ...state, modelSetup: { ...clearModelSetupPreview(setup), stage: "executor" } }, effects: [] } + if (setup.stage === "executor") { + const commanderRecipeId = setup.commanderChoices[setup.commanderSelection]?.id || null + const executorRecipeId = setup.executorChoices[setup.executorSelection]?.id || null + return { + state: { ...state, modelSetup: { ...clearModelSetupPreview(setup), stage: "preview", commandError: undefined } }, + effects: [{ type: "preview-model-setup", commanderRecipeId, executorRecipeId }], + } + } + if (setup.stage === "preview" && setup.candidateHash !== undefined && setup.expectedRevision !== undefined) { + return { state: { ...state, modelSetup: { ...setup, stage: "confirmation" } }, effects: [] } + } + if (setup.stage === "confirmation" && setup.candidateHash !== undefined && setup.expectedRevision !== undefined) { + return { + state: { ...state, modelSetup: { ...setup, stage: "confirming", commandError: undefined } }, + effects: [{ + type: "confirm-model-setup", + commanderRecipeId: setup.commanderChoices[setup.commanderSelection]?.id || null, + executorRecipeId: setup.executorChoices[setup.executorSelection]?.id || null, + expectedRevision: setup.expectedRevision, + candidateHash: setup.candidateHash, + }], + } + } + if (setup.stage === "committed") { + if (setup.pendingRestart) return { state, effects: [] } + const screen = setup.origin + return { + state: { + ...state, + screen, + focus: screen === "main" ? "message-box" : "init-choice", + modelSetup: { ...setup, stage: "loading" }, + }, + effects: [], + } + } + return { state, effects: [] } +} + +function clearModelSetupPreview(setup: ModelSetupState): ModelSetupState { + const { + expectedRevision: _expectedRevision, + candidateHash: _candidateHash, + configurationHash: _configurationHash, + ...current + } = setup + return current +} + export function parseRuntimeCommand(value: string): { command: string; args: string[] } | undefined { const leadingTrimmed = value.trimStart() const commandPrefix = /^\/([a-z][a-z-]*)/i.exec(leadingTrimmed) @@ -182,6 +288,7 @@ export function parseRuntimeCommand(value: string): { command: string; args: str const runtimeCommands = new Set([ "status", + "model-setup", "missions", "resume", "new-session", diff --git a/agentcore/tui/src/launch.ts b/agentcore/tui/src/launch.ts index c1f2f2ad8..056674bcb 100644 --- a/agentcore/tui/src/launch.ts +++ b/agentcore/tui/src/launch.ts @@ -1,4 +1,4 @@ -import { applyKeyCommandWithEffects, type KeyCommand } from "./keyboard" +import { applyKeyCommandWithModelSetupStartupGate, type KeyCommand } from "./keyboard" import { reduceRuntimeEvent } from "./reducer" import { applyRuntimeUiEffect, refreshRuntimeRecords } from "./runtime-effects" import { type RuntimeClient } from "./runtime" @@ -52,16 +52,30 @@ export async function buildHeadlessSnapshot(runtime: RuntimeClient, projectDir: else await close } + if (noStartInspectionScript && state.screen !== "init" && state.screen !== "model-setup") { + state = await applyRuntimeUiEffect(state, runtime, { + type: "load-model-setup", + enterIfMissing: true, + }) + } + if (noStartInspectionScript && !needsExplicitRuntimeResume && state.screen === "resume") { state = { ...state, screen: "main", focus: "message-box" } } - if (!noStartInspectionScript) { + if (!noStartInspectionScript && state.screen !== "init" && state.screen !== "model-setup") { state = await refreshRuntimeRecords(state, runtime) } for (const command of commands) { - const result = applyKeyCommandWithEffects(state, command) + const gate = state.screen === "init" + ? "clear" + : state.modelSetup.startupCheckStatus === "required" + ? "required" + : state.modelSetup.startupCheckStatus === "clear" + ? "clear" + : "blocked" + const result = applyKeyCommandWithModelSetupStartupGate(state, command, gate) state = result.state for (const effect of result.effects) { state = await applyRuntimeUiEffect(state, runtime, effect) diff --git a/agentcore/tui/src/model-setup-view.ts b/agentcore/tui/src/model-setup-view.ts new file mode 100644 index 000000000..d89ef6e00 --- /dev/null +++ b/agentcore/tui/src/model-setup-view.ts @@ -0,0 +1,16 @@ +export type ModelSetupCompletionCopy = Readonly<{ + headline: string + instructions: string +}> + +export function modelSetupCompletionCopy(pendingRestart: boolean): ModelSetupCompletionCopy { + return pendingRestart + ? { + headline: "Selection recorded. Exit and restart NexusLoop to activate it.", + instructions: "This runtime cannot enter the main shell until restart.", + } + : { + headline: "Selection already active. No restart is required.", + instructions: "Enter returns to the main shell.", + } +} diff --git a/agentcore/tui/src/reducer.ts b/agentcore/tui/src/reducer.ts index 7a41c4e3f..b7b4d027f 100644 --- a/agentcore/tui/src/reducer.ts +++ b/agentcore/tui/src/reducer.ts @@ -1,4 +1,5 @@ import type { RuntimeEvent } from "./events" +import type { KeyCommand } from "./keyboard" import { initialState, type UiState } from "./state" import { redactText } from "./redaction" @@ -213,6 +214,39 @@ export function reduceRuntimeEvent(state: UiState, event: RuntimeEvent): UiState } } +export type ModelSetupStartupGate = "pending" | "required" | "clear" | "blocked" + +export function modelSetupStartupGateAllowsInput(gate: ModelSetupStartupGate): boolean { + return gate === "clear" +} + +export function modelSetupStartupGateAllowsCommand(state: UiState, command: KeyCommand, gate: ModelSetupStartupGate): boolean { + if (gate === "clear") return true + if (gate === "required") { + return state.screen === "model-setup" + && !(command.type === "cancel" && (state.modelSetup.stage === "loading" || state.modelSetup.stage === "commander")) + } + return gate === "blocked" + && command.type === "submit" + && state.screen === "model-setup" + && state.modelSetup.stage === "loading" + && state.modelSetup.commandError !== undefined +} + +export function reduceRuntimeEventDuringModelSetupGate( + state: UiState, + event: RuntimeEvent, + gate: ModelSetupStartupGate, +): UiState { + const next = reduceRuntimeEvent(state, event) + if (event.type !== "ProjectInitialized" || gate === "clear") return next + return { + ...next, + screen: gate === "pending" ? "boot" : state.screen, + focus: state.focus, + } +} + export function reduceRuntimeEvents(projectDir: string, events: RuntimeEvent[]): UiState { return events.reduce(reduceRuntimeEvent, initialState(projectDir)) } diff --git a/agentcore/tui/src/runtime-client-factory.ts b/agentcore/tui/src/runtime-client-factory.ts index cd992f91e..094c9f943 100644 --- a/agentcore/tui/src/runtime-client-factory.ts +++ b/agentcore/tui/src/runtime-client-factory.ts @@ -1,6 +1,8 @@ -import { basename } from "path" +import { readFileSync } from "fs" +import { basename, join } from "path" import { createRuntimeServerFromLaunchConfig, + locateProjectRoot, RuntimeServer, RuntimeServerClient, type OpenCodeAdapterFactoryOptions, @@ -8,7 +10,7 @@ import { import { supportedRuntimeEventTypes, type RuntimeEvent } from "./events" import { FakeRuntimeClient, type RuntimeClient, type SubmitUserMessageResult } from "./runtime" -export type TuiRuntimeClientKind = "fake" | "real" +export type TuiRuntimeClientKind = "auto" | "fake" | "real" export interface TuiRuntimeClientFactoryOptions { projectDir: string @@ -30,9 +32,15 @@ export function createTuiRuntimeClient(options: TuiRuntimeClientFactoryOptions): } const env = options.env ?? {} const kind = readRuntimeClientKind(env) - if (kind === "fake") return new FakeRuntimeClient(options.projectDir, options.projectName ?? basename(options.projectDir)) + const projectDir = locateProjectRoot(options.projectDir) + if (kind === "fake") { + return new FakeRuntimeClient(projectDir, options.projectName ?? basename(projectDir)) + } + if (kind === "auto" && !hasApprovedSpec(projectDir)) { + return new FakeRuntimeClient(projectDir, options.projectName ?? basename(projectDir), "restart_required") + } const server = createRuntimeServerFromLaunchConfig({ - projectDir: options.projectDir, + projectDir, env, openCodeAdapterFactoryOptions: options.openCodeAdapterFactoryOptions, }) @@ -45,11 +53,26 @@ export function createTuiRuntimeClient(options: TuiRuntimeClientFactoryOptions): export function readRuntimeClientKind(env: Record): TuiRuntimeClientKind { const raw = env.NXL_RUNTIME_CLIENT - if (raw === undefined || raw.trim() === "") return "fake" + if (raw === undefined || raw.trim() === "") return "auto" if (raw === "fake" || raw === "real") return raw throw new Error(`unknown runtime client kind in NXL_RUNTIME_CLIENT: ${raw}`) } +function hasApprovedSpec(projectDir: string): boolean { + const file = join(projectDir, ".nxl", "spec", "current.json") + try { + const parsed = JSON.parse(readFileSync(file, "utf8")) as unknown + return parsed !== null + && typeof parsed === "object" + && !Array.isArray(parsed) + && Object.getPrototypeOf(parsed) === Object.prototype + && Object.hasOwn(parsed, "status") + && (parsed as { status?: unknown }).status === "approved" + } catch { + return false + } +} + export function isTuiRuntimeEvent(event: unknown): event is RuntimeEvent { if (typeof event !== "object" || event === null) return false const type = (event as { type?: unknown }).type @@ -58,12 +81,20 @@ export function isTuiRuntimeEvent(event: unknown): event is RuntimeEvent { export class TuiRuntimeServerClient implements RuntimeClient { readonly streamMode = "long-lived" as const + readonly modelSetupAuthority = "durable" as const constructor(readonly runtime: RuntimeServerClient) {} async *stream(): AsyncIterable { for await (const event of this.runtime.stream()) { - if (isTuiRuntimeEvent(event)) yield event + if (!isTuiRuntimeEvent(event)) continue + yield event + if (event.type === "RuntimeReady") { + const status = await this.runtime.server.status() + if (!status.specApproved && status.runtimeStatus !== "started") { + yield { type: "ProjectUninitialized", projectDir: status.projectDir } + } + } } } diff --git a/agentcore/tui/src/runtime-effects.ts b/agentcore/tui/src/runtime-effects.ts index 658ef12a8..85b7ae6e8 100644 --- a/agentcore/tui/src/runtime-effects.ts +++ b/agentcore/tui/src/runtime-effects.ts @@ -1,4 +1,6 @@ import { redactText, redactUnknown } from "./redaction" +import { buildModelSetupCandidate, modelSetupCatalog } from "../../runtime/src/model-configuration/model-setup" +import { types as nodeUtilTypes } from "node:util" import { parseRuntimeCommand, type KeySideEffect } from "./keyboard" import { executionCommandFor, @@ -736,6 +738,56 @@ export async function applyRuntimeUiEffect( ): Promise { try { switch (effect.type) { + case "load-model-setup": { + const [catalog, status] = await Promise.all([ + runtime.command("runtime.model_setup_catalog"), + runtime.command("runtime.model_setup_status"), + ]) + const setupDisposition = classifyModelSetupStatus(status) + const next = applyModelSetupCatalogAndStatus(state, catalog, status) + const setupRequired = setupDisposition === "required" + const checked = { + ...next, + modelSetup: { ...next.modelSetup, startupCheckStatus: setupRequired ? "required" as const : "clear" as const }, + } + if (isExternalModelSetupAuthorityStatus(status)) { + if (effect.continueInitializationIfActive) return enterInitializedShell(checked) + if (state.screen === "model-setup") { + const screen = state.modelSetup.origin + return { + ...checked, + screen, + focus: screen === "main" ? "message-box" : "init-choice", + systemActions: [...checked.systemActions, { title: "Model setup unavailable", detail: "Non-setup model authority is active" }], + } + } + return checked + } + if (effect.enterIfMissing && setupRequired) { + return { + ...checked, + screen: "model-setup", + focus: "init-choice", + modelSetup: { ...checked.modelSetup, origin: state.screen === "init" ? "init" : "main", stage: "commander" }, + } + } + if (effect.continueInitializationIfActive && isActiveModelSetupStatus(status)) return enterInitializedShell(checked) + return checked + } + case "preview-model-setup": + return applyModelSetupPreview(state, await runtime.command("runtime.preview_model_setup", { + commander_recipe_id: effect.commanderRecipeId, + executor_recipe_id: effect.executorRecipeId, + })) + case "confirm-model-setup": + return applyModelSetupCommit(state, await runtime.command("runtime.confirm_model_setup", { + commander_recipe_id: effect.commanderRecipeId, + executor_recipe_id: effect.executorRecipeId, + expected_revision: effect.expectedRevision, + candidate_hash: effect.candidateHash, + confirmed_by: "tui-operator", + confirmation: "CONFIRM_MODEL_SETUP", + })) case "load-runtime-status": return applyRuntimeStatus(state, await runtime.command("runtime.status")) case "load-command-authority-summary": @@ -2286,10 +2338,34 @@ export async function applyRuntimeUiEffect( } case "send-command": { const next = await applyNamedRuntimeCommand(state, runtime, effect.command, effect.args ?? []) + if (effect.command === "initialize") { + return await applyRuntimeUiEffect(next, runtime, { + type: "load-model-setup", + enterIfMissing: true, + continueInitializationIfActive: true, + }) + } return shouldRefreshAfterCommand(effect.command) ? await refreshRuntimeRecordsOrRecordError(next, runtime) : next } } } catch (error) { + if (effect.type === "load-model-setup" || effect.type === "preview-model-setup" || effect.type === "confirm-model-setup") { + const message = error instanceof Error ? error.message : String(error) + return { + ...state, + ...(effect.type === "load-model-setup" ? { screen: "model-setup" as const, focus: "init-choice" as const } : {}), + modelSetup: { + ...state.modelSetup, + ...(effect.type === "load-model-setup" ? { + startupCheckStatus: "failed" as const, + origin: state.screen === "init" ? "init" as const : "main" as const, + stage: "loading" as const, + } : {}), + ...(effect.type === "confirm-model-setup" && state.modelSetup.stage === "confirming" ? { stage: "confirmation" as const } : {}), + commandError: redactText(message).slice(0, 240), + }, + } + } if (isOperatorActionEffect(effect)) return recordOperatorActionCommandError(state, error) if (isMissionExecutionEffect(effect)) return recordMissionExecutionCommandError(state, error) if (isReviewEffect(effect)) return recordReviewCommandError(state, error) @@ -2349,6 +2425,8 @@ export async function applyRuntimeUiEffect( export async function refreshRuntimeRecords(state: UiState, runtime: RuntimeClient): Promise { let next = state + next = await applyRuntimeUiEffect(next, runtime, { type: "load-model-setup", enterIfMissing: true }) + if (next.modelSetup.startupCheckStatus !== "clear" || next.screen === "model-setup") return next next = await applyRuntimeUiEffect(next, runtime, { type: "load-runtime-status" }) next = await applyRuntimeUiEffect(next, runtime, { type: "load-recent-missions", limit: 5 }) next = await applyRuntimeUiEffect(next, runtime, { type: "load-reviews", limit: REVIEW_LIMIT }) @@ -2359,6 +2437,223 @@ export async function refreshRuntimeRecords(state: UiState, runtime: RuntimeClie return next } +function classifyModelSetupStatus(value: unknown): "required" | "clear" { + if (!isRecord(value) || !Number.isInteger(value.revision) || typeof value.pending_restart !== "boolean") { + throw new Error("model setup returned malformed status") + } + const authority = value.active_authority_source + if (authority !== undefined && authority !== "explicit" && authority !== "legacy_commander_environment") { + throw new Error("model setup returned malformed status") + } + if (value.status === "missing") { + if (value.revision !== 0 || value.pending_restart !== false || value.setup_hash !== undefined || value.active_setup_hash !== undefined) { + throw new Error("model setup returned malformed status") + } + return authority === undefined ? "required" : "clear" + } + if (value.status !== "ready" || (value.revision as number) < 1 || !isSetupHash(value.setup_hash)) { + throw new Error("model setup returned malformed status") + } + if (!isCurrentModelSetupCandidate(value.candidate)) throw new Error("model setup returned malformed status") + const activeHash = value.active_setup_hash + if (activeHash !== undefined && !isSetupHash(activeHash)) throw new Error("model setup returned malformed status") + if (activeHash === undefined ? value.active_candidate !== undefined : !isCurrentModelSetupCandidate(value.active_candidate)) { + throw new Error("model setup returned malformed status") + } + if (value.pending_restart === false) { + if (activeHash !== value.setup_hash) throw new Error("model setup returned malformed status") + if (!exactSetupValue(value.candidate, value.active_candidate)) throw new Error("model setup returned malformed status") + return "clear" + } + if (activeHash === value.setup_hash) throw new Error("model setup returned malformed status") + return activeHash === undefined ? "required" : "clear" +} + +function isSetupHash(value: unknown): value is string { + return typeof value === "string" && /^[a-f0-9]{64}$/.test(value) +} + +function isCurrentModelSetupCandidate(value: unknown): boolean { + if (!isRecord(value) || nodeUtilTypes.isProxy(value)) return false + const descriptor = Object.getOwnPropertyDescriptor(value, "choices") + if (!descriptor || !("value" in descriptor) || !isRecord(descriptor.value) || nodeUtilTypes.isProxy(descriptor.value)) return false + try { + return exactSetupValue(value, buildModelSetupCandidate(descriptor.value)) + } catch { + return false + } +} + +function exactSetupValue(actual: unknown, expected: unknown): boolean { + if (actual === null || expected === null || typeof actual !== "object" || typeof expected !== "object") return actual === expected + if (nodeUtilTypes.isProxy(actual) || nodeUtilTypes.isProxy(expected)) return false + if (Array.isArray(actual) || Array.isArray(expected)) { + if (!Array.isArray(actual) || !Array.isArray(expected) || actual.length !== expected.length) return false + if (Object.getPrototypeOf(actual) !== Array.prototype || Object.getPrototypeOf(expected) !== Array.prototype) return false + if (!hasDenseDataElements(actual) || !hasDenseDataElements(expected)) return false + for (let index = 0; index < expected.length; index += 1) { + if (!exactSetupValue(actual[index], expected[index])) return false + } + return true + } + const actualRecord = actual as Record + const expectedRecord = expected as Record + if (!hasPlainDataPrototype(actualRecord) || !hasPlainDataPrototype(expectedRecord)) return false + const actualDescriptors = Object.getOwnPropertyDescriptors(actualRecord) + const expectedDescriptors = Object.getOwnPropertyDescriptors(expectedRecord) + if (!hasOnlyEnumerableDataProperties(actualDescriptors) || !hasOnlyEnumerableDataProperties(expectedDescriptors)) return false + const actualKeys = Object.keys(actualDescriptors).sort() + const expectedKeys = Object.keys(expectedDescriptors).sort() + if (actualKeys.length !== expectedKeys.length || actualKeys.some((key, index) => key !== expectedKeys[index])) return false + return expectedKeys.every((key) => exactSetupValue(actualDescriptors[key]!.value, expectedDescriptors[key]!.value)) +} + +function hasDenseDataElements(value: unknown[]): boolean { + const keys = Reflect.ownKeys(value) + if (keys.length !== value.length + 1 || keys.some((key) => typeof key === "symbol")) return false + const descriptors = Object.getOwnPropertyDescriptors(value) + for (let index = 0; index < value.length; index += 1) { + const descriptor = descriptors[String(index)] + if (!descriptor || !("value" in descriptor) || descriptor.enumerable !== true) return false + } + return keys.every((key) => key === "length" || (typeof key === "string" && /^(0|[1-9][0-9]*)$/.test(key) && Number(key) < value.length)) +} + +function hasPlainDataPrototype(value: Record): boolean { + const prototype = Object.getPrototypeOf(value) + return prototype === Object.prototype || prototype === null +} + +function hasOnlyEnumerableDataProperties(descriptors: PropertyDescriptorMap): boolean { + return Reflect.ownKeys(descriptors).every((key) => typeof key === "string" + && "value" in descriptors[key]! + && descriptors[key]!.enumerable === true) +} + +function applyModelSetupCatalogAndStatus(state: UiState, catalogValue: unknown, statusValue: unknown): UiState { + if (!isRecord(catalogValue) || !isRecord(statusValue)) throw new Error("model setup returned malformed state") + if (!exactSetupValue(catalogValue, modelSetupCatalog())) throw new Error("model setup returned malformed catalog") + const choices = (value: unknown, role: string) => { + if (!Array.isArray(value) || value.length > 20) throw new Error(`model setup ${role} recipes are malformed`) + return value.map((item) => { + if (!isRecord(item)) throw new Error(`model setup ${role} recipe is malformed`) + const id = readString(item.recipe_id, "") + const label = readString(item.display_name, "") + if (!id || !label) throw new Error(`model setup ${role} recipe is malformed`) + return { id: id.slice(0, 160), label: label.slice(0, 120) } + }) + } + const commanderChoices = [{ id: "", label: "Leave Commander unconfigured" }, ...choices(catalogValue.commander_recipes, "Commander")] + const executorChoices = [{ id: "", label: "Leave Executor unconfigured" }, ...choices(catalogValue.executor_recipes, "Executor")] + const candidate = isRecord(statusValue.candidate) && isRecord(statusValue.candidate.choices) ? statusValue.candidate.choices : undefined + const activeCandidate = isRecord(statusValue.active_candidate) && isRecord(statusValue.active_candidate.choices) ? statusValue.active_candidate.choices : undefined + const selectedIndex = (items: Array<{ id: string }>, value: unknown) => { + if (value === null) return 0 + if (typeof value !== "string") return 0 + const index = items.findIndex((item) => item.id === value) + return index < 0 ? 0 : index + } + const readiness = (value: unknown) => { + if (!isRecord(value)) return "unconfigured" + const selected = readString(value.selection_status, "unknown") + const connected = readString(value.credential_connection_status, "unknown") + const lifecycle = readString(value.lifecycle_status, "unknown") + return `${selected}; credential ${connected}; lifecycle ${lifecycle}`.slice(0, 180) + } + const label = (items: Array<{ id: string; label: string }>, value: unknown) => items[selectedIndex(items, value)]?.label ?? "Unconfigured" + const pendingCommanderLabel = label(commanderChoices, candidate?.commander_recipe_id) + const pendingExecutorLabel = label(executorChoices, candidate?.executor_recipe_id) + const activeCommanderLabel = label(commanderChoices, activeCandidate?.commander_recipe_id) + const activeExecutorLabel = label(executorChoices, activeCandidate?.executor_recipe_id) + return { + ...state, + modelSetup: { + ...state.modelSetup, + stage: "commander", + commanderChoices, + executorChoices, + commanderSelection: selectedIndex(commanderChoices, candidate?.commander_recipe_id), + executorSelection: selectedIndex(executorChoices, candidate?.executor_recipe_id), + activeCommanderLabel, + activeExecutorLabel, + pendingCommanderLabel, + pendingExecutorLabel, + activeSetupHash: typeof statusValue.active_setup_hash === "string" ? statusValue.active_setup_hash.slice(0, 64) : undefined, + pendingSetupHash: typeof statusValue.setup_hash === "string" ? statusValue.setup_hash.slice(0, 64) : undefined, + pendingRestart: statusValue.pending_restart === true, + commanderReadiness: readiness(statusValue.commander_role_readiness), + executorReadiness: readiness(statusValue.executor_role_readiness), + commandError: undefined, + }, + } +} + +function isActiveModelSetupStatus(value: unknown): boolean { + return isRecord(value) + && value.status === "ready" + && value.pending_restart === false + && isSetupHash(value.setup_hash) + && value.active_setup_hash === value.setup_hash +} + +function isExternalModelSetupAuthorityStatus(value: unknown): boolean { + return isRecord(value) + && value.status === "missing" + && value.revision === 0 + && value.pending_restart === false + && (value.active_authority_source === "explicit" || value.active_authority_source === "legacy_commander_environment") +} + +function enterInitializedShell(state: UiState): UiState { + return { + ...state, + screen: "main", + focus: "message-box", + lastCommand: "initialize", + commander: { ...state.commander, workIntent: "TUI onboarding shell" }, + systemActions: [...state.systemActions, { title: "Initialize selected", detail: "Active model authority verified; entered onboarding shell" }], + } +} + +function applyModelSetupPreview(state: UiState, value: unknown): UiState { + if (!isRecord(value) || !Number.isInteger(value.expected_revision) || typeof value.candidate_hash !== "string" || typeof value.configuration_hash !== "string") { + throw new Error("model setup preview is malformed") + } + return { + ...state, + modelSetup: { + ...state.modelSetup, + stage: "preview", + expectedRevision: Number(value.expected_revision), + candidateHash: value.candidate_hash.slice(0, 64), + configurationHash: value.configuration_hash.slice(0, 64), + commandError: undefined, + }, + } +} + +function applyModelSetupCommit(state: UiState, value: unknown): UiState { + if (!isRecord(value) || (value.status !== "committed" && value.status !== "idempotent") || typeof value.setup_hash !== "string") { + throw new Error("model setup confirmation result is malformed") + } + return { + ...state, + modelSetup: { + ...state.modelSetup, + stage: "committed", + pendingSetupHash: value.setup_hash.slice(0, 64), + pendingRestart: value.restart_required === true, + pendingCommanderLabel: state.modelSetup.commanderChoices[state.modelSetup.commanderSelection]?.label ?? "Unconfigured", + pendingExecutorLabel: state.modelSetup.executorChoices[state.modelSetup.executorSelection]?.label ?? "Unconfigured", + commandError: undefined, + }, + systemActions: [...state.systemActions, { + title: value.restart_required === true ? "Model setup recorded" : "Model setup unchanged", + detail: value.restart_required === true ? "Selection activates after a clean restart" : "The active selection already matches this setup", + }].slice(-12), + } +} + export async function refreshResearchRecords(state: UiState, runtime: RuntimeClient): Promise { let next: UiState = { ...state, research: { ...researchState(state), commandError: undefined } } next = await applyRuntimeUiEffect(next, runtime, { type: "load-research-projection-status" }) @@ -6103,6 +6398,16 @@ function clearCommandErrorFor(command: string, state: UiState): UiState { function applyNamedRuntimeCommand(state: UiState, runtime: RuntimeClient, command: string, args: string[]): Promise { const commandState = { ...state, lastCommand: command } switch (command) { + case "model-setup": + if (args.length > 0) throw new Error("/model-setup accepts no arguments") + if (runtime.modelSetupAuthority === "restart_required") { + throw new Error("Restart NexusLoop after spec approval before opening model setup") + } + return applyRuntimeUiEffect({ + ...commandState, + screen: "model-setup", + modelSetup: { ...commandState.modelSetup, origin: "main", stage: "loading", commandError: undefined }, + }, runtime, { type: "load-model-setup" }) case "stage": return Promise.resolve(stageSuggestedOperatorCommand(commandState, args)) case "stage-command": diff --git a/agentcore/tui/src/runtime-state-merge.ts b/agentcore/tui/src/runtime-state-merge.ts index b5a779526..3b294bfcc 100644 --- a/agentcore/tui/src/runtime-state-merge.ts +++ b/agentcore/tui/src/runtime-state-merge.ts @@ -3,6 +3,9 @@ import type { UiState } from "./state" export function mergeRuntimeEffectState(current: UiState, next: UiState, previousActionCount = 0, baseline?: UiState): UiState { const addedActions = addedSystemActions(next.systemActions, previousActionCount, baseline?.systemActions) + const canUpdateNavigation = baseline !== undefined + && ((current.screen === baseline.screen && current.focus === baseline.focus) + || isStreamInitializedMissingSetupNavigation(current, next, baseline)) const canUpdateRuntimeStatus = baseline === undefined || stableEqual(current.runtimeStatus, baseline.runtimeStatus) const canUpdateAdapterStatus = baseline === undefined || stableEqual(current.adapterStatus, baseline.adapterStatus) const canUpdateResearchProjection = @@ -85,6 +88,7 @@ export function mergeRuntimeEffectState(current: UiState, next: UiState, previou (stableEqual(current.runtimeCommandError, baseline.runtimeCommandError) && stableEqual(current.lastCommand, baseline.lastCommand)) const canUpdateLastCommand = baseline === undefined || stableEqual(current.lastCommand, baseline.lastCommand) + const canUpdateModelSetup = baseline === undefined || stableEqual(current.modelSetup, baseline.modelSetup) const canUpdateProjectName = canUpdateRuntimeStatus || current.header.projectName === baseline?.header.projectName const canUpdateHeaderRuntimeStatus = canUpdateRuntimeStatus || current.header.runtimeStatus === baseline?.header.runtimeStatus @@ -98,6 +102,8 @@ export function mergeRuntimeEffectState(current: UiState, next: UiState, previou return { ...current, + screen: canUpdateNavigation ? next.screen : current.screen, + focus: canUpdateNavigation ? next.focus : current.focus, systemActions: addedActions.length > 0 ? [...current.systemActions, ...addedActions].slice(-12) : current.systemActions, runtimeStatus: canUpdateRuntimeStatus ? next.runtimeStatus : current.runtimeStatus, adapterStatus: canUpdateAdapterStatus ? next.adapterStatus : current.adapterStatus, @@ -144,6 +150,7 @@ export function mergeRuntimeEffectState(current: UiState, next: UiState, previou executorReviewProposalApplyReadiness: canUpdateExecutorReviewProposalApplyReadiness ? next.executorReviewProposalApplyReadiness : current.executorReviewProposalApplyReadiness, executorReviewProposalNarrowApply: canUpdateExecutorReviewProposalNarrowApply ? next.executorReviewProposalNarrowApply : current.executorReviewProposalNarrowApply, minimaxLiveValidation: canUpdateMiniMaxLiveValidation ? next.minimaxLiveValidation : current.minimaxLiveValidation, + modelSetup: canUpdateModelSetup ? next.modelSetup : current.modelSetup, runtimeCommandError: canUpdateRuntimeCommandError ? next.runtimeCommandError : current.runtimeCommandError, lastCommand: canUpdateLastCommand ? next.lastCommand : current.lastCommand, header: { @@ -155,6 +162,17 @@ export function mergeRuntimeEffectState(current: UiState, next: UiState, previou } } +function isStreamInitializedMissingSetupNavigation(current: UiState, next: UiState, baseline: UiState): boolean { + return baseline.screen === "boot" + && current.screen === "resume" + && current.focus === "resume-choice" + && current.resumeSelection === baseline.resumeSelection + && current.lastCommand === baseline.lastCommand + && next.screen === "model-setup" + && next.focus === "init-choice" + && next.modelSetup.stage === "commander" +} + function stableEqual(left: unknown, right: unknown): boolean { return JSON.stringify(left) === JSON.stringify(right) } diff --git a/agentcore/tui/src/runtime.ts b/agentcore/tui/src/runtime.ts index 68c0d15d7..cbe2e53bd 100644 --- a/agentcore/tui/src/runtime.ts +++ b/agentcore/tui/src/runtime.ts @@ -8,6 +8,7 @@ import { ModelCapabilityRegistry } from "../../runtime/src/context/model-capabil import { ContextBudgetService } from "../../runtime/src/context/context-budget-service" import { COMMAND_AUTHORITY_REGISTRY } from "../../runtime/src/authority/command-authority-registry" import { containsConcreteCredentialPayload } from "../../runtime/src/security/redaction" +import { buildModelSetupCandidate, modelSetupCatalog } from "../../runtime/src/model-configuration/model-setup" import type { CommanderApplyPreviewSummary, CommanderApplyResultSummary, CommanderAuditEventSummary, CommanderAuthorityChainSummary, CommanderCyclePreviewSummary, CommanderCycleRecordSummary, CommanderCycleResultSummary, CommanderExecutorReviewPreviewSummary, CommanderExecutorReviewRecordSummary, CommanderExecutorReviewResultSummary, CommanderPlaybookDraftSummary, CommanderPlaybookSummary, CommanderProposalBundleSummary, CommanderProposalSummary, CommanderQueueItemSummary, CommanderQueueKind, CommanderQueueSummary, CommanderTargetContextSummary, CommanderTargetType, CommanderWorkbenchDraftSummary, CommanderWorkbenchReadinessSummary, CommanderWorkbenchStatusSummary, ContextBudgetAllocationSummary, ContextBudgetPreviewSummary, ContextBudgetProfileSummary, ContextBudgetSummaryState, ContinuationPlanPreviewSummary, ContinuationPlanRecordSummary, ContinuationPlanSummary, ContinuationStepResultSummary, ExecutorClaimSummary, ExecutorReviewProposalApplyReadinessPreviewSummary, ExecutorReviewProposalApplyReadinessRecordSummary, ExecutorReviewProposalApplyReadinessSummary, ExecutorReviewProposalCreatePreviewSummary, ExecutorReviewProposalCreateRecordSummary, ExecutorReviewProposalCreateResultSummary, ExecutorReviewProposalDraftCandidateSummary, ExecutorReviewProposalDraftPreviewSummary, ExecutorReviewProposalDraftSummary, ExecutorReviewProposalNarrowApplyPreviewSummary, ExecutorReviewProposalNarrowApplyRecordSummary, ExecutorReviewProposalNarrowApplyResultSummary, ExecutorReviewProposalReviewDecisionPreviewSummary, ExecutorReviewProposalReviewDecisionRecordSummary, ExecutorReviewProposalReviewDecisionResultSummary, ExecutorReviewProposalReviewRequestPreviewSummary, ExecutorReviewProposalReviewRequestRecordSummary, ExecutorReviewProposalReviewRequestResultSummary, ExternalApiAuditRecordSummary, ExternalApiConnectorSummary, ExternalApiResearchIngestionPreviewSummary, ExternalApiResearchIngestionRecordSummary, ExternalApiResearchIngestionResultSummary, ExternalApiRequestPreviewSummary, ExternalApiRequestResultSummary, MiniMaxLiveValidationPreviewSummary, MiniMaxLiveValidationRecordSummary, MiniMaxLiveValidationResultSummary, MiniMaxLiveValidationSurfaceResultSummary, MissionProgressSummary, MissionRecord, MissionResultSummary, ModelCapabilitySummary, OpenCodeHandoffFollowupCounts, OpenCodeHandoffFollowupQueueKind, OpenCodeHandoffFollowupSummary, OpenCodeHandoffPreviewSummary, OpenCodeHandoffReadinessPreviewSummary, OpenCodeHandoffReadinessSummary, OpenCodeHandoffRecordSummary, OpenCodeHandoffResultSummary, OpenCodeProcessSmokePreviewSummary, OpenCodeProcessSmokeRecordSummary, OpenCodeProcessSmokeResultSummary, OpenCodeResultReviewPacketSummary, OpenCodeResultReviewSummary, OpenCodeSessionPlanSummary, OpenCodeSessionPreviewSummary, OpenCodeSessionRecordSummary, OpenCodeSessionSummary, ProposalBundleReadinessSummary, ResearchSynthesisPreviewSummary, ResearchSynthesisRecordSummary, ResearchSynthesisResultSummary, ReviewRequestSummary, RuntimeCheckpointPreviewSummary, RuntimeCheckpointRecordSummary, RuntimeCheckpointScope, RuntimeCheckpointSummary, RuntimeRestorePreviewSummary, RuntimeResumeAnchorSummary, WakeAssessmentPreviewSummary, WakeAssessmentRecordSummary, WakeAssessmentSummary, WakeSchedulePreviewSummary, WakeScheduleRecordSummary, WakeScheduleSummary, WakeSchedulerAuditChainSummary, WakeSchedulerAuditCommandSummary, WakeSchedulerAuditIncidentSummary, WakeSchedulerAuditSummarySummary, WakeSchedulerAuditTimelineEntrySummary, WakeSchedulerBootstrapStatusSummary, WakeSchedulerEventRecordSummary, WakeSchedulerNavigationBoardSummary, WakeSchedulerNavigationCardSummary, WakeSchedulerNavigationCheckpointApprovalUsageSummaryState, WakeSchedulerNavigationCheckpointWriteGroupSummary, WakeSchedulerNavigationCheckpointWriteHistorySummary, WakeSchedulerNavigationCheckpointWritePairComparisonSummary, WakeSchedulerNavigationCheckpointWriteRunPreviewSummary, WakeSchedulerNavigationCheckpointWriteRunRecordSummary, WakeSchedulerNavigationCheckpointWriteRunResultSummary, WakeSchedulerNavigationCheckpointWriteStaleItemSummary, WakeSchedulerNavigationCommandPreviewSummary, WakeSchedulerNavigationStagePreviewSummary, WakeSchedulerNavigationStagedReadGroupSummary, WakeSchedulerNavigationStagedReadHistorySummary, WakeSchedulerNavigationStagedReadPairComparisonSummary, WakeSchedulerNavigationStagedReadStaleItemSummary, WakeSchedulerNavigationStagedRunPreviewSummary, WakeSchedulerNavigationStagedRunRecordSummary, WakeSchedulerNavigationStagedRunResultSummary, WakeSchedulerNavigationStagedCommandRecordSummary, WakeSchedulerNavigationStagedCommandSummary, WakeSchedulerNavigationStagedWriteCommandRecordSummary, WakeSchedulerNavigationStagedWriteCommandSummary, WakeSchedulerNavigationTargetKindSummary, WakeSchedulerNavigationTargetSummary, WakeSchedulerNavigationWriteApprovalRecordSummary, WakeSchedulerNavigationWriteApprovalSummary, WakeSchedulerNavigationWriteReadinessPreviewSummary, WakeSchedulerNavigationWriteBoardSummary, WakeSchedulerNavigationWritePreviewSummary, WakeSchedulerNavigationWriteRunGroupSummary, WakeSchedulerNavigationWriteRunHistorySummary, WakeSchedulerNavigationWriteRunPairComparisonSummary, WakeSchedulerNavigationWriteRunPreviewSummary, WakeSchedulerNavigationWriteRunRecordSummary, WakeSchedulerNavigationWriteRunResultSummary, WakeSchedulerNavigationWriteRunStaleItemSummary, WakeSchedulerNavigationWriteStagePreviewSummary, WakeSchedulerPreviewSummary, WakeSchedulerRecoveryPreviewSummary, WakeSchedulerRecoveryRecordSummary, WakeSchedulerRecoverySummary, WakeSchedulerRecoveryWorkflowPreviewSummary, WakeSchedulerRecoveryWorkflowRecordSummary, WakeSchedulerRecoveryWorkflowStepSummary, WakeSchedulerRecoveryWorkflowSummary, WakeSchedulerRecoveryWorkflowVerificationSummary, WakeSchedulerStateSummary, WakeScheduleTickPreviewSummary, WakeScheduleTickResultSummary } from "./state" import type { CommandAuthorityRecordSummary, CommandAuthoritySummaryState, CommandAuthorityValidationProfileSummary, CommanderGuidanceDeliveryPreviewSummary, CommanderGuidanceDeliveryRecordSummary, CommanderGuidanceDeliveryResultSummary, CommanderGuidanceDeliverySummaryState, CommanderGuidancePreviewSummary, CommanderGuidanceRecordSummary, CommanderGuidanceResultSummary, CommanderGuidanceSummaryState, ContextPacketPreviewSummary, ContextPacketSectionSummary, ContextPacketSummaryState, OpenCodeCommanderQuestionPreviewSummary, OpenCodeCommanderQuestionRecordSummary, OpenCodeCommanderQuestionResultSummary, OpenCodeCommanderQuestionSummaryState, OpenCodeForcedReportRequestSummary, OpenCodeHumanControlPreviewSummary, OpenCodeHumanControlRecordSummary, OpenCodeHumanControlResultSummary, OpenCodeHumanControlSummaryState, OpenCodeLaunchPreviewSummary, OpenCodeLaunchReadinessCheckSummary, OpenCodeLaunchReadinessPreviewSummary, OpenCodeLaunchReadinessSummaryState, OpenCodeLaunchRecordSummary, OpenCodeLaunchResultSummary, OpenCodeProgressPreviewSummary, OpenCodeProgressRecordSummary, OpenCodeProgressResultSummary, OpenCodeProgressSummaryState, OpenCodeResultReportCommandSummary, OpenCodeResultReportPreviewSummary, OpenCodeResultReportRecordSummary, OpenCodeResultReportResultSummary, OpenCodeResultReportSummaryState, OpenCodeResultReviewGateCommandSummary, OpenCodeResultReviewGatePreviewSummary, OpenCodeResultReviewGateRecordSummary, OpenCodeResultReviewGateResultSummary, OpenCodeResultReviewGateSummaryState, OpenCodeSessionInstructionPackFilePreviewSummary, OpenCodeSessionInstructionPackPreviewSummary, OpenCodeSessionInstructionPackRecordSummary, OpenCodeSessionInstructionPackResultSummary, OpenCodeWakeActionExecutionCommandSummary, OpenCodeWakeActionExecutionEvidenceRefSummary, OpenCodeWakeActionExecutionPreviewSummary, OpenCodeWakeActionExecutionRecordSummary, OpenCodeWakeActionExecutionResultSummary, OpenCodeWakeActionExecutionSummaryState, OpenCodeWakeSupervisorBatchPreviewSummary, OpenCodeWakeSupervisorBatchResultSummary, OpenCodeWakeSupervisorCheckSummary, OpenCodeWakeSupervisorContextSectionSummary, OpenCodeWakeSupervisorEvidenceRefSummary, OpenCodeWakeSupervisorExecutionCommandSummary, OpenCodeWakeSupervisorExecutionEvidenceRefSummary, OpenCodeWakeSupervisorExecutionPreviewSummary, OpenCodeWakeSupervisorExecutionRecordSummary, OpenCodeWakeSupervisorExecutionResultSummary, OpenCodeWakeSupervisorExecutionSummaryState, OpenCodeWakeSupervisorPreviewSummary, OpenCodeWakeSupervisorSessionCardSummary, OpenCodeWakeSupervisorSummaryState, OpenCodeWatchdogPreviewSummary, OpenCodeWatchdogRecordSummary, OpenCodeWatchdogResultSummary, OpenCodeWatchdogSummaryState, ResearchIngestionCommandSummary, ResearchIngestionPreviewSummary, ResearchIngestionProvenanceRefSummary, ResearchIngestionRecordSummary, ResearchIngestionResultSummary, ResearchIngestionSummaryState, ResearchMemoryCandidateSummary, ResearchMemoryInspectionPreviewSummary, ResearchMemoryNearDuplicatePreviewSummary, ResearchMemoryRetrievalPreviewSummary, ResearchMemorySearchProfileState, ResearchMemorySummaryState, ResearchNoveltyPreviewSummary } from "./state" import type { CommanderContinuityCommandSummary, CommanderContinuityOpenLoopSummary, CommanderContinuitySectionSummary, CommanderContinuitySourceRefSummary, CommanderContinuitySummaryState, CommanderContinuityThreadCardSummary, CommanderMidMissionContinuityPacketSummary, CommanderProposalContinuityPacketSummary } from "./state" @@ -20,6 +21,7 @@ export interface SubmitUserMessageResult { export interface RuntimeClient { readonly streamMode?: "finite" | "long-lived" + readonly modelSetupAuthority?: "durable" | "fixture" | "restart_required" stream(): AsyncIterable command(name: string, payload?: Record): Promise sendUserMessage(message: string): Promise @@ -39,6 +41,7 @@ const COMMANDER_QUEUE_KINDS: CommanderQueueKind[] = [ ] export class FakeRuntimeClient implements RuntimeClient { + readonly modelSetupAuthority: "fixture" | "restart_required" readonly sentMessages: string[] = [] readonly sentCommands: string[] = [] private readonly missions: MissionRecord[] = [] @@ -105,11 +108,22 @@ export class FakeRuntimeClient implements RuntimeClient { private readonly wakeSchedulerNavigationCheckpointWriteRuns: WakeSchedulerNavigationCheckpointWriteRunResultSummary[] = [] private projectionRebuilds = 0 private sequence = 0 + private fakeModelSetup: { revision: number; setup_hash?: string; active_setup_hash?: string; commander_recipe_id: string | null; executor_recipe_id: string | null; active_commander_recipe_id: string | null; active_executor_recipe_id: string | null } = { + revision: 1, + setup_hash: "4".repeat(64), + active_setup_hash: "4".repeat(64), + commander_recipe_id: null, + executor_recipe_id: null, + active_commander_recipe_id: null, + active_executor_recipe_id: null, + } constructor( private readonly projectDir: string, private readonly projectName: string, + modelSetupAuthority: "fixture" | "restart_required" = "fixture", ) { + this.modelSetupAuthority = modelSetupAuthority if (process.env.NXL_TUI_FAKE_WAKE_SCHEDULER_STALE === "1") { this.wakeSchedulerRecoveryPreviewRecord = fakeWakeSchedulerRecoveryPreview({ recovery_id: "fake-recovery-1", @@ -217,6 +231,23 @@ export class FakeRuntimeClient implements RuntimeClient { async command(name: string, payload: Record = {}): Promise { switch (name) { + case "runtime.model_setup_catalog": + return modelSetupCatalog() + case "runtime.model_setup_status": + return { + status: this.fakeModelSetup.revision === 0 ? "missing" : "ready", + revision: this.fakeModelSetup.revision, + setup_hash: this.fakeModelSetup.setup_hash, + active_setup_hash: this.fakeModelSetup.active_setup_hash, + pending_restart: this.fakeModelSetup.setup_hash !== this.fakeModelSetup.active_setup_hash, + ...(this.fakeModelSetup.revision > 0 ? { candidate: buildModelSetupCandidate({ commander_recipe_id: this.fakeModelSetup.commander_recipe_id, executor_recipe_id: this.fakeModelSetup.executor_recipe_id }) } : {}), + ...(this.fakeModelSetup.active_setup_hash ? { active_candidate: buildModelSetupCandidate({ commander_recipe_id: this.fakeModelSetup.active_commander_recipe_id, executor_recipe_id: this.fakeModelSetup.active_executor_recipe_id }) } : {}), + } + case "runtime.preview_model_setup": + return { preview_version: 1, expected_revision: this.fakeModelSetup.revision, candidate_hash: "2".repeat(64), catalog_hash: "1".repeat(64), configuration_hash: "3".repeat(64), restart_required: true } + case "runtime.confirm_model_setup": + this.fakeModelSetup = { ...this.fakeModelSetup, revision: this.fakeModelSetup.revision + 1, setup_hash: "5".repeat(64), commander_recipe_id: typeof payload.commander_recipe_id === "string" ? payload.commander_recipe_id : null, executor_recipe_id: typeof payload.executor_recipe_id === "string" ? payload.executor_recipe_id : null } + return { status: "committed", revision: this.fakeModelSetup.revision, setup_hash: this.fakeModelSetup.setup_hash, candidate_hash: String(payload.candidate_hash ?? ""), restart_required: true } case "runtime.status": return { runtimeStatus: "fake runtime connected", diff --git a/agentcore/tui/src/snapshot.ts b/agentcore/tui/src/snapshot.ts index 1cf5dddc4..fd8e660e6 100644 --- a/agentcore/tui/src/snapshot.ts +++ b/agentcore/tui/src/snapshot.ts @@ -35,6 +35,26 @@ export function layoutSnapshot(state: UiState): string { return out.join("\n") } + if (state.screen === "model-setup") { + const setup = state.modelSetup + out.push("Model setup") + out.push(` stage=${setup.stage}`) + out.push(` candidate_commander=${setup.commanderChoices[setup.commanderSelection]?.label ?? "unconfigured"}`) + out.push(` candidate_executor=${setup.executorChoices[setup.executorSelection]?.label ?? "unconfigured"}`) + out.push(` active_commander=${setup.activeCommanderLabel}`) + out.push(` active_executor=${setup.activeExecutorLabel}`) + out.push(` active_commander_readiness=${setup.commanderReadiness}`) + out.push(` active_executor_readiness=${setup.executorReadiness}`) + if (setup.pendingRestart) { + out.push(` pending_commander=${setup.pendingCommanderLabel}`) + out.push(` pending_executor=${setup.pendingExecutorLabel}`) + } + out.push(` candidate=${setup.candidateHash?.slice(0, 12) ?? "none"}`) + out.push(` pending_restart=${setup.pendingRestart}`) + if (setup.commandError) out.push(` error=${setup.commandError}`) + return out.join("\n") + } + out.push("Executor") out.push(...lines(state.executor)) out.push("Commander") @@ -51,6 +71,9 @@ export function layoutSnapshot(state: UiState): string { out.push("Live system actions") out.push(...lines(state.systemActions)) out.push("Onboarding") + out.push(` active_model_setup=${state.modelSetup.activeSetupHash?.slice(0, 12) ?? "none"}`) + out.push(` pending_model_setup=${state.modelSetup.pendingSetupHash?.slice(0, 12) ?? "none"}`) + out.push(` model_setup_restart_required=${state.modelSetup.pendingRestart}`) out.push(` provider=${state.providerOnboarding.provider}`) out.push(` model=${state.providerOnboarding.model}`) out.push(` credential=${state.providerOnboarding.credentialSource}`) diff --git a/agentcore/tui/src/state.ts b/agentcore/tui/src/state.ts index 6ac768605..fe9ad063e 100644 --- a/agentcore/tui/src/state.ts +++ b/agentcore/tui/src/state.ts @@ -1,6 +1,6 @@ import type { OperatorCommandExecutionResult, OperatorStagedCommand } from "./operator-actions" -export type Screen = "boot" | "init" | "resume" | "main" +export type Screen = "boot" | "init" | "model-setup" | "resume" | "main" export type FocusTarget = | "init-choice" @@ -57,6 +57,29 @@ export type ProviderOnboardingState = { connectionStatus: string } +export type ModelSetupState = { + origin: "init" | "main" + startupCheckStatus: "pending" | "required" | "clear" | "failed" + stage: "loading" | "commander" | "executor" | "preview" | "confirmation" | "confirming" | "committed" + commanderChoices: Choice[] + executorChoices: Choice[] + commanderSelection: number + executorSelection: number + activeCommanderLabel: string + activeExecutorLabel: string + pendingCommanderLabel: string + pendingExecutorLabel: string + expectedRevision?: number + candidateHash?: string + configurationHash?: string + activeSetupHash?: string + pendingSetupHash?: string + pendingRestart: boolean + commanderReadiness: string + executorReadiness: string + commandError?: string +} + export type ProjectOnboardingState = { plainTextSpec: string gpuQuota: string @@ -5776,6 +5799,7 @@ export type UiState = { search: SearchState approval: ApprovalState providerOnboarding: ProviderOnboardingState + modelSetup: ModelSetupState projectOnboarding: ProjectOnboardingState messageDraft: string submittedMessages: string[] @@ -5901,6 +5925,22 @@ export function initialState(projectDir: string): UiState { localEndpoint: "", connectionStatus: "not tested", }, + modelSetup: { + origin: "init", + startupCheckStatus: "pending", + stage: "loading", + commanderChoices: [{ id: "", label: "Leave Commander unconfigured" }], + executorChoices: [{ id: "", label: "Leave Executor unconfigured" }], + commanderSelection: 0, + executorSelection: 0, + activeCommanderLabel: "Unconfigured", + activeExecutorLabel: "Unconfigured", + pendingCommanderLabel: "Unconfigured", + pendingExecutorLabel: "Unconfigured", + pendingRestart: false, + commanderReadiness: "unconfigured", + executorReadiness: "unconfigured", + }, projectOnboarding: { plainTextSpec: "", gpuQuota: "unset", diff --git a/agentcore/tui/test/keyboard.test.ts b/agentcore/tui/test/keyboard.test.ts index 14cd7df51..178d9191a 100644 --- a/agentcore/tui/test/keyboard.test.ts +++ b/agentcore/tui/test/keyboard.test.ts @@ -1,8 +1,30 @@ import { describe, expect, test } from "bun:test" -import { applyKeyCommand, applyKeyCommandWithEffects, parseRuntimeCommand } from "../src/keyboard" +import { applyKeyCommand, applyKeyCommandWithEffects, applyKeyCommandWithModelSetupStartupGate, parseRuntimeCommand } from "../src/keyboard" import { initialState, type UiState } from "../src/state" describe("TUI keyboard command model", () => { + test("startup setup authority gates direct submit dispatch", () => { + const state = { ...initialState("/tmp/demo"), screen: "main" as const, focus: "message-box" as const, messageDraft: "start mission" } + for (const gate of ["pending", "blocked"] as const) { + expect(applyKeyCommandWithModelSetupStartupGate(state, { type: "submit" }, gate)).toEqual({ state, effects: [] }) + } + expect(applyKeyCommandWithModelSetupStartupGate(state, { type: "submit" }, "required")).toEqual({ state, effects: [] }) + expect(applyKeyCommandWithModelSetupStartupGate(state, { type: "submit" }, "clear").effects).toEqual([ + { type: "send-user-message", message: "start mission" }, + ]) + }) + test("failed startup setup inspection exposes only a no-start retry", () => { + const base = initialState("/tmp/demo") + const state = { + ...base, + screen: "model-setup" as const, + modelSetup: { ...base.modelSetup, stage: "loading" as const, commandError: "setup unavailable" }, + } + expect(applyKeyCommandWithModelSetupStartupGate(state, { type: "submit" }, "blocked")).toMatchObject({ + effects: [{ type: "load-model-setup", enterIfMissing: true, continueInitializationIfActive: true }], + }) + expect(applyKeyCommandWithModelSetupStartupGate(state, { type: "insert", text: "start" }, "blocked")).toEqual({ state, effects: [] }) + }) test("recovery approval slash parsing preserves the terminal raw note suffix", () => { expect(parseRuntimeCommand("/commander-recovery-approve investigation_id=inv confirm=APPROVE human_note=a b ")).toEqual({ command: "commander-recovery-approve", @@ -17,13 +39,14 @@ describe("TUI keyboard command model", () => { args: ["/commander-recovery-approve investigation_id=inv confirm=APPROVE human_note=a b "], }) }) - test("select Initialize enters onboarding shell", () => { + test("select Initialize enters spec onboarding before model setup", () => { const state = { ...initialState("/tmp/demo"), screen: "init" as const, focus: "init-choice" as const } - const next = applyKeyCommand(state, { type: "submit" }) + const result = applyKeyCommandWithEffects(state, { type: "submit" }) - expect(next.screen).toBe("main") - expect(next.lastCommand).toBe("initialize") - expect(next.commander.workIntent).toBe("TUI onboarding shell") + expect(result.state.screen).toBe("main") + expect(result.state.focus).toBe("message-box") + expect(result.state.lastCommand).toBe("initialize") + expect(result.effects).toEqual([{ type: "send-command", command: "initialize" }]) }) test("submit message from message box", () => { @@ -46,24 +69,139 @@ describe("TUI keyboard command model", () => { expect(state.focus).toBe("message-box") }) - test("initialize submit followed by message submit does not resend initialize", () => { - let state: UiState = { ...initialState("/tmp/demo"), screen: "init", focus: "init-choice" } - const effects: string[] = [] + test("model setup keeps Commander and Executor choices independent", () => { + let state: UiState = { + ...initialState("/tmp/demo"), + screen: "model-setup", + modelSetup: { + ...initialState("/tmp/demo").modelSetup, + origin: "main", + stage: "commander", + commanderChoices: [{ id: "", label: "Leave Commander unconfigured" }, { id: "commander-a", label: "Commander A" }], + executorChoices: [{ id: "", label: "Leave Executor unconfigured" }, { id: "executor-b", label: "Executor B" }], + }, + } + state = applyKeyCommand(state, { type: "select-next" }) + state = applyKeyCommand(state, { type: "submit" }) + expect(state.modelSetup.stage).toBe("executor") + expect(state.modelSetup.commanderSelection).toBe(1) + state = applyKeyCommand(state, { type: "submit" }) + expect(state.modelSetup.stage).toBe("preview") + }) - let result = applyKeyCommandWithEffects(state, { type: "submit" }) - state = result.state - effects.push(...result.effects.map((effect) => `${effect.type}:${"command" in effect ? effect.command : effect.message}`)) + test("model setup clears stale preview authority while a changed selection is being previewed", () => { + let state: UiState = { + ...initialState("/tmp/demo"), + screen: "model-setup", + modelSetup: { + ...initialState("/tmp/demo").modelSetup, + stage: "preview", + commanderChoices: [{ id: "", label: "Leave Commander unconfigured" }], + executorChoices: [ + { id: "executor-a", label: "Executor A" }, + { id: "executor-b", label: "Executor B" }, + ], + executorSelection: 0, + expectedRevision: 4, + candidateHash: "a".repeat(64), + configurationHash: "b".repeat(64), + }, + } - result = applyKeyCommandWithEffects(state, { type: "insert", text: "hello runtime" }) - state = result.state - effects.push(...result.effects.map((effect) => `${effect.type}:${"command" in effect ? effect.command : effect.message}`)) + state = applyKeyCommand(state, { type: "cancel" }) + state = applyKeyCommand(state, { type: "select-next" }) + const requested = applyKeyCommandWithEffects(state, { type: "submit" }) + expect(requested.state.modelSetup).toMatchObject({ stage: "preview", executorSelection: 1 }) + expect(requested.state.modelSetup.expectedRevision).toBeUndefined() + expect(requested.state.modelSetup.candidateHash).toBeUndefined() + expect(requested.state.modelSetup.configurationHash).toBeUndefined() + expect(requested.effects).toEqual([{ + type: "preview-model-setup", + commanderRecipeId: null, + executorRecipeId: "executor-b", + }]) + + expect(applyKeyCommandWithEffects(requested.state, { type: "submit" })).toEqual({ + state: requested.state, + effects: [], + }) + }) - result = applyKeyCommandWithEffects(state, { type: "submit" }) - state = result.state - effects.push(...result.effects.map((effect) => `${effect.type}:${"command" in effect ? effect.command : effect.message}`)) + test("committed first-run setup cannot enter the main shell before restart", () => { + const state: UiState = { + ...initialState("/tmp/demo"), + screen: "model-setup", + modelSetup: { ...initialState("/tmp/demo").modelSetup, stage: "committed", pendingRestart: true }, + } + const result = applyKeyCommandWithEffects(state, { type: "submit" }) + expect(result.state.screen).toBe("model-setup") + expect(result.state.modelSetup).toMatchObject({ stage: "committed", pendingRestart: true }) + expect(result.effects).toEqual([]) + expect(applyKeyCommandWithEffects(state, { type: "cancel" })).toEqual({ state, effects: [] }) + }) + + test("unchanged committed setup returns to its originating main screen", () => { + const state: UiState = { + ...initialState("/tmp/demo"), + screen: "model-setup", + modelSetup: { + ...initialState("/tmp/demo").modelSetup, + origin: "main", + stage: "committed", + pendingRestart: false, + }, + } + const result = applyKeyCommandWithEffects(state, { type: "submit" }) + expect(result.state).toMatchObject({ screen: "main", focus: "message-box" }) + expect(result.effects).toEqual([]) + }) - expect(effects).toEqual(["send-command:initialize", "send-user-message:hello runtime"]) - expect(state.submittedMessages).toEqual(["hello runtime"]) + test("model setup cancellation returns to the screen that opened it", () => { + const fromInit: UiState = { + ...initialState("/tmp/demo"), + screen: "model-setup", + modelSetup: { ...initialState("/tmp/demo").modelSetup, origin: "init", stage: "commander" }, + } + expect(applyKeyCommandWithEffects(fromInit, { type: "cancel" }).state).toMatchObject({ screen: "init", focus: "init-choice" }) + + const fromMain: UiState = { + ...initialState("/tmp/demo"), + screen: "model-setup", + modelSetup: { ...initialState("/tmp/demo").modelSetup, origin: "main", stage: "committed" }, + } + expect(applyKeyCommandWithEffects(fromMain, { type: "cancel" }).state).toMatchObject({ screen: "main", focus: "message-box" }) + }) + + test("required first-run setup cannot escape from the Commander stage", () => { + const base = initialState("/tmp/demo") + const state: UiState = { + ...base, + screen: "model-setup", + modelSetup: { + ...base.modelSetup, + origin: "main", + stage: "commander", + startupCheckStatus: "required", + }, + } + expect(applyKeyCommandWithModelSetupStartupGate(state, { type: "cancel" }, "required")).toEqual({ state, effects: [] }) + }) + + test("model setup confirmation owns cancellation until the durable result settles", () => { + const confirmation: UiState = { + ...initialState("/tmp/demo"), + screen: "model-setup", + modelSetup: { + ...initialState("/tmp/demo").modelSetup, + stage: "confirmation", + expectedRevision: 0, + candidateHash: "a".repeat(64), + }, + } + const submitted = applyKeyCommandWithEffects(confirmation, { type: "submit" }) + expect(submitted.state.modelSetup.stage).toBe("confirming") + expect(submitted.effects).toHaveLength(1) + expect(applyKeyCommandWithEffects(submitted.state, { type: "cancel" })).toEqual({ state: submitted.state, effects: [] }) }) test("message box keeps API keys out of TUI state while sending original message", () => { diff --git a/agentcore/tui/test/launch.test.ts b/agentcore/tui/test/launch.test.ts index 21e01c8a5..5ec91ffc7 100644 --- a/agentcore/tui/test/launch.test.ts +++ b/agentcore/tui/test/launch.test.ts @@ -2,9 +2,18 @@ import { afterEach, describe, expect, test } from "bun:test" import { mkdtemp, mkdir, readFile, rm, writeFile } from "fs/promises" import { tmpdir } from "os" import { join } from "path" -import { FakeOpenCodeAdapter, RuntimeServer } from "../../runtime/src/index" +import { + createRuntimeServerFromLaunchConfig, + buildModelSetupCandidate, + EventStore, + FakeOpenCodeAdapter, + ModelSetupService, + RuntimeServer, + RuntimeServerClient, + modelSetupCatalog, +} from "../../runtime/src/index" import type { RuntimeEvent } from "../src/events" -import { runTuiEntrypoint } from "../src/launch" +import { buildHeadlessSnapshot, runTuiEntrypoint } from "../src/launch" import type { RuntimeClient } from "../src/runtime" import { createTuiRuntimeClient } from "../src/runtime-client-factory" @@ -27,6 +36,16 @@ class TestRuntimeClient implements RuntimeClient { async command(name: string, payload?: Record): Promise { this.commandNames.push(name) + if (name === "runtime.model_setup_catalog") return modelSetupCatalog() + if (name === "runtime.model_setup_status") return { + status: "ready", + revision: 1, + setup_hash: "5".repeat(64), + active_setup_hash: "5".repeat(64), + pending_restart: false, + candidate: buildModelSetupCandidate({ commander_recipe_id: null, executor_recipe_id: null }), + active_candidate: buildModelSetupCandidate({ commander_recipe_id: null, executor_recipe_id: null }), + } if (name === "runtime.status") { return { runtimeStatus: "started", @@ -827,6 +846,19 @@ async function makeApprovedProject(dir: string): Promise { ) } +async function commitUnconfiguredModelSetup(dir: string): Promise { + const service = new ModelSetupService({ eventStore: new EventStore(join(dir, ".nxl", "events.jsonl")) }) + const choices = { commander_recipe_id: null, executor_recipe_id: null } as const + const preview = await service.preview(choices) + await service.confirm({ + ...choices, + expected_revision: preview.expected_revision, + candidate_hash: preview.candidate_hash, + confirmed_by: "tui-test-operator", + confirmation: "CONFIRM_MODEL_SETUP", + }) +} + async function readEventKinds(dir: string): Promise { try { return (await readFile(join(dir, ".nxl", "events.jsonl"), "utf8")) @@ -877,6 +909,7 @@ describe("TUI launch boundary", () => { test("real headless runtime client shows status and mission summary", async () => { const dir = await tempProject() await makeApprovedProject(dir) + await commitUnconfiguredModelSetup(dir) const output: string[] = [] await runTuiEntrypoint({ @@ -894,9 +927,202 @@ describe("TUI launch boundary", () => { expect(snapshot).toContain("recent_missions") }) + test("default approved-project setup commits through RuntimeServer persistence", async () => { + const dir = await tempProject() + await makeApprovedProject(dir) + const output: string[] = [] + const keys = [ + { type: "select-next" }, + { type: "submit" }, + { type: "select-next" }, + { type: "submit" }, + { type: "submit" }, + { type: "submit" }, + ] + await runTuiEntrypoint({ + projectDir: dir, + env: { NXL_TUI_HEADLESS: "1", NXL_TUI_KEYS: JSON.stringify(keys), NXL_OPENCODE_ADAPTER: "fake" }, + writeOutput: (snapshot) => output.push(snapshot), + }) + const snapshot = output.join("\n") + expect(snapshot).toContain("stage=committed") + expect(snapshot).toContain("pending_restart=true") + expect(snapshot).not.toContain("screen=main") + const eventKinds = await readEventKinds(dir) + expect(eventKinds).toEqual(["runtime_model_setup_committed"]) + }) + + test("real RuntimeServerClient keeps first-run setup before every runtime lifecycle boundary", async () => { + const dir = await tempProject() + await makeApprovedProject(dir) + const adapter = new SpyOpenCodeAdapter() + const server = createRuntimeServerFromLaunchConfig({ + projectDir: dir, + env: {}, + adapter, + wakeSchedulerBootstrapConfig: { + autostart_enabled: true, + interval_ms: 1_000, + max_due_items: 1, + dry_run: true, + stop_on_error: true, + }, + }) + const client = new RuntimeServerClient({ server, autoStart: true, ownsServer: false }) + const runtime = createTuiRuntimeClient({ projectDir: dir, server }) + const output: string[] = [] + await runTuiEntrypoint({ + projectDir: dir, + env: { NXL_TUI_HEADLESS: "1", NXL_TUI_KEYS: JSON.stringify([{ type: "submit" }, { type: "submit" }, { type: "submit" }, { type: "submit" }]) }, + runtime, + writeOutput: (snapshot) => output.push(snapshot), + }) + + expect(output.join("\n")).toContain("stage=committed") + expect(adapter.startCalls).toBe(0) + expect(await server.status()).toMatchObject({ runtimeStatus: "created", lockHeld: false }) + expect(await readEventKinds(dir)).toEqual(["runtime_model_setup_committed"]) + await client.shutdown() + expect(await readEventKinds(dir)).toEqual(["runtime_model_setup_committed"]) + + const nextAdapter = new SpyOpenCodeAdapter() + const nextServer = createRuntimeServerFromLaunchConfig({ projectDir: dir, env: {}, adapter: nextAdapter }) + const nextClient = new RuntimeServerClient({ server: nextServer, autoStart: true, ownsServer: true }) + await nextClient.command("runtime.status") + expect(nextAdapter.startCalls).toBe(1) + expect(await readEventKinds(dir)).toContain("runtime_started") + await nextClient.shutdown() + }) + + test("headless explicit-resume inspection cannot bypass missing model setup", async () => { + const dir = await tempProject() + await makeApprovedProject(dir) + const adapter = new SpyOpenCodeAdapter() + const server = createRuntimeServerFromLaunchConfig({ + projectDir: dir, + env: {}, + adapter, + wakeSchedulerBootstrapConfig: { + autostart_enabled: true, + interval_ms: 1_000, + max_due_items: 1, + dry_run: true, + stop_on_error: true, + }, + }) + const runtime = createTuiRuntimeClient({ projectDir: dir, server }) + const output: string[] = [] + await runTuiEntrypoint({ + projectDir: dir, + env: { + NXL_TUI_HEADLESS: "1", + NXL_TUI_KEYS: JSON.stringify([ + { type: "submit" }, + { type: "insert", text: "/opencode-session-plan objective=must not start" }, + { type: "submit" }, + ]), + }, + runtime, + writeOutput: (snapshot) => output.push(snapshot), + }) + + expect(output.join("\n")).toContain("screen=model-setup") + expect(adapter.startCalls).toBe(0) + expect(await server.status()).toMatchObject({ runtimeStatus: "created", lockHeld: false }) + expect(await readEventKinds(dir)).toEqual([]) + }) + + test("headless cancel cannot escape required model setup and start Runtime", async () => { + const dir = await tempProject() + await makeApprovedProject(dir) + const adapter = new SpyOpenCodeAdapter() + const server = createRuntimeServerFromLaunchConfig({ + projectDir: dir, + env: {}, + adapter, + wakeSchedulerBootstrapConfig: { + autostart_enabled: true, + interval_ms: 1_000, + max_due_items: 1, + dry_run: true, + stop_on_error: true, + }, + }) + const runtime = createTuiRuntimeClient({ projectDir: dir, server }) + const output: string[] = [] + await runTuiEntrypoint({ + projectDir: dir, + env: { + NXL_TUI_HEADLESS: "1", + NXL_TUI_KEYS: JSON.stringify([ + { type: "cancel" }, + { type: "insert", text: "must not start" }, + { type: "submit" }, + ]), + }, + runtime, + writeOutput: (snapshot) => output.push(snapshot), + }) + + expect(output.join("\n")).toContain("screen=model-setup") + expect(adapter.startCalls).toBe(0) + expect(await server.status()).toMatchObject({ runtimeStatus: "created", lockHeld: false }) + expect(await readEventKinds(dir)).toEqual([]) + }) + + test("real first-run TUI never reaches the configured process spawn seam", async () => { + const dir = await tempProject() + await makeApprovedProject(dir) + let spawnCalls = 0 + const runtime = createTuiRuntimeClient({ + projectDir: dir, + env: { + NXL_RUNTIME_CLIENT: "real", + NXL_OPENCODE_ADAPTER: "process", + NXL_OPENCODE_COMMAND: "opencode", + }, + openCodeAdapterFactoryOptions: { + spawn: (() => { + spawnCalls += 1 + throw new Error("first-run setup must not spawn OpenCode") + }) as never, + }, + }) + const output: string[] = [] + await runTuiEntrypoint({ + projectDir: dir, + env: { NXL_TUI_HEADLESS: "1", NXL_TUI_KEYS: JSON.stringify([{ type: "submit" }, { type: "submit" }, { type: "submit" }, { type: "submit" }]) }, + runtime, + writeOutput: (snapshot) => output.push(snapshot), + }) + + expect(output.join("\n")).toContain("stage=committed") + expect(spawnCalls).toBe(0) + expect(await readEventKinds(dir)).toEqual(["runtime_model_setup_committed"]) + }) + + test("same-process spec approval cannot open model setup through the bootstrap fake client", async () => { + const dir = await tempProject() + const runtime = createTuiRuntimeClient({ projectDir: dir, env: {} }) + await makeApprovedProject(dir) + + const snapshot = await buildHeadlessSnapshot(runtime, dir, { + NXL_TUI_KEYS: JSON.stringify([ + { type: "submit" }, + { type: "insert", text: "/model-setup" }, + { type: "submit" }, + ]), + }) + + expect(snapshot).toContain("Restart NexusLoop after spec approval before opening model setup") + expect(snapshot).not.toContain("stage=committed") + expect(await readEventKinds(dir)).not.toContain("runtime_model_setup_committed") + }) + test("real headless runtime client submits a message and refreshes mission records", async () => { const dir = await tempProject() await makeApprovedProject(dir) + await commitUnconfiguredModelSetup(dir) const output: string[] = [] const keys = [ { type: "submit" }, @@ -926,6 +1152,7 @@ describe("TUI launch boundary", () => { test("status and missions commands update runtime panels", async () => { const dir = await tempProject() await makeApprovedProject(dir) + await commitUnconfiguredModelSetup(dir) const output: string[] = [] const keys = [ { type: "submit" }, @@ -1083,6 +1310,7 @@ describe("TUI launch boundary", () => { test("real headless runtime client loads projection and topics through research command", async () => { const dir = await tempProject() await makeApprovedProject(dir) + await commitUnconfiguredModelSetup(dir) const output: string[] = [] const keys = [ { type: "submit" }, @@ -1110,6 +1338,7 @@ describe("TUI launch boundary", () => { test("real headless runtime client renders empty commander queue surface", async () => { const dir = await tempProject() await makeApprovedProject(dir) + await commitUnconfiguredModelSetup(dir) const output: string[] = [] const keys = [ { type: "submit" }, @@ -1995,8 +2224,9 @@ describe("TUI launch boundary", () => { test("headless executor review on stopped real runtime does not start OpenCode", async () => { const dir = await tempProject() await makeApprovedProject(dir) + await commitUnconfiguredModelSetup(dir) const adapter = new SpyOpenCodeAdapter() - const server = new RuntimeServer({ projectDir: dir, adapter }) + const server = createRuntimeServerFromLaunchConfig({ projectDir: dir, env: {}, adapter }) const runtime = createTuiRuntimeClient({ projectDir: dir, server, env: {} }) const output: string[] = [] const keys = [ @@ -2022,6 +2252,7 @@ describe("TUI launch boundary", () => { test("shutdown command does not report a false post-shutdown refresh error", async () => { const dir = await tempProject() await makeApprovedProject(dir) + await commitUnconfiguredModelSetup(dir) const output: string[] = [] const keys = [ { type: "submit" }, diff --git a/agentcore/tui/test/reducer.test.ts b/agentcore/tui/test/reducer.test.ts index c0257d742..9dc1db345 100644 --- a/agentcore/tui/test/reducer.test.ts +++ b/agentcore/tui/test/reducer.test.ts @@ -1,9 +1,45 @@ import { describe, expect, test } from "bun:test" -import { reduceRuntimeEvent } from "../src/reducer" +import { + modelSetupStartupGateAllowsInput, + modelSetupStartupGateAllowsCommand, + reduceRuntimeEvent, + reduceRuntimeEventDuringModelSetupGate, +} from "../src/reducer" import { layoutSnapshot } from "../src/snapshot" import { initialState } from "../src/state" describe("TUI runtime event reducer", () => { + test("ProjectInitialized cannot expose Resume before model setup authority settles", () => { + const boot = initialState("/tmp/demo") + const initialized = { type: "ProjectInitialized", projectDir: "/tmp/demo" } as const + expect(reduceRuntimeEventDuringModelSetupGate(boot, initialized, "pending")).toMatchObject({ + screen: "boot", + projectDir: "/tmp/demo", + }) + + const setup = { ...boot, screen: "model-setup" as const, focus: "init-choice" as const } + expect(reduceRuntimeEventDuringModelSetupGate(setup, initialized, "required")).toMatchObject({ + screen: "model-setup", + focus: "init-choice", + }) + expect(reduceRuntimeEventDuringModelSetupGate(boot, initialized, "clear")).toMatchObject({ + screen: "resume", + focus: "resume-choice", + }) + expect(modelSetupStartupGateAllowsInput("pending")).toBe(false) + expect(modelSetupStartupGateAllowsInput("blocked")).toBe(false) + expect(modelSetupStartupGateAllowsInput("required")).toBe(false) + expect(modelSetupStartupGateAllowsInput("clear")).toBe(true) + expect(modelSetupStartupGateAllowsCommand(setup, { type: "submit" }, "required")).toBe(true) + expect(modelSetupStartupGateAllowsCommand({ ...setup, screen: "main" }, { type: "submit" }, "required")).toBe(false) + const failed = { + ...initialState("/tmp/demo"), + screen: "model-setup" as const, + modelSetup: { ...initialState("/tmp/demo").modelSetup, stage: "loading" as const, commandError: "setup unavailable" }, + } + expect(modelSetupStartupGateAllowsCommand(failed, { type: "submit" }, "blocked")).toBe(true) + expect(modelSetupStartupGateAllowsCommand(failed, { type: "insert", text: "x" }, "blocked")).toBe(false) + }) test("ProjectUninitialized routes to init screen", () => { const state = reduceRuntimeEvent(initialState("/tmp/demo"), { type: "ProjectUninitialized", projectDir: "/tmp/demo" }) diff --git a/agentcore/tui/test/runtime-client-factory.test.ts b/agentcore/tui/test/runtime-client-factory.test.ts index 262bccd07..0d79b4c8a 100644 --- a/agentcore/tui/test/runtime-client-factory.test.ts +++ b/agentcore/tui/test/runtime-client-factory.test.ts @@ -1,12 +1,12 @@ import { afterEach, describe, expect, test } from "bun:test" import { mkdtemp, rm, writeFile, mkdir } from "fs/promises" -import { join } from "path" +import { basename, join } from "path" import { tmpdir } from "os" import { FakeRuntimeClient } from "../src/runtime" import { reduceRuntimeEvent } from "../src/reducer" import { initialState } from "../src/state" import { createTuiRuntimeClient, isTuiRuntimeEvent, readRuntimeClientKind, TuiRuntimeServerClient } from "../src/runtime-client-factory" -import { EventStore, ExternalApiConnectorRegistry, FakeExternalApiTransport, FakeOpenCodeAdapter, ResearchDb, RuntimeServer, type OpenCodeProcessEventSource, type OpenCodeSpawnedProcess } from "../../runtime/src/index" +import { EventStore, ExternalApiConnectorRegistry, FakeExternalApiTransport, FakeOpenCodeAdapter, ModelSetupService, ResearchDb, RuntimeServer, type OpenCodeProcessEventSource, type OpenCodeSpawnedProcess } from "../../runtime/src/index" import { applyRuntimeUiEffect } from "../src/runtime-effects" const cleanup: string[] = [] @@ -43,6 +43,16 @@ async function makeApprovedProject(dir: string): Promise { 2, ), ) + const setup = new ModelSetupService({ eventStore: new EventStore(join(dir, ".nxl", "events.jsonl")) }) + const choices = { commander_recipe_id: null, executor_recipe_id: null } as const + const preview = await setup.preview(choices) + await setup.confirm({ + ...choices, + expected_revision: preview.expected_revision, + candidate_hash: preview.candidate_hash, + confirmed_by: "tui-runtime-test-operator", + confirmation: "CONFIRM_MODEL_SETUP", + }) } function minimaxConnector() { @@ -139,11 +149,52 @@ async function readFirst(stream: AsyncIterable): Promise { } describe("TUI runtime client factory", () => { - test("no env keeps fake default behavior", async () => { + test("no env keeps legacy fake behavior only before spec approval", async () => { const dir = await tempProject() const client = createTuiRuntimeClient({ projectDir: dir, env: {} }) expect(client).toBeInstanceOf(FakeRuntimeClient) + expect(client.modelSetupAuthority).toBe("restart_required") + }) + + test("no env routes an approved project through the real RuntimeServer client", async () => { + const dir = await tempProject() + await makeApprovedProject(dir) + const client = createTuiRuntimeClient({ projectDir: dir, env: { NXL_OPENCODE_ADAPTER: "fake" } }) + + expect(client).toBeInstanceOf(TuiRuntimeServerClient) + await (client as TuiRuntimeServerClient).runtime.shutdown() + }) + + test("auto mode resolves an approved ancestor before selecting and constructing the real client", async () => { + const dir = await tempProject() + await makeApprovedProject(dir) + const nested = join(dir, "src", "nested") + await mkdir(nested, { recursive: true }) + + const client = createTuiRuntimeClient({ projectDir: nested, env: { NXL_OPENCODE_ADAPTER: "fake" } }) + + expect(client).toBeInstanceOf(TuiRuntimeServerClient) + const runtime = (client as TuiRuntimeServerClient).runtime + await expect(runtime.command("runtime.status")).resolves.toMatchObject({ projectName: basename(dir) }) + await runtime.shutdown() + }) + + test("no env routes a valid large approved spec through the real RuntimeServer client", async () => { + const dir = await tempProject() + await mkdir(join(dir, ".nxl", "spec"), { recursive: true }) + await writeFile(join(dir, ".nxl", "spec", "current.json"), JSON.stringify({ + spec_id: "spec_large_approved", + version: 1, + status: "approved", + objective: "x".repeat(70_000), + success_metrics: ["tests pass"], + })) + + const client = createTuiRuntimeClient({ projectDir: dir, env: { NXL_OPENCODE_ADAPTER: "fake" } }) + + expect(client).toBeInstanceOf(TuiRuntimeServerClient) + await (client as TuiRuntimeServerClient).runtime.shutdown() }) test("NXL_RUNTIME_CLIENT=fake explicitly selects fake", async () => { @@ -152,6 +203,7 @@ describe("TUI runtime client factory", () => { expect(readRuntimeClientKind({ NXL_RUNTIME_CLIENT: "fake" })).toBe("fake") expect(client).toBeInstanceOf(FakeRuntimeClient) + expect(client.modelSetupAuthority).toBe("fixture") }) test("NXL_RUNTIME_CLIENT=real creates RuntimeServer-backed client", async () => { @@ -414,6 +466,20 @@ describe("TUI runtime client factory", () => { await client.runtime.shutdown() }) + test("real runtime stream exposes an uninitialized project without starting it", async () => { + const dir = await tempProject() + const client = createTuiRuntimeClient({ + projectDir: dir, + env: { NXL_RUNTIME_CLIENT: "real", NXL_OPENCODE_ADAPTER: "fake" }, + }) as TuiRuntimeServerClient + const iterator = client.stream()[Symbol.asyncIterator]() + expect((await iterator.next()).value).toMatchObject({ type: "RuntimeReady" }) + expect((await iterator.next()).value).toEqual({ type: "ProjectUninitialized", projectDir: dir }) + expect(await client.runtime.server.status()).toMatchObject({ runtimeStatus: "created", specApproved: false, lockHeld: false }) + await iterator.return?.() + await client.runtime.shutdown() + }) + test("real runtime client maps resume menu commands to runtime commands", async () => { const commands: string[] = [] const shutdownOptions: unknown[] = [] diff --git a/agentcore/tui/test/runtime-effects.test.ts b/agentcore/tui/test/runtime-effects.test.ts index 2bb493dc6..4794005af 100644 --- a/agentcore/tui/test/runtime-effects.test.ts +++ b/agentcore/tui/test/runtime-effects.test.ts @@ -1,9 +1,12 @@ import { describe, expect, test } from "bun:test" +import { buildModelSetupCandidate, modelSetupCatalog } from "../../runtime/src/index" import { commanderRecoveryApprovalDisplay, commanderRecoveryAuthorityValues, commanderRecoveryPreviewDiagnostics } from "../src/commander-recovery-view" import type { RuntimeEvent } from "../src/events" -import { applyRuntimeUiEffect } from "../src/runtime-effects" -import { parseRuntimeCommand } from "../src/keyboard" +import { applyRuntimeUiEffect, refreshRuntimeRecords } from "../src/runtime-effects" +import { modelSetupCompletionCopy } from "../src/model-setup-view" +import { applyKeyCommandWithEffects, parseRuntimeCommand } from "../src/keyboard" import { commandTypeFromSlash } from "../src/operator-actions" +import { mergeRuntimeEffectState } from "../src/runtime-state-merge" import { FakeRuntimeClient, orderQueueItems, type RuntimeClient } from "../src/runtime" import { layoutSnapshot } from "../src/snapshot" import { snapshotUiState } from "../src/state-snapshot" @@ -42,8 +45,18 @@ class RejectingRuntime implements RuntimeClient { this.sendCommandCalls += 1 throw new Error("runtime should not receive init command") } - async command(): Promise { + async command(name: string): Promise { this.commandCalls += 1 + if (name === "runtime.model_setup_catalog") return modelSetupCatalog() + if (name === "runtime.model_setup_status") return { + status: "ready", + revision: 1, + setup_hash: "5".repeat(64), + active_setup_hash: "5".repeat(64), + pending_restart: false, + candidate: buildModelSetupCandidate({ commander_recipe_id: null, executor_recipe_id: null }), + active_candidate: buildModelSetupCandidate({ commander_recipe_id: null, executor_recipe_id: null }), + } throw new Error("runtime should not receive init command") } } @@ -56,7 +69,17 @@ class RefreshFailAfterSubmitRuntime implements RuntimeClient { async sendCommand(): Promise { return { ok: true } } - async command(): Promise { + async command(name: string): Promise { + if (name === "runtime.model_setup_catalog") return modelSetupCatalog() + if (name === "runtime.model_setup_status") return { + status: "ready", + revision: 1, + setup_hash: "5".repeat(64), + active_setup_hash: "5".repeat(64), + pending_restart: false, + candidate: buildModelSetupCandidate({ commander_recipe_id: null, executor_recipe_id: null }), + active_candidate: buildModelSetupCandidate({ commander_recipe_id: null, executor_recipe_id: null }), + } throw new Error("refresh failed after accepted mission") } } @@ -546,6 +569,18 @@ class OpenCodeHandoffRuntime implements RuntimeClient { async command(name: string, payload?: Record): Promise { this.calls.push({ name, payload }) switch (name) { + case "runtime.model_setup_catalog": + return modelSetupCatalog() + case "runtime.model_setup_status": + return { + status: "ready", + revision: 1, + setup_hash: "5".repeat(64), + active_setup_hash: "5".repeat(64), + pending_restart: false, + candidate: buildModelSetupCandidate({ commander_recipe_id: null, executor_recipe_id: null }), + active_candidate: buildModelSetupCandidate({ commander_recipe_id: null, executor_recipe_id: null }), + } case "runtime.preview_opencode_handoff": return { proposal_id: payload?.proposalId, @@ -801,6 +836,18 @@ class MissionExecutionRuntime implements RuntimeClient { this.calls.push(`${name}:${JSON.stringify(payload)}`) const missionId = String(payload.missionId ?? payload.mission_id ?? "mission-1") switch (name) { + case "runtime.model_setup_catalog": + return modelSetupCatalog() + case "runtime.model_setup_status": + return { + status: "ready", + revision: 1, + setup_hash: "5".repeat(64), + active_setup_hash: "5".repeat(64), + pending_restart: false, + candidate: buildModelSetupCandidate({ commander_recipe_id: null, executor_recipe_id: null }), + active_candidate: buildModelSetupCandidate({ commander_recipe_id: null, executor_recipe_id: null }), + } case "runtime.status": return { runtimeStatus: "started", @@ -976,6 +1023,319 @@ class FailingMissionExecutionRuntime extends MissionExecutionRuntime { } describe("runtime UI effects", () => { + test("model setup loads exact recipes, previews, and confirms through RuntimeClient", async () => { + const calls: Array<{ name: string; payload?: Record }> = [] + const runtime = new FakeRuntimeClient("/tmp/demo", "demo") + runtime.command = (async (name: string, payload: Record = {}) => { + calls.push({ name, payload }) + if (name === "runtime.model_setup_catalog") return modelSetupCatalog() + if (name === "runtime.model_setup_status") return { status: "missing", revision: 0, pending_restart: false } + if (name === "runtime.preview_model_setup") return { expected_revision: 0, candidate_hash: "a".repeat(64), configuration_hash: "b".repeat(64), restart_required: true } + if (name === "runtime.confirm_model_setup") return { status: "committed", revision: 1, setup_hash: "c".repeat(64), candidate_hash: "a".repeat(64), restart_required: true } + throw new Error(`unexpected ${name}`) + }) as RuntimeClient["command"] + + let state = await applyRuntimeUiEffect({ ...initialState("/tmp/demo"), screen: "model-setup" }, runtime, { type: "load-model-setup" }) + expect(state.modelSetup.commanderChoices.map((item) => item.id)).toEqual(["", "commander-anthropic-claude-sonnet-4-5", "commander-google-gemini-2-5-flash", "commander-openai-gpt-4-1-mini-responses"]) + state = await applyRuntimeUiEffect(state, runtime, { type: "preview-model-setup", commanderRecipeId: "commander-anthropic-claude-sonnet-4-5", executorRecipeId: "executor-anthropic-claude-sonnet-4-5" }) + expect(state.modelSetup.candidateHash).toBe("a".repeat(64)) + expect(state.modelSetup.pendingRestart).toBe(false) + state = await applyRuntimeUiEffect(state, runtime, { type: "confirm-model-setup", commanderRecipeId: "commander-anthropic-claude-sonnet-4-5", executorRecipeId: "executor-anthropic-claude-sonnet-4-5", expectedRevision: 0, candidateHash: "a".repeat(64) }) + expect(state.modelSetup).toMatchObject({ stage: "committed", pendingRestart: true, pendingSetupHash: "c".repeat(64) }) + expect(calls.map((item) => item.name)).toEqual([ + "runtime.model_setup_catalog", + "runtime.model_setup_status", + "runtime.preview_model_setup", + "runtime.confirm_model_setup", + ]) + }) + test("model setup slash command enters the async screen with main-shell origin", async () => { + const runtime = new FakeRuntimeClient("/tmp/demo", "demo") + runtime.command = (async (name: string) => { + if (name === "runtime.model_setup_catalog") return modelSetupCatalog() + if (name === "runtime.model_setup_status") return { status: "missing", revision: 0, pending_restart: false } + throw new Error(`unexpected ${name}`) + }) as RuntimeClient["command"] + const state = await applyRuntimeUiEffect( + { ...initialState("/tmp/demo"), screen: "main", focus: "message-box" }, + runtime, + { type: "send-command", command: "model-setup" }, + ) + expect(state).toMatchObject({ screen: "model-setup", modelSetup: { origin: "main", stage: "commander" } }) + }) + test("unchanged model setup confirmation stays active without requesting restart", async () => { + const runtime = new FakeRuntimeClient("/tmp/demo", "demo") + runtime.command = (async (name: string) => { + if (name === "runtime.confirm_model_setup") return { + status: "idempotent", + revision: 1, + setup_hash: "c".repeat(64), + candidate_hash: "a".repeat(64), + restart_required: false, + } + throw new Error(`unexpected ${name}`) + }) as RuntimeClient["command"] + const state = await applyRuntimeUiEffect( + { ...initialState("/tmp/demo"), screen: "model-setup" }, + runtime, + { type: "confirm-model-setup", commanderRecipeId: "commander-a", executorRecipeId: "executor-b", expectedRevision: 1, candidateHash: "a".repeat(64) }, + ) + expect(state.modelSetup).toMatchObject({ stage: "committed", pendingRestart: false, pendingSetupHash: "c".repeat(64) }) + expect(state.systemActions.at(-1)).toEqual({ title: "Model setup unchanged", detail: "The active selection already matches this setup" }) + expect(modelSetupCompletionCopy(state.modelSetup.pendingRestart)).toEqual({ + headline: "Selection already active. No restart is required.", + instructions: "Enter returns to the main shell.", + }) + }) + test("model setup durable confirmation survives an Escape key race", async () => { + let release!: () => void + const gate = new Promise((resolve) => { release = resolve }) + const runtime = new FakeRuntimeClient("/tmp/demo", "demo") + runtime.command = (async (name: string) => { + if (name === "runtime.confirm_model_setup") { + await gate + return { status: "committed", revision: 1, setup_hash: "c".repeat(64), candidate_hash: "a".repeat(64), restart_required: true } + } + throw new Error(`unexpected ${name}`) + }) as RuntimeClient["command"] + const confirmation: UiState = { + ...initialState("/tmp/demo"), + screen: "model-setup", + modelSetup: { + ...initialState("/tmp/demo").modelSetup, + stage: "confirmation", + expectedRevision: 0, + candidateHash: "a".repeat(64), + }, + } + const submitted = applyKeyCommandWithEffects(confirmation, { type: "submit" }) + const baseline = submitted.state + const effectResult = applyRuntimeUiEffect(baseline, runtime, submitted.effects[0]!) + const current = applyKeyCommandWithEffects(baseline, { type: "cancel" }).state + release() + const merged = mergeRuntimeEffectState(current, await effectResult, baseline.systemActions.length, baseline) + expect(merged.modelSetup).toMatchObject({ stage: "committed", pendingRestart: true, pendingSetupHash: "c".repeat(64) }) + }) + test("failed model setup confirmation releases the owned confirmation state", async () => { + const runtime = new FakeRuntimeClient("/tmp/demo", "demo") + runtime.command = (async () => { throw new Error("confirmation failed") }) as RuntimeClient["command"] + const state = await applyRuntimeUiEffect( + { + ...initialState("/tmp/demo"), + screen: "model-setup", + modelSetup: { ...initialState("/tmp/demo").modelSetup, stage: "confirming" }, + }, + runtime, + { type: "confirm-model-setup", commanderRecipeId: null, executorRecipeId: null, expectedRevision: 0, candidateHash: "a".repeat(64) }, + ) + expect(state.modelSetup).toMatchObject({ stage: "confirmation", commandError: "confirmation failed" }) + }) + test("model setup keeps active readiness attached to active selection while a replacement is pending", async () => { + const runtime = new FakeRuntimeClient("/tmp/demo", "demo") + runtime.command = (async (name: string) => { + if (name === "runtime.model_setup_catalog") return modelSetupCatalog() + if (name === "runtime.model_setup_status") return { + status: "ready", + revision: 2, + setup_hash: "2".repeat(64), + active_setup_hash: "1".repeat(64), + pending_restart: true, + active_candidate: buildModelSetupCandidate({ commander_recipe_id: "commander-anthropic-claude-sonnet-4-5", executor_recipe_id: "executor-anthropic-claude-sonnet-4-5" }), + candidate: buildModelSetupCandidate({ commander_recipe_id: "commander-google-gemini-2-5-flash", executor_recipe_id: "executor-google-gemini-2-5-flash" }), + commander_role_readiness: { selection_status: "selected", credential_connection_status: "connected", lifecycle_status: "ready" }, + executor_role_readiness: { selection_status: "selected", credential_connection_status: "connected", lifecycle_status: "ready" }, + } + throw new Error(`unexpected ${name}`) + }) as RuntimeClient["command"] + const state = await applyRuntimeUiEffect({ ...initialState("/tmp/demo"), screen: "model-setup" }, runtime, { type: "load-model-setup" }) + expect(state.modelSetup).toMatchObject({ + activeCommanderLabel: "Anthropic Claude Sonnet 4.5", + activeExecutorLabel: "Anthropic Claude Sonnet 4.5", + pendingCommanderLabel: "Google Gemini 2.5 Flash", + pendingExecutorLabel: "Google Gemini 2.5 Flash", + commanderReadiness: "selected; credential connected; lifecycle ready", + executorReadiness: "selected; credential connected; lifecycle ready", + pendingRestart: true, + }) + }) + test("initialization skips setup only for an exact active durable setup", async () => { + const runtime = new FakeRuntimeClient("/tmp/demo", "demo") + let state = await applyRuntimeUiEffect( + { ...initialState("/tmp/demo"), screen: "model-setup" }, + runtime, + { type: "load-model-setup", continueInitializationIfActive: true }, + ) + expect(state.screen).toBe("main") + expect(state.systemActions.at(-1)?.detail).toContain("Active model authority verified") + + runtime.command = (async (name: string) => { + if (name === "runtime.model_setup_catalog") return modelSetupCatalog() + if (name === "runtime.model_setup_status") return { + status: "ready", + revision: 2, + setup_hash: "5".repeat(64), + active_setup_hash: "4".repeat(64), + pending_restart: true, + candidate: buildModelSetupCandidate({ commander_recipe_id: null, executor_recipe_id: null }), + active_candidate: buildModelSetupCandidate({ commander_recipe_id: null, executor_recipe_id: null }), + } + throw new Error(`unexpected ${name}`) + }) as RuntimeClient["command"] + state = await applyRuntimeUiEffect( + { ...initialState("/tmp/demo"), screen: "model-setup" }, + runtime, + { type: "load-model-setup", continueInitializationIfActive: true }, + ) + expect(state.screen).toBe("model-setup") + }) + test("runtime refresh enters the existing onboarding screen when durable setup is missing", async () => { + const runtime = new FakeRuntimeClient("/tmp/demo", "demo") + const commands: string[] = [] + runtime.command = (async (name: string) => { + commands.push(name) + if (name === "runtime.model_setup_catalog") return modelSetupCatalog() + if (name === "runtime.model_setup_status") return { status: "missing", revision: 0, pending_restart: false } + throw new Error(`unexpected ${name}`) + }) as RuntimeClient["command"] + const state = await refreshRuntimeRecords({ ...initialState("/tmp/demo"), screen: "resume" }, runtime) + expect(state).toMatchObject({ screen: "model-setup", modelSetup: { origin: "main", stage: "commander" } }) + expect(commands.sort()).toEqual(["runtime.model_setup_catalog", "runtime.model_setup_status"]) + }) + test("runtime refresh keeps a first committed setup pre-start until fresh reconstruction", async () => { + const runtime = new FakeRuntimeClient("/tmp/demo", "demo") + const commands: string[] = [] + runtime.command = (async (name: string) => { + commands.push(name) + if (name === "runtime.model_setup_catalog") return modelSetupCatalog() + if (name === "runtime.model_setup_status") return { + status: "ready", + revision: 1, + setup_hash: "5".repeat(64), + pending_restart: true, + candidate: buildModelSetupCandidate({ commander_recipe_id: null, executor_recipe_id: null }), + } + throw new Error(`unexpected ${name}`) + }) as RuntimeClient["command"] + const state = await refreshRuntimeRecords({ ...initialState("/tmp/demo"), screen: "resume" }, runtime) + expect(state).toMatchObject({ screen: "model-setup", modelSetup: { pendingRestart: true } }) + expect(commands.sort()).toEqual(["runtime.model_setup_catalog", "runtime.model_setup_status"]) + }) + test("runtime refresh fails closed when model setup authority cannot be inspected", async () => { + for (const failure of ["reject", "malformed_catalog", "malformed_status", "malformed_candidate", "divergent_active_candidate"] as const) { + const runtime = new FakeRuntimeClient("/tmp/demo", "demo") + const commands: string[] = [] + runtime.command = (async (name: string) => { + commands.push(name) + if (name === "runtime.model_setup_catalog") { + if (failure === "reject") throw new Error("setup status unavailable") + return failure === "malformed_catalog" + ? { commander_recipes: "malformed", executor_recipes: [] } + : modelSetupCatalog() + } + if (name === "runtime.model_setup_status") { + if (failure === "malformed_status") return { status: "missing", revision: 1, pending_restart: false } + if (failure === "malformed_candidate") return { + status: "ready", + revision: 1, + setup_hash: "5".repeat(64), + active_setup_hash: "5".repeat(64), + pending_restart: false, + candidate: { choices: { commander_recipe_id: "commander-forged", executor_recipe_id: null } }, + active_candidate: { choices: { commander_recipe_id: "commander-forged", executor_recipe_id: null } }, + } + if (failure === "divergent_active_candidate") return { + status: "ready", + revision: 1, + setup_hash: "5".repeat(64), + active_setup_hash: "5".repeat(64), + pending_restart: false, + candidate: buildModelSetupCandidate({ commander_recipe_id: "commander-anthropic-claude-sonnet-4-5", executor_recipe_id: null }), + active_candidate: buildModelSetupCandidate({ commander_recipe_id: "commander-google-gemini-2-5-flash", executor_recipe_id: null }), + } + return { status: "missing", revision: 0, pending_restart: false } + } + throw new Error(`unexpected ${name}`) + }) as RuntimeClient["command"] + const state = await refreshRuntimeRecords({ ...initialState("/tmp/demo"), screen: "resume" }, runtime) + expect(state).toMatchObject({ screen: "model-setup", modelSetup: { startupCheckStatus: "failed", stage: "loading" } }) + expect(state.modelSetup.commandError).toBeDefined() + expect(layoutSnapshot(state)).toContain(" error=") + expect(commands.sort()).toEqual(["runtime.model_setup_catalog", "runtime.model_setup_status"]) + } + }) + test("runtime refresh preserves the shell when mutually exclusive registry authority is already active", async () => { + for (const authoritySource of ["explicit", "legacy_commander_environment"] as const) { + const runtime = new FakeRuntimeClient("/tmp/demo", "demo") + runtime.command = (async (name: string) => { + if (name === "runtime.status") return {} + if (name === "runtime.model_setup_catalog") return modelSetupCatalog() + if (name === "runtime.model_setup_status") return { + status: "missing", + revision: 0, + pending_restart: false, + active_authority_source: authoritySource, + } + if (name.startsWith("runtime.")) return [] + throw new Error(`unexpected ${name}`) + }) as RuntimeClient["command"] + const state = await refreshRuntimeRecords({ ...initialState("/tmp/demo"), screen: "main" }, runtime) + expect(state.screen).toBe("main") + } + }) + test("setup inspection retry continues when mutually exclusive registry authority is active", async () => { + for (const authoritySource of ["explicit", "legacy_commander_environment"] as const) { + const runtime = new FakeRuntimeClient("/tmp/demo", "demo") + runtime.command = (async (name: string) => { + if (name === "runtime.model_setup_catalog") return modelSetupCatalog() + if (name === "runtime.model_setup_status") return { + status: "missing", + revision: 0, + pending_restart: false, + active_authority_source: authoritySource, + } + throw new Error(`unexpected ${name}`) + }) as RuntimeClient["command"] + const failed = { + ...initialState("/tmp/demo"), + screen: "model-setup" as const, + modelSetup: { + ...initialState("/tmp/demo").modelSetup, + stage: "loading" as const, + startupCheckStatus: "failed" as const, + commandError: "model setup status unavailable", + }, + } + const state = await applyRuntimeUiEffect(failed, runtime, { + type: "load-model-setup", + enterIfMissing: true, + continueInitializationIfActive: true, + }) + expect(state).toMatchObject({ screen: "main", modelSetup: { startupCheckStatus: "clear", commandError: undefined } }) + expect(state.systemActions.at(-1)?.detail).toContain("Active model authority verified") + } + }) + test("explicit model setup command does not expose onboarding under non-setup authority", async () => { + for (const authoritySource of ["explicit", "legacy_commander_environment"] as const) { + const runtime = new FakeRuntimeClient("/tmp/demo", "demo") + runtime.command = (async (name: string) => { + if (name === "runtime.model_setup_catalog") return modelSetupCatalog() + if (name === "runtime.model_setup_status") return { + status: "missing", + revision: 0, + pending_restart: false, + active_authority_source: authoritySource, + } + throw new Error(`unexpected ${name}`) + }) as RuntimeClient["command"] + const state = await applyRuntimeUiEffect( + { ...initialState("/tmp/demo"), screen: "model-setup", modelSetup: { ...initialState("/tmp/demo").modelSetup, origin: "main" } }, + runtime, + { type: "load-model-setup" }, + ) + expect(state).toMatchObject({ screen: "main", focus: "message-box", modelSetup: { startupCheckStatus: "clear" } }) + expect(state.systemActions.at(-1)).toEqual({ title: "Model setup unavailable", detail: "Non-setup model authority is active" }) + } + }) test("Commander recovery staged commands classify mutations as writes", () => { for (const command of [ "/commander-recovery-approve", @@ -1030,7 +1390,7 @@ describe("runtime UI effects", () => { ]) }) - test("init-only commands are handled locally without runtime dispatch", async () => { + test("init-only commands inspect setup without dispatching an init command", async () => { const runtime = new RejectingRuntime() const state = initialState("/tmp/demo") @@ -1038,7 +1398,7 @@ describe("runtime UI effects", () => { expect(next.lastCommand).toBe("initialize") expect(next.runtimeCommandError).toBeUndefined() - expect(runtime.commandCalls).toBe(0) + expect(runtime.commandCalls).toBe(2) expect(runtime.sendCommandCalls).toBe(0) }) diff --git a/agentcore/tui/test/runtime-state-merge.test.ts b/agentcore/tui/test/runtime-state-merge.test.ts index 23a1d063d..c59d22211 100644 --- a/agentcore/tui/test/runtime-state-merge.test.ts +++ b/agentcore/tui/test/runtime-state-merge.test.ts @@ -4,6 +4,53 @@ import { mergeRuntimeEffectState } from "../src/runtime-state-merge" import { initialState, type UiState } from "../src/state" describe("interactive runtime effect state merge", () => { + test("applies async navigation only while the initiating screen and focus remain current", () => { + const baseline: UiState = { ...initialState("/tmp/demo"), screen: "main", focus: "message-box" } + const effectResult: UiState = { + ...baseline, + screen: "model-setup", + focus: "init-choice", + modelSetup: { ...baseline.modelSetup, origin: "main", stage: "commander" }, + } + expect(mergeRuntimeEffectState(baseline, effectResult, 0, baseline)).toMatchObject({ + screen: "model-setup", + focus: "init-choice", + }) + + const moved: UiState = { ...baseline, screen: "resume", focus: "resume-choice" } + expect(mergeRuntimeEffectState(moved, effectResult, 0, baseline)).toMatchObject({ + screen: "resume", + focus: "resume-choice", + }) + }) + + test("lets missing setup supersede the stream-driven boot to resume transition", () => { + const baseline = initialState("/tmp/demo") + const streamInitialized: UiState = { + ...baseline, + screen: "resume", + focus: "resume-choice", + } + const missingSetup: UiState = { + ...baseline, + screen: "model-setup", + focus: "init-choice", + modelSetup: { ...baseline.modelSetup, origin: "main", stage: "commander" }, + } + + expect(mergeRuntimeEffectState(streamInitialized, missingSetup, 0, baseline)).toMatchObject({ + screen: "model-setup", + focus: "init-choice", + }) + + const operatorMoved = { ...streamInitialized, resumeSelection: 1 } + expect(mergeRuntimeEffectState(operatorMoved, missingSetup, 0, baseline)).toMatchObject({ + screen: "resume", + focus: "resume-choice", + resumeSelection: 1, + }) + }) + test("rebases async runtime effect fields without dropping newer stream state", () => { const base: UiState = { ...initialState("/tmp/demo"), diff --git a/docs/ARCHITECTURE.md b/docs/ARCHITECTURE.md index 24c81fdb2..d7a15676b 100644 --- a/docs/ARCHITECTURE.md +++ b/docs/ARCHITECTURE.md @@ -157,8 +157,22 @@ selection assertion, consults only bounded schema-validated local OpenCode catal configuration, and authentication state, and returns enum-only readiness evidence. Dynamic or remote authority is reported as unknown without plugin, network, provider, or mutation activity. Runtime source execution under -`agentcore/upstream` is not a supported boundary. See ADR-039. First-run setup -and Runtime integration remain 9W4E work. +`agentcore/upstream` is not a supported boundary. See ADR-039. + +Branch 9W4E adds the code-owned six-recipe setup catalog and one append-only +`runtime_model_setup_committed` transition. Runtime reconstructs the existing +ADR-035 configuration and ADR-036 registry only at the next process +construction; commits never hot-reload an active registry. Persisted setup, +explicit registry authority, and legacy Commander environment authority are +pairwise exclusive. + +Executor readiness invokes the exact configured OpenCode launch executable +with the fixed packaged arguments `nexusloop executor-readiness-v1`. Runtime +validates the exact projection/provider/model/binding echo and tri-state +evidence under fixed process, byte, timeout, concurrency, cancellation, and +shutdown limits. The observation remains evidence only. OpenTUI stages +independent Commander/Executor choices, previews exact hashes, requires +confirmation, and renders active versus pending-next-start state. See ADR-040. ```text NexusLoop domain control plane @@ -533,3 +547,5 @@ The target architecture is **not**: - `agentcore/adr/ADR-034-commander-model-provider-protocols.md` - `agentcore/adr/ADR-035-unified-model-profiles-and-role-bindings.md` - `agentcore/adr/ADR-036-runtime-model-profile-registry-and-role-readiness.md` +- `agentcore/adr/ADR-039-opencode-owned-executor-readiness-command.md` +- `agentcore/adr/ADR-040-first-run-model-setup-and-role-selection.md` diff --git a/docs/COMMANDER_PROVIDERS.md b/docs/COMMANDER_PROVIDERS.md index 0364d9f37..92d754519 100644 --- a/docs/COMMANDER_PROVIDERS.md +++ b/docs/COMMANDER_PROVIDERS.md @@ -6,8 +6,8 @@ provider settings do not authorize Commander. ## Internal Model-Profile Contract Branch 9W4B0 defines the internal configuration vocabulary, and 9W4B1 activates -validated snapshots in an immutable RuntimeServer registry. It is still not a -CLI, TUI, persistent configuration, or discovery feature. A model connection contains only bounded +validated snapshots in an immutable RuntimeServer registry. Branch 9W4E adds a +credential-free append-only setup surface over that same vocabulary. A model connection contains only bounded provider/account authority identifiers, a model profile selects an exact model, and independent role bindings select profiles for `commander` and `executor`. The Executor role means the primary tactical OpenCode model only, not small, @@ -39,6 +39,26 @@ the connection's provider kind. Its policy identity is part of the Executor projection hash. OpenCode discovery and authentication cannot populate this registry. +## First-run Recipes + +The setup catalog contains exactly three Commander and three primary Executor +recipes: Anthropic Claude Sonnet 4.5, Google Gemini 2.5 Flash, and OpenAI +GPT-4.1 mini. Commander uses the existing native Messages, Generative AI, and +Responses conformance entries. Executor uses the existing static provider +mapping. Either role may be explicitly unconfigured. + +Setup stores recipe identities and semantic hashes only. It never stores a +credential, environment name, endpoint, header, provider object, OpenCode auth +record, or catalog payload. Connector and credential readiness remain +role-owned. A setup commit is pending until restart and cannot mutate the +active registry. + +Executor readiness is observed by the packaged +`opencode nexusloop executor-readiness-v1` command using the exact executable +configured for Executor launch. Availability and credential connection remain +independent, and unknown is not ready. The command cannot select a model, +create provider mapping, authorize Commander, or change either projection. + The kernel accepts no endpoint, base-URL, header, package, plugin, provider option, environment-variable, OAuth, credential, or catalog authority field. An exact model ID is inert external identifier data and is never parsed as an diff --git a/docs/TUI_UX.md b/docs/TUI_UX.md index 52db881f1..b51fb18f2 100644 --- a/docs/TUI_UX.md +++ b/docs/TUI_UX.md @@ -29,6 +29,19 @@ When the user opens NexusLoop without an initialized project: The init flow should feel like entering a system that is already alive, not a detached setup script. +### Model Setup + +After project/spec onboarding, a missing durable setup automatically enters +the existing in-shell model setup view. Commander and primary Executor choices are independent and each +supports an explicit unconfigured state. The view separates selected, +connected, and ready status; it never asks for or displays credentials. + +OpenTUI requests a Runtime-owned preview, displays bounded role selections and +hashes, and requires a distinct confirmation bound to the displayed revision +and candidate hash. A successful commit is append-only and reports restart +required. Until restart, the active selection and pending-next-start selection +are rendered separately. Cached TUI state is display evidence only. + ## Resume Flow When existing runtime state is present: diff --git a/phases/9W4E/BASELINE.md b/phases/9W4E/BASELINE.md new file mode 100644 index 000000000..06c05f637 --- /dev/null +++ b/phases/9W4E/BASELINE.md @@ -0,0 +1,43 @@ +# 9W4E Clean Baseline + +- Base: `5c114669ebe8b6fd3ddd50a912a6ff821b882edd` +- Branch: `redesign/branch9w4e-first-run-model-setup-v2` +- Local and `origin/main` divergence before branch creation: `0 0` +- Worktree before branch creation: clean + +## Runtime + +The current-session clean-base Runtime run completed successfully. Its verbose +output was truncated by the terminal capture; the exact count is rerun and +recorded during final validation rather than copied from the obsolete branch. + +```text +0 fail +$ tsc --noEmit +``` + +## TUI + +```text +bun install v1.3.11 (af24e281) + +Checked 179 installs across 188 packages (no changes) [34.00ms] + +323 pass +0 fail +4149 expect() calls +Ran 323 tests across 7 files. [5.05s] +$ tsc --noEmit +``` + +## CLI Integration + +```text +....... [100%] +7 passed in 1.93s +``` + +The base worktree was clean, ancestry was `0 0`, PR #132 ancestry was absent, +and `origin/main` was exactly the required base. Frozen paths are the fourteen +entries in `FROZEN.lock`; this branch does not modify any frozen path, +`agentcore/upstream`, manifest, or lockfile. diff --git a/phases/9W4E/FORBIDDEN.md b/phases/9W4E/FORBIDDEN.md new file mode 100644 index 000000000..d9d1562b1 --- /dev/null +++ b/phases/9W4E/FORBIDDEN.md @@ -0,0 +1,23 @@ +# 9W4E Forbidden Scope + +- Dynamic Commander provider/model discovery or user-created conformance. +- Arbitrary endpoints, headers, packages, plugins, callbacks, fetch objects, or credentials. +- Credential storage, OpenCode `auth.json` mutation, or cross-role credential authority. +- Provider model-execution probes, fallback, failover, retry, streaming, or registry hot reload. +- OpenCode observations as selection, mapping, Commander conformance, or credential-sharing authority. +- Environment-, CLI-, or user-selected readiness executables, arguments, source modules, preloads, or assertions. +- Implicit role fallback or auxiliary OpenCode model selection. +- New Commander transports, hosted tools, retained provider state, or response retrieval. +- External MCP/research, proposal, governance, research mutation, or replay capability. +- Browser/dashboard setup, `agentcore/upstream`, dependencies, manifests, lockfiles, or frozen files. +- Changes to `resume_supported=false`, `provider_tool_loop_enabled=false`, or + `external_read_execution_enabled=false`. + +Setup state is credential-free selection authority only. Connector and +credential readiness remain separate existing runtime boundaries. + +The packaged OpenCode Executor command may inspect pinned OpenCode +availability and credential-source state only after NexusLoop has selected one +exact Executor projection. It may emit only bounded tri-state evidence for that +identity and may not expose a catalog, auth source, path, environment name, or +raw provider/config/plugin data. diff --git a/phases/9W4E/IMPLEMENTATION_BRIEF.md b/phases/9W4E/IMPLEMENTATION_BRIEF.md new file mode 100644 index 000000000..75bc35a1f --- /dev/null +++ b/phases/9W4E/IMPLEMENTATION_BRIEF.md @@ -0,0 +1,39 @@ +# 9W4E Implementation Brief + +Migration compatibility: when an explicit or legacy Commander registry is +already active, missing durable setup remains visible but does not replace that +mutually exclusive authority with first-run onboarding. Historical user-flow +fixtures establish an explicit unconfigured setup through the real TUI before +testing unrelated post-setup behavior; the dedicated first-run scenario keeps +the missing-setup path intact. + +1. Add versioned setup catalog/types/service/projection under + `agentcore/runtime/src/model-configuration/`. +2. Define three exact Commander recipes and three exact Executor recipes from + existing native protocol evidence; permit either role to be unconfigured. +3. Reconstruct candidates from recipe IDs, validate with ADR-035, project both + roles, and return bounded safe hashes/readiness fields. +4. Confirm via expected revision/candidate hash and EventStore + `appendIfLatestKind`; identical writes are idempotent and all others are stale or + conflicting. +5. Add a bounded Runtime resolver for the packaged 9W4E0 command. It accepts + only the exact immutable projection and invokes the exact process-adapter + executable used for launch with fixed `nexusloop executor-readiness-v1` + arguments. It is lifecycle-owned, capped at 2,048 request bytes, 4,096 + response bytes, five seconds, concurrency two, and zero retries. +6. Project persisted setup during launch and construct the immutable ADR-036 + registry plus Commander provider assertions. Reject explicit/legacy/persisted + source conflicts. Construct the production resolver from the process + adapter's immutable executable/cwd/environment authority; + package-internal resolver injection remains test-only and no production + environment input can select it. +7. Add canonical RuntimeServer/client commands for catalog, status, preview, + and confirm. All are pre-start safe; confirmation temporarily acquires or + reuses the run lock and is lifecycle-owned through durable settlement. +8. Replace the placeholder onboarding panel with keyboard-driven independent + role selection, preview, explicit confirmation, pending/current selection, + blocked readiness, and restart-required state. +9. Add red authority/runtime/TUI tests and one real headless CLI/OpenTUI E2E + proving commit, clean restart, active projections, exact Executor launch + model, and no secret durability. +10. Add ADR-040 and update implemented architecture/TUI/provider facts. diff --git a/phases/9W4E/MODEL_SETUP_SURFACE_AUDIT.md b/phases/9W4E/MODEL_SETUP_SURFACE_AUDIT.md new file mode 100644 index 000000000..71a8c83e4 --- /dev/null +++ b/phases/9W4E/MODEL_SETUP_SURFACE_AUDIT.md @@ -0,0 +1,69 @@ +# 9W4E Model Setup Surface Audit + +## Current Ingresses + +- Project initialization: TUI `ProjectUninitialized` -> init choice -> the + existing non-authoritative `initialize` shell action. +- Model configuration: ADR-035 pure validators and projections under + `agentcore/runtime/src/model-configuration/`. +- Runtime registry: `RuntimeServerOptions.modelProfileRuntimeRegistry`, deeply + revalidated and immutable by `ModelProfileRuntimeRegistry`. +- Commander authority: legacy validated + `NXL_COMMANDER_INVESTIGATION_*` configuration, static conformance, connector + registry, and provider readiness. +- Executor authority: static provider mapping and exact primary OpenCode + `--model provider/model` launch projection. +- Connector construction: `NXL_EXTERNAL_API_CONNECTORS_JSON`; connector URLs, + headers, credential references, and environment names remain outside model + configuration. +- Credential resolution: Commander through `ExternalApiRequestService`; + Executor through its independent readiness resolver. +- OpenCode observations: availability/connection evidence only; no selection + or Commander conformance authority. +- Runtime startup: `createRuntimeServerFromLaunchConfig` and + `readRuntimeServerLaunchOptionsFromEnv`. +- Runtime client: canonical `RuntimeServer.command`, `RuntimeServerClient`, and + `TuiRuntimeServerClient` routing. +- TUI: `providerOnboarding` is currently static display state; init selection + enters the main shell without durable model authority. +- Durability: `.nxl/events.jsonl` through `EventStore.appendIfLatest`; no model + setup event exists at the base. + +## 9W4E Seam + +9W4E adds one model-setup domain under runtime model configuration. The launch +factory projects committed setup before RuntimeServer construction. Runtime +commands expose catalog/status/preview/confirm only. OpenTUI displays and stages +those safe DTOs and sends exact hashes back for confirmation. + +No other ingress may override the selection. Existing explicit registry and +legacy environment inputs conflict with persisted authority. Executor launch +continues to inject exactly one primary `--model` argument and leaves every +auxiliary model untouched. + +## Pinned OpenCode Observation Boundary + +The pinned OpenCode `provider.list` route remains unsuitable as a Runtime DTO: +it exposes broad provider/model records and its connected list is not exact +credential evidence. 9W4E instead consumes the command packaged by 9W4E0: +`opencode nexusloop executor-readiness-v1`. + +The packaged command owns the audited OpenCode-side catalog, configuration, +plugin, and authentication semantics. It accepts one exact selected identity, +discards unrelated providers/models and raw state, and emits only the echoed +identity, independent availability/credential enums, and deterministic +evidence ID. It performs no provider request, discovery refresh, mutation, +fallback, or retry. Partial or dynamic authority remains unknown. + +Runtime never imports OpenCode provider, auth, config, plugin, or catalog +modules. It invokes the exact validated process-adapter command used for +Executor launch, replacing launch arguments with only the fixed readiness +subcommand. The same cwd and parent-plus-configured environment policy is used. +No readiness-specific executable, Bun config, preload, module path, dependency +path, or environment assertion exists. + +Runtime caps request/output at 2,048/4,096 bytes, timeout at five seconds, and +concurrency at two. It requires exactly one newline-terminated flat JSON +observation, rejects identity/version/evidence drift, performs zero retries, +and owns cancellation plus shutdown drain. Failures become unknown readiness; +they never alter selection. diff --git a/phases/9W4E/SCOPE_QUESTIONS.md b/phases/9W4E/SCOPE_QUESTIONS.md new file mode 100644 index 000000000..1151a3e43 --- /dev/null +++ b/phases/9W4E/SCOPE_QUESTIONS.md @@ -0,0 +1,118 @@ +# 9W4E Scope Questions + +All authority-affecting questions are resolved from the base source. + +## What is durable authority? + +One append-only `runtime_model_setup_committed` event in `.nxl/events.jsonl`. +The event contains exact built-in recipe references, semantic hashes, a +contiguous revision, and the prior hash. Startup reconstructs and fully +revalidates the ADR-035 configuration from those code-owned recipes. Projection +rejects malformed, duplicate, truncated, +unknown-version, or hash-invalid records. + +## How is startup activation possible without hot reload? + +Launch construction projects setup records before constructing RuntimeServer. +The resulting 9W4B1 registry is immutable for that process. A commit made by a +running server is pending for the next process start; the active registry is +never replaced in place. + +## How are connector and credential authority kept separate? + +Commander recipes contain only code-owned non-secret execution assertions: +transport, provider, connector ID, model, capability booleans, and hard limits. +Connector URL/header/credential source remains in +`ExternalApiConnectorRegistry` and `ExternalApiRequestService`. Setup events do +not contain those values. Missing connector credentials produce blocked +readiness without invalidating a safe selection. + +## Which Commander recipes are selectable? + +Only exact native-model contracts already proven on the base: + +- Anthropic Messages: `claude-sonnet-4-5-20250929`; +- Google Generative AI: `gemini-2.5-flash`; +- OpenAI Responses: `gpt-4.1-mini`. + +OpenAI-compatible Chat Completions remains a verified protocol family but has +no single code-owned endpoint/model onboarding recipe, so it is not offered. +The compatibility matrix cannot create setup authority. + +## Which Executor choices are selectable? + +The same three exact provider/model pairs are offered as bounded built-in +choices under the static Executor provider mapping. This branch does not add +catalog discovery. Executor means only the primary tactical `--model` seam. + +## What happens when both roles choose the same model? + +The candidate records the same provider/model intent through two explicit role +choices and two independent role bindings. Role-owned credential authority and +readiness remain separate; equal opaque credential-binding IDs, if introduced +by a future code-owned recipe, would not share a resolver or secret. + +## How do authority sources interact? + +Persisted setup, an explicitly injected runtime registry, and legacy Commander +environment authority are pairwise exclusive. Any simultaneous presence +blocks startup. They are never merged or prioritized. + +## Does the default fake/real client policy change? + +Yes, narrowly. An unset `NXL_RUNTIME_CLIENT` is now code-owned `auto`: the +legacy fake client remains available only before an approved spec so existing +spec onboarding can complete, while an approved project always constructs the +real RuntimeServer client. Explicit `fake` remains a fixture/operator override; +it is not production setup evidence. This prevents the normal model-setup flow +from reporting an in-memory fake commit. RuntimeServer still revalidates the +spec and owns every durable setup preview and confirmation. + +## Is a controlled restart automatic? + +No active workload is silently restarted. Commit reports +`restart_required=true`. First-run headless/user flow cleanly shuts down and a +subsequent invocation activates the committed setup. Later changes follow the +same staged-next-start rule. + +## How is production Executor readiness observed? + +The prior injected-only resolver was insufficient for production launch. 9W4E +consumes the process-isolated OpenCode-owned protocol packaged by 9W4E0. +Runtime sends the already-selected exact Executor projection to `opencode +nexusloop executor-readiness-v1`; that command consults pinned OpenCode +model/configuration/plugin/authentication semantics and emits only exact identity plus +`available|unavailable|unknown` and `connected|disconnected|unknown`. +Runtime validates identity and computes the credential-free evidence ID. + +The public `provider.list` route is not used because its `connected` list can +include configured providers without proving a usable credential source. The +child checks model availability and credential-source presence independently; +only built-in `anthropic`, `google`, and `openai` credential observations can +be definitive. Missing or incomplete catalog state remains `unknown`. + +The executable is exactly the validated process adapter `command` used for +Executor launch. Readiness arguments are fixed internally to `nexusloop +executor-readiness-v1`; launch arguments never become readiness arguments. +No readiness-specific executable, preload, source module, dependency path, or +environment assertion exists. Process injection remains package-internal test +machinery for protocol and lifecycle adversaries. + +The observer is not discovery authority. It cannot return choices, change a +profile, create a mapping, authorize Commander, or select fallback. Its process +is registered before the first await, bounded by timeout/output/concurrency, +cancelled and drained during shutdown, and freshly invoked by the launch gate. +An injected resolver remains a package-internal test seam and is never +constructed from production environment configuration. + +## How does a shared exact model remain usable by both roles? + +The Runtime capability registry derives role eligibility from the two immutable +selection projections, not from readiness evidence. If Commander and Executor +select the same exact provider kind and model, Runtime still creates separate +role-specific capability records. The Commander record retains only its static +Commander conformance, tool, and context authority; the Executor record retains +its independently conservative primary-tactical-model capability evidence. +Sharing provider/model selection intent never combines contexts, tools, limits, +or readiness. This does not change either projection hash, conformance, +mapping, credential authority, or readiness result. diff --git a/phases/9W4E/THREAT_MODEL.md b/phases/9W4E/THREAT_MODEL.md new file mode 100644 index 000000000..890d94eff --- /dev/null +++ b/phases/9W4E/THREAT_MODEL.md @@ -0,0 +1,26 @@ +# 9W4E Threat Model + +| Threat | Control | +| --- | --- | +| TUI state becomes authority | Server reconstructs candidates from exact recipe IDs and revalidates every preview/commit. | +| Matrix/catalog/user claim creates Commander support | Exact built-in setup catalog and conformance registry only. | +| Secret or endpoint enters setup | Exact credential-free event allowlist; configuration grammar; recursive forbidden-value tests. | +| Stale/concurrent confirmation wins | Expected revision, candidate hash, and EventStore expected-tail compare-and-append. | +| Partial write becomes active | Strict JSON/event/hash projection fails closed. | +| Active registry mutates | Startup-only construction; commit is pending-next-start. | +| Role fallback shares authority | Independent explicit bindings; unconfigured remains unconfigured. | +| Equal credential IDs share secrets | Commander and Executor readiness resolvers remain independent. | +| Legacy/explicit/persisted authority merges | Pairwise startup conflict checks. | +| Setup changes auxiliary OpenCode models | Existing exact single primary `--model` projection only. | +| Caller object methods/proxies execute | Setup parser delegates semantic state to hardened ADR-035 validators and reconstructs exact primitive recipe inputs. | +| Shutdown races append | Commit registers owned work before its first await, acquires or reuses the run lock, and shutdown drains it before `runtime_shutdown`. | +| OpenCode catalog/auth becomes selection authority | Observer receives an exact immutable projection and emits only matching tri-state evidence; unrelated entries are discarded in the child. | +| Raw OpenCode state crosses process boundary | Strict one-object protocol and output cap; no lists, auth records, paths, URLs, headers, plugin data, or raw errors are accepted. | +| Partial OpenCode state becomes definitive absence | Child/process/timeout/truncation failures map to `unknown`; only a complete observation may emit `unavailable` or `disconnected`. | +| Observer outlives RuntimeServer | Runtime owns each subprocess before its first await; shutdown terminates and drains all observations before `runtime_shutdown`. | +| Fixture or caller forges readiness identity | Runtime validates projection/provider/model/binding/version and recomputes the evidence ID from safe semantic results. | +| Alternate readiness executable or source preload replaces authority | Runtime derives readiness from the same validated process-adapter command used for launch and supplies only the fixed packaged subcommand; no readiness executable/args/preload input exists. | +| OpenCode state produces unsafe or partial certainty | The packaged command owns pinned OpenCode semantics and emits bounded tri-state evidence; Runtime maps malformed, timed-out, cancelled, mismatched, or failed observations to unknown. | + +GitHub/provider content and OpenCode observations remain evidence, not setup +authority. No provider request is performed during setup. diff --git a/phases/9W4E/VALIDATION.md b/phases/9W4E/VALIDATION.md new file mode 100644 index 000000000..77d2fcdea --- /dev/null +++ b/phases/9W4E/VALIDATION.md @@ -0,0 +1,79 @@ +# 9W4E Validation Record + +This file records the commands used for the local gate. Development failures +and invalid invocations remain separate from final release evidence; the final +handoff reports the fresh detached exact-head rerun. + +## Required commands + +```text +cd agentcore/upstream +npx -y bun@1.3.13 install --frozen-lockfile +npx -y bun@1.3.13 install --frozen-lockfile +npx -y bun@1.3.13 run --cwd packages/opencode build:nexusloop-readiness +npx -y bun@1.3.13 run --cwd packages/opencode test:nexusloop-readiness-package +cd packages/opencode +npx -y bun@1.3.13 test test/cli/nexusloop-executor-readiness.test.ts test/cli/nexusloop-executor-readiness-state.test.ts + +cd agentcore/runtime +bun install --frozen-lockfile +bun test src/model-configuration/model-setup.test.ts src/model-configuration/opencode-executor-readiness-resolver.test.ts src/runtime.test.ts +bun test +bun run typecheck + +cd agentcore/tui +bun install --frozen-lockfile +bun test test/keyboard.test.ts test/launch.test.ts test/runtime-client-factory.test.ts test/runtime-effects.test.ts test/runtime-state-merge.test.ts +bun test +bun run typecheck + +uv run pytest tests/integration/cli -q +uv run pytest tests/e2e_user/scenarios/test_model_setup_executor_readiness_tui.py -q +uv run pytest tests/e2e_user -q +git diff --check +``` + +## Superseded candidate historical gate + +The earlier implementation candidate `26de3cc0afb0254b3d81737e16761c96a9f4f9c1` +completed the historical suite, but its harness excluded the setup prerequisite +event prefix from six scenario assertions. That evidence is superseded by the +startup-gate repair and must not be used as final-head validation: + +```text +........................................................................ [ 80%] +.................. [100%] +90 passed in 2097.99s (0:34:57) +real 2114.57 +user 1395.14 +sys 489.74 +``` + +## Invalid and failed development runs + +- A detached-worktree command containing recursive removal was rejected before + execution; it changed nothing and is excluded. +- A simplistic pre-install `node_modules` check reported the tracked + `agentcore/server-fork/node_modules` fixture. A follow-up proved it was the + only such path and that no untracked dependency directory was inherited. +- An upstream test invocation from the workspace root ran no tests because the + root intentionally points test discovery at `do-not-run-tests-from-root`. + The supported package invocation subsequently passed all 50 readiness tests. +- `python -m py_compile` was invalid because `python` was unavailable. The + supported `uv run python -m py_compile` exposed one f-string syntax error; + the error was fixed and the identical supported command passed. +- Two early targeted E2E invocations were not polled to an exit code and are + excluded. The identical targeted command subsequently passed. +- The first Runtime full run found one stale authority-test expectation. The + assertion was corrected to the truthful non-process setup classification and + the identical suite subsequently passed. +- The first TUI full run found six established-project fixture assumptions. + Those fixtures were updated to commit explicit unconfigured setup rather + than weakening onboarding, and the identical suite subsequently passed. +- The first full historical run produced `47 failed, 43 passed` because + established scenarios lacked durable setup. The second produced + `6 failed, 84 passed` because six no-mutation assertions included the real + setup prerequisite lifecycle. The temporary event-prefix filtering that made + the next run pass concealed the first-run startup defect and has been removed. + The repaired scenarios inspect the complete journal, including the setup + commit, and assert that the prerequisite creates no runtime lifecycle events. diff --git a/phases/9W4E/checklist.md b/phases/9W4E/checklist.md new file mode 100644 index 000000000..91f8de504 --- /dev/null +++ b/phases/9W4E/checklist.md @@ -0,0 +1,12 @@ +# 9W4E Checklist + +- [x] 1. Verify exact base, clean worktree, frozen boundaries, and baseline verifiers. +- [x] 2. Audit setup, registry, packaged readiness, launch, client, TUI, durability, and PR #132 regressions. +- [x] 3. Resolve authority questions and freeze threat/implementation boundaries. +- [x] 4. Add red setup, packaged resolver, transaction, startup, TUI, and E2E tests. +- [x] 5. Implement immutable setup catalog and durable authority. +- [x] 6. Activate persisted setup only during next Runtime construction. +- [x] 7. Integrate packaged Executor readiness and exact launch gating. +- [x] 8. Add canonical runtime/client commands and OpenTUI role selection. +- [x] 9. Add ADR-040 and update implemented canonical documentation. +- [ ] 10. Run complete final validation, detached historical E2E, and mechanical guards for the startup-gate repair head. diff --git a/tests/e2e_user/sandbox.py b/tests/e2e_user/sandbox.py index a1daff203..3f5a06cd8 100644 --- a/tests/e2e_user/sandbox.py +++ b/tests/e2e_user/sandbox.py @@ -35,6 +35,8 @@ def __init__(self, repo_root: Path, recorded_dir: Path) -> None: self._create_venv() self.env = self._build_env() self.runner = CliRunner(self.nxl_executable, self.env, recorded_dir) + self._automatic_model_setup = True + self._model_setup_bootstrap_active = False @property def python(self) -> Path: @@ -100,14 +102,65 @@ def run_cli( timeout: int = 300, transcript_name: str | None = None, ) -> CommandResult: + target = cwd or self.root + self._bootstrap_unconfigured_model_setup(target, args) return self.runner.run( args, - cwd=cwd or self.root, + cwd=target, stdin=stdin, timeout=timeout, transcript_name=transcript_name, ) + def require_manual_model_setup(self) -> None: + """Keep a scenario on the genuine missing-setup first-run path.""" + self._automatic_model_setup = False + + def _bootstrap_unconfigured_model_setup(self, project: Path, args: Sequence[str]) -> None: + """Give pre-9W4E scenarios explicit setup through the real TUI boundary.""" + if not self._automatic_model_setup or self._model_setup_bootstrap_active or args: + return + if self.env.get("NXL_TUI_HEADLESS") != "1" or self.env.get("NXL_RUNTIME_CLIENT") != "real": + return + if any( + key.startswith("NXL_COMMANDER_INVESTIGATION_") and value not in (None, "0") + for key, value in self.env.items() + ): + return + spec_path = project / ".nxl" / "spec" / "current.json" + try: + spec = json.loads(spec_path.read_text(encoding="utf-8")) + except (FileNotFoundError, json.JSONDecodeError, OSError): + return + if not isinstance(spec, dict) or spec.get("status") != "approved": + return + events_path = project / ".nxl" / "events.jsonl" + if events_path.exists() and '"kind":"runtime_model_setup_committed"' in events_path.read_text(encoding="utf-8"): + return + prior_keys = self.env.get("NXL_TUI_KEYS") + prior_runner_keys = self.runner.env.get("NXL_TUI_KEYS") + keys = json.dumps([{"type": "submit"}] * 4) + self.env["NXL_TUI_KEYS"] = keys + self.runner.env["NXL_TUI_KEYS"] = keys + self._model_setup_bootstrap_active = True + try: + result = self.runner.run([], cwd=project, timeout=300) + finally: + self._model_setup_bootstrap_active = False + if prior_keys is None: + self.env.pop("NXL_TUI_KEYS", None) + else: + self.env["NXL_TUI_KEYS"] = prior_keys + if prior_runner_keys is None: + self.runner.env.pop("NXL_TUI_KEYS", None) + else: + self.runner.env["NXL_TUI_KEYS"] = prior_runner_keys + if result.exit_code != 0: + raise AssertionError(result.stdout + result.stderr) + events = events_path.read_text(encoding="utf-8") if events_path.exists() else "" + if '"kind":"runtime_model_setup_committed"' not in events: + raise AssertionError("real TUI did not commit the historical scenario model setup prerequisite") + def run_cli_background( self, args: Sequence[str], diff --git a/tests/e2e_user/scenarios/test_commander_executor_review_tui.py b/tests/e2e_user/scenarios/test_commander_executor_review_tui.py index 6c6876fb5..b5c1681e4 100644 --- a/tests/e2e_user/scenarios/test_commander_executor_review_tui.py +++ b/tests/e2e_user/scenarios/test_commander_executor_review_tui.py @@ -76,13 +76,9 @@ def test_user_runs_commander_executor_review_without_live_execution(sandbox) -> assert "executor-review-secret-abc123" not in result.stdout assert "abc123" not in result.stdout - events_path = project / ".nxl" / "events.jsonl" - events = [ - json.loads(line) - for line in events_path.read_text(encoding="utf-8").splitlines() - if line.strip() - ] if events_path.exists() else [] + events = sandbox.list_events(project) event_kinds = [event["kind"] for event in events] + assert event_kinds.count("runtime_model_setup_committed") == 1 forbidden = { "opencode_handoff_started", "opencode_handoff_created", @@ -119,7 +115,7 @@ def test_user_runs_commander_executor_review_without_live_execution(sandbox) -> assert forbidden.isdisjoint(event_kinds) assert "runtime_started" not in event_kinds assert not any(kind.startswith("commander_executor_review_") for kind in event_kinds) - assert all(kind == "runtime_shutdown" for kind in event_kinds) + assert event_kinds == ["runtime_model_setup_committed"] serialized_events = json.dumps(events) assert "executor-review-secret" not in serialized_events assert "executor-review-secret-abc123" not in serialized_events diff --git a/tests/e2e_user/scenarios/test_context_budget_registry_tui.py b/tests/e2e_user/scenarios/test_context_budget_registry_tui.py index 8d18e1958..ec2647103 100644 --- a/tests/e2e_user/scenarios/test_context_budget_registry_tui.py +++ b/tests/e2e_user/scenarios/test_context_budget_registry_tui.py @@ -90,17 +90,9 @@ def test_user_inspects_context_budget_registry_without_launching_or_mutating(san assert "vendor-secret" not in result.stdout assert "model-secret" not in result.stdout - events_path = project / ".nxl" / "events.jsonl" - events = ( - [ - json.loads(line) - for line in events_path.read_text(encoding="utf-8").splitlines() - if line.strip() - ] - if events_path.exists() - else [] - ) + events = sandbox.list_events(project) event_kinds = [event["kind"] for event in events] + assert event_kinds.count("runtime_model_setup_committed") == 1 assert event_kinds.count("runtime_started") == 0 assert event_kinds.count("opencode_session_planned") == 0 forbidden = { diff --git a/tests/e2e_user/scenarios/test_context_packet_compiler_tui.py b/tests/e2e_user/scenarios/test_context_packet_compiler_tui.py index 27d4368b4..6fdda457f 100644 --- a/tests/e2e_user/scenarios/test_context_packet_compiler_tui.py +++ b/tests/e2e_user/scenarios/test_context_packet_compiler_tui.py @@ -89,17 +89,9 @@ def test_user_previews_context_packet_compiler_without_launching_or_mutating(san assert "vendor-secret" not in result.stdout assert "model-secret" not in result.stdout - events_path = project / ".nxl" / "events.jsonl" - events = ( - [ - json.loads(line) - for line in events_path.read_text(encoding="utf-8").splitlines() - if line.strip() - ] - if events_path.exists() - else [] - ) + events = sandbox.list_events(project) event_kinds = [event["kind"] for event in events] + assert event_kinds.count("runtime_model_setup_committed") == 1 assert event_kinds.count("runtime_started") == 0 assert event_kinds.count("opencode_session_planned") == 0 forbidden = { diff --git a/tests/e2e_user/scenarios/test_executor_review_proposal_draft_tui.py b/tests/e2e_user/scenarios/test_executor_review_proposal_draft_tui.py index 94f549484..ee4faceb4 100644 --- a/tests/e2e_user/scenarios/test_executor_review_proposal_draft_tui.py +++ b/tests/e2e_user/scenarios/test_executor_review_proposal_draft_tui.py @@ -74,13 +74,9 @@ def test_user_previews_executor_review_proposal_drafts_without_mutation(sandbox) assert "executor-draft-secret-abc123" not in result.stdout assert "abc123" not in result.stdout - events_path = project / ".nxl" / "events.jsonl" - events = [ - json.loads(line) - for line in events_path.read_text(encoding="utf-8").splitlines() - if line.strip() - ] if events_path.exists() else [] + events = sandbox.list_events(project) event_kinds = [event["kind"] for event in events] + assert event_kinds.count("runtime_model_setup_committed") == 1 forbidden = { "opencode_handoff_started", "opencode_handoff_created", @@ -120,7 +116,7 @@ def test_user_previews_executor_review_proposal_drafts_without_mutation(sandbox) } assert forbidden.isdisjoint(event_kinds) assert "runtime_started" not in event_kinds - assert all(kind == "runtime_shutdown" for kind in event_kinds) + assert event_kinds == ["runtime_model_setup_committed"] serialized_events = json.dumps(events) assert "executor-draft-secret" not in serialized_events assert "executor-draft-secret-abc123" not in serialized_events diff --git a/tests/e2e_user/scenarios/test_model_setup_executor_readiness_tui.py b/tests/e2e_user/scenarios/test_model_setup_executor_readiness_tui.py new file mode 100644 index 000000000..3ea289d10 --- /dev/null +++ b/tests/e2e_user/scenarios/test_model_setup_executor_readiness_tui.py @@ -0,0 +1,302 @@ +from __future__ import annotations + +import json +import platform +import re +import shutil +import time +from pathlib import Path + +import pytest + + +@pytest.mark.phase_m4 +def test_first_run_model_setup_activates_exact_executor_through_production_observer(sandbox) -> None: + sandbox.require_manual_model_setup() + install = sandbox.install_from_current_repo() + assert install.exit_code == 0, install.stdout + install.stderr + + project = sandbox.make_empty_project_dir("model_setup_executor_readiness_project") + secret = "executor-observer-secret-never-published" + bun = shutil.which("bun") + assert bun is not None + inherited_authority = [ + "ANTHROPIC_API_KEY", + "GOOGLE_GENERATIVE_AI_API_KEY", + "OPENAI_API_KEY", + "OPENCODE_AUTH_CONTENT", + "OPENCODE_CONFIG", + "OPENCODE_CONFIG_CONTENT", + "OPENCODE_CONFIG_DIR", + "OPENCODE_DISABLE_MODELS_FETCH", + "OPENCODE_ENABLE_EXPERIMENTAL_MODELS", + "OPENCODE_MODELS_PATH", + "OPENCODE_PURE", + ] + for key in inherited_authority: + sandbox.env.pop(key, None) + sandbox.runner.env.pop(key, None) + sandbox.env.update({"NXL_TUI_HEADLESS": "1"}) + sandbox.runner.env.update({"NXL_TUI_HEADLESS": "1"}) + + def run_tui(keys: list[dict[str, str]], *, timeout: int = 300): + encoded = json.dumps(keys) + sandbox.env["NXL_TUI_KEYS"] = encoded + sandbox.runner.env["NXL_TUI_KEYS"] = encoded + result = sandbox.run_cli([], cwd=project, timeout=timeout) + assert result.exit_code == 0, result.stdout + result.stderr + return result + + approved = run_tui([ + {"type": "submit"}, + {"type": "insert", "text": "Verify first-run model setup and exact Executor launch authority"}, + {"type": "submit"}, + {"type": "insert", "text": "approve spec"}, + {"type": "submit"}, + {"type": "insert", "text": "/model-setup"}, + {"type": "submit"}, + ]) + assert "screen=main" in approved.stdout + assert "Restart NexusLoop after spec approval before opening model setup" in approved.stdout + current_spec_path = project / ".nxl" / "spec" / "current.json" + assert current_spec_path.exists(), approved.stdout + current_spec = json.loads(current_spec_path.read_text(encoding="utf-8")) + assert current_spec["status"] == "approved" + approved_events = (project / ".nxl" / "events.jsonl").read_text(encoding="utf-8") + assert "runtime_model_setup_committed" not in approved_events + + setup = run_tui([ + {"type": "select-next"}, + {"type": "submit"}, + {"type": "select-next"}, + {"type": "submit"}, + {"type": "submit"}, + {"type": "submit"}, + ]) + assert "stage=committed" in setup.stdout + assert "pending_restart=true" in setup.stdout + assert "Anthropic Claude Sonnet 4.5" in setup.stdout + + platform_name = "windows" if platform.system().lower() == "windows" else platform.system().lower() + architecture = {"x86_64": "x64", "amd64": "x64", "aarch64": "arm64"}.get( + platform.machine().lower(), platform.machine().lower() + ) + executable_name = "opencode.exe" if platform_name == "windows" else "opencode" + packaged_opencode = ( + sandbox.repo_root + / "agentcore" + / "upstream" + / "packages" + / "opencode" + / "dist" + / f"opencode-{platform_name}-{architecture}" + / "bin" + / executable_name + ) + assert packaged_opencode.is_file(), "build:nexusloop-readiness must produce the packaged OpenCode executable" + packaged_env = { + "NXL_OPENCODE_ADAPTER": "process", + "NXL_OPENCODE_COMMAND": str(packaged_opencode), + "NXL_OPENCODE_ARGS_JSON": "[]", + "OPENCODE_AUTH_CONTENT": json.dumps({"anthropic": {"type": "api", "key": secret}}), + "HOME": str(sandbox.root / "packaged-home"), + "XDG_CONFIG_HOME": str(sandbox.root / "packaged-config"), + "XDG_DATA_HOME": str(sandbox.root / "packaged-data"), + "XDG_STATE_HOME": str(sandbox.root / "packaged-state"), + "XDG_CACHE_HOME": str(sandbox.root / "packaged-cache"), + } + sandbox.env.update(packaged_env) + sandbox.runner.env.update(packaged_env) + packaged_status = run_tui([ + {"type": "submit"}, + {"type": "insert", "text": "/model-setup"}, + {"type": "submit"}, + ]) + assert "executor=Anthropic Claude Sonnet 4.5" in packaged_status.stdout + assert "credential connected" in packaged_status.stdout + assert "lifecycle ready" in packaged_status.stdout + assert secret not in packaged_status.stdout + + models_path = sandbox.root / "opencode-models.json" + models = { + "anthropic": { + "id": "anthropic", + "name": "Anthropic", + "env": ["ANTHROPIC_API_KEY"], + "models": { + "claude-sonnet-4-5-20250929": { + "id": "claude-sonnet-4-5-20250929", + "name": "Claude Sonnet 4.5", + "release_date": "2025-09-29", + "attachment": False, + "reasoning": True, + "temperature": True, + "tool_call": True, + "limit": {"context": 200000, "output": 8192}, + } + }, + } + } + models_path.write_text(json.dumps(models), encoding="utf-8") + auth_content = json.dumps({"anthropic": {"type": "api", "key": secret}}) + capture = sandbox.root / "opencode-launch-args.json" + opencode = sandbox.root / "opencode-fixture" + opencode.write_text( + f"""#!{bun} +import {{ createHash }} from "node:crypto"; +import {{ readFileSync, writeFileSync }} from "node:fs"; +const args = process.argv.slice(2); +if (args[0] === "nexusloop" && args[1] === "executor-readiness-v1") {{ + const input = JSON.parse(await new Response(Bun.stdin.stream()).text()); + let availability = "unknown"; + try {{ + const catalog = JSON.parse(readFileSync(process.env.OPENCODE_MODELS_PATH, "utf8")); + availability = catalog?.[input.provider_id]?.models?.[input.model_id] ? "available" : "unavailable"; + }} catch {{}} + const connected = Boolean(process.env.OPENCODE_AUTH_CONTENT); + const semantic = {{ + policy_version: "nexusloop_opencode_executor_readiness_policy_v1", + selection_projection_hash: input.selection_projection_hash, + provider_id: input.provider_id, + model_id: input.model_id, + credential_binding_id: input.credential_binding_id, + provider_availability_status: availability, + credential_connection_status: connected ? "connected" : "disconnected", + }}; + process.stdout.write(JSON.stringify({{ + observation_version: 1, + selection_projection_hash: input.selection_projection_hash, + provider_id: input.provider_id, + model_id: input.model_id, + credential_binding_id: input.credential_binding_id, + provider_availability_status: availability, + credential_connection_status: connected ? "connected" : "disconnected", + evidence_id: "opencode-readiness-v1-" + createHash("sha256").update(JSON.stringify(semantic)).digest("hex"), + }}) + "\\n"); + process.exit(0); +}} +if (args.includes("--model")) {{ + writeFileSync(process.env.NXL_E2E_LAUNCH_CAPTURE, JSON.stringify(args)); + process.exit(0); +}} +for await (const _chunk of Bun.stdin.stream()) {{}} +""", + encoding="utf-8", + ) + opencode.chmod(0o755) + + configured = { + "NXL_TUI_HEADLESS": "1", + "NXL_OPENCODE_ADAPTER": "process", + "NXL_OPENCODE_COMMAND": str(opencode), + "NXL_OPENCODE_ARGS_JSON": json.dumps(["--stdio"]), + "HOME": str(sandbox.root / "opencode-home"), + "XDG_CONFIG_HOME": str(sandbox.root / "opencode-config"), + "XDG_DATA_HOME": str(sandbox.root / "opencode-data"), + "OPENCODE_MODELS_PATH": str(models_path), + "OPENCODE_AUTH_CONTENT": auth_content, + "NXL_REAL_OPENCODE_LAUNCH": "1", + "NXL_E2E_LAUNCH_CAPTURE": str(capture), + } + sandbox.env.update(configured) + sandbox.runner.env.update(configured) + + def command_keys(commands: list[str]) -> list[dict[str, str]]: + keys: list[dict[str, str]] = [{"type": "submit"}] + for command in commands: + keys.extend([{"type": "insert", "text": command}, {"type": "submit"}]) + return keys + + planned = run_tui(command_keys([ + "/opencode-session-plan objective=verify exact persisted model setup", + "/opencode-sessions", + ])) + session_match = re.search(r"latest=(opencode_session_[A-Za-z0-9._-]+)", planned.stdout) + assert session_match, planned.stdout + session_id = session_match.group(1) + + packed = run_tui(command_keys([ + f"/opencode-session-instruction-pack-write session={session_id}", + "/opencode-session-instruction-packs", + ])) + pack_match = re.search(r"latest=(opencode_instruction_pack_[A-Za-z0-9._-]+)", packed.stdout) + if not pack_match: + pack_match = re.search(r"- (opencode_instruction_pack_[A-Za-z0-9._-]+) status=written", packed.stdout) + assert pack_match, packed.stdout + pack_id = pack_match.group(1) + + sandbox.env.pop("OPENCODE_AUTH_CONTENT") + sandbox.runner.env.pop("OPENCODE_AUTH_CONTENT") + disconnected = run_tui(command_keys([ + f"/opencode-launch-preview session={session_id} pack={pack_id}", + "/model-setup", + ])) + assert "commander=Anthropic Claude Sonnet 4.5" in disconnected.stdout + assert "executor=Anthropic Claude Sonnet 4.5" in disconnected.stdout + assert "pending_restart=false" in disconnected.stdout + assert "credential disconnected" in disconnected.stdout + assert "lifecycle unknown" in disconnected.stdout + assert not capture.exists() + + sandbox.env["OPENCODE_AUTH_CONTENT"] = auth_content + sandbox.runner.env["OPENCODE_AUTH_CONTENT"] = auth_content + models_path.write_text("{", encoding="utf-8") + malformed_launch = run_tui(command_keys([f"/opencode-launch-preview session={session_id} pack={pack_id}"])) + assert "Executor role readiness is not ready for the selected model profile" in malformed_launch.stdout + malformed = run_tui(command_keys(["/model-setup"])) + assert "credential connected" in malformed.stdout + assert not capture.exists() + + models_path.write_text(json.dumps(models), encoding="utf-8") + sandbox.env["OPENCODE_AUTH_CONTENT"] = auth_content + sandbox.runner.env["OPENCODE_AUTH_CONTENT"] = auth_content + conflicting_args = json.dumps([str(opencode), "--stdio", "--model=wrong/provider"]) + sandbox.env["NXL_OPENCODE_ARGS_JSON"] = conflicting_args + sandbox.runner.env["NXL_OPENCODE_ARGS_JSON"] = conflicting_args + conflicting = run_tui(command_keys([f"/opencode-launch-preview session={session_id} pack={pack_id}"])) + assert "preconfigured OpenCode primary model conflicts with runtime model-profile authority" in conflicting.stdout + assert not capture.exists() + + exact_args = json.dumps([str(opencode), "--stdio"]) + sandbox.env["NXL_OPENCODE_ARGS_JSON"] = exact_args + sandbox.runner.env["NXL_OPENCODE_ARGS_JSON"] = exact_args + launched = run_tui([ + {"type": "submit"}, + {"type": "insert", "text": "start runtime for exact model launch verification"}, + {"type": "submit"}, + {"type": "insert", "text": f"/opencode-launch-readiness session={session_id} pack={pack_id}"}, + {"type": "submit"}, + {"type": "insert", "text": f"/opencode-launch session={session_id} pack={pack_id}"}, + {"type": "submit"}, + {"type": "insert", "text": "/opencode-launches"}, + {"type": "submit"}, + ]) + deadline = time.monotonic() + 2 + while not capture.exists() and time.monotonic() < deadline: + time.sleep(0.02) + assert capture.exists(), launched.stdout + args = json.loads(capture.read_text(encoding="utf-8")) + assert args.count("--model") == 1 + model_index = args.index("--model") + assert args[model_index + 1] == "anthropic/claude-sonnet-4-5-20250929" + assert not any(value.startswith("--model=") or value.startswith("-m") for value in args) + assert not any(name in " ".join(args) for name in ["small_model", "title", "summary", "compaction", "subagent"]) + + events_text = (project / ".nxl" / "events.jsonl").read_text(encoding="utf-8") + durable = setup.stdout + approved.stdout + planned.stdout + packed.stdout + disconnected.stdout + malformed_launch.stdout + malformed.stdout + launched.stdout + events_text + for forbidden in [ + secret, + "NXL_E2E_LAUNCH_CAPTURE", + "OPENCODE_AUTH_CONTENT", + "OPENCODE_MODELS_PATH", + "authorization", + "api-key", + "auth.json", + ]: + assert forbidden not in durable + events = [json.loads(line) for line in events_text.splitlines() if line.strip()] + assert sum(event.get("kind") == "runtime_model_setup_committed" for event in events) == 1 + assert sum(event.get("kind") == "spec_approved" for event in events) == 1 + assert sum(event.get("kind") == "opencode_session_launch_started" for event in events) == 1 + assert sum(event.get("kind") == "opencode_session_launch_succeeded" for event in events) == 1 + assert events[-1]["kind"] == "runtime_shutdown" diff --git a/tests/e2e_user/scenarios/test_opencode_process_smoke_tui.py b/tests/e2e_user/scenarios/test_opencode_process_smoke_tui.py index 85bed65ec..0bb00efe4 100644 --- a/tests/e2e_user/scenarios/test_opencode_process_smoke_tui.py +++ b/tests/e2e_user/scenarios/test_opencode_process_smoke_tui.py @@ -77,13 +77,9 @@ def test_user_inspects_opencode_process_smoke_without_live_process(sandbox) -> N assert "opencode-smoke-secret" not in result.stdout assert "opencode-smoke-secret-abc123" not in result.stdout - events_path = project / ".nxl" / "events.jsonl" - events = [ - json.loads(line) - for line in events_path.read_text(encoding="utf-8").splitlines() - if line.strip() - ] + events = sandbox.list_events(project) event_kinds = [event["kind"] for event in events] + assert event_kinds.count("runtime_model_setup_committed") == 1 assert "runtime_started" not in event_kinds assert "opencode_process_smoke_blocked" in event_kinds assert "opencode_process_smoke_started" not in event_kinds diff --git a/tests/e2e_user/scenarios/test_research_memory_novelty_tui.py b/tests/e2e_user/scenarios/test_research_memory_novelty_tui.py index 1ab6e864b..499660847 100644 --- a/tests/e2e_user/scenarios/test_research_memory_novelty_tui.py +++ b/tests/e2e_user/scenarios/test_research_memory_novelty_tui.py @@ -85,17 +85,9 @@ def test_user_previews_research_memory_and_novelty_without_execution_or_mutation assert "research-memory-secret-abc123" not in result.stdout assert "token=abc123" not in result.stdout - events_path = project / ".nxl" / "events.jsonl" - events = ( - [ - json.loads(line) - for line in events_path.read_text(encoding="utf-8").splitlines() - if line.strip() - ] - if events_path.exists() - else [] - ) + events = sandbox.list_events(project) event_kinds = [event["kind"] for event in events] + assert event_kinds.count("runtime_model_setup_committed") == 1 assert event_kinds.count("runtime_started") == 0 forbidden = { "opencode_session_planned",