Skip to content

fix(cli,metadata-protocol): os migrate resume completes an interrupted recorded-by run, and os serve reports interrupted migration runs at boot - #21527

Merged
objectstack-fleet[bot] merged 6 commits into
mainfrom
claude/issue-21498-compose-migration-recovery
Oct 3, 2026
Merged

objectstack-fleet[bot] merged 6 commits into
mainfrom
claude/issue-21498-compose-migration-recovery

Conversation

@objectstack-fleet

Copy link
Copy Markdown
Contributor

Fixes #21498

Clause-②: no

What changed

MigrationRecoveryPlugin (@objectstack/runtime) owns the migration-plans registry and the ADR-0119 D2 boot scan. Nothing composed it. This PR composes it once in every stack that can host a plan. It also makes the plan's owner register the plan, because composing the registry alone does not make resume work.

  • The os migrate data boot (packages/cli/src/utils/data-migration-plugins.ts) composes new MigrationRecoveryPlugin() beside PlatformObjectsPlugin. This is the one composition point for recorded-by, resume, value-shapes, summary-nulls, files-to-references, meta --stored, audit-metadata-bodies and os storage orphans.
  • The os serve boot (packages/cli/src/commands/serve.ts, step 5c-bis) composes it beside the PlatformObjectsPlugin auto-registration. os start and os dev spawn serve, so they get it too. The block uses the same presence guard as the block above it, so a config that composes its own instance keeps that one. Measured with a host config that composes new MigrationRecoveryPlugin(): Plugin registered: com.objectstack.migration-recovery ×1, Plugin superseded ×0, Service 'migration-plans' registered ×1.
  • The plan's owner registers its plan (packages/metadata-protocol/src/plugin.ts). assembleMetadataProtocol is the code path that both ObjectQLPlugin and MetadataProtocolPlugin run. It now hands metadata.recorded-by-sentinel-to-null to migration-plans at kernel:ready, but only when a registry is composed.
    • Why kernel:ready. The registry is registered in the recovery plugin's init(), and the kernel runs that after the engine's init(). At kernel:ready, every init() has run, so a missing registry is really missing.
    • Why it lands before the scan. This hook is registered in Phase 1. The scan's hook is registered in the recovery plugin's start(), which is Phase 2. dispatchHookPropagating runs handlers in registration order, so the scan sees the plan.
    • Not gated on runPlatformMigrations. Registering a plan runs nothing.
  • Changesets: @objectstack/cli patch and @objectstack/metadata-protocol patch. packages/runtime gains only a test, so it ships nothing.
  • Docs that this change made false: the checklist item platform-core.interrupted-migration-boot-report (rev 2) and row D14 in docs/qa/platform-checklist/FOLLOW-UPS.md. Both said that no boot composes the plugin.

Why the registry alone was not enough (the ruling's mechanism, measured)

The ruling expected recorded-by's plans.register() and resume's plans.get() to "meet one registry". They cannot. recorded-by and resume are separate processes, and an in-memory registry does not survive a process boundary. The third ablation below measures exactly that state: the composition is present and the owner registration is removed. resume still refuses with the original sentence, and the serve scan calls the run unresumable.

The ruling's intent is that a run resumes once its owning package is loaded. That holds only if the owner registers the plan in every process. So the owner, @objectstack/metadata-protocol, now does. The refusal sentence in resume.ts is unchanged. It is now true exactly when it appears: no loaded package registers the plan.

The public door, before and after

Fixture: three sys_metadata_history rows hold recorded_by = 'system'. A real process ran the recorded-by plan under runMigrationJournal and was SIGKILLed inside chunk 0's transaction. That leaves run_started and chunk_started(0) in the journal, with no chunk committed.

reading before (25797a16e1) after
os migrate resume --json resumable: false resumable: true
os migrate resume --run RUN_ID --yes --json exit 1: "belongs to plan 'metadata.recorded-by-sentinel-to-null', which no loaded package registers … Load the package that owns this migration and re-run." exit 0, status: completed, chunksCommitted 1/1, 0 sentinel rows left, journal run_started, chunk_started, chunk_started, chunk_done, run_done
os serve over the run 0 lines mention it Interrupted migration run 'RUN_ID' (plan 'metadata.recorded-by-sentinel-to-null', …) … Resume with: os migrate resume --run RUN_ID, plus 1 interrupted migration run(s) found in sys_migration_journal. They are NOT resumed automatically …

The boot scan's cost and noise (A4)

  • Log level and silence on a clean database. os serve was booted over a journal whose only run had concluded, and over a fresh database. Neither boot printed a scan line. The fresh-database boot listed 3 boot-diagnostic warnings, and none of them came from the scan. The scan logs at warn, and only when it has an interrupted run to report or when its read fails.

  • Cost. The scan is one findInterruptedRuns call. I timed it on SQLite, median of 20: 0.26 ms with an empty journal, 0.65 ms with 1 concluded run, 11.2 ms with 51 concluded runs. It re-reads each run's events, so the cost grows linearly with journal history (see the acceptance notes). These are shared-box numbers.

  • In the one-shot os migrate boot the scan runs too. I measured two effects:

    • A data command booted over an interrupted run warns about that run on stderr first. os migrate resume therefore lists the run twice in human mode.
    • A read-only boot of a database with no journal table yet logs "Migration journal scan failed …". For example, os migrate value-shapes --json on a fresh project went from 5 to 6 WARN lines. Those commands already exit 1 on that database (see the findings).

    A registry-only constructor option would remove both effects. I wrote one, measured it, and withdrew it. It would have widened @objectstack/runtime's public surface (making Clause-② yes) for a cosmetic gain.

Pins

  • packages/cli/src/commands/migrate/resume.recorded-by.integration.test.ts (integration tier). Every boot runs in a hook.
    • The crash comes from a real SIGKILLed child process.
    • The real os migrate resume command lists the run resumable: true, and --run … --yes completes it. The test then reads the rows and the journal on its own connection.
    • The real os serve reports the run with Resume with: os migrate resume --run RUN_ID.
    • The control: os serve over a fresh database prints no scan line.
  • packages/runtime/src/migration-recovery-plugin.plan-owner.test.ts. A real ObjectKernel runs ObjectQLPlugin, PlatformObjectsPlugin and the recovery plugin over in-memory SQLite, with the journal rows a chunk-0 crash leaves. The scan reports the run as resumable, and the registry returns the plan after boot.

Reverse verification

Each mutation went through scripts/ablation-replace.mjs, which proves the anchor hit and that the restore matches the HEAD blob. Every pin turned red as expected.

ablation (at) what went red stayed green
serve 5c-bis kernel.use deleted (fe23beb97e, src via tsx, no build) serve scan pin: expected [] to have a length of 1 resume ×2, fresh control
data-boot plugins.push(new MigrationRecoveryPlugin()) deleted (fe23beb97e) list expected false to be true; act exit 1 with the original refusal serve ×2
owner registration replaced (e5cdb086b0, metadata-protocol rebuilt, ablation-dist-preflight marker present in 2 dist files) runtime pin 2/2; cli list, act and serve scan (… to contain 'Resume with: …') fresh control

For the third ablation, the restore leg rebuilt the package. ablation-dist-preflight --absent passed, the tree was clean against HEAD, and the registration call was back in dist/. ⚠️ My first attempt at this ablation was a no-op. It used a void 'marker' statement, which esbuild drops, so the marker never reached dist/. The preflight refused before any test ran. I re-anchored the marker on an assignment and re-ran.

Local verification

  • Gates. node scripts/pm/dispatch-gates.mjs --commands --repo objectstack-ai/objectstack derived 69 commands at e5cdb086b0, and all 69 exit 0 at e5cdb086b0. --ran reconciliation: "69 derived, 69 run, 0 NOT-MEASURED, 0 UNRUN (a DERIVED zero)".
    • check:driver-memory-census went red on the first version of the runtime pin, which bound the frozen @objectstack/driver-memory. The pin now uses SQLite.
    • check:dual-build-cjs-loads and check:i18n-coverage first answered PREREQUISITE NOT MET (no dist). I re-ran them after a full turbo build (72/72).
  • Lint. pnpm lint (the whole repo, eslint . --no-inline-config) exits 0 at e5cdb086b0.
  • Pins at e5cdb086b0.
    • runtime (2 files): 13/13 pass.
    • cli integration (the new pin plus the 5 files that boot the data composition: preview-read-only, meta.stored-flow-resolution, platform-migrations-arming, schema-migrate.one-shot-family, schema-migrate.teardown): 98 pass, 1 skipped (the live-PG cell).
  • Package suites at fe23beb97e. The only change since then is the runtime test's driver.
    • @objectstack/cli unit tier: 251 files, 3611 pass, 29 skipped. Two files first failed on a missing packages/cli/dist and passed after pnpm --filter @objectstack/cli build.
    • @objectstack/metadata-protocol full suite: 205 files pass, 3 skipped.
    • typecheck passes for cli and metadata-protocol. runtime's typecheck was re-run at e5cdb086b0 and passes.

Acceptance notes

  • Out of scope, reported for filing. resume now reaches the runner, which refuses two kinds of recorded-by run with PLAN_CHANGED:

    • a run that committed a chunk before it was interrupted (load() only selects rows that still hold the sentinel, so the chunk plan recomputed on resume hashes differently);
    • a run started with a non-default --chunk-size.

    Measured at the public door: 203 rows killed in chunk 1, and 3 rows at chunk size 2. In both cases the list says resumable: true and --run … --yes exits 1 Refused (PLAN_CHANGED). For these runs, re-running --apply remains the recovery. The pin above interrupts in chunk 0 and deliberately does not pin around this.

  • Out of scope, reported for filing. On a project whose database does not exist yet, os migrate resume --json, recorded-by --json and value-shapes --json already exit 1 with an opaque "The database refused to run this query for object …" at 25797a16e1. The scan now adds one warning to those runs.

  • The boot scan's cost grows linearly with journal history: one query per run ever recorded. 11 ms at 51 runs. This is an observation and has no carrier.

  • recorded-by's in-process plans.register(plan) now re-registers a plan the owner already registered. The last registration wins, and it carries the flag's chunk size. I left it in place because the ruling names it.


Generated by Claude Code

claude added 6 commits October 3, 2026 01:49
…he migrate data boot; owner registers its plan

Claude-Session: https://claude.ai/code/session_016GiHYRmLSNWTfbX9gVQkpz
Co-authored-by: Claude <noreply@anthropic.com>
…ugin beside the journal

Claude-Session: https://claude.ai/code/session_016GiHYRmLSNWTfbX9gVQkpz
Co-authored-by: Claude <noreply@anthropic.com>
… (no new runtime option); checklist item follows the composition

Claude-Session: https://claude.ai/code/session_016GiHYRmLSNWTfbX9gVQkpz
Co-authored-by: Claude <noreply@anthropic.com>
…lans composition

Claude-Session: https://claude.ai/code/session_016GiHYRmLSNWTfbX9gVQkpz
Co-authored-by: Claude <noreply@anthropic.com>
…e frozen memory driver

Claude-Session: https://claude.ai/code/session_016GiHYRmLSNWTfbX9gVQkpz
Co-authored-by: Claude <noreply@anthropic.com>
@github-actions github-actions Bot added size/l documentation Improvements or additions to documentation tests tooling labels Oct 3, 2026
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

📓 Docs Drift Check

This PR changes 2 package(s): @objectstack/cli, @objectstack/metadata-protocol, touching 2 documentable anchor(s).

1 hand-written doc(s) NAME something this change touched and may need an implementation-accuracy re-verification:

  • content/docs/kernel/services-checklist.mdx (via assembleMetadataProtocol (symbol, a top-level function))
What this run could not see
  • 1 anchor(s) matched too much of the corpus to be a work list: os serve (command, 30 pages)
  • 1 name(s) were too generic to anchor anything (single lowercase words)
  • the SDK route bridge reached 54 of 206 client-bound route-ledger rows — the other 152 have no registrar path: tail to select them, so pages documenting THEIR client methods cannot appear above, on this or any run. Of those 152: 0 are remediable by widening that discovery convention (an in-repo file declares the path; the convention did not scan it); 55 are structural — on a ledger where NOT ONE row is declared in-repo, so no discovery change reaches them at any price; 97 are undecided (no in-repo declaration, on a ledger that has other in-repo registrars — absence and an unreadable spelling are not distinguishable here). The rows themselves: node scripts/docs-audit/affected-docs.mjs --bridge-coverage
  • a page that states a rule by its inputs shares no identifier with the emitter that implements the rule, so an emitter-only diff cannot list it — not on this run and not on any run. Measured on fix(driver-sql): emit varchar(maxLength) for a text field a declared index keys on #11430: content/docs/protocol/objectql/types.mdx documents the text-family column mapping by the ObjectQL type names it maps FROM (text / textarea / html) while the diff changed createColumn; it went unlisted, and it was the page that diff falsified, in four places. No shared token exists to detect this on, so a rule your change carries has to be re-read by hand in the pages that restate it.
  • a key NAME is not a key, so the hand re-read the line above prescribes can land on the wrong schema. The same spelling is authorable on one governed type and a [REMOVED] tombstone on another for each of active, aria, joins, objects, template, tools and version (censused on [finding] tools is a key on BOTH AgentSchema (tombstoned, dead) and SkillSchema (live, cloud-attested), so a name-based search attributes skill examples to the agent key — it produced a false stop-the-line alarm on PR #19059 #19093 over the liveness ledger's governed types, top-level keys); nothing in a search result distinguishes the two, so a grep hit on a LIVE example reads as evidence about the DEAD key. Measured on fix(spec): the agent.tools liveness row says dead — it claimed live on a key the schema tombstoned #19059: content/docs/ai/agents.mdx was reported as contradicting the agent.tools tombstone over its tools: example at :161, which is inside the defineSkill({ block opened at :155 — the page was already correct. Settle ownership by PARSING the value against both schemas, never by the name: that literal PASSES SkillSchema, and as an AgentSchema it FAILS at tools with the tombstone prescription. ⛔ These names are not the whole class — a key retired through a .strict() guidance map leaves no tombstone in the walked shape and none of them here (tool.category, live as AIToolDefinition.category).

Coarse fallback — 35 page(s) merely mention a changed package (the pre-#9192 predicate, kept for the deliberately-wide backstop): node scripts/docs-audit/affected-docs.mjs --json ad7c3518983a1bb63fd4601954ac92d055124e42 → packageMentionDocs.

Which tree this was computed on

This run read content/docs from 37cce35f749bdedeeed7bc9ecb8c01ce1a8826d8 — the merge of head e5cdb086b0ff9b90965d408e67c9e69d434e8b07 into base ad7c3518983a1bb63fd4601954ac92d055124e42, which is what actions/checkout gives a pull_request run. Not the PR head.

A worktree cut from an older main holds a different content/docs, so re-deriving there can legitimately return a different list — that is a different tree, not a wrong row. To answer on the same tree:

# while this PR is open — GitHub drops the merge commit once it closes
git fetch origin 37cce35f749bdedeeed7bc9ecb8c01ce1a8826d8 && git checkout 37cce35f749bdedeeed7bc9ecb8c01ce1a8826d8
# afterwards, rebuild it from the two parents, which stay fetchable
git fetch origin ad7c3518983a1bb63fd4601954ac92d055124e42 e5cdb086b0ff9b90965d408e67c9e69d434e8b07 && git checkout -B drift-repro ad7c3518983a1bb63fd4601954ac92d055124e42 && git merge --no-ff e5cdb086b0ff9b90965d408e67c9e69d434e8b07

node scripts/docs-audit/affected-docs.mjs --json ad7c3518983a1bb63fd4601954ac92d055124e42

⚠️ That checkout carried uncommitted changes, so the commit above does not fully identify what was read.

Advisory only, and a precision-first one (#9192): a page is listed because it names a
symbol, wire route or SDK method this diff touched — not because it mentions a changed
package. Each row says which anchor put it there, so a wrong row is reportable rather than
merely annoying. To re-verify, run the docs-accuracy-audit workflow scoped to these files:
node scripts/docs-audit/affected-docs.mjs ad7c3518983a1bb63fd4601954ac92d055124e42 → pass the list as
args.docs, on the commit named under Which tree this was computed on.

@objectstack-fleet
objectstack-fleet Bot marked this pull request as ready for review October 3, 2026 03:35
@objectstack-fleet
objectstack-fleet Bot enabled auto-merge October 3, 2026 03:35
@objectstack-fleet
objectstack-fleet Bot added this pull request to the merge queue Oct 3, 2026
Merged via the queue into main with commit 550f4cc Oct 3, 2026
36 checks passed
@objectstack-fleet
objectstack-fleet Bot deleted the claude/issue-21498-compose-migration-recovery branch October 3, 2026 04:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation size/l tests tooling

Projects

None yet

2 participants