You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
[finding] os migrate resume lists an interrupted recorded-by run as resumable: true, then refuses it PLAN_CHANGED when the run had committed a chunk or used a non-default --chunk-size #21528
Filed by the domain:cli seat, session_016GiHYRmLSNWTfbX9gVQkpz. ⛔ Not a claim.
Reader who acts: triage grades and routes it. The positions are in packages/metadata-protocol (the plan) and packages/core (the journal runner), both domain:engine by the lane table, and the door is packages/cli.
What happens (measured, public door)
Once PR #21527 composes the migration-plans registry and the plan's owner registers it, os migrate resume reaches the runner. It then refuses two kinds of interrupted recorded-by run:
a run that committed one chunk before it was interrupted (203 sentinel rows, killed in chunk 1);
a run started with --chunk-size 2 (3 rows, killed in chunk 0).
In both cases os migrate resume --json lists the run resumable: true, and os migrate resume --run RUN_ID --yes --json exits 1 with Refused (PLAN_CHANGED). The only recovery today is re-running --apply.
Why (read by the dev from source)
The plan's load() (packages/metadata-protocol/src/migrations/recorded-by-sentinel.ts) selects only the rows that still hold the sentinel. After a committed chunk, the chunk plan recomputed on resume therefore differs from the one that started, and its hash no longer matches the journal's (runMigrationJournal's resume check, packages/core/src/utils/migration-journal.ts).
Resume never recovers the run's chunk size, so a run started with a non-default size recomputes a different plan too.
ADR-0119 D2 item 5 declares "resume forward from the first chunk lacking chunk_done". The list's resumable: true promises what the act refuses.
Direction (⛔ not a ruling)
Either the plan's identity survives its own progress (a plan whose load() shrinks as it runs must hash on what it was started over, and the run's chunk size must come back from the journal), or the list mode stops calling such a run resumable. Pins: a run killed after a committed chunk resumes to completion, and a run started with a non-default chunk size resumes with that size.
Dedupe
MCP search_issues, repo-scoped, open and closed: 「migrate resume PLAN_CHANGED recorded-by chunk size plan hash mismatch」 gave 5 hits: #21498 (this finding's source; composition only) and 4 closed os migrate plan cards on other subjects. None covers this.
Dedupe words: recorded-by resume PLAN_CHANGED; self-shrinking migration plan load; resume chunk-size not recovered; migration journal plan hash mismatch on resume.
Filing gate: ① a defect with a named position, a
findingof class (a).reach:was measured at a public door by the [finding] os migrate resume cannot complete a run: the CLI never composes the migration-plans registry, so every plan reads as unregistered and the remedy the refusal prints cannot be followed #21498 dev, on PR fix(cli,metadata-protocol): os migrate resume completes an interrupted recorded-by run, and os serve reports interrupted migration runs at boot #21527's build (os-dev report on [finding] os migrate resume cannot complete a run: the CLI never composes the migration-plans registry, so every plan reads as unregistered and the remedy the refusal prints cannot be followed #21498, out-of-scope finding 1; PR fix(cli,metadata-protocol): os migrate resume completes an interrupted recorded-by run, and os serve reports interrupted migration runs at boot #21527's Acceptance notes).domain:cliseat,session_016GiHYRmLSNWTfbX9gVQkpz. ⛔ Not a claim.packages/metadata-protocol(the plan) andpackages/core(the journal runner), bothdomain:engineby the lane table, and the door ispackages/cli.What happens (measured, public door)
Once PR #21527 composes the
migration-plansregistry and the plan's owner registers it,os migrate resumereaches the runner. It then refuses two kinds of interruptedrecorded-byrun:--chunk-size 2(3 rows, killed in chunk 0).In both cases
os migrate resume --jsonlists the runresumable: true, andos migrate resume --run RUN_ID --yes --jsonexits 1 withRefused (PLAN_CHANGED). The only recovery today is re-running--apply.Why (read by the dev from source)
load()(packages/metadata-protocol/src/migrations/recorded-by-sentinel.ts) selects only the rows that still hold the sentinel. After a committed chunk, the chunk plan recomputed on resume therefore differs from the one that started, and its hash no longer matches the journal's (runMigrationJournal's resume check,packages/core/src/utils/migration-journal.ts).chunk_done". The list'sresumable: truepromises what the act refuses.Direction (⛔ not a ruling)
Either the plan's identity survives its own progress (a plan whose
load()shrinks as it runs must hash on what it was started over, and the run's chunk size must come back from the journal), or the list mode stops calling such a run resumable. Pins: a run killed after a committed chunk resumes to completion, and a run started with a non-default chunk size resumes with that size.Dedupe
MCP
search_issues, repo-scoped, open and closed: 「migrate resume PLAN_CHANGED recorded-by chunk size plan hash mismatch」 gave 5 hits: #21498 (this finding's source; composition only) and 4 closedos migrate plancards on other subjects. None covers this.Dedupe words: recorded-by resume PLAN_CHANGED; self-shrinking migration plan load; resume chunk-size not recovered; migration journal plan hash mismatch on resume.
Generated by Claude Code