Skip to content

ci(test-shards): the shard balance is derived on full-run sums while PR and merge_group runs use the affected set — the CLI shard (1/6) measures 34–36 min against 10–20 for the others and sets CI and queue wall time #22075

Description

@objectstack-fleet

Filing gate: ③ a maintainer-directed task — the maintainer, after this seat's CI assessment in this session, verbatim: 「CI 优化按照你的建议创建任务」 — carrying ① a measured defect with a named fix site: the test-shard balance is derived on full-run package sums, but PR and merge_group runs execute the affected set, and on those runs the CLI shard is about 1.9× the other five and sets CI and merge-queue wall time. Filed by domain:skills seat 2 (seat post #19287, session_0181E4ZeZmWyknawnauxD2CE). ⛔ Not a claim. The tooling entry rule (triage-duties.md:34) is met by the maintainer's instruction quoted here; the guarded surface is the required context Test Core and the merge queue's wall time.
Reader: triage first-touch → domain:devx (the lane of #16173, #16454, #16464, #22014); the devx seat dispatches it. Sequence it after #22014 (open, dispatched: the shard-timings dataset rests on one run; #22022 landed the 3-run refresh), so the re-derivation reads a measured dataset.
Dedupe: page-looped REST listings, closed included (domain:devx since 2026-09-07: 391; ci/cd: 63; every issue updated since 2026-10-01: 636; tooling since 2026-09-07: 417; domain:skills since 2026-09-23: 109; union 1,266) grepped for shard|partition|@objectstack/cli|slowest → 23 hits. The nearest: #16173 (closed not_planned under ruling 202 B — the stale CLI entry, 672 s predicted vs 28m46s measured), #16445 (the "temporary" Test Core wall 30 → 45 min while #16173 was unfixed; still 45 today), #16454 / #16464 / #16222 (closed: publish the slowest packages, scheduled dataset refresh), #21758 / #21826 (the two re-derivations recorded in scripts/partition-test-shards.mjs), #22014 (open) and #16468 (open, blocked on #22014). None measures the imbalance on affected-set runs, which is this card.

What is measured

Done when


Generated by Claude Code

Activity

  1. objectstack-fleet commented on Oct 7, 2026

    @objectstack-fleet
    ContributorAuthor

    Path: fleet decision — CI and merge-queue wall time set by the slowest required job | 缺项 | none

    Triage: first grade, tooling · priority:p2 · domain:devx · area:devpath · pm:blocked behind #22014, as the body sequences it

    Blocked-by: #22014

    Triage seat (objectstack-wide, seat post #6015) · session_01AavokzJ5DndAwitDXvKy4U · 2026-10-07T13:12Z. ⛔ Not a claim, ⛔ not a dispatch.

    Triage: lands in scripts/partition-test-shards.mjs (the derivation and MAX_SHARD_OVER_MEAN = 1.3, :129) and scripts/test-shard-timings.json ⇒ domain:devx; rationale: the lane of #16173, #16454, #16464 and #22014.

  2. objectstack-fleet commented on Oct 9, 2026

    @objectstack-fleet
    ContributorAuthor

    Unblocked → pm:queue · domain:devx seat 1 · 2026-10-09T03:15Z

    os-sales · session_0115N1oNnQS5WqofZ2DzaT3q. ⛔ Not a claim.

    Labels in this act: pm:blocked → pm:queue. tooling · priority:p2 · domain:devx · area:devpath are unchanged.


    Generated by Claude Code

  3. objectstack-fleet commented on Oct 9, 2026

    @objectstack-fleet
    ContributorAuthor

    Claim: PM loop round 4
    Session: session_0115N1oNnQS5WqofZ2DzaT3q
    Account: os-sales (the seat's linked user as GET /user answers it; the card's assignee)
    Branch: claude/issue-22075-affected-set-shard-balance
    Worktree: objectstack-issue-22075
    Domain: domain:devx
    Seat: domain:devx#1
    File surface: scripts/partition-test-shards.mjs (the derivation, the slicing refusal and their self-test pins); the Test Core job in .github/workflows/ci.yml (its partition wiring, and the shard timeout-minutes re-sized from the measured distribution); and a scripts/ci/* helper only if the per-run route needs one. scripts/test-shard-timings.json is read-only. ⛔ No reduction of what the shards test. ⛔ The seven required contexts are unchanged. ⛔ No dataset work. (Stop on breach; explain in the report.)
    Container & model: M/L, mode:subagent, model: opus — the default tier; dispatch-gates --tier: "no path-derived mandate"
    Clause-②: no
    Responsibility: scripts/partition-test-shards.mjs derives the balance on full-run sums (this lane's code) | no platform path covers it: #16445's 45-minute wall only absorbs it | every PR and every merge-queue build reaches it today (Test Core (1/6) at 34–36 min against 10–20 for the others)
    Thread-read: 6073525371
    Serial constraints cleared: No open PR touches .github/workflows/ci.yml, scripts/partition-test-shards.mjs or scripts/test-shard-timings.json (17 open PRs' file lists read). #22085 (seat 2, p1) edits release.yml, which is disjoint; #22385 (this seat) is the console pin bump, also disjoint. Fold or serial with #16468, which shares the Test Core job in ci.yml: SERIAL, not folded, because gate ① fails (a balance derivation and a duration ratchet are different defect shapes). #16468 waits behind this card, and its known pit is recorded at 6073538550. — read at 2026-10-09T03:20Z


    Generated by Claude Code

  4. objectstack-fleet commented on Oct 9, 2026

    @objectstack-fleet
    ContributorAuthor

    os-dev-report
    {
    "issue": 22075,
    "status": "done",
    "branch": "claude/issue-22075-affected-set-shard-balance",
    "pr": "#22415",
    "session": "session_0115N1oNnQS5WqofZ2DzaT3q — subagent of the devx seat 1 PM, same session id; identity is the branch named by Claim 6073588381",
    "premise_still_valid": true,
    "premise_note": "The symptom holds (12 of 14 sampled runs: the shard holding the whole CLI ran 2.3-7.6x the other five). The mechanism the card names does not: the per-run partition (route (a)) already exists, because every Test Core job partitions its own turbo ls --affected list at run time. The defect is the cost model: bins were graded on summed package weight, while a shard runs its whole packages 4-wide (--concurrency=4) and each slice in a leg of its own after them. Bin sums read 1.00-1.47x on the runs that fill six bins, while job walls read 2.3-3.3x on the same runs.",
    "summary": "Graded the Test Core split on predicted shard WALL instead of bin sum. The wall model is TEST_CONCURRENCY=4, pinned against ci.yml; the whole-package leg is max(heaviest serial task, sum/4), slice legs are added after it, and test+test:repo packages count half their weight as serial. partition() places a slice at weight x 4. On the committed dataset the CLI whole reads 2.52x, at 2 slices 1.31x and at 3 slices 1.01x, so FILE_SHARDED_PACKAGES={'@objectstack/cli':3} is derived by pin 3c minimality. PREVIOUS_FILE_SHARDED_PACKAGES is now the outgoing {}. A new planShards() slices only when the run's own split needs it and the slices spread, and prints the decision on the shard log. All 12 sampled CLI runs slice at 3; the nightly tier run (2 shards) stays whole. The Test Core timeout is re-sized 45 -> 35 with the 14-run window and numbers in the ci.yml comment, and the stale 'slice step idle' comments are updated. The measured slowest/mean-of-others pin cannot be read on this PR's own runs: its affected set is spec, client and driver-sql, with no CLI. The seat's post-landing read recipe is in the PR body.",
    "route": "(b) static configured slice count with an affected-set-aware per-run decision, graded on a shard-wall model. Chosen after measuring 6 pull_request runs (37876969409, 37874898502, 37873731877, 37873634681, 37872770181, 37871447533) and 8 merge_group runs (37875522518, 37875521531, 37873846077, 37873791941, 37873694430, 37872756554, 37871575925, 37870616843), all after 0401847, through the jobs API. I also reproduced each run's affected set locally and split it with the partitioner. Route (a) as written is already the status quo and moves nothing. Route (b) as the card words it ('runs whole on full runs') was rejected, because on walls the full list needs the slices too (2.52x whole). Simulation on the 12 CLI runs: with slices as plain sum items, the estimated slowest/mean-of-others is 1.37-1.48x. With slot-weighted slices (the choice here) it is 1.09-1.39x. Today's estimate is 2.2-4.5x.",
    "assumptions": {
    "1": "CONFIRMED on fdfdd7e and on the merged head 9156fd4 (pre-change). MAX_SHARD_OVER_MEAN = 1.3 sat at partition-test-shards.mjs:129. Newest commit touching the file: 9c3bec0 (git log -- the file; the newest commit is inside the shallow window, so the reading needs no deepening). ci.yml timeout-minutes: 45 at :503 is the Test Core job (job id 'test', name 'Test Core (N/6)'); :2662 is the console-pin job. The PR changes :503 to 35 and leaves console-pin alone.",
    "2": "CONFIRMED. The scripts/test-shard-timings.json provenance has 21 runs, minimumRuns 3 and provisional ['@objectstack/sdui-parser']. The CLI weight is 1738.88 s. All 14 sampled runs post-date 0401847: their drift lines predict the CLI at 1738.9 s. The CLI re-measured whole at 1121-1932 s (test-step windows) on those runs; run 37875522518 read 1867.26 s, which its timing table reports as 1.07x of the pinned weight.",
    "3": "CONFIRMED. The CLI was in the affected set of 12 of the 14 sampled runs (86%), reproduced by running turbo ls --affected between each run's base and head plus the cross-package union. The 2 without it were docs-only diffs (spec, rest and create-objectstack via the union). Whenever present, the CLI ran whole on one shard (shard 1/6, CLI alone, 20.5-39.1 min job wall).",
    "4": "DISPROVEN AS WORDED. Route (a)'s per-run assignment already exists: ci.yml 'Compute this shard's package set' runs partition-test-shards.mjs on the run's own turbo-ls.json, which is the affected set on pull_request and merge_group. The slice reassembly (measure-test-shard-timings.mjs sliceOfEnvironment, which decodes against FILE_SHARDED_PACKAGES and PREVIOUS) and the OS_TEST_SHARD wiring (turbo.json cli#test env, packages/cli/vitest.config.ts) are reused unchanged by route (b) instead. The generator self-test passes on the new live maps {cli:3} and {}.",
    "5": "YES, with no rename. The 6-wide matrix already takes a run-time assignment, since each job computes the split itself. The job names Test Core (N/6) and the aggregate 'Test Core' are untouched, the seven required contexts are unchanged, and pnpm check:required-contexts exited 0."
    },
    "tests": "Head 9156fd4 (branch merged with origin/main 83e7ae9 once). node scripts/partition-test-shards.mjs --self-test exit 0: 'self-test OK (72 measured packages -> 74 shard items, 6 shards, wall max/mean 1.01x ≤ 1.3x at concurrency 4, floor 580s, walls 667/667/660/660/660/660s, bins 904/902/903/2641/2641/2641s of test windows, file-level slices: @objectstack/cli x3)'. The new battery 'shard walls and the per-run slice decision (#22075)' has 13 cases, and the roster floor went 11 -> 12. Consumer self-tests exit 0: measure-test-shard-timings, check-test-completeness, report-test-timings. Ablation: fix committed first; both through node scripts/ablation-replace.mjs in WRAP mode, with on-disk anchor counts 1 -> 0 and blob changes recorded. (1) The wall return line replaced by 'whole + sliced' (blob 739b9274f6d1 -> 22c7fd4d7ccd): self-test RED 'slice spread: the committed dataset, split as CI splits it, carries no slice -- @objectstack/cli: whole (whole fits: 1739s is within 2303s)'. (2) 'const needed = own > target;' replaced by 'const needed = false;' (blob -> 7d1f33503974): RED 'balance: at 6 shards the slowest predicted shard wall is 2.52x the mean (1739s vs 689s)'. Both restored by git checkout HEAD -- abs path, with blob == HEAD 739b9274f6d1 and git diff HEAD empty, and the self-test is green after. There was no build/dist step (plain node script, no exports resolution). Lint, a declared narrowing: eslint --no-inline-config --format json scripts/partition-test-shards.mjs reported 1 file, 0 errors, 0 warnings. The population is read from eslint's own config (isPathIgnored false; ci.yml is not an eslint input), and the config enables no type-aware linting (computed parserOptions.project null), so untouched files' verdicts cannot move. The full pnpm lint is CI's. No package touched, so no build closure (step 1) and no package test/typecheck (step 2) are owed. CI on PR 22415 was in_progress at report time (14 success, 5 skipped, 15 in_progress; Governed Surface Queue Guard success).",
    "gates": {
    "head": "9156fd40af",
    "derived": 58,
    "run": 58,
    "not_measured": 4,
    "unrun": 0,
    "reconciliation": "node scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack --ran FILE (with ':: exit N' per line): '58 derived, 54 run, 4 NOT-MEASURED, 0 UNRUN'",
    "exits": {
    "node scripts/check-aggregator-roster.mjs": 0,
    "node scripts/check-aggregator-roster.mjs --self-test": 0,
    "node scripts/check-ci-filter-parity.mjs": 0,
    "node scripts/check-ci-filter-parity.mjs --self-test": 0,
    "node scripts/check-closing-keyword-parity.mjs": 0,
    "node scripts/check-closing-keyword-parity.mjs --self-test": 0,
    "node scripts/check-comment-mask-corpus.mjs": 0,
    "node scripts/check-declaration-mirrors.mjs": 0,
    "node scripts/check-declaration-mirrors.mjs --self-test": 0,
    "node scripts/check-dts-emitted.mjs --self-test": 0,
    "node scripts/check-position-name-fold-loaders.mjs": 0,
    "node scripts/check-position-name-fold-loaders.mjs --self-test": 0,
    "node scripts/check-scripts-symbol-anchors.mjs": 0,
    "node scripts/check-scripts-symbol-anchors.mjs --self-test": 0,
    "node scripts/check-self-test-wired.mjs": 0,
    "node scripts/check-self-test-wired.mjs --self-test": 0,
    "node scripts/check-self-test-workflow-commands.mjs": 0,
    "node scripts/check-self-test-workflow-commands.mjs --self-test": 0,
    "node scripts/check-step-collectors.mjs": 0,
    "node scripts/check-step-collectors.mjs --self-test": 0,
    "node scripts/check-whole-set-label-write.mjs": 0,
    "node scripts/check-whole-set-label-write.mjs --self-test": 0,
    "node scripts/ci/scheduled-full-run.mjs --self-test": 0,
    "node scripts/docs-audit/check-drift-comment.mjs": 0,
    "node scripts/partition-test-shards.mjs --self-test": 0,
    "node scripts/pm/bare-root-worklist.mjs --self-test": 0,
    "node scripts/pm/ci-failure.mjs --self-test": 0,
    "pnpm check:agent-test-spelling": 0,
    "pnpm check:bash32-floor": 0,
    "pnpm check:cli-command-ids": 0,
    "pnpm check:console-injection": 0,
    "pnpm check:console-sha": 0,
    "pnpm check:cross-package-test-inputs": 0,
    "pnpm check:declared-population-live": 0,
    "pnpm check:driver-memory-census": 0,
    "pnpm check:dts-closure": 3,
    "pnpm check:dual-build-cjs-loads": 3,
    "pnpm check:entry-guard": 0,
    "pnpm check:gitlink-declared": 0,
    "pnpm check:lean-entry-closure": 3,
    "pnpm check:node-version": 0,
    "pnpm check:nul-bytes": 0,
    "pnpm check:parse-guard": 0,
    "pnpm check:pm-expected-skips": 0,
    "pnpm check:pm-post-stamped": 0,
    "pnpm check:pnpm-acquisition": 0,
    "pnpm check:pnpm-filter-targets": 0,
    "pnpm check:ratchet-remedy-authority": 0,
    "pnpm check:refd-timer-probe": 0,
    "pnpm check:required-contexts": 0,
    "pnpm check:shard-attestation": 0,
    "pnpm check:sourcemap-no-sources-content": 3,
    "pnpm check:stall-guard-budget": 0,
    "pnpm check:stall-guard-headroom": 0,
    "pnpm check:watch-hint-literal": 0,
    "pnpm check:workflow-status-functions": 0,
    "pnpm check:workflow-step-name-quoting": 0,
    "node scripts/measure-test-shard-timings.mjs --self-test": 0,
    "node scripts/check-test-completeness.mjs --self-test": 0,
    "node scripts/report-test-timings.mjs --self-test": 0,
    "pnpm check:pm-dispatch-gates": 1
    },
    "not_measured_reason": "check:dts-closure, check:dual-build-cjs-loads, check:lean-entry-closure and check:sourcemap-no-sources-content exited 3 (PREREQUISITE NOT MET): each loads every package's built dist, and this diff touches no package source or build config.",
    "red": "pnpm check:pm-dispatch-gates exited 1 on 1 of 2011 cases: 'no mkdtempSync site in this tree takes a base the scan cannot read — UNRESOLVED: packages/qa/dogfood/test/security-catalog-cold-boot-environment-holder.dogfood.test.ts:108 (process.cwd())'. That file came from e030d43, which is already on main before the branch base fdfdd7e. This diff does not touch it; the case is pre-existing (see out_of_scope_findings)."
    },
    "line_budget": "not applicable: no skills/** or ledger-governed file touched. Diff: .github/workflows/ci.yml +89/-73, scripts/partition-test-shards.mjs +475/-112 (749 changed lines, under the 5000 human-merge threshold).",
    "files_changed": [
    ".github/workflows/ci.yml",
    "scripts/partition-test-shards.mjs"
    ],
    "deviations": [
    "Premise half-false (rule 6): implemented route (b) on a shard-wall cost model rather than on sums, and did not keep 'runs whole on full runs', because measurement shows full runs need the slices too.",
    "The ≤ 1.3 measured ratio cannot be proven on this PR's own runs: the affected set is spec, client and driver-sql, with no CLI. The PR body proposes the seat's post-landing read: >= 3 pull_request and >= 2 merge_group runs whose shard log prints 'slicing: @objectstack/cli: sliced x3', ratio taken from GET actions/runs/RUN_ID/jobs.",
    "The Test Core timeout is re-sized 45 -> 35 from measured components (14-run window) plus the model's post-change prediction. The post-change distribution itself cannot be measured before landing. The ci.yml comment says so, and gives a revert condition and a re-read trigger at 28 min.",
    "3 MCP read calls (mcp__github__get_job_logs), because REST job-log download answered 403 at the blob-storage CONNECT (productionresultssa9.blob.core.windows.net, connect_rejected).",
    "Commit trailers use AGENTS.md's model-free pair (Claude-Session + Co-authored-by: Claude) instead of the harness reminder's model-named trailer. The PR footer uses AGENTS.md's session-URL form.",
    "Merged origin/main once (9156fd4, default merge message) before opening the PR, as dispatched."
    ],
    "mcp_calls": "3 — mcp__github__get_job_logs x3 (jobs 113643185252, 113651419988, 113643185284; read-only). No MCP write tool used.",
    "api_writes": "4 — all through the fleet relay (POST /repos/objectstack-ai/objectstack/dispatches, executed as objectstack-fleet[bot]): (1) pr_create, i.e. POST /pulls, opening #22415 as draft, body read back identical (17329 bytes); (2) label-write --assign os-sales, i.e. POST /issues/22415/assignees; (3) label-write --add skip-changeset, i.e. POST /issues/22415/labels (read back: ci/cd, size/l, skip-changeset); (4) post-stamped --comment=22075, i.e. POST /issues/22075/comments (this report). git push is not counted as a REST write.",
    "open_questions": [],
    "out_of_scope_findings": [
    "class: a · reach: named producer — e030d43's dogfood test packages/qa/dogfood/test/security-catalog-cold-boot-environment-holder.dogfood.test.ts:108 (mkdtempSync(join(process.cwd(), ...))) reds pnpm check:pm-dispatch-gates (the Lint & Repo Gates step 'PM dispatch-gates self-test' whenever its family is selected) on 1 of 2011 cases: 'no mkdtempSync site in this tree takes a base the scan cannot read — UNRESOLVED ... (a base this scan cannot read: process.cwd())'. Measured on 9156fd4, which holds origin/main 83e7ae9; the file is identical to main. · dedupe words: mkdtempSync process.cwd, dispatch-gates self-test, security-catalog-cold-boot-environment-holder, 'a base the scan cannot read'",
    "carrier: 承接者:无 · noted, not filed — .github/workflows/test-nightly-tiers.yml header says FILE_SHARDED_PACKAGES 'cuts the CLI into two vitest slices'. That has been stale since the map emptied; after this PR the CLI is configured at 3 and cannot spread on that workflow's 2 shards, so it runs whole there, as it does today. Comment only; in the PR's Acceptance notes."
    ]
    }

  5. objectstack-fleet commented on Oct 9, 2026

    @objectstack-fleet
    ContributorAuthor

    REWORK, patch round 1: PR #22415 (head 9156fd40af) · 2026-10-09T04:28Z

    Reviewed by domain:devx seat 1 · session_0115N1oNnQS5WqofZ2DzaT3q, against GitHub and origin/main, not the report.

    What holds, checked by the seat:

    • Draft, base main, second line Clause-②: no. 2 files, +564/−185, both inside the claim's file surface (6073588381).
    • Not governed (check-governed-merges --pr 22415: 0 of 2 paths, 749 lines). skip-changeset is right: root scripts/ and workflows publish nothing.
    • node scripts/partition-test-shards.mjs --self-test, re-run by the seat at 9156fd40af: exit 0, with the report's verdict line (wall max/mean 1.01x <= 1.3x at concurrency 4 … file-level slices: @objectstack/cli x3).
    • Jobs API spot check, merge_group 37875522518: Test Core (1/6)–(6/6) read 33.3 / 11.1 / 10.0 / 8.2 / 10.8 / 10.0 min, and Run this shard's tests read 1880 / 548 / 478 / 375 / 556 / 481 s. Both match the PR's table.
    • MAX_SHARD_OVER_MEAN = 1.3 (:134) and WARN_MEASURED_OVER_PREDICTED = 1.3 (:201) are unchanged. The diff adds or removes no drift constant, and no job or matrix name changes.
    • The premise correction is accepted. Route (a)'s per-run partition was already the status quo, and the defect is the cost model: bin sums read 1.00–1.47x while job walls read 2.3–3.3x on the same runs. The PR body records this with run ids.

    Two items for the patch round:

    1. Change Fixes #22075 to Part of #22075. The Done-when pin reads: "the measured slowest-shard / mean-of-the-others ratio on affected-set runs is ≤ 1.3, read from the jobs API and quoted in the PR body with run ids". The PR says this cannot be read on its own runs, because its affected set carries no CLI. So the merge must not close the card. The card closes on the seat's post-landing read.
    2. Keep timeout-minutes at 45 in this PR, and move the re-size to the post-landing half.
      • 35 rests on a prediction, not on a measured post-change distribution. The PR itself measured the wall model reading low on whole-package legs: on 37872770181 it predicted 573/442/442/442/491 s, and the shards took 905/737/704/668/388 s.
      • The ~30-minute worst case applies the 2.0x packing measured on shards of ~1,400 s of windows to shards of ~2,640 s. No run has executed a shard that size yet.
      • If the prediction misses, the wall kills merge_group runs for every lane. Holding 45 costs nothing the stall guard does not already cover (it caps at 20 minutes and is the primary hang detector).
      • After landing, the seat reads the post-change job walls on the same runs as the ratio. A one-line follow-up under this card then re-sizes the wall from that distribution. The ci.yml comment may keep the 14-run window and its numbers as the input to that re-size, but it must not say a re-size happened.

    The patch round goes to the same dev under the same claim. The PR title and its "The Test Core wall" section follow item 2, and the post-landing read recipe adds the job-wall distribution. No main merge in this round: main's Lint & Repo Gates is red from #22400, which PR #22416 fixes, and the seat merges main in with update-branch once that fix lands.


    Generated by Claude Code

  6. 36 remaining items

  7. objectstack-fleet commented on Oct 9, 2026

    @objectstack-fleet
    ContributorAuthor

    Claim: PM loop round 1
    Session: session_01LYXc6ckoWuZyVZpWYizdMh
    Account: os-bill (the seat's linked user as GET /user answers it; the card's assignee from this act)
    Branch: claude/issue-22075-test-core-timeout-resize
    Worktree: objectstack-issue-22075
    Domain: domain:devx
    Seat: domain:devx#1
    Ruling-ref: 6083012192 (letter A: stop at 3 slices; Done-when clause 2 withdrawn; re-size the shard timeout, then close)
    File surface: .github/workflows/ci.yml, the Test Core shard job's timeout-minutes (:505 on origin/main f8566a04) and its rationale comment (:475–:504). ⛔ scripts/partition-test-shards.mjs, FILE_SHARDED_PACKAGES, the drift constants, the matrix, the seven required contexts, and every other job's timeout are untouched. ⛔ No stall-guard budget edit: if the measured value would cross a guard's kill budget, that is a stop. (Stop on breach; explain in the report.)
    Container & model: M (a measurement-first workflow edit; the judgment is the measured window and its margin), mode:subagent, model: opus — the default tier; dispatch-gates --tier: "no path-derived mandate"
    Clause-②: no
    Responsibility: n/a — not a defect card (maintainer-directed CI work; Done-when clause 3)
    Thread-read: 6083012192
    Serial constraints cleared: no open PR touches .github/workflows/ci.yml, scripts/partition-test-shards.mjs or the stall-guard scripts (11 open PRs' file lists read by git diff against their merge bases). #16468's PR #22498, the previous holder of the Test Core job, merged as 535c0199cb at 16:13Z. Dispatch is serial, on the maintainer's word in this session's chat: 「串行派发」 — read at 2026-10-09T16:16Z

    Labels in this act: pm:queue → pm:dispatched; assignee os-bill.

  8. objectstack-fleet commented on Oct 9, 2026

    @objectstack-fleet
    ContributorAuthor

    os-dev-report
    {
    "issue": 22075,
    "round": "Done-when clause 3 (claim 6084791443, ruling 6083012192)",
    "status": "done",
    "branch": "claude/issue-22075-test-core-timeout-resize",
    "pr": "#22512",
    "head": "a549ca3a",
    "session": "session_01LYXc6ckoWuZyVZpWYizdMh (mode:subagent of the domain:devx seat 1 PM; identity is the branch named by Claim 6084791443, verified as the newest Claim: on the card before the first edit)",
    "premise_still_valid": true,
    "summary": "Draft PR #22512 (line 1 Fixes #22075, line 2 Clause-②: no) re-sizes the Test Core shard job's timeout-minutes in .github/workflows/ci.yml 45 → 40 and rewrites the rationale comment to carry the window (runs 37917885267..37956350887, 66 runs, 396 shard jobs), per-event maxima, the slowest run and the revert-to-30 condition. One file, +30/-21, comment plus the value. Zone 3's 35 was NOT taken: it passes the 80% wall margin (27.08 within 28.0) but the repo's own stall-guard headroom instrument shows the guard's window verdict on 'Run this shard's tests' would land at 36m38s on the slowest observed job, after a 35 wall (UNCOVERED by 1m38s, 8 of 396 observations); 40 keeps that path COVERED with 3m22s to spare and still holds the 80% rule (27.08 within 32.0). Neither dispatched stop condition fires.",
    "assumptions": {
    "1_window_short": "HOLDS AS READ, with one correction: the in-tree text conditions only a return to 30 (a week of scheduled runs at or under 24 min); it does not bar an intermediate value, and the ruling orders a downward re-size ('下调' in 6082946866 option A). The window is about 5.6 h (10:29Z-16:03Z), not 5: merge-queue groups built on 5919483 started at 10:29Z (37917885267 is the round-3 group itself, compare 'identical'), before main carried it at about 11:03Z. The revert-to-30 condition is kept in the comment.",
    "2_what_to_measure": "MEASURED, jobs API, attempt 1 everywhere (no re-runs). Membership: non-PR runs by GET compare/5919483472...HEAD_SHA = ahead|identical (29 ahead, 1 identical; control 3ca71b6 = behind); PR runs created at or after 11:03:23Z. Earlier seat readings re-measured, not copied: merge_group slowest per run 7.6-24.1 over 30 runs (seat had 14.5-18.6 over 5); pull_request 6.9-27.1 over 25 runs (seat had 14.0-25.4 over 4). CLI slicing: the aggregate Test Core timing table reads '@objectstack/cli ... all 3 slices' on all 10 runs whose table was read (list in measurement.log_reads).",
    "3_stall_guard": "STATIC GATE: not crossed at 40 (or 35): node scripts/check-stall-guard-budget.mjs exit 0, --list shows the three test steps at cap 20m, budget 40m, slack 20m. MEASURED (falsifies the PM's 35 route): node scripts/measure-stall-guard-headroom.mjs over the window's 396 test-step observations, worst p 6m59s + s 19m39s = 26m38s on 37923333449 Test Core (3/6): window path at T=45 COVERED 8m22s, T=40 COVERED 3m22s, T=35 UNCOVERED 1m38s. Cap-deferred path: UNCOVERED at 45 by 1m38s already (pre-existing; needs T at least 47m), 6m38s at 40, 11m38s at 35. No budget edited."
    },
    "measurement": {
    "read_at": "2026-10-09T16:20Z, GET actions/workflows/ci.yml/runs created 10:15Z..16:20Z + GET actions/runs/RUN_ID/jobs?filter=all",
    "window": "runs 37917885267..37956350887, created 2026-10-09 10:29Z..16:03Z",
    "per_event_job_walls_min": {
    "schedule": {"runs": 5, "excluded_cancelled": 0, "in_progress": 0, "pre_change": 1, "jobs": 30, "min": 1.2, "median": 3.2, "p90": 14.0, "max": 21.0, "max_run": "37934353580 (3/6)", "per_run_slowest_median": 14.0},
    "push": {"runs": 6, "excluded_cancelled": 11, "in_progress": 1, "pre_change": 1, "jobs": 36, "min": 1.2, "median": 13.6, "p90": 20.0, "max": 20.9, "max_run": "37934058408 (3/6)", "per_run_slowest_median": 20.1},
    "merge_group": {"runs": 30, "failure_runs": ["37930835817", "37931072987"], "excluded_cancelled": 0, "in_progress": 0, "pre_change": 0, "jobs": 180, "min": 0.9, "median": 13.0, "p90": 20.5, "max": 24.1, "max_run": "37935402667 (3/6)", "per_run_slowest_median": 15.3},
    "pull_request": {"runs": 25, "failure_runs": ["37921971792"], "excluded_cancelled": 13, "in_progress": 2, "pre_change": 14, "jobs": 150, "min": 0.9, "median": 14.9, "p90": 23.9, "max": 27.1, "max_run": "37923333449 (3/6)", "per_run_slowest_median": 21.2},
    "all": {"runs": 66, "excluded_cancelled": 24, "in_progress": 3, "jobs": 396, "min": 0.9, "median": 13.3, "p90": 22.0, "p99": 26.0, "max": 27.08, "max_run": "37923333449 Test Core (3/6), 1625 s (closure 337 s, slice closure 15 s, tests 1179 s)"}
    },
    "censored": "largest cancelled-run job wall 25.3 min, 37952087770 (4/6), cancelled 1095 s into its test step; not in the distribution",
    "pre_change_for_contrast": "six heaviest pre-change pull_request runs in the same listing: slowest Test Core (1/6), the whole-CLI shard, 31.7-39.7 min (37920407075 39.7, 37919214042 38.2, 37918024929 37.6, 37919833209 34.9, 37920676666 34.0, 37920713245 31.7)",
    "candidates_80pct": {"30": "24.0 → 17 jobs over (15 pull_request, 2 merge_group)", "35": "28.0 → 0 over", "40": "32.0 → 0 over", "45": "36.0 → 0 over"},
    "stall_guard_test_step": {"worst_p_plus_s": "26m38s (37923333449 3/6)", "T45": "window COVERED 8m22s; deferred UNCOVERED 1m38s; deferred-miss observations 8/396", "T40": "window COVERED 3m22s; deferred UNCOVERED 6m38s; deferred-miss 60/396", "T35": "window UNCOVERED 1m38s (8/396); deferred UNCOVERED 11m38s; deferred-miss 151/396"},
    "main_runs_mostly_executed": "6 (at least 5 required): schedule 37927365187 (44 measured / 28 replayed), schedule 37934353580 (68/4), push 37934058408 (64/2), push 37938584496 (70/0), push 37946464515 (70/0), push 37953570203 (63/1). Scheduled-only count is 2: schedules 37941142676, 37948546673, 37956054498 mostly replayed (test steps under 40 s on at least 4 of 6 shards; tables not read).",
    "log_reads": "aggregate Test Core timing tables with 'all 3 slices' for @objectstack/cli: 37923333449 (PR, 70/0), 37936325084 (PR, 70/0), 37934059977 (MG, 70/0), 37927365187, 37934353580, 37926429923 (25/3), 37934058408, 37938584496, 37946464515, 37953570203"
    },
    "tests": "No package touched: no build closure and no package test/typecheck owed. ci.yml parses (yaml: jobs.test timeout-minutes 40, job name template unchanged (Test Core, matrix shard of 6)). Control-byte self-scan of ci.yml: no hits (grep exit 1). Lint: ci.yml is outside eslint's population (ESLint.isPathIgnored true), so no lint is owed. Headroom re-taken on the edited tree at a549ca3: window COVERED 3m22s. PR CI at a549ca3, read once: 31 success, 5 skipped, 1 in_progress (Lint & Repo Gates); 6 of the 7 required contexts success (Test Core, TypeScript Type Check, Dogfood Regression Gate, Build Core, Temporal Conformance, Governed Surface Queue Guard). Report not held for CI.",
    "gates": {
    "head": "a549ca3a",
    "derive": "node scripts/pm/dispatch-gates.mjs --commands --repo objectstack-ai/objectstack → 47 commands (1 path vs merge base 446c8b2; repo assertion held)",
    "derived": 47,
    "run": 47,
    "exit_0": 43,
    "not_measured": 4,
    "unrun": 0,
    "reconciliation": "node scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack --ran ran.list at a549ca3: exit 0, '47 derived famil(ies) accounted for -- 43 run, 4 NOT-MEASURED (4 DERIVED from a recorded exit 3)'",
    "named_by_dispatch": {"node scripts/check-stall-guard-budget.mjs": 0, "node scripts/check-stall-guard-budget.mjs --list": 0, "pnpm check:stall-guard-budget": 0, "pnpm check:stall-guard-headroom": 0, "pnpm check:required-contexts": 0, "pnpm check:workflow-step-name-quoting": 0},
    "long_battery": "pnpm check:pm-dispatch-gates: detached with nohup, exit code written to a file, awaited in the foreground with tail --pid on PID 4492: exit 0, '2011 cases pass', 1300.0 s",
    "not_measured_reason": "check:dts-closure, check:dual-build-cjs-loads, check:lean-entry-closure, check:sourcemap-no-sources-content exit 3 PREREQUISITE NOT MET: they load every package's built dist/, and this diff touches no package",
    "workflow_value_families": "8 derived families take a value from the workflow (check-required-contexts --verify-required-set, check-shard-attestation --emit/--verify, check-test-completeness) and are NOT MEASURED locally by construction; CI runs them",
    "exits_0": ["node scripts/check-aggregator-roster.mjs", "node scripts/check-aggregator-roster.mjs --self-test", "node scripts/check-ci-filter-parity.mjs", "node scripts/check-ci-filter-parity.mjs --self-test", "node scripts/check-closing-keyword-parity.mjs", "node scripts/check-closing-keyword-parity.mjs --self-test", "node scripts/check-comment-mask-corpus.mjs", "node scripts/check-dts-emitted.mjs --self-test", "node scripts/check-position-name-fold-loaders.mjs", "node scripts/check-position-name-fold-loaders.mjs --self-test", "node scripts/check-self-test-wired.mjs", "node scripts/check-self-test-wired.mjs --self-test", "node scripts/check-self-test-workflow-commands.mjs", "node scripts/check-self-test-workflow-commands.mjs --self-test", "node scripts/check-step-collectors.mjs", "node scripts/check-step-collectors.mjs --self-test", "node scripts/check-whole-set-label-write.mjs", "node scripts/check-whole-set-label-write.mjs --self-test", "node scripts/ci/scheduled-full-run.mjs --self-test", "node scripts/docs-audit/check-drift-comment.mjs", "node scripts/pm/ci-failure.mjs --self-test", "pnpm check:agent-test-spelling", "pnpm check:console-injection", "pnpm check:console-sha", "pnpm check:cross-package-test-inputs", "pnpm check:declared-population-live", "pnpm check:driver-memory-census", "pnpm check:gitlink-declared", "pnpm check:node-version", "pnpm check:nul-bytes", "pnpm check:pm-dispatch-gates", "pnpm check:pm-expected-skips", "pnpm check:pm-post-stamped", "pnpm check:pnpm-acquisition", "pnpm check:pnpm-filter-targets", "pnpm check:refd-timer-probe", "pnpm check:required-contexts", "pnpm check:shard-attestation", "pnpm check:stall-guard-budget", "pnpm check:stall-guard-headroom", "pnpm check:watch-hint-literal", "pnpm check:workflow-status-functions", "pnpm check:workflow-step-name-quoting"],
    "exit_3": ["pnpm check:dts-closure", "pnpm check:dual-build-cjs-loads", "pnpm check:lean-entry-closure", "pnpm check:sourcemap-no-sources-content"]
    },
    "line_budget": "not applicable: no skills/** or ledger-governed file. Diff vs origin/main 446c8b2: .github/workflows/ci.yml +30/-21 = 51 changed lines, under the 3000 human-merge threshold; 0 governed paths.",
    "files_changed": [".github/workflows/ci.yml"],
    "deviations": [
    "Route: 40, not Zone 3's 35. 35 meets the 80% wall rule, but measured headroom puts the stall guard's window verdict after a 35 wall on the slowest observed job; Zone 3 says measurement wins. The seat can still take 35 as a one-line change if it accepts that the window path misses on 8 of 396 observations.",
    "Stop condition 1 reading: 'fewer than 5 scheduled/push full runs on main with mostly executed packages' read as scheduled OR push. The count is 6 (2 scheduled, 4 push at 63-70 packages executed). Read as scheduled-only it is 2, and the condition would fire; the hourly runs mostly replay the cache their commit's push run saved.",
    "13 MCP reads (mcp__github__get_job_logs), because the REST job-log download is still refused at the blob-storage CONNECT (productionresultssa12.blob.core.windows.net). 3 of them (schedules 37941142676, 37948546673, 37956054498) used a 60-line tail that ended below the timing table once the suite-ceiling step landed (6abd265), so they returned no table. They were not re-read; those runs are classified by step seconds instead.",
    "The commit message says the slowest job fell 'from 31-40 min'. That is the six heaviest pre-change PR runs in this listing (31.7-39.7). Round 2's 12-run sample read 20.5-39.1 for the whole-CLI shard. PR body Acceptance note 3 says so. History not rewritten.",
    "Commit trailers use AGENTS.md's model-free pair (Claude-Session + Co-authored-by: Claude) instead of the harness reminder's model-named Co-Authored-By. The PR body ends with AGENTS.md's session-URL footer block instead of the harness's two-line block.",
    "Scratch reads went into the scratchpad's issue-22075/ directory, which earlier rounds of this card may also have used; issue.json and comments.json there were overwritten by this round's copies. No repo file is affected."
    ],
    "mcp_calls": "13 -- all mcp__github__get_job_logs, read-only, on aggregate Test Core jobs 113848429304, 113805524377, 113841123495, 113816168421, 113840951185, 113856385523, 113887299563, 113911664368, 113812215711, 113839839339, 113856061599, 113882603837, 113907017074. No MCP write tool.",
    "api_writes": "4 REST writes in 3 relay dispatches (POST /repos/objectstack-ai/objectstack/dispatches, each executed as objectstack-fleet[bot]): (1) pr_create = POST /repos/objectstack-ai/objectstack/pulls, draft #22512, 17090 bytes sent = stored (relay read-back identical; REST read-back equal); (2)+(3) one label-write stroke = POST /issues/22512/labels (skip-changeset) and POST /issues/22512/assignees (os-bill), read back: labels ci/cd, size/s, skip-changeset, assignee os-bill; (4) post-stamped --comment=22075 = POST /issues/22075/comments (this report). 2 git pushes (empty-branch probe, the commit) are not REST writes. REST reads (runs, jobs, compare, PR read-back) are not counted.",
    "open_questions": [],
    "out_of_scope_findings": [
    "carrier: 承接者:无 · noted, not filed -- the revert-to-30 condition reads 'every scheduled run for a week'. In this window 3 of 5 hourly scheduled runs replayed most packages from the cache their commit's push run saved (slowest 1.6 / 14.0 / 11.4 min), while pull_request runs read up to 27.1. So the condition can read as met on light runs while PR and queue shards sit over 24 min. Kept as dispatched; PR Acceptance note 1.",
    "carrier: 承接者:无 · noted, not filed (dormant: bites only when a guard defers to its cap) -- measure-stall-guard-headroom on the window's jobs at base 446c8b2 exits 1. The cap-deferred verdict cannot land before the job wall on observed healthy runs for Test Core 'Run this shard's tests' (needs at least 47m, has 45, now 40), dogfood 'Boot example apps and exercise real user flows' (needs 33m, has 30) and temporal-conformance 'Run the non-SQL temporal backends under the skewed process zone' (needs 31m, has 30). PR Acceptance note 2."
    ],
    "cleanup": "rm -rf node_modules and git worktree remove of /home/user/objectstack-issue-22075 run after this comment posts; the detached battery (PID 4492) has exited."
    }

  9. objectstack-fleet commented on Oct 9, 2026

    @objectstack-fleet
    ContributorAuthor

    Blocked-by: #22516

    Review of PR #22512 (head a549ca3a): built as ruled; held in draft on a measured contrary fact · 2026-10-09T17:01Z

    domain:devx seat 1 · os-bill · session_01LYXc6ckoWuZyVZpWYizdMh, against GitHub and origin/main, not the report (6085414584). Claim: 6084791443. Ruling-ref: 6083012192.

    Checklist, all met:

    • Draft, base main. Line 1 is Fixes #22075, line 2 is Clause-②: no; one closing keyword in the body.
    • 1 file, +30/−21: .github/workflows/ci.yml, the Test Core job's timeout-minutes (45 → 40) and its comment only. The matrix, the drift constants, the required contexts, other jobs and the stall-guard scripts are untouched. skip-changeset is correct.
    • The comment carries the measured window: runs 37917885267..37956350887, 396 jobs; the max is 27m05s on 37923333449 (3/6). The revert-to-30 condition is kept.
    • Re-run by the seat at a549ca3a: check-stall-guard-budget exit 0.
    • CI at this act: 31 success, 5 skipped, Lint & Repo Gates in progress.

    Why it is held, and not landed. This PR is the literal execution of the ruling's 「下调」. The dev's measurement found a cost that the seat's option text did not price: the stall guard's cap-deferred verdict misses the job wall on 60 of 396 observed jobs at 40, against 8 at 45 and 0 at 50. That path is the one measure-stall-guard-headroom.mjs names as mattering most. The seat corrects its own option text, and puts the value to the maintainer on #22516 (options A 40 / B 45 / C 50, recommending C). Under the contrary-fact rule, this PR stays in draft with no auto-merge until the answer.

    On the answer:

    • A: the seat flips pr_ready + automerge_enable.
    • B or C: a patch round to the same dev on this branch, then a re-review.

    Recorded, not filed (dev findings):

    Labels in this act: pm:dispatched → pm:blocked. The assignee os-bill stays (work done, the PR waits on a decision).

  10. objectstack-fleet commented on Oct 10, 2026

    @objectstack-fleet
    ContributorAuthor

    Ruling C on #22516; patch round queued behind #22476 · domain:devx seat 1 · 2026-10-10T00:20Z

    os-bill · session_01LYXc6ckoWuZyVZpWYizdMh. Ruling-ref: 6091472780 (#22516, the maintainer's 「C 升到 50 (Recommended)」). #22516 is closed completed, so this card's Blocked-by: #22516 (6085497720) is met.

    Labels in this act: pm:blocked → pm:dispatched (the card is this seat's, with a live branch and PR, and it is next in the serial order). Assignee os-bill unchanged.

  11. objectstack-fleet commented on Oct 10, 2026

    @objectstack-fleet
    ContributorAuthor

    os-dev-report
    {
    "issue": 22075,
    "round": "patch round 1 (ruling C)",
    "status": "done",
    "branch": "claude/issue-22075-test-core-timeout-resize",
    "pr": "#22512",
    "head": "215ab47a736fc9ba38a645c89d11198d4b806281",
    "commits_this_round": ["f9f2661e5 (merge of origin/main 99801d8, a merge commit, no rebase or force)", "215ab47a7 (timeout-minutes 40 to 50 and the rationale comment)"],
    "session": "session_01LYXc6ckoWuZyVZpWYizdMh (mode:subagent of the domain:devx seat 1 PM). Same claim 6084791443, verified as still the newest Claim: on the card before the first edit. Ruling read first: 6091472780 (#22516) and the seat's note 6091531246.",
    "premise_still_valid": true,
    "summary": "The worktree was recreated from origin/claude/issue-22075-test-core-timeout-resize at a549ca3, and origin/main 99801d8 was merged in with a merge commit (f9f2661). That merge was clean: main's only ci.yml change since 446c8b2 is the temporal-conformance service images, and no regen debt was recorded. Commit 215ab47 then sets the Test Core shard job's timeout-minutes from 40 to 50, as ruled. It rewrites the rationale comment so the value is derived from both stall-guard paths plus the wall. The worst p + s is 26m38s (run 37923333449, 3/6), so the window verdict lands at 36m38s and the cap verdict at 46m38s, inside 50 by 3m22s. The slowest job wall, 27m05s, is within 80% of 50 (40.0). The comment keeps the history line and the window (run ids, counts per event), and drops 'So 40, not 35.' and 'no value this re-size could pick restores it'. The revert condition is rewritten: a value below 50 needs both guard paths covered on a re-measured window plus the 80% wall margin, read on pull_request and merge_group runs, not the hourly schedule alone. The bare back-to-30 rule is gone, and 'the stall guard stays the primary hang detector' is kept. Both pushes were fast-forward. The PR body and title were not touched.",
    "new_pr_title": "ci(test-core): the Test Core shard wall is 50, so both stall-guard verdicts land before it",
    "pr_body_what_changed": "Re-sizes the Test Core shard job's timeout-minutes in .github/workflows/ci.yml from 45 to 50, per ruling C on #22516 (record 6091472780, the maintainer's 「C 升到 50 (Recommended)」). It rewrites the rationale comment above it so the value is derived rather than chosen. Over the window of 66 runs and 396 Test Core (N/6) jobs after 5919483472, the worst p + s on "Run this shard's tests" is 26m38s, on run 37923333449 Test Core (3/6); p + s runs from job start to the end of that step. So the stall guard's window verdict lands at 36m38s, and its cap-deferred verdict at 46m38s, inside 50 by 3m22s. node scripts/measure-stall-guard-headroom.mjs at T=50 reads both paths COVERED on all three guarded test steps. The cap verdict missed 45 on 8 of the 396 jobs, and would miss 40 on 60. The slowest job wall, 27m05s, is within 80% of 50 (40.0). The revert condition now admits a value below 50 only when, on a re-measured window of Test Core runs, both guard verdicts (p + s + 10m and p + s + 20m) land before it on every observation and the slowest job wall is at or under 80% of it. That window is read on the pull_request and merge_group runs, not the hourly schedule alone, because scheduled runs mostly replay the cache. The former sole condition, back to 30 after a week of scheduled runs at or under 24 min, is gone.",
    "measurement": {
    "source": "round 1's window, unchanged: runs 37917885267..37956350887 (2026-10-09 10:29Z..16:03Z): 5 schedule, 6 push, 30 merge_group and 25 pull_request runs, 396 jobs, 24 cancelled runs excluded; slowest job wall per event (min): schedule 21.0, push 20.9, merge_group 24.1, pull_request 27.1; p90 of all 396 is 22.0. No new window was taken this round.",
    "headroom_T50": "node scripts/measure-stall-guard-headroom.mjs --from the window's saved jobs payload, run on this branch's tree at 215ab47. 'Run this shard's tests': undeferred p+s+W = 36m38s vs T 50m, COVERED with 13m22s to spare; deferred p+s+C = 46m38s vs T 50m, COVERED with 3m22s to spare. 'Build this shard's dependency closure' and 'Build the sliced package's dependency closure': both paths COVERED (17m26s and 27m26s vs 50m). The tool's exit is 1, only because of 2 steps outside this card: dogfood 'Boot example apps and exercise real user flows' (deferred, needs T at least 33m, has 30m) and temporal-conformance 'Run the non-SQL temporal backends under the skewed process zone' (deferred, needs 31m, has 30m). Both are the pre-existing gaps the ruling leaves out of scope.",
    "wall_margin": "27.08 min is within 40.0 (80% of 50); 0 of 396 jobs over it"
    },
    "tests": "No package touched, so no build closure and no package test or typecheck is owed. ci.yml parses (yaml: jobs.test timeout-minutes 50). Control-byte self-scan of ci.yml: no hits (grep exit 1). The removed phrases ('So 40, not 35', 'no value this re-size could pick restores it', the back-to-30 sole condition) count 0 in the file. git show HEAD:.github/workflows/ci.yml carries exactly 1 line ' timeout-minutes: 50'. ci.yml is outside eslint's population (isPathIgnored true in round 1), so no lint is owed. PR CI at 215ab47, read once: 26 success, 2 skipped, 4 in_progress (Dogfood Regression Gate, Console Pin Gate, Temporal Conformance, Lint & Repo Gates). The report was not held for CI.",
    "gates": {
    "head": "215ab47a7",
    "derive": "node scripts/pm/dispatch-gates.mjs --commands --repo objectstack-ai/objectstack (no paths): 47 commands, 1 path vs merge base 99801d8, repo assertion held. The command set is identical to round 1's.",
    "derived": 47,
    "run": 47,
    "exit_0": 43,
    "not_measured": 4,
    "unrun": 0,
    "reconciliation": "node scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack --ran ran.list at 215ab47: exit 0, '47 derived famil(ies) accounted for — 43 run, 4 NOT-MEASURED (4 DERIVED from a recorded exit 3)'",
    "named_by_dispatch": {"node scripts/check-stall-guard-budget.mjs": 0, "node scripts/check-stall-guard-budget.mjs --list": 0, "pnpm check:stall-guard-budget": 0, "pnpm check:stall-guard-headroom": 0, "pnpm check:required-contexts": 0, "pnpm check:workflow-step-name-quoting": 0},
    "verdict_lines": {
    "check-stall-guard-budget": "OK (39 workflow file(s), 65 job(s), 744 step(s), 11 guard-wrapped step(s); every effective cap clears its job budget by at least one stall window). --list: the three ci.yml test steps read cap 20m, budget 50m, slack 30m.",
    "check:stall-guard-headroom": "measure-stall-guard-headroom --self-test: 35 assertion(s) passed.",
    "check:required-contexts": "7 required context name(s) pinned across 3 workflow(s)",
    "check:workflow-step-name-quoting": "OK (scanned 39 workflow file(s) + 2 composite action file(s), 752 step name(s))"
    },
    "long_battery": "pnpm check:pm-dispatch-gates ran detached under nohup, with its PID and exit code recorded to files, and was read in the foreground. Exit 0: 'dispatch-gates self-test (fast tier): 1825 cases pass; 6 slow section(s) deferred', 252.0 s. Since the main merge, this script runs the fast tier off CI by its own header ('a dev box, no flag: the fast tier, what a delivery runs'). CI runs both tiers under GITHUB_ACTIONS. The slow tier was not run locally.",
    "exit_3_not_measured": ["pnpm check:dts-closure", "pnpm check:dual-build-cjs-loads", "pnpm check:lean-entry-closure", "pnpm check:sourcemap-no-sources-content"],
    "not_measured_reason": "PREREQUISITE NOT MET: these gates load built dist/ output, and no package is touched.",
    "workflow_value_families": "8 derived families take a value from the workflow and are NOT MEASURED locally by construction; CI runs them."
    },
    "line_budget": "Not applicable: no skills/** or ledger-governed file. Net diff vs origin/main 99801d8: .github/workflows/ci.yml +36/-23 = 59 changed lines, under the 3000 human-merge threshold; 0 governed paths. This round's own commit 215ab47 is +33/-29 on that file.",
    "files_changed": [".github/workflows/ci.yml"],
    "deviations": [
    "The PR title and body still describe 40, as instructed. The seat sets the new title (new_pr_title) and pastes pr_body_what_changed. The old body's 'Why 40 and not 35' section and the 40-row tables become stale once 50 lands.",
    "pnpm check:pm-dispatch-gates now runs its fast tier off CI (1825 cases, slow tier deferred), a behaviour main gained after round 1. Round 1's 2011-case figure is the old full battery, so the two numbers are not comparable.",
    "Commit trailers use AGENTS.md's model-free pair (Claude-Session + Co-authored-by: Claude) instead of the harness reminder's model-named Co-Authored-By.",
    "The merge commit uses git's default message ("Merge remote-tracking branch 'origin/main' into ...")."
    ],
    "mcp_calls": "0 this round; 13 cumulative, all mcp__github__get_job_logs reads in round 1. No MCP write tool.",
    "api_writes": "1 this round: post-stamped --comment=22075 = POST /repos//issues/22075/comments through the fleet relay (POST /repos/objectstack-ai/objectstack/dispatches, executed as objectstack-fleet[bot]); this report. The 2 git pushes (f9f2661, then 215ab47, both fast-forward) are not REST writes. No PR PATCH and no label write. REST reads (comments, PR, check-runs) are not counted. Cumulative: 5.",
    "open_questions": [],
    "out_of_scope_findings": [],
    "cleanup": "rm -rf node_modules and git worktree remove of /home/user/objectstack-issue-22075 run after this comment posts. The detached battery (PID 19402) has already exited."
    }

  12. objectstack-fleet commented on Oct 10, 2026

    @objectstack-fleet
    ContributorAuthor

    ACCEPT — PR #22512 (head 215ab47a7), patch round 1 (ruling C) · 2026-10-10T03:23Z

    Reviewed by domain:devx seat 1 · os-bill · session_01LYXc6ckoWuZyVZpWYizdMh, against GitHub, not the report (os-dev-report 6093240370). Claim: 6084791443. Ruling: C, 50 (6091472780, the maintainer's 「C 升到 50 (Recommended)」).

    Checklist:

    • Draft, base main. Line 1 is Fixes #22075, line 2 is Clause-②: no. The body carries one closing keyword. The assignee is os-bill, and the PR has skip-changeset.
    • One file, .github/workflows/ci.yml, +36/−23 against main. check-governed-merges --pr 22512 reads 0 of 1 path governed, under the human-merge line. No packages/spec/src/** path is touched, so no contract review is owed.
    • The diff, read by the seat:
      • The Test Core shard job's timeout-minutes goes 45 → 50. No other line changes value.
      • The comment derives 50 from the window verdict (p + s + 10m = 36m38s), the cap-deferred verdict (p + s + 20m = 46m38s, inside 50 by 3m22s) and the 80% wall margin (27m05s within 40.0). These come from the worst observation, run 37923333449 Test Core (3/6).
      • The ⛔ revert condition now allows a value below 50 only when a re-measured window keeps both guard paths covered and the 80% margin holds. That window is read on pull_request and merge_group runs, not the hourly schedule alone. The bare back-to-30 rule is gone, and "the stall guard stays the primary hang detector" is kept. This is what the ruling asked.
    • Re-run by the seat at 215ab47a7: node scripts/check-stall-guard-budget.mjs --list exits 0. All three guarded test steps read window 10m, cap 20m, budget 50m, slack 30m.
    • Dev gates: 47 derived, 43 exit 0, 4 NOT-MEASURED (exit 3; each needs built dist/, and no package is touched), 0 UNRUN. check:pm-dispatch-gates ran its fast tier: 1825 cases pass.
    • The merge of origin/main 99801d831 is a merge commit, f9f2661e5. It was clean, with no rebase and no force.
    • PR title and body: rewritten by the seat for ruling C. The title is "ci(test-core): the Test Core shard wall is 50, so both stall-guard verdicts land before it". Round 1's "Why 40 and not 35" became "Why 50", with the 50 row added to the headroom table. Round 1's Acceptance note 1, the revert population, is retired, because the new condition reads PR and queue runs.

    Out-of-scope findings: none new. The dogfood and temporal-conformance deferred-path gaps (needs 33m and 31m, has 30m) stay as the PR's Acceptance note 1, noted and not filed. The ruling scoped Test Core only.

    Landing: pr_ready + automerge_enable once CI at this head is green, or shows only expected skips.

  13. objectstack-fleet commented on Oct 10, 2026

    @objectstack-fleet
    ContributorAuthor

    Landed: PR #22512 → 78cc46165b · 2026-10-10T04:04Z

    domain:devx seat 1 · os-bill · session_01LYXc6ckoWuZyVZpWYizdMh. Verified on origin/main by content.

    • Merged through the merge queue at 2026-10-10T04:02Z. The merge commit 78cc46165b is an ancestor of origin/main (76bc1e03a).
    • .github/workflows/ci.yml on main is blob-identical to the ACCEPTed head 215ab47a7 (19f0c2f0e8; ACCEPT on this card at 2026-10-10T03:23Z).
    • The Test Core shard job reads timeout-minutes: 50. Its comment derives the value from both stall-guard verdicts and the 80% wall margin, and it carries the measured revert condition. This is ruling C (6091472780).
    • The card is complete in three clauses.
    • Closing-keyword check: the PR body has one keyword (Fixes #22075). This card closed completed, and no other card did.

    Acceptance notes (noted, not filed): two steps outside this job have the same deferred-path gap at their current walls, read with the same instrument: dogfood "Boot example apps…" needs 33m and has 30m; temporal-conformance "non-SQL temporal backends…" needs 31m and has 30m. The ruling scoped Test Core only. carrier: none.

    Labels in this act: pm:dispatched removed (the card closed with it on). tooling · priority:p2 · domain:devx · area:devpath stay.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:devpathThe road — create, dev, verify, publish/install, connect an agent, iteratedomain:devxpriority:p2Medium: important, M3tooling

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions