Skip to content

ci(test-core): the Test Core shard wall is 50, so both stall-guard verdicts land before it - #22512

Merged
objectstack-fleet[bot] merged 3 commits into
mainfrom
claude/issue-22075-test-core-timeout-resize
Oct 10, 2026
Merged

objectstack-fleet[bot] merged 3 commits into
mainfrom
claude/issue-22075-test-core-timeout-resize

Conversation

@objectstack-fleet

@objectstack-fleet objectstack-fleet Bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #22075
Clause-②: no

What this does

This PR re-sizes the Test Core shard job's timeout-minutes in .github/workflows/ci.yml from 45 to 50, per ruling C on #22516 (record 6091472780, the maintainer's 「C 升到 50 (Recommended)」). It also rewrites the rationale comment above the value, so the value is derived and the revert condition is stated against measurement. This is Done-when clause 3 of the card. One file changes: a comment plus one value. No step, matrix, job name, required context, stall-guard setting or other job's timeout changes.

Rulings.

The window

The window covers every CI run whose tested tree carries 5919483472. Read from the jobs API at 2026-10-09T16:20Z.

  • Membership. For push, schedule and merge_group, the run's head commit must answer ahead or identical to GET compare/5919483472...HEAD_SHA. That is 29 ahead and 1 identical (the round-3 queue group itself). The control 3ca71b6e05, its parent, answers behind.
  • Pull requests. For pull_request, the run must be created at or after 11:03:23Z, when main first carried the commit. A PR run tests the PR merged onto main as of that moment.
  • Range. Runs 37917885267 to 37956350887, created 10:29Z to 16:03Z.
  • Wall. completed_at − started_at of each Test Core (N/6) job, on attempt 1. No run in the window was re-run.
event runs included cancelled runs excluded in progress excluded pre-change excluded jobs job wall min / median / p90 / max (min) slowest job per-run slowest, median / max
schedule 5 0 0 1 30 1.2 / 3.2 / 14.0 / 21.0 37934353580 (3/6) 14.0 / 21.0
push 6 11 1 1 36 1.2 / 13.6 / 20.0 / 20.9 37934058408 (3/6) 20.1 / 20.9
merge_group 30 (2 failure) 0 0 0 180 0.9 / 13.0 / 20.5 / 24.1 37935402667 (3/6) 15.3 / 24.1
pull_request 25 (1 failure) 13 2 14 150 0.9 / 14.9 / 23.9 / 27.1 37923333449 (3/6) 21.2 / 27.1
all 66 24 3 16 396 0.9 / 13.3 / 22.0 / 27.08 37923333449 (3/6), 27m05s
  • Failure runs are kept. Their Test Core jobs ran to a conclusion under the same wall. 37921971792 shard 4/6 failed at 24.9 min, 37930835817 shard 2/6 failed at 19.6 min, and 37931072987 shard 2/6 failed at 20.7 min.
  • Cancelled runs are excluded. Their jobs are censored by supersession. The largest cancelled job wall is 25.3 min, on 37952087770 (4/6), cancelled 1095 s into its test step. That is below the maximum above.
  • Before the change, the six heaviest pre-change pull_request runs in the same listing had slowest jobs of 31.7–39.7 min: 37920407075 39.7, 37919214042 38.2, 37918024929 37.6, 37919833209 34.9, 37920676666 34.0, 37920713245 31.7. Each time the slowest was Test Core (1/6), the shard that ran the whole CLI alone.

Did the CLI run as slices? I read the aggregate Test Core job's timing table on 10 runs. On each, @objectstack/cli reads all 3 slices.

run event packages measured / cache-replayed slowest job
37923333449 pull_request (the window's slowest) 70 / 0 27.1
37936325084 pull_request 70 / 0 20.8
37934059977 merge_group 70 / 0 24.0
37927365187 schedule 44 / 28 18.9
37934353580 schedule 68 / 4 21.0
37926429923 push 25 / 3 15.4
37934058408 push 64 / 2 20.9
37938584496 push 70 / 0 20.7
37946464515 push 70 / 0 19.6
37953570203 push 63 / 1 20.5
  • The other three scheduled runs (37941142676, 37948546673, 37956054498) were not log-read. Their test steps read under 40 s on at least four of six shards, so they mostly replayed the cache that the push run on the same commit saved. Three more log reads, cut short above the table, are not counted.
  • The dispatch's main-run floor holds. Six main runs executed most of their packages: 2 scheduled (37927365187, 37934353580) and 4 push (37934058408, 37938584496, 37946464515, 37953570203). That is at least the 5 required. Push runs are affected-set runs by design, but these four executed 63–70 packages. Counting the hourly scheduled runs alone gives 2; see Acceptance notes.

Every included run (66), with its six job walls

run event created (UTC) conclusion Test Core (1/6)–(6/6) job walls (min) slowest
37927365187 schedule 12:01 success 11.4 / 12.8 / 10.9 / 18.9 / 7.0 / 4.5 18.9 (4/6)
37934353580 schedule 13:05 success 13.2 / 13.9 / 21.0 / 20.4 / 9.0 / 13.7 21.0 (3/6)
37941142676 schedule 14:01 success 1.2 / 1.6 / 1.3 / 1.4 / 1.3 / 1.4 1.6 (2/6)
37948546673 schedule 15:01 success 1.4 / 14.0 / 1.4 / 13.0 / 1.5 / 1.7 14.0 (2/6)
37956054498 schedule 16:01 success 11.4 / 1.8 / 1.7 / 1.8 / 1.6 / 1.7 11.4 (1/6)
37926429923 push 11:52 success 8.7 / 12.0 / 15.4 / 14.2 / 6.7 / 4.2 15.4 (3/6)
37928884094 push 12:15 success 8.6 / 3.2 / 1.4 / 1.5 / 1.8 / 1.2 8.6 (1/6)
37934058408 push 13:02 success 13.5 / 18.2 / 20.9 / 15.6 / 14.0 / 14.7 20.9 (3/6)
37938584496 push 13:40 success 12.2 / 16.4 / 20.7 / 20.0 / 11.3 / 12.9 20.7 (3/6)
37946464515 push 14:44 success 17.1 / 17.9 / 13.3 / 19.6 / 9.9 / 13.6 19.6 (4/6)
37953570203 push 15:41 success 9.0 / 17.4 / 20.5 / 19.6 / 14.9 / 13.5 20.5 (3/6)
37917885267 merge_group 10:29 success 9.1 / 3.5 / 2.0 / 1.3 / 1.5 / 1.2 9.1 (1/6)
37918648139 merge_group 10:36 success 18.4 / 18.5 / 22.8 / 14.4 / 11.3 / 10.1 22.8 (3/6)
37919164432 merge_group 10:41 success 17.1 / 17.7 / 17.2 / 20.4 / 12.8 / 14.2 20.4 (4/6)
37919744896 merge_group 10:47 success 11.5 / 2.4 / 0.9 / 1.1 / 1.3 / 1.1 11.5 (1/6)
37920849417 merge_group 10:58 success 8.4 / 2.4 / 1.1 / 1.1 / 1.3 / 1.2 8.4 (1/6)
37921059083 merge_group 11:00 success 18.8 / 15.6 / 23.4 / 15.4 / 13.0 / 16.5 23.4 (3/6)
37922022714 merge_group 11:09 success 12.3 / 15.9 / 15.3 / 16.6 / 13.6 / 11.8 16.6 (4/6)
37922106947 merge_group 11:10 success 18.6 / 16.7 / 18.6 / 15.9 / 16.9 / 16.6 18.6 (1/6)
37924560197 merge_group 11:34 success 13.1 / 14.5 / 13.6 / 10.1 / 12.0 / 11.2 14.5 (2/6)
37924620787 merge_group 11:34 success 13.8 / 14.7 / 10.7 / 10.4 / 8.0 / 11.3 14.7 (2/6)
37924680433 merge_group 11:35 success 14.1 / 13.6 / 16.0 / 14.4 / 9.7 / 6.9 16.0 (3/6)
37925822821 merge_group 11:46 success 13.1 / 1.6 / 1.0 / 0.9 / 1.2 / 1.3 13.1 (1/6)
37930060454 merge_group 12:26 success 15.0 / 13.9 / 22.5 / 21.1 / 11.8 / 16.1 22.5 (3/6)
37930189782 merge_group 12:27 success 8.4 / 4.7 / 1.3 / 1.2 / 0.9 / 1.1 8.4 (1/6)
37930835817 merge_group 12:33 failure 15.5 / 19.6 / 16.0 / 18.9 / 14.9 / 16.7 19.6 (2/6)
37931072987 merge_group 12:35 failure 14.8 / 20.7 / 23.4 / 23.0 / 14.5 / 12.6 23.4 (3/6)
37934059977 merge_group 13:02 success 12.8 / 20.5 / 24.0 / 16.1 / 17.2 / 16.3 24.0 (3/6)
37934063662 merge_group 13:02 success 14.7 / 11.1 / 14.5 / 12.1 / 11.8 / 11.4 14.7 (1/6)
37934410546 merge_group 13:05 success 19.9 / 20.7 / 21.2 / 23.3 / 17.9 / 14.3 23.3 (4/6)
37934475276 merge_group 13:06 success 12.8 / 4.1 / 4.4 / 1.2 / 1.1 / 1.1 12.8 (1/6)
37935402667 merge_group 13:14 success 20.7 / 20.4 / 24.1 / 21.4 / 14.7 / 16.9 24.1 (3/6)
37942731232 merge_group 14:14 success 10.4 / 9.5 / 13.4 / 7.7 / 1.8 / 1.5 13.4 (3/6)
37942773589 merge_group 14:15 success 7.6 / 2.6 / 1.2 / 1.3 / 1.4 / 1.4 7.6 (1/6)
37942824803 merge_group 14:15 success 19.0 / 14.6 / 23.9 / 18.3 / 16.8 / 17.0 23.9 (3/6)
37943908618 merge_group 14:24 success 13.4 / 13.8 / 10.4 / 9.5 / 10.1 / 10.6 13.8 (2/6)
37950612548 merge_group 15:17 success 15.2 / 18.1 / 22.9 / 14.9 / 15.8 / 11.1 22.9 (3/6)
37953132138 merge_group 15:37 success 7.8 / 4.2 / 2.1 / 1.1 / 1.2 / 1.1 7.8 (1/6)
37953459452 merge_group 15:40 success 15.0 / 19.0 / 22.7 / 21.8 / 16.1 / 15.5 22.7 (3/6)
37953831319 merge_group 15:43 success 10.7 / 11.7 / 13.8 / 12.5 / 6.7 / 7.0 13.8 (3/6)
37956350887 merge_group 16:03 success 10.5 / 5.6 / 1.4 / 1.2 / 1.0 / 1.1 10.5 (1/6)
37921971792 pull_request 11:09 failure 22.7 / 17.7 / 26.0 / 24.9 / 15.5 / 10.6 26.0 (3/6)
37922028318 pull_request 11:09 success 16.3 / 22.8 / 25.2 / 16.1 / 15.4 / 19.4 25.2 (3/6)
37922045966 pull_request 11:09 success 11.1 / 2.3 / 1.0 / 1.6 / 1.2 / 1.4 11.1 (1/6)
37922427884 pull_request 11:13 success 16.4 / 22.8 / 26.1 / 15.7 / 13.6 / 12.1 26.1 (3/6)
37923333449 pull_request 11:22 success 13.5 / 23.1 / 27.1 / 20.4 / 16.1 / 11.5 27.1 (3/6)
37924699846 pull_request 11:35 success 12.7 / 14.1 / 16.1 / 23.6 / 15.1 / 18.4 23.6 (4/6)
37926531514 pull_request 11:53 success 6.9 / 5.6 / 1.2 / 1.2 / 1.1 / 1.3 6.9 (1/6)
37926958001 pull_request 11:57 success 20.9 / 15.4 / 25.4 / 25.2 / 15.8 / 12.9 25.4 (3/6)
37927225061 pull_request 12:00 success 15.9 / 22.1 / 24.1 / 22.9 / 15.8 / 16.1 24.1 (3/6)
37930004341 pull_request 12:26 success 10.6 / 8.3 / 10.0 / 12.9 / 2.8 / 3.3 12.9 (4/6)
37930169155 pull_request 12:27 success 10.5 / 3.7 / 2.2 / 1.1 / 0.9 / 1.0 10.5 (1/6)
37932302950 pull_request 12:46 success 10.4 / 11.6 / 15.4 / 14.4 / 10.3 / 5.5 15.4 (3/6)
37934454127 pull_request 13:06 success 10.8 / 11.8 / 14.0 / 10.5 / 5.4 / 7.5 14.0 (3/6)
37934496824 pull_request 13:06 success 8.2 / 2.0 / 1.1 / 1.1 / 1.2 / 1.0 8.2 (1/6)
37936325084 pull_request 13:22 success 17.5 / 18.1 / 20.8 / 20.1 / 11.5 / 13.7 20.8 (3/6)
37939267019 pull_request 13:46 success 16.8 / 23.1 / 25.5 / 24.7 / 20.0 / 18.9 25.5 (3/6)
37943716557 pull_request 14:22 success 17.0 / 21.2 / 17.9 / 18.4 / 18.7 / 17.1 21.2 (2/6)
37944266225 pull_request 14:27 success 22.0 / 23.9 / 26.2 / 25.0 / 19.8 / 18.7 26.2 (3/6)
37947173860 pull_request 14:50 success 12.3 / 13.2 / 14.7 / 17.7 / 16.9 / 12.9 17.7 (4/6)
37948151438 pull_request 14:58 success 14.1 / 3.7 / 2.2 / 1.0 / 1.1 / 1.1 14.1 (1/6)
37948626652 pull_request 15:01 success 15.6 / 23.6 / 21.9 / 22.8 / 20.0 / 13.0 23.6 (2/6)
37950504865 pull_request 15:16 success 11.1 / 12.1 / 11.1 / 9.8 / 1.6 / 2.3 12.1 (2/6)
37951404817 pull_request 15:23 success 21.1 / 19.2 / 25.9 / 25.1 / 20.0 / 12.6 25.9 (3/6)
37951544850 pull_request 15:24 success 16.2 / 23.1 / 25.5 / 22.9 / 16.0 / 19.7 25.5 (3/6)
37953898847 pull_request 15:43 success 8.2 / 6.0 / 1.3 / 1.2 / 1.1 / 1.2 8.2 (1/6)

Why 50

Three measured readings set the value, and the largest answer wins. All are taken over the window above, on the slowest observation: run 37923333449 Test Core (3/6). Its "Run this shard's tests" step ends 26m38s after the job starts (p 6m59s plus s 19m39s).

  1. The window verdict. A silent freeze is killed 10 min after the last output, at p + s + 10m = 36m38s.
  2. The cap-deferred verdict. When the liveness probe sees the suite still writing, the guard waits until p + s + 20m = 46m38s.
  3. The 80% wall margin. The slowest job wall, 27m05s, must be within 80% of the wall.

Readings 1 and 2 come from node scripts/measure-stall-guard-headroom.mjs --from over the window's saved jobs payloads. The 50 row was taken on this branch's tree at 215ab47a7. The other rows are round 1's readings over the same window.

timeout-minutes window verdict (p+s+10m) cap-deferred verdict (p+s+20m) window misses (of 396) deferred misses (of 396) 80% of wall jobs over it
50 COVERED, 13m22s to spare COVERED, 3m22s to spare 0 0 40.0 0
45 (before) COVERED, 8m22s to spare UNCOVERED by 1m38s 0 8 36.0 0
40 (round 1) COVERED, 3m22s to spare UNCOVERED by 6m38s 0 60 32.0 0
35 UNCOVERED by 1m38s UNCOVERED by 11m38s 8 151 28.0 0

50 is the smallest whole-five value at which both guard verdicts land before the job timeout on every observation. The deferred path needs a wall of at least 47 min.

check-stall-guard-budget --list at 215ab47a7 reads, on all three guarded test steps: window 10m, cap 20m, budget 50m, slack 30m. The seat re-ran it at this head (exit 0).

The comment

  1. The backstop rationale paragraph is unchanged.
  2. The raise history becomes one line: 30 → 45 (CI: the shard-timings file is stale for the CLI package — 672s predicted vs 28m46s measured against a 30-minute timeout, so Test Core shard 1/6 is one slow run from being killed on any PR touching the CLI #16173), held at 45 until this card cut the CLI into 3 slices (5919483472).
  3. Re-sized 45 → 50. This part carries the window: the run ids, the counts per event, the per-event maxima and the p90. It also carries the worst p + s with its run, and the three readings above.
  4. ⛔ The revert condition. A value below 50 is allowed only on a re-measured window of Test Core runs, read with measure-stall-guard-headroom over every completed pull_request, merge_group, push and schedule run in it. On that window, both guard verdicts (p + s + 10m and p + s + 20m) must land before the new value on every observation. The slowest Test Core (N/6) job wall must also be at or under 80% of the new value.
    • The window is read on pull_request and merge_group runs, not the hourly schedule alone, because scheduled runs mostly replay the cache.
    • The former sole condition is gone. It said: back to 30 after a week of scheduled runs at or under 24 min.
    • "The stall guard stays the primary hang detector" is kept.

Gates

Derived with node scripts/pm/dispatch-gates.mjs --commands --repo objectstack-ai/objectstack against the merge base 99801d831, after origin/main was merged in (merge commit f9f2661e5, clean). It names 1 path and 47 commands. Each command ran at head 215ab47a7, with its exit code written before any pipe. --ran reconciles: 47 accounted for, 43 run, 4 NOT-MEASURED.

  • 43 exit 0. These include:
    • pnpm check:stall-guard-budget;
    • pnpm check:stall-guard-headroom (35 assertion(s) passed);
    • pnpm check:required-contexts (7 required context name(s) pinned across 3 workflow(s));
    • pnpm check:workflow-step-name-quoting (752 step name(s)).
  • 4 exit 3, PREREQUISITE NOT MET. These are check:dts-closure, check:dual-build-cjs-loads, check:lean-entry-closure and check:sourcemap-no-sources-content. Each loads every package's built dist/, and this diff touches no package. They are NOT MEASURED.
  • pnpm check:pm-dispatch-gates ran detached and exited 0, its fast tier only: 1825 cases pass, 6 slow sections deferred. Off CI that is now the script's default. CI runs both tiers.
  • Outside the derived list: node scripts/check-stall-guard-budget.mjs exits 0: "every effective cap clears its job budget by at least one stall window".
  • Families that take a value from the workflow (--verify-required-set, check-shard-attestation --emit/--verify, check-test-completeness) are NOT MEASURED locally by construction. CI runs them.

No package and no published file is touched, so this PR carries skip-changeset.

Acceptance notes

  1. The deferred stall-guard path outside this job (noted, not filed). At their current walls, two other steps have the same gap this re-size closes for Test Core, per the same instrument:
    • dogfood "Boot example apps and exercise real user flows" (needs 33m, has 30m);
    • temporal-conformance "Run the non-SQL temporal backends under the skewed process zone" (needs 31m, has 30m).
    • The ruling covers Test Core only. These two are recorded so the trade is visible.
  2. What the round-1 commit message's "31-40 min" covers. It is the band of the six heavy pre-change pull_request runs above (31.7–39.7). Round 2's 12-run sample read 20.5–39.1 min for the shard running the whole CLI. The tables above are the reading.

Generated by Claude Code

…sured post-slicing job walls

The 3-slice CLI split (5919483) took the slowest Test Core job from
31-40 min to at most 27m05s across 66 completed runs (396 shard jobs) whose
tested tree carries it. Two measured rules set the wall: the slowest job
within 80% of it (35 and 40 both pass), and the stall guard's window
verdict landing before it: the worst observed "Run this shard's tests"
ends 26m38s into its job, so the verdict lands at 36m38s, after a 35 wall
and inside 40. The comment carries the window, the per-event maxima and
the revert-to-30 condition.

Claude-Session: https://claude.ai/code/session_01LYXc6ckoWuZyVZpWYizdMh
Co-authored-by: Claude <noreply@anthropic.com>
…rdicts land before it

Ruling C on the timeout value: 40 -> 50. The worst observed p + s on
"Run this shard's tests" is 26m38s (run 37923333449, Test Core 3/6), so
the guard's window verdict lands at 36m38s and its cap-deferred verdict
at 46m38s, inside 50 by 3m22s; the slowest job wall, 27m05s, is within
80% of 50. The revert condition now admits a lower value only when a
re-measured window keeps both guard paths covered and the slowest job
wall at or under 80% of it.

Claude-Session: https://claude.ai/code/session_01LYXc6ckoWuZyVZpWYizdMh
Co-authored-by: Claude <noreply@anthropic.com>
@objectstack-fleet objectstack-fleet Bot changed the title ci(test-core): re-size the Test Core shard wall 45 -> 40 from the measured post-slicing job walls ci(test-core): the Test Core shard wall is 50, so both stall-guard verdicts land before it Oct 10, 2026
@objectstack-fleet
objectstack-fleet Bot marked this pull request as ready for review October 10, 2026 03:39
@objectstack-fleet
objectstack-fleet Bot enabled auto-merge October 10, 2026 03:39
@objectstack-fleet
objectstack-fleet Bot added this pull request to the merge queue Oct 10, 2026
Merged via the queue into main with commit 78cc461 Oct 10, 2026
42 checks passed
@objectstack-fleet
objectstack-fleet Bot deleted the claude/issue-22075-test-core-timeout-resize branch October 10, 2026 04:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci/cd size/s skip-changeset PR has no user-facing published change; bypasses the changeset gate

Projects

None yet

2 participants