Skip to content

ci(test-core): wire the shard timing-drift check: red past 1.5x measured/predicted, warning past 1.3x - #21998

Merged
objectstack-fleet[bot] merged 3 commits into
mainfrom
claude/issue-16465-wire-shard-drift-check
Oct 6, 2026
Merged

objectstack-fleet[bot] merged 3 commits into
mainfrom
claude/issue-16465-wire-shard-drift-check

Conversation

@objectstack-fleet

Copy link
Copy Markdown
Contributor

Fixes #16465
Clause-②: no

Wires the Test Core shard timing-drift check the tree has carried unwired since #16173, adds the card's 1.3x as a warning tier under its red, and restates the ci.yml comments that the wiring and the #21487 / #21826 changes left stale. Root scripts/ and one workflow, nothing published: skip-changeset.

The ruling this follows

Triage 5925054826, as the unlock 6013064225 restates it, verbatim:

  • "one red rule, the in-tree partition-test-shards.mjs --check-drift at MAX_MEASURED_OVER_PREDICTED = 1.5;"
  • "the card's 1.3× is a ::warning:: only;"
  • "⛔ no second red rule, so no 70%-of-timeout-minutes red."

The maintainer's authority on the card: 「同意你的建议,你负责执行派发所有可行的优化」.

What changed

.github/workflows/ci.yml, the Test Core shard job only

  • New step Check this shard's timing drift: --check-drift over .turbo/runs/*.json with --label "Test Core (N/6)" (the matrix shard), with no if: and no continue-on-error. It sits between the run-summary upload and the completeness guard, above the attestation pair, so a drift red also withholds that shard's attestation. A shard with no packages writes no summary; the step reports that as NOT MEASURED and exits 0, because the script itself treats zero inputs as a usage error.
  • The "⛔ THE DRIFT STEP IS DELIBERATELY NOT WIRED HERE YET" block is replaced by a description of what the step does. The stale CLI figures ("458.15s recorded, 1231.52s measured") and "the CLI is halved across two runners" are gone. The comment now cites the dataset as chore(ci): refresh the Test Core shard-timings dataset #21826 refreshed it (provenance run 37262126122, CLI 1702.69s) and the second sample below.
  • WHY SIX now names @objectstack/cli (1702.69s) as the heaviest indivisible suite, ahead of @objectstack/spec (1134.86s). The 6/7/8/10-shard table is re-derived on the current dataset with the partitioner's own partition() / balanceOf(): 1.04x / 1.22x / 1.39x / 1.74x, with the max at 1703s at every count. The old claim that the self-test "pins this arithmetic" is corrected: what it pins is pin 3, the heaviest package within 1.3x of the mean at SHARD_COUNT.
  • The 45-minute timeout comment. ⛔ The value stays 45. The revert condition is now stated against measured wall time instead of "once CI: the shard-timings file is stale for the CLI package — 672s predicted vs 28m46s measured against a 30-minute timeout, so Test Core shard 1/6 is one slow run from being killed on any PR touching the CLI #16173 lands": back to 30 only when the slowest shard's job wall time stays at or under 24 minutes (80% of 30) on every scheduled run for a week. On the 16 main runs after the chore(ci): refresh the Test Core shard-timings dataset #21826 refresh, shard 1/6 (the CLI alone) read 6m14s to 35m43s of job wall time, and was over 30 minutes on 9 of them (for example 34m39s in run 37415122516 and 35m43s in run 37453598388). Two present-tense "30-minute wall" phrases in the same job now say "timeout-minutes wall".
  • Slice-leg prose. The build-closure step and the test step now say that FILE_SHARDED_PACKAGES has been empty since ci(test-core): retire the CLI's file-level slicing by the partitioner's own slice-count derivation #21487, so no shard runs a slice today.

scripts/partition-test-shards.mjs

  • WARN_MEASURED_OVER_PREDICTED = 1.3 sits beside MAX_MEASURED_OVER_PREDICTED = 1.5, which is unchanged. It is the same ratio over the same executed windows; it is not MAX_SHARD_OVER_MEAN, and its docblock says why the two are kept separate.

  • driftReport() gains warned, the band strictly between the two bounds and exclusive of the red, so one shard gets one verdict.

  • renderDriftVerdict() returns the four verdicts (OK, WARN, DRIFT, NOT MEASURED) without printing them. WARN emits exactly one ::warning title=Test Core shard timing drift::… line on stdout and exits 0; DRIFT exits 1, with the same remedy text as before. escapeWorkflowCommandMessage() keeps the message on one line.

  • New self-test battery drift warning tier under the red (#16465), with 12 cases, registered in SELF_TEST_BATTERIES (floor 12). The roster floor goes from 10 to 11. The cases cover:

    • both edges of the band, the red's exclusivity and a healthy 1.18x reading;
    • replay-only input;
    • band non-empty (WARN below the red) and not below the balance tolerance;
    • the rendered annotation line, the red's exit code, and command escaping.

    Nothing in the self-test prints a workflow command (check:self-test-workflow-commands green).

Not touched: the attestation schema, scripts/test-shard-timings.json, measure-test-shard-timings.mjs, and the required-check names. Test Core and Test Core (N/6) are unchanged, and check:required-contexts is green.

The second sample (carried item 2): executed windows only

The test-core-run-summary-* artifacts could not be read from this container: the egress policy denies productionresultssa17.blob.core.windows.net. Job logs reach only their last 5000 lines through get_job_logs. So the sample comes from each run's Test Core timing table, which the aggregator echoes into its log. The table carries turbo's own execution windows through samplesFromSummary(), which is the same reader and the same exclusions as the gate. Replayed packages are listed there and excluded here.

shard reading source runs
1/6 (CLI alone) exact, 0.61x, 1.05x, 0.65x, 0.78x, 1.02x, 1.03x, 1.01x, 0.79x the 8 scheduled runs 37416453417, 37427862594, 37433381795, 37440129185, 37446949708, 37453598388, 37460325624, 37467795959
2/6 to 6/6 NOT MEASURED whole. The table prints only the 10 slowest packages, so a shard's lighter members have no window here same runs
(small-set push runs) exact on every shard that ran: 0.90x (spec alone), 0.48x (rest alone), 0.67x (create-objectstack alone); 0.18x (client alone) 37422473456; 37413386379

On shards 2 to 6, the measured heavy members and the turbo --concurrency=4 capacity bound (the sum of a shard's task windows is at most 4x its test step) give these results:

  • Proven under 1.5x: 8 shard-run pairs (for example shard 3 in 37427862594, at most 0.82x).
  • Undecided everywhere else. No measured package read 1.5x or more in any of the 9 near-full runs.
  • Highest package reading: @objectstack/spec, 1.45x / 1.39x / 1.43x on the three runs that executed it (37433381795, 37453598388, 37467882762).
  • Nearest the bound: shard 2/6. spec carries 70% of its prediction, so its 11 unmeasured members would have to average 1.61x to 1.77x before the shard reached 1.5x.

Stop condition: not met. No executed-window reading put any shard at or over 1.5x, so the red is wired as ruled. The per-shard numbers, the bounds and this PR's own run are on #16465 in the os-dev-report.

Verification (at 430859c2f)

  • Derived gates: node scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack --ran … printed "59 derived famil(ies) accounted for — 59 run, 0 NOT-MEASURED (a DERIVED zero — all 59 recorded an exit code and none of them is 3)". Every gate exited 0.
    • Four of them (check:dts-closure, check:dual-build-cjs-loads, check:lean-entry-closure, check:sourcemap-no-sources-content) first answered exit 3 (PREREQUISITE NOT MET). They exited 0 once a workspace build had run through the verify lock: 73 tasks, VERDICT command-exit 0.
  • node scripts/partition-test-shards.mjs --self-test → self-test OK (72 measured packages -> 72 shard items, 6 shards, max/mean 1.04x …).
  • --check-drift on synthetic summaries: 1.06x gives OK, exit 0; 1.34x gives WARN plus one ::warning line, exit 0; 1.54x gives DRIFT, exit 1. A replay sits beside each, and they are excluded.
  • The step's bash, run against an empty directory, gave NOT MEASURED and exit 0. Against one summary it gave the verdict.
  • check-governed-merges --test on the final two paths: "0 of 2 path(s) hit the register … NOT governed", with 494 changed lines.

Acceptance notes

  • Shards 2 to 6 still lack a whole-shard reading from main. This step is what produces one. The first scheduled runs after landing print it per shard, and the WARN annotations show where.
  • This PR's own Test Core run is a thin sample.
    • Its affected set, simulated with turbo ls --affected and the cross-package union at the merge base, is @objectstack/spec, @objectstack/client and @objectstack/driver-sql.
    • spec's test leg is expected to replay, which makes it NOT MEASURED. client and driver-sql run alone.
    • So it exercises the wiring rather than the dataset.
  • Expect the warning on shard 2/6 of cold full runs. spec read 1.39x to 1.45x against its single-run 1134.86s entry. That is the warning doing its job, and the remedy is the refresh lane, not this step.
  • Small-sample hazard, observed and not filed. driftReport() has no floor on predicted seconds. A PR whose shard executes only a few-second package gets a ratio of that package alone. Every lone-package reading found here was under 1.0x, so this is noted, not filed.
  • The 24-minute (80%) revert criterion is this PR's choice, stated with its reason in the comment. The current 45 is already at 79% on shard 1/6's slowest reading (35m43s).

Generated by Claude Code

claude added 3 commits October 6, 2026 13:39
….5x red

WARN_MEASURED_OVER_PREDICTED sits beside MAX_MEASURED_OVER_PREDICTED: the same
ratio over the same executed windows, an annotation-only tier under the red.
renderDriftVerdict() returns the four verdicts without printing them, so the
self-test reads the ::warning line a runner receives.

Claude-Session: https://claude.ai/code/session_01VF48aw8RPG6wzDnMgp6rtw
Co-authored-by: Claude <noreply@anthropic.com>
…n pair

Runs partition-test-shards.mjs --check-drift over .turbo/runs/*.json on every
shard, with no if: and no continue-on-error, so a red past 1.5x measured/
predicted withholds the attestation; past 1.3x it annotates. Restates the
comments the wiring made stale: the drift block's CLI figures, WHY SIX (the
CLI is the heaviest indivisible suite on the refreshed dataset), the 45-minute
timeout's revert condition against measured shard wall time, and the idle
slice-leg prose.

Claude-Session: https://claude.ai/code/session_01VF48aw8RPG6wzDnMgp6rtw
Co-authored-by: Claude <noreply@anthropic.com>
@objectstack-fleet objectstack-fleet Bot added the skip-changeset PR has no user-facing published change; bypasses the changeset gate label Oct 6, 2026
@objectstack-fleet objectstack-fleet Bot added the skip-changeset PR has no user-facing published change; bypasses the changeset gate label Oct 6, 2026
@objectstack-fleet
objectstack-fleet Bot marked this pull request as ready for review October 6, 2026 15:01
@objectstack-fleet
objectstack-fleet Bot added this pull request to the merge queue Oct 6, 2026
Merged via the queue into main with commit 9c3bec0 Oct 6, 2026
43 checks passed
@objectstack-fleet
objectstack-fleet Bot deleted the claude/issue-16465-wire-shard-drift-check branch October 6, 2026 15:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci/cd size/m skip-changeset PR has no user-facing published change; bypasses the changeset gate

Projects

None yet

2 participants