Skip to content

Commit 4fe161b

Browse files
committed
ci: accumulate runs until the workspace is covered, instead of hunting for one that covers it
The lane's first live run (34083991141) measured the fact the documented refresh procedure does not state: NO single green run measures the whole workspace. Of the seven retained green push runs on main, the best measured 52 of the 71 packages the committed dataset holds; the rest measured 2, 3, 13, 18, 22 and 49. All seven were partial cache replays, so every candidate was rejected and the job exited 1 having regenerated nothing. That is the cache design rather than luck: turbo's key is namespaced per shard and only main pushes write it, so a package whose inputs have not changed is a HIT -- and the generator refuses hits rather than recording a replayed ~0.1s window as a suite's cost. "Download six artifacts from any green run" therefore measures a SLICE of the workspace. Runs are now accumulated until coverage is satisfied, each fenced by its own `--run <id>` group. That grouping is exactly what the #16473 fix on this branch added: slices are summed WITHIN a run and the per-run sums are medianed ACROSS runs, so a package measured by three of the accumulated runs gets the median of three observations -- the property the dataset's merge rule always claimed and could not previously deliver. Feeding several runs without that grouping is refused by name, so this path could not have been taken before the rider. The argument list is rebuilt from the accepted set each round rather than appended to, so a run the generator refuses is dropped cleanly instead of poisoning every later attempt. The PR body and commit message now name every contributing run; the newest is the one the refresh is dated from. When even the whole accumulation falls short the job still refuses loudly, and its message names the remedy that would be a DECISION rather than a tuning knob: keeping the last measured weight for a package this refresh did not measure. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019RfFHiRCSs3JXLK4cwcfox
1 parent ea96c27 commit 4fe161b

2 files changed

Lines changed: 95 additions & 31 deletions

File tree

‎.github/workflows/shard-timings-refresh.yml‎

Lines changed: 90 additions & 27 deletions
Original file line numberDiff line numberDiff line change
@@ -50,13 +50,29 @@
5050
# regeneration happens where the data already is, on a timer, and no seat needs
5151
# reachability it does not have.
5252
#
53-
# CHOOSING THE RUN IS THE HARD PART — see scripts/ci/select-shard-timings-run.mjs
53+
# CHOOSING THE RUNS IS THE HARD PART — see scripts/ci/select-shard-timings-run.mjs
5454
# ------------------------------------------------------------------------------
5555
# "The newest run" is wrong three different ways here (cancelled runs, cache
5656
# replays, expired artifacts), and a refresh built on the wrong run is worse than
5757
# no refresh because it stamps a fresh `measuredAt` on numbers nobody measured.
5858
# That script carries the argument and the self-test; this file only drives it.
5959
#
60+
# ⚠ AND IT IS "RUNS", PLURAL, WHICH THE DOCUMENTED PROCEDURE DOES NOT SAY.
61+
# Measured on this lane's first live run (34083991141): of the seven retained
62+
# green push runs on main, the BEST measured 52 of the 71 packages the committed
63+
# dataset holds, and the others measured 2, 3, 13, 18, 22 and 49. All seven were
64+
# partial cache replays. That follows from the cache design rather than from luck
65+
# — turbo's key is namespaced per shard and only main pushes write it, so a
66+
# package whose inputs have not changed is a HIT, and the generator refuses hits
67+
# rather than recording a replayed ~0.1s window as a suite's cost. "Download six
68+
# artifacts from any green run" therefore measures a SLICE of the workspace.
69+
#
70+
# So the regeneration step accumulates runs, each under its own `--run <id>`
71+
# group — the grouping #16473 added, which sums a sliced package's slices within
72+
# a run and medians the per-run sums across runs. A package seen in three of the
73+
# accumulated runs gets the median of three observations, which is the property
74+
# the dataset's own merge rule always claimed and could not previously deliver.
75+
#
6076
# THE PINS, AND THE ONE THING A MACHINE MUST NOT DECIDE
6177
# ----------------------------------------------------
6278
# partition-test-shards.mjs `--self-test` grades the dataset against the
@@ -214,11 +230,30 @@ jobs:
214230
for (const r of runs) console.log(` ${r.run_id} ${r.created_at} ${r.head_sha.slice(0, 10)}`);
215231
'
216232
217-
# Each eligible run is tried in turn: download, generate, and then judge
218-
# whether it MEASURED the workspace. A run that passed selection can still
219-
# be a cache replay, and the only honest test for that is the resulting
220-
# dataset's package coverage — so the loop is the point, not a retry.
221-
- name: Regenerate the dataset from the first run that actually measured
233+
# ACCUMULATE runs until the workspace is covered — do not look for one run
234+
# that covers it, because there is no such run.
235+
#
236+
# MEASURED on the first live run of this lane (run 34083991141), and it is
237+
# the fact that shapes this step: of the seven retained green push runs on
238+
# main, the BEST measured 52 of the 71 packages the committed dataset
239+
# holds, and the rest measured 2, 3, 13, 18, 22 and 49. Every one of them
240+
# was a partial cache replay. That is not bad luck, it is the cache design:
241+
# turbo's key is namespaced per shard and only main pushes write it, so a
242+
# package whose inputs have not changed is a HIT — and the generator
243+
# refuses hits rather than recording a replayed ~0.1s window as a suite's
244+
# cost. "Download six artifacts from any green run" therefore measures a
245+
# SLICE of the workspace, never all of it.
246+
#
247+
# So runs are accumulated. Each contributes its six summaries under its own
248+
# `--run <id>` group, which is exactly the grouping #16473 added: slices are
249+
# summed within a run and the per-run sums are medianed across runs, so a
250+
# package measured by three of these runs gets a median of three
251+
# observations rather than whichever run happened to be read last. Feeding
252+
# several runs WITHOUT that grouping is refused by the generator by name.
253+
#
254+
# The loop stops at the first accumulation that covers the workspace, so a
255+
# quiet week costs one run's download and a busy one costs a few.
256+
- name: Regenerate the dataset, accumulating runs until the workspace is covered
222257
id: generate
223258
env:
224259
GITHUB_TOKEN: ${{ github.token }}
@@ -227,6 +262,7 @@ jobs:
227262
WORK="$RUNNER_TEMP/refresh"
228263
mkdir -p "$WORK"
229264
CHOSEN=''
265+
ACCEPTED=()
230266
231267
RUN_COUNT=$(node -e 'console.log(JSON.parse(require("fs").readFileSync(process.env.RUNNER_TEMP + "/candidates.json", "utf8")).length)')
232268
for i in $(seq 0 $((RUN_COUNT - 1))); do
@@ -265,20 +301,29 @@ jobs:
265301
echo "::endgroup::"; continue
266302
fi
267303
268-
# One run, so no `--run` grouping is needed: every summary here came
269-
# from run $RUN_ID. Feeding SEVERAL runs would require `--run <id>`
270-
# before each run's six files — the generator refuses the ambiguity
271-
# rather than taking the last value (#16473).
272-
# An array, not `xargs`: a split invocation would run the generator
273-
# twice and the second would overwrite the first's output with a
274-
# partial dataset.
275-
mapfile -t SUMMARY_FILES < "$SUMDIR/summaries.txt"
276-
if ! node scripts/measure-test-shard-timings.mjs "${SUMMARY_FILES[@]}" \
304+
ACCEPTED+=("$RUN_ID")
305+
306+
# The argument list is REBUILT from the accepted set every round
307+
# rather than appended to, so a run the generator refuses can be
308+
# dropped cleanly instead of poisoning every later attempt. Each run
309+
# is fenced by its own `--run <id>`; an array, not `xargs`, because a
310+
# split invocation would run the generator twice and the second would
311+
# overwrite the first's output with a partial dataset.
312+
ARGS=()
313+
for R in "${ACCEPTED[@]}"; do
314+
ARGS+=( --run "$R" )
315+
mapfile -t RUN_FILES < "$WORK/$R/summaries.txt"
316+
ARGS+=( "${RUN_FILES[@]}" )
317+
done
318+
319+
if ! node scripts/measure-test-shard-timings.mjs "${ARGS[@]}" \
277320
--out "$WORK/refreshed.json"; then
278-
echo "::warning::The generator refused run $RUN_ID's summaries; skipping this run."
321+
echo "::warning::The generator refused the set including run $RUN_ID; dropping that run and continuing."
322+
unset 'ACCEPTED[-1]'
279323
echo "::endgroup::"; continue
280324
fi
281325
326+
echo "Accumulated ${#ACCEPTED[@]} run(s): ${ACCEPTED[*]}"
282327
if node scripts/ci/select-shard-timings-run.mjs --check-coverage \
283328
--committed scripts/test-shard-timings.json \
284329
--refreshed "$WORK/refreshed.json" \
@@ -292,11 +337,18 @@ jobs:
292337
done
293338
294339
if [ -z "$CHOSEN" ]; then
295-
echo "::error::No eligible run produced a complete measurement of the workspace. Every candidate was a cache replay, refused by the generator, or lost its artifacts. NOTHING was regenerated and no PR was opened — this is a refusal, not a quiet success."
340+
echo "::error::The ${#ACCEPTED[@]} eligible run(s) on main, accumulated together, still do not measure every package the committed dataset holds — the shortfall above names what is missing. A package can stay unmeasured across every retained run: turbo's cache key is namespaced per shard and only main pushes write it, so a package whose inputs have not changed is a HIT in all of them, and the generator refuses hits rather than recording a replay as a duration. NOTHING was regenerated and no PR was opened — this is a refusal, not a quiet success. If this is the steady state rather than a quiet week, the dataset needs a merge rule (keep the last measured weight for a package this refresh did not measure) rather than a wider run window; that is a decision, not a tuning knob, and it is tracked on the card."
296341
exit 1
297342
fi
298-
echo "run_id=$CHOSEN" >> "$GITHUB_OUTPUT"
299-
echo "head_sha=$(node -e 'const rs=JSON.parse(require("fs").readFileSync(process.env.RUNNER_TEMP+"/candidates.json","utf8"));console.log(rs.find(r=>String(r.run_id)===process.argv[1]).head_sha)' "$CHOSEN")" >> "$GITHUB_OUTPUT"
343+
# Candidates arrive newest-first, so ACCEPTED[0] is the most recent run
344+
# in the set and its commit is the one the refresh is dated from. The
345+
# full list travels alongside it: every run in it contributed
346+
# measurements, and the PR body names them all rather than implying one.
347+
NEWEST="${ACCEPTED[0]}"
348+
echo "run_id=$NEWEST" >> "$GITHUB_OUTPUT"
349+
echo "runs=${ACCEPTED[*]}" >> "$GITHUB_OUTPUT"
350+
echo "run_count=${#ACCEPTED[@]}" >> "$GITHUB_OUTPUT"
351+
echo "head_sha=$(node -e 'const rs=JSON.parse(require("fs").readFileSync(process.env.RUNNER_TEMP+"/candidates.json","utf8"));console.log(rs.find(r=>String(r.run_id)===process.argv[1]).head_sha)' "$NEWEST")" >> "$GITHUB_OUTPUT"
300352
301353
# Byte comparison, and it decides everything downstream. `cmp -s` rather
302354
# than a diff of parsed JSON: the committed artefact is the file, so the
@@ -343,6 +395,8 @@ jobs:
343395
if: steps.compare.outputs.changed == 'true'
344396
env:
345397
RUN_ID: ${{ steps.generate.outputs.run_id }}
398+
RUNS: ${{ steps.generate.outputs.runs }}
399+
RUN_COUNT: ${{ steps.generate.outputs.run_count }}
346400
HEAD_SHA: ${{ steps.generate.outputs.head_sha }}
347401
PARTITIONER_EXIT: ${{ steps.bins.outputs.partitioner_exit }}
348402
USED_PAT: ${{ secrets.RELEASE_PUSH_TOKEN != '' }}
@@ -364,14 +418,23 @@ jobs:
364418
echo
365419
echo "## Source"
366420
echo
367-
echo "- Run: https://github.com/$GITHUB_REPOSITORY/actions/runs/$RUN_ID"
368-
echo "- Commit measured: \`$HEAD_SHA\`"
369-
echo "- Selected because its six \`Test Core (N/6)\` jobs all concluded \`success\`, its six"
370-
echo " run-summary artifacts were still retained, and the dataset it produced measures every"
371-
echo " package the committed one measured that is still in the workspace (the check that"
372-
echo " rejects a cache replay)."
421+
echo "Measured across $RUN_COUNT accumulated run(s). No single green run measures the whole"
422+
echo "workspace — turbo's cache is namespaced per shard and only main pushes write it, so a"
423+
echo "package whose inputs have not changed is a HIT and the generator refuses hits rather"
424+
echo "than recording a replay as a duration. Runs are therefore accumulated, each fenced by"
425+
echo "its own \`--run\` group, until every package the committed dataset holds is measured"
426+
echo "again; a package seen in several of them gets the median of those observations."
427+
echo
428+
for R in $RUNS; do
429+
echo "- https://github.com/$GITHUB_REPOSITORY/actions/runs/$R"
430+
done
431+
echo
432+
echo "- Newest run in the set: \`$RUN_ID\`, commit \`$HEAD_SHA\` — the date this refresh carries."
433+
echo "- Every run above had all six \`Test Core (N/6)\` jobs conclude \`success\` with its six"
434+
echo " run-summary artifacts still retained; runs that were cancelled, failed or had lost"
435+
echo " their artifacts were rejected by name in the log before any of these were used."
373436
echo
374-
echo "## Measured per-shard suite time on that run"
437+
echo "## Measured per-shard suite time on the newest run in the set"
375438
echo
376439
echo "\`\`\`"
377440
echo "$SHARDS"
@@ -431,7 +494,7 @@ jobs:
431494
# heredoc would carry its own leading whitespace into the message body.
432495
git commit \
433496
-m "chore(ci): refresh the Test Core shard-timings dataset" \
434-
-m "Regenerated by .github/workflows/shard-timings-refresh.yml from the six test-core-run-summary artifacts of run $RUN_ID ($HEAD_SHA). Generated, never hand-edited."
497+
-m "Regenerated by .github/workflows/shard-timings-refresh.yml from the test-core-run-summary artifacts of $RUN_COUNT accumulated run(s) ($RUNS), newest $RUN_ID at $HEAD_SHA. Generated, never hand-edited."
435498
git push origin "$BRANCH"
436499
437500
PR_URL=$(gh pr create --base main --head "$BRANCH" \

‎scripts/ci/select-shard-timings-run.mjs‎

Lines changed: 5 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -549,10 +549,11 @@ async function main() {
549549
if (!report.ok) {
550550
console.error(
551551
`select-shard-timings-run: COVERAGE SHORTFALL -- ${report.lost.length} package(s) the committed ` +
552-
'dataset measured, and the workspace still contains, are NOT measured by this refresh: ' +
553-
`${report.lost.join(', ')}. Every one of them would silently fall back to the test-file-count ` +
554-
'ESTIMATE, so this run measured a warm cache rather than the workspace. Rejecting it and trying ' +
555-
'an older run.'
552+
'dataset measured, and the workspace still contains, are NOT measured by the runs accumulated ' +
553+
`so far: ${report.lost.join(', ')}. Every one of them would silently fall back to the ` +
554+
'test-file-count ESTIMATE, so what has been read so far is a warm cache rather than the ' +
555+
'workspace. Not a verdict on any one run: no single run measures everything (turbo caches per ' +
556+
'shard, and the generator refuses hits), so the caller adds the next older run and asks again.'
556557
);
557558
process.exit(1);
558559
}

0 commit comments

Comments
 (0)