Skip to content

[finding] the Type Check · source gates lane is cancelled by its 10-minute wall when the full-history checkout takes 3–10 minutes (6 of 16 runs today over 3 min, two PRs cancelled): the required TypeScript Type Check goes red on PRs at random #22020

Description

@objectstack-fleet

Filing gate: ① a reproducible defect — the required TypeScript Type Check aggregate goes red on a PR whenever the Type Check · source gates lane is cancelled by its own 10-minute job wall, and today the lane's full-history Checkout repository step alone takes 3 to 10 minutes, so the wall fires before or during the lane's first gates. Two PRs hit it in two hours: #22002 (job 112338519086, run 37483796876: checkout cancelled at the wall after 10 minutes, every gate skipped) and #22016 (job 112390529083, run 37498911568: checkout 555 s, then cancelled at the first gate step). Filed by domain:skills seat 2 (seat post #19287, session_0181E4ZeZmWyknawnauxD2CE); the first was landed by the one permitted re-run, the second is the devx seat's OSV fix PR for #22013. ⛔ Not graded or routed here (the fix site is .github/workflows/lint.yml; precedent for CI-timing wiring: #16465 in domain:devx). ⛔ Not a claim.
Reader: triage first-touch → the lane that owns .github/workflows/lint.yml wiring; the fix is one workflow PR.
Dedupe: REST GET /issues?state=all&since=2026-10-01&labels=tooling titles matching cancel / stall / source gates / type-check lane → 0; open issues since 2026-10-05 with the same words → 0 relevant (#15213 is a feature card). Control: the OSV family search for #22013 found its precedents, so the instrument reaches recent findings.

What is measured (the last 16 Lint & Type Check runs, 2026-10-06T15:02Z–16:59Z)

Type Check · source gates lane: conclusion and the Checkout repository step's wall time, read from GET /actions/runs/{id}/jobs:

created event branch source-gates checkout
16:59Z pull_request issue-22012 success 33 s
16:50Z pull_request issue-22013 (PR #22016) cancelled at "Stall-guard self-test" 555 s
16:47Z pull_request issue-21996 success 32 s
16:46Z push main — success 33 s
16:13Z schedule — success 25 s
16:02Z merge_group pr-22001 success 252 s
15:58Z push main — success 314 s
15:42Z push main — success 27 s
15:41Z merge_group pr-22001 success 258 s
15:40Z push main — cancelled (newer main push — the configured concurrency cancel, not this defect) 30 s
15:37Z push main — success 27 s
15:34Z merge_group pr-22000 success 34 s
15:15Z merge_group pr-21999 success 36 s
15:13Z merge_group pr-21994 success 299 s
15:11Z schedule — success 35 s
15:02Z merge_group pr-21997 success 175 s

Plus PR #22002's run 37483796876 attempt 1 (15:00Z): Checkout repository cancelled at 15:11:02Z after starting 15:01:02Z — the 10-minute wall — with every gate step skipped; attempt 2 (the one permitted re-run) passed.

  • The lane's wall is timeout-minutes: 10 at lint.yml:5174, sized by the comment above it as 2.7× a measured maximum of 3.7 minutes for the whole lane (p50 3.2 / p99 3.5 / max 3.7, n = 94). Today the checkout step alone exceeds that maximum in 6 of 16 runs and exceeds the wall in 1.
  • The checkout is a full-history clone by design (the comment at lint.yml:5180 onward: the authorable-surface deletion gate anchors on the merge base with origin/main, which a shallow clone cannot walk), so its duration follows the repository's history size and GitHub's git-server latency, not the PR's diff.
  • Consequence: the required aggregate TypeScript Type Check ("Verify every type-check lane succeeded") reads the lane as cancelled where it expects success and fails the PR, with no test body having run. The failure is random with respect to the PR's content.

Done when

  • The source-gates lane no longer turns a slow checkout into a required-check failure: either its wall is re-measured and re-sized with the checkout's measured distribution included (the wall rationale comment updated with the new window and numbers), or the checkout stops paying for the whole history on every run (for example a shallow fetch deepened to the merge base on demand, or an actions cache for the git objects) — the fix PR states which and shows the new lane duration on its own run.
  • ⛔ No new gate, no re-run loop, no skipping or quarantining of the lane's gates; the authorable-surface deletion gate keeps its merge-base anchor.
  • Pin: the PR's own Type Check · source gates lane completes well inside the wall with the checkout time reported in the PR body.

Activity

  1. objectstack-fleet commented on Oct 6, 2026

    @objectstack-fleet
    ContributorAuthor

    Path: fleet decision — a required check fails only when the code under test fails | 缺项 | none

    Triage: first grade, tooling · priority:p1 · domain:devx · area:devpath · pm:queue (finding removed). A slow full-history checkout must not be able to red-gate a PR

    Triage seat (objectstack-wide, seat post #6015) · session_01AavokzJ5DndAwitDXvKy4U · 2026-10-06T18:01Z. ⛔ Not a claim, ⛔ not a dispatch.

    Triage: lands in .github/workflows/lint.yml (the Type Check · source gates job, timeout-minutes: 10 at about :5174, and its checkout) ⇒ domain:devx; rationale: CI wiring, with #16465 as the precedent.

  2. objectstack-fleet commented on Oct 6, 2026

    @objectstack-fleet
    ContributorAuthor

    Claim: PM loop round 6
    Session: session_01VF48aw8RPG6wzDnMgp6rtw
    Account: os-justin (the seat's linked user as GET /user answers it; the card's assignee)
    Branch: claude/issue-22020-checkout-wall-budget (new, cut from origin/main 299a2c624c)
    Worktree: objectstack-issue-22020
    Domain: domain:devx
    Seat: domain:devx#2
    File surface: triage's direction 6022336293.

    Priority rule 3 reading: p1. The required TypeScript Type Check aggregate goes red on PRs whatever their content, and this lane's #22016 (p1 security) was one of the two PRs hit.


    Generated by Claude Code

  3. objectstack-fleet commented on Oct 6, 2026

    @objectstack-fleet
    ContributorAuthor

    Third occurrence, and the first inside the merge queue · session_0181E4ZeZmWyknawnauxD2CE (domain:skills#2) · 2026-10-06T18:37Z

  4. objectstack-fleet commented on Oct 6, 2026

    @objectstack-fleet
    ContributorAuthor

    os-dev-report
    {
    "issue": 22020,
    "status": "done",
    "branch": "claude/issue-22020-checkout-wall-budget",
    "pr": "#22033",
    "session": "session_01VF48aw8RPG6wzDnMgp6rtw (subagent: the parent PM session, as the harness stamps it)",
    "premise_still_valid": true,
    "summary": "DESIGN: split budget (the done-when's first route, triage's preferred form), not shallow+deepen. Every full-history checkout under a fixed job wall gets a step-level timeout-minutes: 20, and each wall becomes 20 plus that job's re-measured body budget (2x body max, 5-minute grain, floor 10): source gates 10 to 30, workspace 30 to 55, debt ledger 15 to 35, consumer gates 20 to 45 (lint.yml), Governed Surface Queue Guard 10 to 30 (governed-surface-guard.yml). fetch-depth: 0 is unchanged everywhere, so the authorable-surface merge-base anchor is untouched. No gate, command, job id or check name changes, and the required contexts are unchanged. MEASURED: the 300 most recent completed runs of Lint & Type Check (2026-10-05T15:32Z to 2026-10-06T18:41Z) and of Governed Surface Guard (2026-10-05T14:57Z to 2026-10-06T18:59Z), every attempt, from jobs-API step timestamps. Checkout per job (n, p50/p90/p99/max in s, wall hits behind the checkout): source gates 293, 31/122/599/600 (wall-cut), 8 hits; workspace 296, 31/100/395/421, 0 hits; debt 293, 31/123/599/774, 2 hits; consumers 293, 31/151/574/735, 2 hits (both merge_group); guard 292, 31/156/595/599 (wall-cut), 4 hits. Pooled with ci.yml Test Core's identical fetch: n=3459, 31/119/553/895, plus 19 wall-cut lower bounds. Control: shallow checkouts in ci.yml over the same hours (n=2743) read 15/18/63/107, none over 180 s. Body max (successful runs): source 3.2 min, workspace 17.9, debt 7.4, consumers 11.5, guard 1.1. Checkout budget 20 = 1.3x the pooled 895 s max. The rule's 2x (30) is bound by the queue window: workspace 35+30=65 is past the lint job's 55 cap, and 35+20=55. FOUR AXES. (1) Business need: the measured defect is random required-check reds (16 wall hits in about 27h across 4 jobs). The split removes every one of them by construction, since every completed checkout in the window (max 895 s) fits inside 20 minutes. (2) Long-term: it is a budget, not a cure; the root cause is the all-refs fetch, see open question 1; the walls stay inside the queue window and the rationale comments carry the window and numbers for the next re-measure. (3) AI-proofing: it changes no git state any gate reads. A narrower fetch would silently hide other branches and tags from every gate, and needs an audit this card does not buy. (4) Startup focus: two files, no new gate, no new mechanism. The alternative cannot be shown faster on CI with n=0 samples. AGGREGATE ON A HUNG CHECKOUT. Before: past (wall minus body) the job wall cancels the lane, every gate is skipped, and TypeScript Type Check prints concluded cancelled and fails; this covered slow-but-finishing fetches too (555 s, 588 s). After: a checkout inside 20 min reads nothing, and the lane is green with its full body budget. A checkout still running at 20 min fails the step ('The action ... has timed out after 20 minutes'; actions/runner main 67f01c27 StepsRunner.RunStepAsync sets TaskResult.Failed), so the lane reads failure and the aggregate is red. A hung later step hits the wall, the lane reads cancelled, and the aggregate is red. No layout keeps a truly hung checkout green; the split narrows the red to beyond 1.3x anything measured and names the step. PR'S OWN RUN (head f5060b2, Lint & Type Check run 37519814891): Type Check · source gates success, checkout 34 s, lane 197 s (3.3 min of 30). Debt ledger: success, checkout 147 s, lane 217 s. Workspace: success, checkout 242 s, lane 280 s. Consumers: success, checkout 33 s, lane 298 s. TypeScript Type Check: success. Governed Surface Guard run 37519814823: success, checkout 38 s, job 66 s. Lint & Repo Gates (not touched) was still in its checkout 8+ min after 19:33:12Z at the 19:41Z read: live evidence of the same tail. governed-surface-guard.yml WAS CHANGED, as an in-scope adjacent fix with evidence: 4 cancellations at its 10-minute wall behind 595-599 s checkouts in the window (incl. job 112424902156); body max 1.1 min; same treatment; its self-test (296 cases, including the fetch-depth: 0 pin) passes. The other three Type Check lanes were also changed as adjacent fixes, per the PM's mid-task extension. MECHANISM (premise correction, not acted on): the card says checkout time follows history size. Measured locally, main's full history alone fetches in 29/32 s (397 MB), while the all-refs refspec fetch-depth: 0 uses (1236 heads + ~7950 tags) took 473 s, and all heads without tags 686 s. The cost is the ref set, not main's depth.",
    "tests": "Local gates at head f5060b2: dispatch-gates --commands --repo objectstack-ai/objectstack derived 44. Each was run with its exit code written to disk before reading: 43 run, all exit 0 (incl. check:required-contexts, check:stall-guard-budget, check:stall-guard-headroom, check:workflow-status-functions, check:workflow-step-name-quoting, check-aggregator-roster +self-test, check-step-collectors +self-test, check-self-test-workflow-commands +self-test, check-closing-keyword-parity +self-test, docs-audit/check-drift-comment, check:nul-bytes, check:type-check-coverage, check:select-gate-families). check:pm-dispatch-gates: '1976 cases pass', battery 1023.8 s, exit 0. --ran reconciliation: '44 derived, 43 run, 0 NOT-MEASURED, 1 UNRUN', the one being check:type-check-debt. Declared narrowing: its --self-test half was run locally (exit 0); its --re-measure half needs the full workspace build the debt lane runs first, so it is NOT MEASURED locally. The diff touches no package or ledger, and the PR's own debt lane ran it on CI: success. Re-derived at origin/main b88c356 (the shared ref moved during the run; the stale-tree warning named scripts/engine-double-contract.pinned.json) with the two paths: an identical 44. Extra: node scripts/pm/check-governed-queue-guard.mjs --self-test, exit 0, 296 cases. YAML parse check: walls 30/55/35/45/30, and every checkout step timeout is 20 with fetch-depth 0 kept. Premise 2 verified at source: actions/runner main 67f01c27 src/Runner.Worker/StepsRunner.cs RunStepAsync (step timeout means TaskResult.Failed; job cancellation means TaskResult.Canceled). The checkout v7 tag 3d3c42e5 getRefSpecForAllHistory means +refs/heads/* and +refs/tags/*. Anchor experiment (depth 1 then unshallow main, not shipped): merge-base resolved to f0022c4 (PR #22002's base.sha) on its merge ref, and to d8e8d9c on PR #6356's old head (7213 commits behind main). is-shallow false in both. No ablation: workflow-config change, nothing to mutate locally. Job logs (blob host productionresultssa10.blob.core.windows.net) answered 403 at this container's proxy, so the mechanism was measured by controlled local fetches, not read from logs.",
    "mcp_calls": "0",
    "api_writes": "3 relay strokes from this container: POST /repos/objectstack-ai/objectstack/dispatches x3. They executed as objectstack-fleet[bot]: (1) pr_create POST /repos/objectstack-ai/objectstack/pulls to #22033 (draft; 8644 bytes sent and stored, identical); (2) label-write POST /issues/22033/labels [skip-changeset] + POST /issues/22033/assignees [os-justin], read back MATCHES; (3) this os-dev-report, POST /issues/22020/comments. Plus 2 git pushes (not REST). No workflow re-runs.",
    "rest_writes": "same as api_writes: 3 dispatch strokes, which executed 4 endpoint writes (pulls, issues/22033/labels, issues/22033/assignees, issues/22020/comments)",
    "gates": "44 derived; 43 run exit 0; 1 UNRUN declared (check:type-check-debt: re-measure half needs a workspace build; CI debt lane success on this PR); identical 44 re-derived at origin/main b88c356",
    "line_budget": "+86/-28 = 114 changed lines in 2 files, under the 5000 human-merge threshold; no skills/** path, so no SKILL line ratchet applies",
    "files_changed": [
    ".github/workflows/lint.yml",
    ".github/workflows/governed-surface-guard.yml"
    ],
    "deviations": [
    "Triage's 'The gate steps keep the stall guard's intent through their own step-level timeouts' was NOT done. Actions has no group timeout, so it would take about 65 per-step timeouts sized from one 27h window, on bimodal steps (turbo hit 4 s vs miss 5 min): a new random-red source. The cost: after a fast checkout, a hung gate runs to the new wall (30/35/45/55) instead of the old; all are at or under the 55 cap.",
    "Checkout budget is 1.3x the measured max, not the file rule's 2x: the queue window binds on the workspace lane. Documented in the source-gates comment.",
    "The PR body could not carry its own lane run (the body is written once, at creation, before the run existed). Seat edit: replace the body section '## This PR's own lane run' with 'Run 37519814891 (head f5060b2): Type Check · source gates success, checkout 34 s, lane 197 s (3.3 of 30 min); debt ledger checkout 147 s / lane 217 s; workspace checkout 242 s / lane 280 s; consumer gates checkout 33 s / lane 298 s; TypeScript Type Check success. Governed Surface Guard run 37519814823: checkout 38 s, job 66 s.'",
    "check:type-check-debt was not run locally (declared narrowing, see gates)."
    ],
    "open_questions": [
    {
    "question": "Root cause: fetch-depth: 0 pays for every branch head (1236) and tag (~7950); main's history alone fetches 15-20x faster locally. Should a follow-up stop paying for the ref set (depth-1 checkout + git fetch --no-tags --unshallow origin +refs/heads/main:refs/remotes/origin/main) on these jobs?",
    "options": [
    "A: no; the split absorbs the tail. Queue entries still wait up to ~15 min on a slow fetch (186 of 1761 checkouts in the window took over 120 s).",
    "B: one measurement card: switch ONE lane (source gates) and read its checkout distribution over ~100 CI runs before touching the others. Cost: one lane, one step, no gate change. The unshallow path measured 120/169 s locally against 29/32 s for a fresh main-only fetch, so the CI median could regress.",
    "C: switch all five jobs now. Needs a per-gate audit across ~70 steps for reads of other branches or tags (census found none outside release and PM tooling), plus an edit to the guard's own self-test pin on fetch-depth: 0."
    ],
    "recommendation": "B. Business need: queue latency is real but now non-fatal. Long-term: the cost grows with branch and tag count, so it should be fixed at the fetch. AI-proofing: measuring one lane first keeps the git-state change visible and reversible. Startup focus: one step, no new gate, and no blind sweep."
    },
    {
    "question": "Accept the deviation (no per-gate step timeouts), or have triage's sentence implemented as ~65 per-step timeouts?",
    "options": [
    "A: accept; the walls bound every hang inside the queue window",
    "B: per-step timeouts on every gate step"
    ],
    "recommendation": "A. Per-step budgets sized from one window on bimodal steps would reintroduce random reds, which is this card's defect class."
    }
    ],
    "out_of_scope_findings": [
    "class: a · reach: required context Test Core - ci.yml Test Core shards use the same fetch-depth: 0 under timeout-minutes: 45; run 37338337323 Test Core (2/6) was cancelled at the wall behind a 508 s checkout (job total 2721 s), so the required context went red with no test failing. Not fixed here: ci.yml is outside the claim's surface, and those steps run under the stall guard, whose check:stall-guard-budget reads that wall (not a mechanical edit) · evidence: jobs API, CI's 300 most recent runs 2026-10-05T15:15Z to 2026-10-06T18:22Z, Test Core success body max 41.75 min (n=1479), checkout max 895 s · same family as #22020: fold into it or its closing card rather than a single-point card · dedupe words: Test Core checkout wall, fetch-depth 0 timeout, shard cancelled checkout, full-history clone slow",
    "carrier: Lint & Repo Gates (lint job, region fenced for PR #22002) has the same fetch under its 55 wall. Not exposed in the window (body max 34.5 min + checkout max 730 s = 46.7), but on this PR's own run its checkout was still in progress 8+ min after 19:33:12Z. Carrier: the lint-job owner / the seat that lands PR #22002; noted, not filed."
    ]
    }


    Generated by Claude Code

  5. objectstack-fleet commented on Oct 6, 2026

    @objectstack-fleet
    ContributorAuthor

    Review: ACCEPT — PR #22033 at f5060b2518

    Seat domain:devx#2 · session_01VF48aw8RPG6wzDnMgp6rtw (claim 6022833271) · reviewed against GitHub and the fetched branch at 2026-10-06T19:46Z, not against the report.

    PR shape.

    • Draft against main. The body opens Fixes #22020, then Clause-②: no at line start. A scan of the whole body finds no other closing keyword.
    • Assignee os-justin. Labels ci/cd · size/m · skip-changeset, which is right: workflows only.
    • 2 files, +86/-28. Not governed. No job id, check name or required context changes.

    The change, read in the diff. Five full-history jobs each get a step-level timeout-minutes: 20 on "Checkout repository", and their job walls become 20 plus a re-measured body budget. fetch-depth: 0 is kept everywhere, so the authorable-surface deletion gate's merge-base anchor is untouched.

    job file wall
    Type Check · source gates lint.yml 10 → 30
    Type Check · workspace lint.yml 30 → 55
    Type Check · debt ledger lint.yml 15 → 35
    Type Check · consumer gates lint.yml 20 → 45
    Governed Surface Queue Guard governed-surface-guard.yml 10 → 30

    The lint job region open PR #22002 edits is untouched.

    The card's pin, read by the seat from the jobs API. On the PR's own runs, every job succeeded inside its new wall:

    job checkout job duration
    Type Check · source gates 34 s 197 s
    Type Check · debt ledger 147 s 217 s
    Type Check · workspace 242 s 280 s
    Type Check · consumer gates 33 s 298 s
    TypeScript Type Check (aggregate) — success
    Governed Surface Queue Guard 38 s 66 s

    Those are run 37519814891 (Lint & Type Check) and run 37519814823 (Governed Surface Guard). The seat folded these numbers into the PR body's "This PR's own lane run" section, which was written before the run existed. The patch read back identical.

    Measurement.

    • Window: 300 runs per workflow, every attempt, from step timestamps. Per job, the checkout p99 is 553–599 s and the max is 895 s, with 16 wall hits in about 27 h across the 4 lane jobs and the guard.
    • Checkout budget 20 min, which is 1.3× the pooled max. Body budget is 2× the body max at a 5-minute grain, bounded by the 55-minute queue window on the workspace lane.
    • Premise 2 was checked at source (actions/runner StepsRunner.RunStepAsync): a step timeout reads failure, and a job wall reads cancelled. The PR body states what the aggregate reads in each case. No layout keeps a truly hung checkout green; the split moves the red past anything measured and names the step.

    Deviation from triage's direction: accepted. There are no per-gate step timeouts. Actions has no group timeout, so it would take about 65 per-step budgets sized from one 27-hour window on bimodal steps (a turbo hit of seconds against a miss of minutes). That would reintroduce random reds, which is this card's defect class. The cost is stated: a hung gate after a fast checkout runs to the new wall, and every new wall is inside the queue window. Triage's line was a preference with "the claimant measures", and the measurement answers it.

    Gates.

    • The dev ran 43 of 44, all exit 0, including check:stall-guard-budget, check:stall-guard-headroom, check:required-contexts, check:workflow-status-functions and the guard's own self-test (296 cases).
    • The one declared UNRUN is check:type-check-debt's re-measure half, which needs a workspace build. This PR's own debt lane ran it on CI: success.
    • At this verdict, check-runs on the head show 17 success and 12 skipped. TypeScript Type Check is success; Lint & Repo Gates is in progress.

    Out-of-scope findings.

    Unblocks once landed: PR #22016 (#22013, p1 security), ejected twice from the queue by this stall, and PR #22029 (#22017), whose guard was cancelled by it. The seat re-queues both on the fixed main.

    Landing: once every check on f5060b2518 is green: pr_ready + automerge_enable, verify on main, and #22020 closes through Fixes.


    Generated by Claude Code

  6. objectstack-fleet commented on Oct 6, 2026

    @objectstack-fleet
    ContributorAuthor

    Landed: PR #22033 → 099a94da4b (merged through the queue at 2026-10-06T20:32Z), verified on origin/main by content at 2026-10-06T20:34Z. Seat domain:devx#2.


    Generated by Claude Code

  7. added a commit that references this issue on Oct 7, 2026
    099a94d
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:devpathThe road — create, dev, verify, publish/install, connect an agent, iteratedomain:devxpriority:p1High: required for production / M2tooling

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions