Repository navigation
π Estate CI-health backlog: public repos with independently-red CIΒ #464
Description
Activity
π Status β 2026-06-21: the Hypatia-lane root cause is a central
actions/cacheSHA corruption (already fixed & merged)Actioning the CI-health findings surfaced 2026-06-20/21. I have
hypatia+standardswrite access (the reporting session did not, so it could only diagnose). Cross-ref nextgen-typing#69 ("out of scope β central").(1) ROOT CAUSE β central
actions/cacheSHA corruption β β FIXED & MERGEDThe estate-wide
scan / Hypatia Neurosymbolic Analysisfailure at "Prepare all required actions":Unable to resolve action actions/cache@d4373f267a887d77f9eb0683a479ec60b1fe5b2b (unable to find version d4373f267a887d77f9eb0683a479ec60b1fe5b2b)β¦is, as diagnosed, not in any consumer workflow. It was pinned once, centrally, in both Hypatia reusables in
standards:.github/workflows/hypatia-scan-reusable.yml.github/workflows/governance-reusable.yml
d4373fβ¦resolves to nothing upstream β it's a corruption of v4.2.2's real commitd4323d4β¦.It was already repaired and merged before I reached it: standards#394 (merged 2026-06-21 10:52Z, commit
d72fe5a) re-pinned both reusables to the genuinev4.2.0commit1bd1e32aβ¦, preserving the# v4.2.0comment.Independently verified this session via
git ls-remote https://github.com/actions/cache:SHA upstream ref resolves? d4373fβ¦(the corrupt pin)(none) β bogus 1bd1e32aβ¦(the repair)refs/tags/v4.2.0β 0057852bβ¦("most common")v4+v4.3.0β 27d5ce7fβ¦(used across hypatia)main+v5+v5.0.5β git grep d4373fβ¦acrossstandards+hypatiaβ 0 matches. hypatia's own workflows are clean (they pinactions/cacheto the valid27d5ce7fβ¦ # v5.0.5).β οΈ Propagation caveat β necessary but not yet sufficientConsumers pin these reusables by standards commit SHA, not
@main(hypatia uses@5eb28d7dβ¦; most repos@861b5e91β¦). The repair landed as a new standards HEAD (d72fe5a), so a consumer still pinned at a pre-#394 SHA keeps dereferencing the broken cache pin until it is re-enrolled tod72fe5a+. A consumer's lane therefore only goes green after re-enrollment (gitbot-fleetenroll-repos) β not automatically.(2)
governance / Check Workflow Stalenessred β EXPECTED drift, not a new defectstandards/scripts/check-workflow-staleness.shfails any consumer whose pinned reusable SHA β current standards HEAD ("Workflow pins Hypatia reusable before cache/baseline-delay fix. Refresh to current standards SHA."). Because #394 advanced HEAD tod72fe5a, every consumer is stale by definition right now β this is precisely the signal that the re-enrollment pass is pending, the same root cause as (1) viewed from the propagation side.Remediation = gitbot-fleet
enroll-reposre-pin of consumers tod72fe5a+. That's out of scope forstandards/hypatia(consumer repos and gitbot-fleet aren't in my access), so per the task I'm recording it here as expected post-#394 drift.(3)
nextgen-databases, pre-existing & repo-internal β recorded (out of scope to fix)Both live in
nextgen-databases(notstandards/hypatia), so flagging with grounded fixes rather than touching them:- K9 pedigree β
verisimdb/connectors/test-infra/deploy.k9.nclβ "Pedigree block missing 'name'". Confirmed against the schema: instandards/k9-svc/pedigree.ncl,Metadata.name | Stringis the only metadata field with nodefault, so it is mandatory. Fix: addmetadata.name; per the canonicalk9-svc/pandoc/container/deploy.k9.nclsample, ideally alsometadata.version+validation.pedigree_versionand a leash leveltrust_level/security_levelβ'Kennel | 'Yard | 'Hunt(a shell-runningdeploy.k9.nclis'Hunt). governance / Trusted-base reduction policyred β perstandards/docs/TRUSTED-BASE-REDUCTION-POLICY.adoc+scripts/check-trusted-base.sh: an undocumented soundness-relevant escape hatch in a proof-bearing file. Disposition is per-repo β discharge, budget (// TRUSTED:), axiom (// AXIOM:), or a dated debt entry innextgen-databases'docs/proof-debt.md.
Durable record
Full audit with the verification table + propagation analysis:
standards/docs/audits/audit-hypatia-cache-sha-corruption-2026-06-21.adoc(+.a2mlcompanion), filed as draft standards#396.
Net. The Hypatia-lane root cause (1) is fixed & merged (
standards#394), independently verified. (2) is expected post-fix staleness drift awaiting a gitbot-fleetenroll-reposre-pin tod72fe5a+ (out of my scope). (3) is two pre-existingnextgen-databases-internal items for that repo's maintainers. Nothing here required (or received) an unverified fix.
Generated by Claude Code
βΆοΈ Remaining action β propagate the fix via gitbot-fleetenroll-reposThe central root cause is fixed + merged (standards#394). The only step left to turn the estate green is to re-pin stale consumers β they pin the reusables by
standardscommit SHA (not@main), so they keep dereferencing the old workflow and showCheck Workflow Stalenessred until refreshed.Target pin: current
standards/mainHEAD =4ddc926(containsactions/cache@1bd1e32aβ¦ # v4.2.0; re-resolve to current HEAD at run time).Operation (same as the 2026-05-27 orphan re-pin sweep β Contents API + auto-merge squash):
- Run gitbot-fleet
enroll-reposto re-pin each stale consumer's.github/workflows:hypatia-scan.ymlβhypatia-scan-reusable.yml@4ddc926governance.ymlβgovernance-reusable.yml@4ddc926
- Prioritise the observed-failing repos: nextgen-databases, KnotTheory.jl, nextgen-typing, plus the repo list above.
- Verify on 2β3 samples: scan lane green + staleness check green; post the result here.
Not covered by this sweep (separate per-repo fixes in
nextgen-databases):- K9 pedigree β
verisimdb/connectors/test-infra/deploy.k9.nclneedsmetadata.name(+version/pedigree_version+trust_level 'Hunt). governance / trusted-basered β undocumented escape hatch β entry in that repo'sdocs/proof-debt.md.
Note: this step cannot be driven from a session scoped only to
hypatia/standards(which is where the diagnosis + audit #396 were done). It needs a session/agent scoped to gitbot-fleet (and the consumer repos). Cross-ref nextgen-typing#69.
Generated by Claude Code
- Run gitbot-fleet
β enroll-repos re-pin is blocked β the propagation actuator (.git-private-farm Actions) is down
Picked this up from a session scoped to gitbot-fleet + .git-private-farm. The re-pin is fully scoped and ready, but it cannot run: the farm workflow that performs the Contents-API + auto-merge-squash sweep is failing at the infrastructure level. Reporting before any half-done state is created.
Target pin (confirmed)
standards/mainHEAD =4ddc926(carriesactions/cache@1bd1e32aβ¦ # v4.2.0, the #394 repair).stapeln+panoplyare already pinned here β proof the target resolves.Estate consumers enumerated (GitHub code search, both reusables)
SHA-stale = pinned SHA β
4ddc926, excluding@mainauto-trackers (already get the fix) and the 2 already at HEAD:Reusable refs @main(skip)already 4ddc926SHA-stale to bump governance-reusable.yml225 62 2 ~143 hypatia-scan-reusable.yml197 0 2 ~195 Old-SHA clusters (one farm run keys on one OLD_SHA, so this is ~17 dispatches):
- governance:
5a93d9d5β102,861b5e91β30,d72fe5aβ4,5eb28d7dβ2,3b3549e2β2, +3 singletons (echidna,ephapax,valence-shell) - hypatia-scan:
5a93d9d5β102,915139d7β54,97df762β16 (the 2026-05-27 orphan SHA still live),6cd37728β12,5eb28d7dβ5,d72fe5aβ3, +3 singletons (bunsenite,eclexia,vcl-ut)
Priority repos located:
nextgen-typingboth reusables@5a93d9d5;nextgen-databasesgovernance@main(no re-pin) + hypatia-scan@915139d7;KnotTheory.jlboth@5a93d9d5(β οΈ this one was missed by paginated code-search and only found by targeted lookup β code-search pagination is lossy, so the final sweep needs a per-repo verification pass to catch stragglers).The blocker β farm GitHub Actions are not allocating runners
I fired one validation dispatch into
.git-private-farmsha-bump-propagate.yml(workflow_dispatch,nextgen-typinghypatia-scan5a93d9d5β4ddc926). It failed in 6 s with no runner assigned (runner_id: 0, run27914137962) β before any step ran. Inputs passed every validation gate; this is not an input/logic failure.It is not isolated. Every
.git-private-farmActions run today dies the same way:- 14+
Propagate(repository_dispatch) runs β allfailurein 7β19 s - scheduled
Well-Known Standards,Estate CI-deadlock Audit,RustSec Audit,Forge Drift Detectionβfailurein 4β5 s OSSF Scorecardβstartup_failurein 0 s
GitHub-hosted (
ubuntu-latest) jobs failing in <20 s with no runner /startup_failureacross unrelated workflows = Actions are administratively or financially blocked on the private farm repo (spending-limit / billing exhaustion, or Actions disabled). This is the "billing (separate)" blocker already noted on this issue β and it sits upstream of the re-pin: the estate's own propagation automation has been failing against it all day.Net
- 0 consumers re-pinned, 0 verified. The single sample failed on infra, not on the change.
- The farm is the only estate-wide write path (consumers are out of this session's repo scope), so the sweep is blocked until farm Actions billing/runners are restored.
- Once restored, the sweep is ~17
workflow_dispatchcalls intosha-bump-propagate.yml(one per reusable Γ old-SHA, allnew_sha=4ddc926) + a re-search pass for stragglers + sample verification. I can re-fire immediately on the word that Actions are live again.
Out-of-scope items unchanged (K9 pedigree on
nextgen-databases, trusted-base proof-debt) β not part of this sweep.
Posted by Claude Code (gitbot-fleet session). No changes were pushed; this is a blocker report.
Generated by Claude Code
- governance:
β 3 priority repos re-pinned + admin-merged β and the real
scanblocker is a fresh Hypatia compile regressionDid these directly (owner-admin token), not via the farm:
.git-private-farmActions are still billing-blocked (every run dies <20s with no runner).bag-of-actionsoffers no propagation path (Zig/Elixir continuation runtime; its only workflows arecodeql+push-email-notify).Target corrected β pinned to current HEAD, not
4ddc926Re-resolved
standards/mainHEAD at run time (as the task instructs): it has advanced tod7c22711(#432, 2026β06β26), well past4ddc926(06β21). The samples were already atd135b05b(#416, 06β24 β postβcacheβfix), so their cache lane was already resolved; the live-open items were staleness + the newer reusable fixes (#424--exit-zero, #429GITHUB_TOKEN). All re-pinned tod7c22711.Done (all merged)
repo PR reusables β d7c2271extra Check Workflow Stalenessnextgen-typing #74 β hypatia-scan, governance, scorecard β pins all @Head nextgen-databases #52 β hypatia-scan, governance, scorecard governance @mainβpinned; removed retiredscorecard-enforcer.ymlπ’ success KnotTheory.jl #37 β hypatia-scan, governance, scorecard β π’ success Notes: I bumped scorecard-reusable too β the staleness checker fails on any of {scorecard, hypatia-scan, governance} pin β HEAD, not just the two named.
KnotTheory.jlwas missing from code-search enumeration (the index is lossy) β found by direct lookup.Verification
governance / Check Workflow Stalenessβ success on nextgen-databases + KnotTheory.jl (post-mergemain).- In CI,
actions/cache@27d5ce7fβ¦resolves fine β the central cache corruption is fully closed.
π΄
scan / Hypatia Neurosymbolic Analysisstill RED β estate-wide β and it is NOT the pinThe reusable
git cloneshypatia.gitand runsmix escript.build; hypatiamainno longer compiles as of #545 (7df53b1b, 06β26 20:29):== Compilation error in file lib/rules/structural_drift.ex == ** (CompileError) undefined variable "real_subdirs"lib/rules/structural_drift.ex:1019still referencesreal_subdirs, but fix(scanner): make SD022 path-drift rule precise + fix real doc/workflow driftΒ #545 renamed that binding toreal_basenames(src_dir_index/1, line 963).real_subdirsis now defined nowhere βescript.buildfails β the scan job dies at "Build Hypatia scanner" on every repo that calls the reusable.- Latent second defect: line 1004 regex has one capture group, but line 1006 destructures
[_, prefix, dir](two) βFunctionClauseErroronce the compile error is fixed.
I did not push a fix β there's no Elixir/mix in this session to compile-verify it, and you're actively iterating on this file (the bash fallback can't substitute). Flagging for a tested fix. Until hypatia
maincompiles, no repo's scan lane can go green β this is the real "scan green" blocker now (cache is closed).Still open
- Full estate re-pin (~200 consumers): needs either
.git-private-farmActions billing restored (farm sweep), or the same direct-token approach I used here, repo-by-repo. Happy to run the rest the same way on your go. - Hypatia
maincompile fix (fix(scanner): make SD022 path-drift rule precise + fix real doc/workflow driftΒ #545 follow-up) to clear the estate-widescanred.
Posted by Claude Code (gitbot-fleet session).
Generated by Claude Code
- added 5 commits that reference this issue
on Jun 27, 2026 143 remaining items
- added 6 commits that reference this issue
on Jun 27, 2026 - added a commit that references this issue
on Jul 7, 2026 - addedtri/controlSafety triangle β gate it so it cannot regressSafety triangle β gate it so it cannot regress
on Aug 3, 2026 - addedcicdCI/CD: workflows, actions, lockfiles, pins, runners, release gatesCI/CD: workflows, actions, lockfiles, pins, runners, release gatesand removed
on Aug 26, 2026 - addedpriority:p1High - schedule nextHigh - schedule nextscope:estateAffects many or all repos across the estateAffects many or all repos across the estatestatus:readyFully specified and ready to be picked upFully specified and ready to be picked up
on Sep 30, 2026 Status (2026-09-30 issue sweep): reviewed and kept open as a live meta-tracker. Feature/RFC issues were consolidated into the roadmap #886. Root-cause fixes in flight: #883 (scanner build + npx precision), #888 (workflow_audit path anchoring, 0-AI-MANIFEST.deed), standards#1084 (harden-runner in standards-only workflows). Ledger:
dev-notes/audits/hypatia-issue-sweep/2026-09-30.md.
After the allow-list fix (27 repos) and billing (separate), the remaining stuck PRs are blocked by each repo's own red CI β real build/test/governance failures, not the dependency bump. Rebasing won't help; each needs per-repo repair (home: hypatia ci-health-sweep, #461).
Sampled failures:
bofig(Build and test, Hypatia) Β·project-wharf(Analyze rust/actions, build, workflow linter) Β·double-track-browser(TypeScript Tests, language-policy) Β·coq-jr(language-policy, Well-Known, Validate Security Files).Repos with stuck PRs (per-repo CI repair needed):
Note: ~33 of the burn-cut PRs opened this session are also gated on these same repo-CI issues (or billing).