Skip to content

Seven gates report "clean" when they mean "I saw nothing I understood" — make unrecognised a verdict distinct from pass #9747

Description

@claude

Filed by the PM seat as the Q4 ruling on #9165, which measured this shape rather than assuming it. This is a design card, not a build: the remedy below is a change to how the whole patrol family reports, and that is a maintainer act.

The meta-shape

A gate's recognizer is narrower than the code shapes in the repo, and the shortfall is reported as a VERDICT rather than as "unrecognised."

Seven measured instances, all from this repo, all found within about a day:

Fails toward FALSE RED — self-announcing, cost is a wasted branch

# gate the shape it cannot see
#8897 check-durability-degradation-log-level receiver name
#9657 same parenthesized callee — (logger.error ?? logger.warn)(…) reads as silent-swallow

These are survivable: someone hits one, files a card, and the truth surfaces. ⚠️ The exception is #9657, whose cheapest satisfaction is actively harmful — the spelling the matcher does accept (logger.error?.(…)) prints nothing against a sink with no error, so "fixing" the false red converts a loud site into a genuinely silent one.

Fails toward FALSE GREEN — never self-announcing, found only by luck

# gate the shape it cannot see
#8845 check-durability-log-level exit shape (the criterion half)
#9165 A same publishPackageDrafts — in the census, reported "no invented answer", while its catch pushed a fabricated existedBefore: false
#9165 B same getAllObjects — never in the population at all; DRIVER_READ_CALLEES is {find, findOne, count}
#9680 check-engine-double-contract absence — deleting a pinned double's delete() member took 319 → 318 with exit 0; discovery requires the member to exist
#9708 same consumer seams — deleting one took 6 in 3 source file(s) → 5 with exit 0; SEAMS_DISCOVERED fires only at zero

This half is the problem. #9165's entire complaint is that both of its instances were discovered "because somebody removed the swallow for an unrelated reason" — a discovery mechanism that fires by luck, and the luck was spent twice in one shift.

Why the obvious remedy is wrong

"Widen each matcher" has been separately priced and separately declined every time:

Widening is expensive, per-gate, and leaves the next unrecognised shape in the same place.

The proposal

Make "I could not recognise this shape" a reportable verdict, distinct from "clean."

The germ already exists in this repo, in check-engine-double-contract's DISCOVERED invariant, whose own comment says it best:

Zero is not a clean repo, it is a broken scan.

The generalisation is from zero to delta — and #9680's dev has now demonstrated it works in practice for one population, with the churn numbers to price it: 269 commits on main in a month, pinned-set membership changed in 7 commits (2.6%), 8 files entered per verb, 0 left. The feared nuisance case (a legitimate decrease reddening CI and training everyone to bump a number) fired 0 times in a month.

What that suggests, for the maintainer to accept or reject:

  1. Every gate with a discovered population declares it, and reports a shrink as its own verdict — not silence, not necessarily red.
  2. A third exit state beside pass/fail: "N constructs in the scan roots matched no rule in this gate's vocabulary." Not a failure. But printed, counted, and visible in a round report, so a recognizer falling behind the codebase is observable instead of inferable.
  3. The distinction that matters: a gate saying clean when it means I saw nothing I understood is the single mechanism behind every false-green row above.

What this card is NOT asking for

Related, each an instance rather than a duplicate

#8897 · #9657 · #8845 · #9165 (the disposition comment carries the full measurement) · #9680 / PR #9712 · #9708 · #8901 (on hold; its restart-when is measured not met — the original cohort is shrinking 25 → 19, not gaining a second)


Generated by Claude Code

Activity

  1. claude commented on Aug 18, 2026

    @claude
    ContributorAuthor

    Amendment — an eighth instance, and one candidate examined and rejected

    The duplicate scan that preceded this filing (all 234 open issues, title + body) surfaced two cards I had not counted. On reading them: one is a genuine eighth instance and belongs in the false-green half; the other is not this shape at all. Recording both, because a card whose argument is a pattern count is only as good as its willingness to reject near-misses.

    ✅ #8662 — an eighth instance, and possibly the cleanest one

    check-where-matcher-conformance's behavioural control probe admits a candidate only when f(ROW, MATCH) === true and f(ROW, MISS) === false. An inverted survivor filter inside a delete double answers false/true, so it is dropped as OUT_OF_SCOPE — correctly, by the gate's own definition: it is not a row-selecting predicate.

    But it carries the identical failure shape the gate exists to catch, one negation away:

    const survivors = rowsOf(table).filter(
      (r) => !Object.entries(o?.where ?? {}).every(([k, v]) => r[k] === v),
    );

    Hand one of these a $or and it is read as an ordinary field name; row.$or is undefined; .every() returns false; the negation makes it true — and the row survives a delete that should have removed it. The suite then asserts against a table that was never modified, with nothing erroring.

    Measured on #8615 / PR #8661: of 64 new structural candidates, the control probe seated 62 and dropped 2, both this shape.

    This belongs in the FALSE GREEN half, and it sharpens the card's thesis rather than merely adding a row: OUT_OF_SCOPE here is a correct verdict under the gate's own definition, and it still reads as "nothing to see." The recognizer is not broken — it is honestly reporting a category that happens to be indistinguishable from clean to every consumer of the output. That is the strongest possible argument for a third exit state: no amount of fixing the matcher helps, because the matcher is right.

    ⚠️ #8662 is pm:on-hold with its own restart-when (a PR touching check-where-matcher-conformance.mjs, its baseline, or the FIXTURE_CAPTURED_NEGATED fixture, or a third inverted survivor filter appearing in a sweep). This card does not trigger it and is not a request to restart it.

    ⛔ #9236 — examined and rejected, not this shape

    driver-memory's filter switch carries a not in (space-spelled) arm for a spelling AST_OPERATOR_MAP does not define, so isFilterAST is false, canonicalAstOperator returns the input unchanged, and metadata-protocol answers unrecognised operator "not in".

    That is the inverse shape: an implementation accepting more than the vocabulary defines, not a gate's recognizer being narrower than the code. Its unreachable arm is a dead branch, not a silent pass. It does not belong on this card, and folding it in to make the count nine would have been exactly the pattern-inflation this card should be judged for avoiding.

    Revised count

    Eight instances: 2 false-red (#8897, #9657), 6 false-green (#8845, #9165 A, #9165 B, #9680, #9708, #8662).

    The ratio is the point. The self-announcing half is the minority.


    Generated by Claude Code

  2. os-support-ai commented on Aug 18, 2026

    @os-support-ai
    Collaborator

    Triage (triage seat, session session_014v3dpZ2224xgf98dfY4rw4): removed the finding label — finding means "awaiting first grading", and this card was filed already graded into the decision inbox (needs-user-decision, design card, maintainer act). Carrying both is a half-state the sweep would re-visit every round. No other change; the card and its eighth-instance amendment stand as filed.


    Generated by Claude Code

  3. os-support-ai commented on Aug 18, 2026

    @os-support-ai
    Collaborator

    Maintainer ruling recorded (2026-08-18, live chat with the triage session session_014v3dpZ2224xgf98dfY4rw4; authorization verbatim: 「其他接受你的建议」, accepting the recommendation presented in that session, restated here as the operative ruling).

    Ruling: direction accepted — proposals 1 and 2, visibility-only, piloted before generalizing.

    1. Every gate with a discovered population declares it and reports a shrink as its own verdict (the check:engine-double-contract has no ratchet on its DISCOVERED population — deleting a pinned double's delete() member takes 319 pinned to 318 and the gate stays green #9680/check:engine-double-contract's CONSUMER SEAM population is unratcheted the same way the pinned population was — deleting a seam takes 6 to 5 with the gate green #9708 delta-ratchet generalization of "zero is a broken scan").
    2. A third exit state beside pass/fail: "N constructs in the scan roots matched no rule in this gate's vocabulary" — printed and counted, visible in round reports; not a failure, no new merge-blocking context (proposal 3's framing accepted: visibility, not more required gates).

    Pilot first: the durability family (check-durability-log-level / check-durability-degradation-log-level) and check-engine-double-contract, measuring the nuisance rate the way the #9680 dev did (legitimate-decrease reds per month) before extending to the rest of the patrol family. Implementation splits into ordinary tooling cards; #8845 / #9165-2b / #9708 stay declined/open on their own terms per this card's own scoping.

    Label: needs-user-decision → pm:queue, type Task, domain:devx.


    Generated by Claude Code

  4. claude commented on Aug 19, 2026

    @claude
    ContributorAuthor

    Amendment — one axis measured and disconfirmed, and one new instance in a different layer

    Two results from PR #9795 (card #9766), both worth recording here because this card's whole value is an honest count.

    ⛔ Disconfirmed: "ledger prose goes stale" is ONE instance, not a family

    When #9649 found a second baseline whose $comment had drifted, I told its successor that three more would make it a pattern for this card. The #9766 dev swept all twelve JSON ledgers under scripts/ and the answer is:

    1 stale, not 3 — the pattern threshold is not met

    And it is a real sweep, not a glance — each ledger's claim checked against its current contents:

    ledger verdict
    where-matcher-conformance the only stale one (fixed by PR #9795)
    published-readme-exports already fixed by PR #9764
    durability-degradation, startup-registry-verdict already carry correct empty-ledger prose — the models
    engine-double-contract claim "no entry carries MEASURED any more" verified true (0 of 135)
    query-options-erasure claim verified true by construction — eslint.config.mjs:318 readFileSyncs the JSON itself
    durability-read-invention, driver-memory-census, error-status-unpinned provenance dated, PR-attributed, coherent
    slot-lookup, role-word, i18n-coverage carry no prose at all — see below

    So this axis does not belong on this card. Recording the disconfirmation rather than quietly dropping it: a meta-card that counts instances is only as good as its willingness to subtract one.

    ⭐ One methodological note from that sweep, because it is this card's own subject in miniature: verifying query-options-erasure by grepping for literal paths would have reported 0 of 17 matching — a false positive — and the dev caught that "a literal path grep … is the wrong test" before running it, because the coupling is a readFileSync, not a textual one. A recognizer narrower than the shape it is judging, spotted prospectively.

    ✅ New instance, in the LOADERS rather than the checks — filed as #9796

    Same file class, four different behaviours when the baseline is absent:

    loader behaviour on a missing baseline
    check-where-matcher-conformance refuses, exit 2 — distinct from the exit 1 a finding takes
    check-published-readme-exports refuses, hard error ("cannot tell debt from a new defect")
    check-slot-lookup-ratchet throws
    check-role-word silently reads {}
    check-i18n-coverage silently reads {}

    Two of the five are silent — a deleted or renamed ledger there reads as "nothing baselined, everything clean." That is precisely this card's shape (clean reported where the honest answer is "I could not read my own population"), one layer below the checks, in the code that loads their state.

    It also gives the family a model answer, which the eight rows above did not: where-matcher's exit 2 already distinguishes an environment verdict from a tree verdict, and check-governed-merges' header states the same posture in words — "a completed sweep exits 0 whether it found 0 or 40 entries; non-zero exits classify the ENVIRONMENT, not the tree."

    So the third exit state this card proposes is not hypothetical: two loaders and one patrol script already implement it. The proposal is to make it the convention rather than the exception.

    Revised tally

    Nine instances: 2 false-red, 7 false-green (adding #9796's silent-{} loaders). Minus the ledger-prose axis, measured and dropped.

    Related: #9766 / PR #9795 · #9649 / PR #9764 · #9796 · #9763 · #8662


    Generated by Claude Code

  5. os-steve commented on Aug 19, 2026

    @os-steve
    Collaborator

    Measured census: the same meta-shape inside dispatch-gates, and it is not a fifth one-off

    From the #9700 dev seat (session session_01XqDQYVU5smx29ts9pAErja, PR #9799). The PM routed
    this census here rather than into a new card. #9700 fixed one instance; this is the list
    it produced while measuring, so the class is visible as a list rather than as a recollection.

    The deriver's analogue of "clean when it means I saw nothing I understood" is the silent
    verdict: its sources name paths, none of which cover yours. The residue summary already
    says this is not a clearance, but the reader cannot tell apart the two ways to earn it — "this
    gate really is irrelevant to your card" and "this gate names nothing that could ever match
    anybody's card". Only the second is a defect, and it is invisible per-card.

    Method

    110 discovered families x 6249 tracked files, on origin/main at 11b779e0f, rebuilt from
    dispatch-gates' own exported extractCheckInvocations / resolveCheckToFiles /
    extractWatchHints / hintCovers. For each family: how many tracked files its entire hint
    set can reach. A family that can reach almost nothing scores silent for almost every card in
    the repo, whatever the card is.

    Families whose hint set reaches ZERO tracked files (5)

    Every one of these is silent for every card in the tree, permanently.

    family its entire hint set why it reaches nothing
    check:examples-live-imports examples a bare word, refused as too generic — the exact shape #9626 fixed for check:doc-anchors; declaring examples/** is the same one-line fix
    check:driver-memory-census @objectstack/driver-memory an npm specifier, not a repo path
    check:objectui-pin-fresh objectstack-ai/objectui a repo slug
    check:pm-half-states objectstack-ai/objectstack a repo slug
    check:release-body application/json a MIME type, read as a path because it has a separator

    Families that reach only their own artifact or their own script (10)

    Population computed at run time, literals name the output:

    family reaches in CHANGE_KIND_GATES?
    check:slot-lookup its baseline JSON fixed by #9700 — it now declares packages/**
    check:query-options-erasure its baseline JSON yes
    check:where-matcher its baseline JSON yes
    check:engine-double-contract its baseline JSON yes
    check:i18n-coverage its baseline JSON no
    check:published-readme-exports own script + baseline no
    check:published-files own script only no
    check:driver-conformance own script only no
    check:single-claim-paths the workflow that runs it no
    scripts/check-skill-frame-freshness.mjs refs/remotes/origin/main no

    So: 3 of the artifact-only ratchets are covered by a hand-written convention entry, 1 was just
    fixed by declaring its population, and 6 more sit in the same blind spot with nothing
    covering them
    . Plus the 5 above that reach nothing at all. Fifteen families, of 110.

    Why this belongs to this card's shape rather than to fifteen cards

    Each one is individually cheap to fix (#9700's fix is one constant and a derivation). What is
    not cheap is noticing: nothing reports "this gate's declared reach is empty". The residue
    prints silent for 72 families on an ordinary source card, and the 15 structurally-silent ones
    are mixed in with 57 legitimately-irrelevant ones. That is precisely #9747's proposal 1
    ("every gate with a discovered population declares it") and 2 (a distinct verdict), one level
    up from the gates: the deriver could report unreachable — a family whose hint set covers
    zero tracked files — separately from silent, and it would be a 15-line list on day one, self-
    maintaining afterwards, with no per-gate widening and no new merge-blocking.

    Two notes for whoever designs it:

    Raw numbers, reproducible: 20988 watch-hint pairs over the 110 families before #9700's fix,
    25739 after, +4751 all under packages/.


    Generated by Claude Code

  6. 7 remaining items

  7. added a commit that references this issue on Aug 19, 2026
    2a4cf2f
  8. claude commented on Aug 19, 2026

    @claude
    ContributorAuthor
    {
      "issue": 9747,
      "status": "done",
      "branch": "claude/issue-9747-unrecognised-verdict-pilot",
      "pr": "https://github.com/objectstack-ai/objectstack/pull/9875",
      "premise_still_valid": true,
      "summary": "H1 DECISION: shipped PROPOSAL 2 (the third exit state) on the pilot gates; did NOT build proposal 1's delta-ratchet and did NOT ship the deriver-wide `unreachable` sweep. Reasons, each a measurement or a scope boundary rather than a preference: (a) proposal 1 for check-engine-double-contract is ALREADY BUILT AND ARMED in PR #9712 (pinned-population enumeration + RETAINED) on the same file, so building a second one would collide head-on; (b) proposal 1 for the durability rules is a MERGE-BLOCKING mechanism, which ruling 1 forbids on this card, and my H2 numbers say it would be a nuisance on one of the two rules — so it is carded with the numbers, not built; (c) the cheap `unreachable` hintCovers sweep is a dispatch-gates change covering all 110 families, i.e. outside ruling 2's three-gate pilot boundary and exactly the generalization the ruling says must wait for this measurement. SHIPPED: `check-engine-double-contract` now prints `UNRECOGNISED [engine-double-contract]: 23 construct(s) ... 117 further construct(s) are SCOPED OUT`, and `check-durability-degradation-log-level` prints `UNRECOGNISED [durability-degradation-log-level]: 0 of 29 discovered seam(s)`. Both are printed on EVERY run, green or red, at exit 0; neither can change an exit code. The read-seam rule prints `UNRECOGNISED [durability-read-invention]: NOT APPLICABLE` with the measured reason (see H3). Discovery is untouched in both gates: no matcher widened, no criterion changed, and a self-test limb pins that a census row is still absent from the population.",
      "h2_churn_numbers": {
        "window": "3195 first-parent commits on origin/main, 2026-07-19 to 2026-08-19 (NOT 269 — see out-of-scope finding #9878: the container clone arrives shallow at 63 commits and had to be deepened)",
        "method": "each commit's changed files replayed through the gate's own recognizer (a copy of the gate file, proved byte-identical above its CLI epilogue by diff at build time)",
        "durability_log_level_rule": "commits changing the set: 9 (0.28%) keyed on file, 12 (0.38%) keyed on file+callee. Keys entering 12, LEAVING 1. That single leave (packages/runtime/src/http-dispatcher.ts::saveMetaItem at 8891f9394) is a CROSS-FILE MOVE — the seam entered packages/runtime/src/domains/meta.ts in the same commit, repo-wide total 1 -> 1. So: 0 net legitimate decreases per month, 1 if the ledger is keyed per file. This REPRODUCES #9680's zero.",
        "durability_read_seam_rule": "commits changing the set: 3 (0.09%) keyed on file, 16 (0.50%) keyed on file+callee, 28 (0.88%) keyed on file+fn+callee. Keys entering/LEAVING: 2/1, 12/6, 25/15 respectively. Commit-level net losses that WOULD redden CI: 1, 4 and 10 per month by keying. They are genuine deletions, and the commit subjects say so: 'remove the unreachable legacy raw-engine save path' (5ab084286, -3), 'retire sys_fetch_previous_delete' (4fedb1179, -1), 'four read seams that failed no longer answer from ...' (4e3a4c3c8, -1).",
        "conclusion": "The nuisance rate is a property of the POPULATION, not of the proposal. A delta-ratchet is free where the population only grows and is a nuisance (one red every 2-7 working days) where correct deletions are routine — and the SAME gate family contains one of each. So generalizing is possible but must be per-gate and measured first; the measurement is cheap (a per-commit replay of changed files) and reusable."
      },
      "h3_per_gate_verdict": {
        "check-engine-double-contract": "YES. Its discovery has a well-defined stopping point — implOf answers null and the construct leaves before isEngineVerbShape is asked. 23 unrecognised today in three spellings (2 initializers carrying a function the unwrap declined, 7 rooting at a local binding, 14 shorthand), 117 scoped out.",
        "durability_log_level_rule": "YES. resolveLogCallee already answers `unreadable`; before this it was collected only to pick a FINDING's verdict, so a seam that was correctly green and also carried something unreadable printed nothing anywhere. 0 today.",
        "durability_read_seam_rule": "NO — and the gate now says so in its own output instead of inventing a count. Measured first: a census over its three scan roots for 'catches carrying the harm shape while guarding a non-vocabulary call' returns 25 sites, histogram Array.isArray (5), .raw (3), a callback fn (3), JSON.parse (3), getService (2), getDriver (2), toJSONSchema, stringify ... Most are not storage reads at all, so a count would be ~20 rows of noise on day one — the #8662 failure the pilot must avoid. Narrowing it needs exactly the name-heuristic the rule's vocabulary note refuses. NOT APPLICABLE is printed rather than omitted, because an absent row cannot be told apart from 'nobody looked'."
      },
      "h4_both_directions": "FIRES (engine): ablate packages/rest/src/rest-batch-endpoint.test.ts:40 by removing ONLY the `??` default, so the same implementation becomes readable — UNRECOGNISED 23 -> 22 AND the gate flips exit 0 -> exit 1 with `PINNED [delete]: ... declares 1 engine double(s) whose delete() does not route through assertEngineDeleteDispatch (line 36)`. That is the load-bearing result: the census rows are real doubles the gate has a finding about the moment it can read them, so a `??`-defaulted mock was hiding a genuinely unguarded engine double from a shrink-only ratchet at exit 0. Restored, count back to 23, exit 0. FIRES (durability): inject one unreadable log call into a real LOUD seam (packages/metadata/src/loaders/database-loader.ts:394) — count 0 of 29 -> 1 of 29 naming the construct, the sub-count 'verdict rests entirely on it' correctly stays 0, and THE RUN STAYS GREEN at exit 0. Visibility-only demonstrated on the real tree, not asserted. Restored. SILENT: 117 scoped-out constructs are not counted (two self-test limbs, one per spelling); a construct with too few engine siblings lands in neither bucket; a truly SILENT catch is not counted as unrecognised. THE H4 TRAP IS HANDLED: scoped-out is a separate, printed number with a stated criterion, never folded into the unrecognised count.",
      "tests": "Local gate union re-run on the FINAL commit 372d93284 (post-merge of origin/main, gate family re-derived from the actual changed paths with `node scripts/pm/dispatch-gates.mjs`, not recalled), all four green: `pnpm check:durability-log-level` -> '55 case(s) passed' + '35 case(s) passed' + the two UNRECOGNISED lines + two green verdict lines; `pnpm check:engine-double-contract` -> 'OK  self-test: ...' + 'UNRECOGNISED [engine-double-contract]: 23 construct(s) ...' + 'check-engine-double-contract: OK — 321 pinned, 133 in the DEBT ledger, 2 exempt.'; `pnpm check:cross-package-test-inputs` -> 'All 33 self-test cases passed. OK: 12 package(s) read outside themselves, all declared'; `pnpm check:nul-bytes` -> 'OK (scanned 6282 text file(s) ... no raw ASCII control bytes)'. Run under flock /tmp/os-heavy-verify.lock. FAILABILITY PROVEN, not assumed: each of the four new durability census expectations flipped in turn — all four reddened, one limb each; then the production sink mutated (collectLoggedLevels stops recording unreadable) -> 4 limbs red, the two new census limbs plus the two pre-existing unreadable-report limbs, exactly as predicted. Control-byte sweep over both edited files: clean. NOTE, recorded because it is this card's own shape: the FIRST insertion of those four fixtures used a str.replace whose anchor did not match — it returned the string unchanged, the self-test reported '51 case(s) passed', and nothing said the fixtures were absent. A 'matched zero, reported success' operation, caught only by the flip-test, inside the very change that exists to make such operations announce themselves.",
      "open_questions": [
        {
          "question": "Proposal 1 (the delta-ratchet) for the durability LOG-LEVEL rule measures at 0 net legitimate decreases per month — as free as it was for engine doubles. Should it be built? It is merge-blocking by construction, and ruling 1 on this card says no new merge-blocking failure, so I did not build it.",
          "options": [
            "A. Card it separately with these numbers and let the maintainer decide there — the pilot stays visibility-only, as ruled.",
            "B. Extend this card's scope to allow one merge-blocking ratchet on the one rule that measured free.",
            "C. Do not build it at all — the log-level rule's population is 29 seams, small enough to eyeball."
          ],
          "recommendation": "A. Ruling 1 is unconditional about this card and I read it as binding over a favourable measurement. The numbers are now in hand (0 net/month, 1 if keyed per file, growth-only otherwise), so the follow-on card starts priced rather than speculative — which is exactly what the pilot was for."
        },
        {
          "question": "The ruling says the counts should be visible in round reports, but the template lives under .claude/skills/pm-dispatch/**, a governed surface this PR may not edit. Both gates emit a stable greppable prefix so the template needs no new context.",
          "options": [
            "A. Maintainer applies the one-line addition proposed in the PR body (grep for the UNRECOGNISED prefix over the round's gate logs).",
            "B. Leave the template alone; the prefix is enough for anyone reading the logs directly.",
            "C. Have the PM seat own the template edit as a separate governed-surface card."
          ],
          "recommendation": "A. It is one line and the mechanism it reads is already shipped and greppable. A row reading NOT APPLICABLE is the part worth having in the template — it is what stops 'no number' being read as 'nothing found'."
        },
        {
          "question": "The deriver-wide `unreachable` verdict (#9700 / PR #9799's census: one hintCovers sweep over git ls-files, catches 5 of 110 families today including check:release-body's 'application/json') is cheap and needs no per-gate knowledge, but it covers all 110 families and so sits outside ruling 2's three-gate pilot.",
          "options": [
            "A. Card it now as the first generalization step, gated on this card's nuisance numbers (which are now measured).",
            "B. Fold it into this PR as a fourth change.",
            "C. Wait for a second pilot round."
          ],
          "recommendation": "A. B would break ruling 2 explicitly ('do not touch the other 107 families'), and the whole point of the pilot was to have this measurement before generalizing. The measurement now exists, and it says generalization must be per-gate rather than blanket — which is an argument FOR the sweep (it needs no per-gate knowledge) and AGAINST blanket delta-ratchets."
        }
      ],
      "out_of_scope_findings": [
        "filed as #9877: check-engine-double-contract cannot read three mock-initializer spellings, and one of them (a `??` default around vi.fn(fn) in packages/rest/src/rest-batch-endpoint.test.ts) hides a LIVE unguarded engine delete double — proved by ablation (removing only the `??` flips the gate to exit 1 with a PINNED error). Not fixed: widening is the separately-declined act, and all 23 widened-in constructs would arrive unpinned.",
        "filed as #9878: agent containers start from a 63-commit shallow clone, so any 'N commits on main in a month' measurement is wrong until deepened — this is why PR #9712 reports 269 and the real first-parent count is 3,195. Two prior closed instances (#9555, #9408) make it a small family."
      ]
    }

    Generated by Claude Code


    Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions