Merging #133 and #138 turned main red, and neither PR was at fault.
Both were green on their own head SHA. Both were branched 31 commits back, and both added tests. tests/test_doc_claims.py compares a hard-coded server-suite count in AGENTS.md against a live collect — neither PR's CI run ever saw the other's tests, so each was individually consistent and the merge of the two was not. 1040 claimed, 1047 collected. Tests (server) and Tests (server · py3.13) are both required checks, so main went red and stayed blocking until #152.
This is the same count-drift family as the earlier playbook-count problem, but with a new trigger: it is not reachable from any single PR. A per-PR gate structurally cannot see a number that only two merges together move, so no amount of care on the contributor side prevents it.
Options
- Run the doc-claims gate on
main after merge — catches it, but only after main is already red. Cheap, and at least the signal is immediate and unambiguous.
- Have the PR job collect against the merge result rather than the PR head. GitHub already builds a merge commit for
pull_request events; collecting there would have caught this before either merge. Better, and probably the real fix.
- Stop hard-coding the count. The number in
AGENTS.md is documentation of a fact the test suite already knows. A generated line, or a tolerance, removes the whole class — at the cost of losing a genuine tripwire that has caught real drift before.
Option 2 plus keeping the exact count is my current preference, but this is worth deciding rather than patching again.
Also worth noting
The same structural blind spot applies to every other hard-coded count the doc-claims checker guards (playbooks, detectors, flags, providers). None of them are safe from a two-PR interaction, only from a one-PR one.
Merging #133 and #138 turned
mainred, and neither PR was at fault.Both were green on their own head SHA. Both were branched 31 commits back, and both added tests.
tests/test_doc_claims.pycompares a hard-coded server-suite count inAGENTS.mdagainst a live collect — neither PR's CI run ever saw the other's tests, so each was individually consistent and the merge of the two was not. 1040 claimed, 1047 collected.Tests (server)andTests (server · py3.13)are both required checks, somainwent red and stayed blocking until #152.This is the same count-drift family as the earlier playbook-count problem, but with a new trigger: it is not reachable from any single PR. A per-PR gate structurally cannot see a number that only two merges together move, so no amount of care on the contributor side prevents it.
Options
mainafter merge — catches it, but only aftermainis already red. Cheap, and at least the signal is immediate and unambiguous.pull_requestevents; collecting there would have caught this before either merge. Better, and probably the real fix.AGENTS.mdis documentation of a fact the test suite already knows. A generated line, or a tolerance, removes the whole class — at the cost of losing a genuine tripwire that has caught real drift before.Option 2 plus keeping the exact count is my current preference, but this is worth deciding rather than patching again.
Also worth noting
The same structural blind spot applies to every other hard-coded count the doc-claims checker guards (playbooks, detectors, flags, providers). None of them are safe from a two-PR interaction, only from a one-PR one.