3.8.0 scope: the deep-gate fixes, made reviewable - #139
Merged
Conversation
…rfaces The gate already agreed pyproject/__init__/CITATION. It did not check the two places that state the CURRENT version in prose (RELEASE.md, readiness_pack/ PROGRESS.md), nor PyPI, nor the project page — and PR 575 upstream carried a stale 3.6.2 for weeks without anything noticing. Extended in place rather than as a second script: a second measuring station for the same quantity is the next drift. Historical statements (since vX, as of vX, old changelog headings) are excluded by design and covered by a test — bumping them would turn a fact into a lie. External surfaces have three states. Unreachable is NICHT MESSBAR: it does not fail the run, and it never counts as green; the summary says what was actually verified. --require-external turns it into a failure for the release checklist.
The audit recommended a stability window in weeks. The rule here is the
opposite shape: no calendar, no cadence. A release happens when it can answer a
checkable list, so what gets slowed down is vagueness rather than speed.
The scope list answers the stash question by measurement instead of assumption.
stash@{0} applies cleanly to main and is out precisely because of that: it adds
a targetSubjectDigest key to the per-edge result, and that result is signed into
a relation statement. New output at a public interface is a MINOR, not a PATCH.
The line claimed 87 findings on this tree from ruff 0.16's expanded default set. Both halves were wrong, and measuring properly showed why: the local ruff is 0.15.10, and all 87 findings sit in an untracked scratchpad/ directory that 'ruff check .' happens to walk. Over the 258 tracked .py files the tree is clean. I had read the output of a command as an answer about the repository, when it was an answer about the working directory. The real numbers were already in pyproject.toml, measured by the change itself.
Cut-list item 1 for 3.7.1. Measured drift was larger than the list assumed: not only the harness digest and the absence rule were missing, but also today's three corrections. Both docs still carried anchors[] as a predicate field, which PR 575 never had. docs/upstream/eval-result.md is now a mirror of the submitted file and says so: when the two differ, the PR is the source of truth and this copy is the one that is wrong. It also stopped claiming the PR is unopened. IN_TOTO_PROFILE.md drops the anchors row and says plainly that earlier revisions were wrong there, rather than quietly deleting it. Documentation only, no src/ change.
Measured on 2026-08-07: git describe --tags returned corpus-review-2026-07-25-iter10, _semver_tuple read it as (0, 0, 0), every real version compared as bumped past it, and check 3 stopped applying. Under that blind spot one non-trivial commit sat undelivered since v3.7.0 with no [Unreleased] section, and the gate reported OK. The check did not fail. It stopped checking, and silence looked exactly like agreement — the same shape as the vanished-anchor case in check 4. Three states: a release tag, no tags at all, or tags that exist but none is a release. The third is reported in its own words.
The corrected check 3 immediately found a real gap: one non-trivial commit had been sitting since v3.7.0 with no changelog trace. This is that trace, and it states plainly what the release gate asks for — semantics unchanged.
Aligning the two docs today fixed the instance and left the class open: nothing checked that they stay aligned, which is how they drifted apart in the first place. Both had listed anchors as a predicate field, neither carried the absence rule, and the mirror still claimed the PR was unopened. Deliberately not a byte comparison against the upstream file — that lives in another repository, so the test would pass or fail depending on what happens to be checked out next to this one. These are the sentences the submission makes, each one droppable by a future edit without anyone noticing.
Measured at 19:19: the list named four items while the branch carried seven commits. The two it omitted were the two that were found rather than planned — the post-tag drift fix and the invariant tests. A scope list that does not contain what actually ships is the thing it was built against, so the omission is recorded in the list rather than quietly filled in.
The submitted spec gained a third Non-claims paragraph today (35c83da). The mirror was pinned at 6fdf5bf and the profile page never carried it, so both described a predicate that says less than the one actually submitted. The mirror takes the paragraph verbatim; the profile page states it in its own voice and keeps the 'can be' hedge, because the predicate never fixes which class a detector counts as positive. Measured, not assumed: the other five upstream changes of today were already present in both files via 871453c. Only this one had drifted.
…umber's object Check 4 watches the places somebody entered into _TRACKED_PLACES. The place that goes stale is the one nobody entered, so check 6 sweeps tracked files for current-release claim shapes outside the declared set and asks for a decision: declare it, or reword it. It fires even when the number is right today — agreeing now is not the property. Measured, not assumed: it caught its own author. The first version of docs/version_truth_list.md quoted the anchor form literally, and a page about version places became one. Every number now carries its object and the file it was read from. _source_version() returns value and origin together, because deriving them separately at two call sites is how one run reports two different values for 'the version' — measured today, three times, in this repo's neighbours. The truth list itself is measured (git ls-files + line scan): one source, no derived places at all (nothing pulls the version automatically), four checked copies plus two external surfaces, 17 historical lines that must stay old. Tests +8 (28 total in this file): undeclared claim in README, undeclared claim that matches today's version, historical forms must not trip, declared places not double-reported, untracked file is not a repo claim, plus three that pin the object-and-source wording.
…d assurance state
SUPPORT.md and COMPATIBILITY.md add no promise. SUPPORT repeats SECURITY.md's
supported-version sentence instead of softening it — two rules would mean the
stricter one is only true on paper — and says plainly that there is no response-time
commitment, because inventing one would be worth less. COMPATIBILITY writes down what
the existing SemVer commitment already implies, including the part that is easy to
miss: adding a key to a signed structure is breaking, which is why stash@{0} stayed
out of 3.7.1.
The self-assessment states the Scorecard value as measured and unsoftened: 6.5/10
(v5.5.0, api.securityscorecards.dev, 2026-08-07T18:37:54Z), with the four zeros named
and their three different causes separated. No application was filed — that is
outward-facing and needs its own GO.
It also caught an overclaim in scorecard.yml: the comment said Pinned-Dependencies was
maxed because actions are SHA-pinned. Measured 3/10. Actions are one input among
several. The comment now carries its measurement date, because a score nobody
re-measured is the same defect class this repo reports about version numbers.
Scope list extended by the five items this and the previous round produced.
…e check looks Owner decision 2026-08-07: publish, do not curate. The badge is live, so it moves; the README states the measured number next to it (6.5/10, v5.5.0, 2026-08-07) and explains all four zeros in one sentence each, including the two that a single maintainer cannot fix. Signed-Releases was measured, not guessed. ossf/scorecard docs/checks.md: the check reads GitHub RELEASE ASSETS for *.sig / *.sigstore.json / *.intoto.jsonl. Measured the same day: v3.7.0's assets are the wheel, the sdist and SHA256SUMS. The provenance was never missing — attest-build-provenance puts it in GitHub's attestation store and on PyPI, which is not where the check looks. The bundle is now copied, unchanged, to a second location and attached to the release. Nothing is re-signed. On the asset name: .intoto.jsonl describes the content. The check does not verify signatures, so a name alone can buy points — picking one for that reason would be the exact defect this project exists to make visible. Not verified: the release workflow was not run. The next release shows whether it takes effect.
…es an allowlist L6-01 from the deep gate on d3401a7 (P1, jury-confirmed): the step I added three hours earlier copied the provenance bundle into dist/ — the same directory handed to packages-dir. Measured with twine 7.0.0: InvalidDistribution, before any network I/O. The GitHub Release would have gone out and the PyPI leg would have failed, splitting the release in half. My own commit message conceded 'Not verified: the release workflow was not run'; the gate ran it instead. Two changes, and the second is the point. (a) The provenance and SHA256SUMS now go to release-assets/, a directory with one consumer. A directory serving two consumers with different admissible contents IS the coupling; a second directory removes it. (b) The publish gate carried a hand-maintained removal list ('rm -f dist/SHA256SUMS # not a distributable'). It did not fail on my new file — it simply never mentioned it, and silence read as approval. A blocklist can only name what someone already thought of. The gate now asserts the property: after removals, dist/ holds EXACTLY one wheel and one sdist; anything else stops the publish loudly, on the leg where it is still cheap. That is finding L6-03's shape, which says the same thing about a sibling list and notes it is 'exactly like L4-01' — a class closed at one call site. Third instance of that pattern in one night. Measured bidirectionally: wheel+sdist pass; SHA256SUMS caught; the d3401a7 case caught; two wheels caught. Not verified: the release workflow still has not run. The next release shows whether the split is really gone.
…ipt's own edge Deep-gate finding L4-01 (P1, jury-confirmed, wf_1c023644-953). verify_relationship_edges checked at the receipt's OWN edge whether the target verifies standalone and whether a declared targetSubjectDigest binds. _walk_chain did neither: it adjudicated ancestor SYNTAX and skipped ancestor CRYPTO entirely. Measured against the pre-fix file straight out of git: a forged ancestor gives FAIL at hop 1 and VERIFIED at hops 2, 3, 4 and 5. lineage=VERIFIED with safeForAutomation=true, and the presenter picks the distance by inserting one self-signed hop. The invariant now enforced is distance invariance: verdict(D at hop 1) == verdict(D at hop n) over byte-identical evidence, for every defect class. Three gates in _walk_chain mirroring the direct arm one to one — attached-but-not-an-object, attached-but-unverified, and the subject pin on every traversed ancestor edge. Same three in the Rust walker, minus the first: TargetInfo is a typed struct, so that case cannot arise there. The asymmetry is stated in the code rather than silently absent. One thing found while fixing, same shape one level down: the walk returned the first error in DFS order, so an unverified sibling listed BEFORE a back-edge masked the cycle and the reported code depended on the order the issuer wrote the edges in. A cycle is structural and is now decided before any descent, in both languages. Two test changes, both because the tests encoded the defect. test_relation_property's reachability helper walked THROUGH unverified nodes — it modelled exactly the gap. And its assertion demanded the cycle code even where a stronger ancestor gate legitimately fires first; FAIL stays mandatory, but the reason must now be a NAMED one from a closed set, never 'FAIL without a code'. New durable regression: tests/test_relation_gate_distance_invariance.py — property over a generator varying hop depth, five defect classes, plus the two counter-directions that keep it honest (a clean chain still VERIFIES at every depth; an ancestor beyond the attached horizon stays declared-only). 5 tests, 40 subtests. Full suite 2012 tests OK, ruff clean, cargo check/test/fmt/clippy clean. NOT MEASURED, and it matters: no behavioural vector was run against the Rust binary in this round. The Rust change compiles and lints, and verify-relation / verify-relation-statement do exist as subcommands — contrary to the finding's note that the parity oracle waits for one — but Python/Rust parity for these vectors is unverified here. The relation_signer check on ancestor edges is also NOT included: it needs ancestor edges in the emitted lineage result, and that structure is signed, so adding to it is breaking under COMPATIBILITY.md. Own increment.
…verity und Status Deep-gate Fund L5-01, P0, jury-bestaetigt (wf_cfe249d0-ee8). _norm() — NFKC plus Entfernen der Kategorien Cc/Cf — lief in _resolve_current ueber 'severity' und 'status' und NICHT ueber 'id' und 'superseded_by'. Nachbar-Felder in derselben Funktion, ungleich behandelt. Der Angriff, gemessen gegen die Vor-Fix-Fassung aus git: ein offener P0 mit superseded_by = "PB-X<unsichtbar>" plus ein geschlossener Koeder mit id = "PB-X<unsichtbar>". Roh sind die Zeichenketten verschieden, der Verweis liest sich als legitime Supersession auf einen vorhandenen ANDEREN Eintrag, der P0 faellt aus der Zaehlung, und das signierte Register meldet 0 offene P0/P1. Fuer einen menschlichen Pruefer sehen beide Kennungen gleich aus. ALLE SECHS geprueften unsichtbaren Zeichen kamen durch (U+200B U+200C U+200D U+FEFF U+00AD U+2060); nach dem Fix wird jedes gefangen. KEINE SPERRLISTE. Der Fund sagt es ausdruecklich: eine Liste verbotener Zeichen ist die Bauart, die im Nachbarbefund L5-02 versagt hat. Der Kennungsraum wird EINMAL beim Eintritt normalisiert; 'id', 'superseded_by' und die Praesenzmenge laufen durch dasselbe _norm() wie alles andere, was diese Funktion adjudiziert. Eine normalisierte Kennungs-Kollision ist fail-closed, und der Kanal haengt an den Status, damit keiner der beiden typisierten Gruende zur Dekoration wird: verschiedene Status sind ein WIDERSPRUCH (der bestehende Kanal behaelt seine Bedeutung), gleiche eine ANOMALIE. Zusaetzlich: eine Kennung, die zu nichts normalisiert, ist selbst eine Anomalie — sonst kollabierten mehrere still auf denselben leeren Schluessel. KEIN UEBERBLOCKEN, gemessen: das echte Register traegt 17 Findings, deren Kennungen roh UND normalisiert eindeutig sind und sich beim Normalisieren nicht aendern; C12.2 meldet weiter PASS mit 17 ausgewerteten. Neue Regression tests/test_findings_register_identity_axis.py: 8 Tests, 24 Subtests, ueber einen Korpus statt ueber ein Beispiel, mit dem vom Fund verlangten META-Test auf der Severity-Achse (eine Suite, die nur die Kennungs-Achse faengt, hat die Klasse neu aufgezaehlt statt sie zu schliessen) und zwei Gegenrichtungen. Gegen die Vor-Fix-Fassung aus git schlagen 13 Tests fehl, gegen die neue keiner. Eine eigene Schwaeche beim Bauen gefunden und entfernt: die erste Fassung fuehrte ein Kollisions-Set, das befuellt und nie gelesen wurde — eine Variable, die wie ein Riegel aussieht und keiner ist. Das ist die Form des Nachbarbefunds L1-03, und sie faellt bei einem Fix gegen genau diese Klasse doppelt auf. Volle Suite 2020 Tests OK, ruff sauber. NICHT GEMESSEN: der Rust-Verifier kennt diesen Pfad nicht; das Register ist ein Python-seitiges Gate-Artefakt. Die uebrigen Achsen des Fund-Korpus (ASCII-Homoglyphen wie kyrillisches 'a') sind NICHT abgedeckt — NFKC vereinheitlicht sie nicht, und eine Homoglyphen-Tabelle waere wieder die Sperrlisten-Form. Als offener Rest benannt statt still gelassen. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… P1)
loads_strict owns the input_bytes cap, but that cap is a FILE proxy: on the
direct-dict path there are no bytes to measure, so it is inert. Every public
surface that takes an already-parsed structure and then decodes or expands
something proportional to its size was therefore unbounded.
Measured on verify_evidence_pack against the pre-fix file, with a payload ONE
unit over the string_len limit: 2.6 MB peak in 0.095 s before, 0.1 MB in
0.004 s after. (The first measurement used "A"*1000001, which is not valid
base64 — the old path bailed for an unrelated reason and the number said
nothing. Only valid base64 measures the amplification.)
THE CLASS, not the instance. The finding is explicit that wiring a 6th, 7th,
... call site re-opens on the next added surface, so the population is DERIVED
in tests/test_structural_budget_reachability.py and the property is
REACHABILITY in the static call graph. I had measured the family three times
and got three numbers (8/6, the finding's 10/8, 24/22 wide) — that is the
symptom of an enumerated population.
The most important finding of this round is against my own gate. After wiring
anchors.verify_anchor, renewal.verify_sequence dropped off the list without
being touched: renewal.py binds a LOCAL variable named verify_anchor, and my
name-based call graph merged it with anchors.verify_anchor — FALSE COVERAGE. I
had even defended the wide graph in a comment ("a too-narrow edge produces a
false finding"). True, but the other error is the expensive one. The graph is
now module-qualified and resolves conservatively: own module first, else only a
repo-unique definition, never a locally bound name, and no edge at all when
ambiguous.
Six surfaces wired, each mapped to ITS OWN documented failure form (the
precedent both covered members already set: anchors -> BundleFormatError,
policy -> PolicyError):
evidence_pack.verify_evidence_pack result dict over_budget b64decode of proof
dsse._payload_bytes (2 members) BundleFormatError signatures[i].sig per entry
persample.verify_sample_opening BundleFormatError whole proof list decoded, then capped
anchors_chia.verify_offline_merkle result dict bytes.fromhex over unbounded key/value
anchors.verify_anchor BundleFormatError canonicalRoot + proof base64
relation.verify_relationship_edges fail-closed result list walked before its own cap reports
Three times the same shape: ONE dimension was bounded and the surface looked
covered. dsse bounded the signature COUNT, not its SIZE. persample bounded the
disclosure SEGMENT, not the container. anchors bounded the FIELD SET, not the
value sizes. relation bounds the edge count, but only after walking the list,
and `reason` is unbounded free text.
renewal.verify_sequence is excluded WITH PROOF, not wired: it takes
list[list[ArchiveTimeStamp]] — typed objects, not parsed JSON —, the walker
would step over them and measure only list length, which renewal_ats_chain
already bounds more sharply. Wiring it would have turned this gate green
without protecting anything. The exclusion requires evidence: a test asserts
the named budget dimension exists and is actually used in that module.
2099 tests pass, 7 skipped, ruff clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…L5-02, P1)
The gate granted PASS from a discipline MARKER minus a NEGATION blocklist.
That shape cannot terminate against natural language (Ranum; CWE-183
inverted), and measured against the old file it does not even come close:
32 phrasings of "the audit did not happen" -> 20 passed --strict
6 unrelated mentions of the vocabulary -> 3 passed --strict
"the adversarial audit was dropped for this release" passed. So did
"refactored the adversarial fixture loader into its own module". Every fix of
that shape is one more word in the list.
This is the SAME shape as L5-01 one file over — there a blocklist of invisible
characters, here a blocklist of negation words. Both are replaced by
establishing the property instead of forbidding the form.
THE INVERSION, as the finding requires: one canonical attesting line, matched
as a WHOLE line, carrying the version it attests:
pre-tag-adversarial-audit: RUN | version=3.7.0
A negation cannot live inside a closed full-line form, so no vocabulary has to
be enumerated. Two things change:
* PROSE CANNOT MOVE THE VERDICT IN EITHER DIRECTION. The CHANGELOG is
presentational; it neither grants a pass nor withholds one. Measured first:
for 3.7.0 the CHANGELOG already granted nothing (changelog_records_audit
was False), so removing that path blocks nothing real.
* THE RECORD MUST SAY WHICH VERSION IT ATTESTS. Until now any marker-carrying
file under audit_artifacts/<token>/ granted the pass, so a record copied
over from an earlier release attested the new one by sitting in the right
folder.
audit_records_for stays marker-based for its existing consumers (the candidate
matrix, the C12.2 scan); only the new attesting_records_for feeds the verdict.
A MISSING verdict now names the marker-carrying record that failed to attest —
that is the likeliest cause of a surprising refusal, and staying mute about it
is how a correct gate gets called broken and then loosened.
Two existing tests encoded the old behaviour and are updated, not deleted:
test_positive_marker_still_passes asserted that PROSE grants a pass — exactly
the defect — so its INTENT (the gate must not blanket-reject) is kept and its
input becomes a real attestation, plus a new counterpart asserting that the
same prose alone no longer passes.
HONEST LIMIT, in the gate's own comment too: this is provenance-SHAPED, not
provenance. The finding's end state is a runner-signed record whose subject
digest equals the artifact being tagged; this repo has no signing path for that
yet. What is closed is that prose no longer decides.
MIGRATION, declared: I added the canonical line to the existing 3.6.0 and 3.7.0
records. Both already state in prose that the audit ran; the line transcribes
that. It is a transcription by me, not a new audit.
2109 tests pass, 7 skipped, ruff clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…, P1) tests/conftest.py carried a frozenset of 39 test ids that skip outside a git checkout. That list IS the defect: commit 2c5e7a5 had already appended ids to it once, and the gate still measured six MORE tests failing from an extracted sdist at HEAD. A list cannot know about the method somebody adds tomorrow to a module already on it. Measured against the REAL sdist (python -m build --sdist, extracted, run outside the checkout) rather than against an approximation: old conftest (enumerated): 7 failed, 2060 passed, 49 skipped new conftest (derived): 0 failed, 2057 passed, 66 skipped Six of the seven are the ones the finding names (test_intoto_spec_diff). The seventh came from a test file written in this same session — nobody had added it to any list, and the derivation covered it anyway. The question is now answered by measurement: does this module read a ROOT-relative path that does not exist here? Then we are outside a checkout, its assertions are about pruned material, and it SKIPs honestly. TWO DEFECTS IN MY OWN FIX, both found by measuring instead of trusting green: 1. The first version decomposed path CHAINS, so `parents[1] / "src" / "proofbundle"` was read as a root-level `proofbundle`. It flagged 23 modules even in a complete checkout. I then narrowed the rule to the FIRST SEGMENT, which removed the false positives — and also stopped catching the six tests the finding is about, because the sdist prunes LEAVES under shipped directories (docs/ is grafted, docs/IN_TOTO_PROFILE.md is not). The narrowing would have traded a real defect for a comfortable green. The actual bug was the decomposition; with chains joined, the full-path rule is precise. 2. I trimmed the fallback list to three modules because I BELIEVED the rest redundant. Seven tests failed. The set is measurable — empty the list in the extracted sdist, run, and what falls belongs in it. It is six modules, 15 ids, each with a documented reason. _REPO_CONTEXT_TESTS is now what the finding allows it to be: a documented explicit-exclusion list for modules whose repo dependency is not visible as a path literal (import from scripts/, dynamic loading, a glob over .github). Two tests of my own from this session are marked at the point of use with skipUnless rather than growing the central list. OPEN QUESTION, not explained away: published-artifact-gate.yml ALREADY runs the extracted-sdist suite and fails the job on a failure. So the invariant being false at HEAD means that job was either red or not triggered on the measured tree. I did not determine which. 2116 tests pass in the checkout, ruff clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…2-02, P2)
merkle_path (256) WAS enforced — in merkle.verify_inclusion, i.e. AFTER the
line that base64-decodes the entire proof list. Wrong order: for a proof that
breaks the cap and therefore can never be valid, the full work was done first
and the rejection came second. The budget module names "cap before work" as the
pattern in its own docstring.
The L2-01 structural bound added earlier already closed the unbounded case —
8,000,000 entries now fail at json_nodes instead of after ~10 s. But a window
remained between the cap and json_nodes: up to 200,000 entries were still
decoded in full. Two bounds, two different quantities.
Measured (order oracle from the finding):
n= 256 0.001 s ok=False (proof does not bind)
n= 257 0.000 s ok=False "refused before decoding"
n>= 200000 fail-closed at the structural budget
2116 tests pass, ruff clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…uple (L1-01, P2)
evaluation_card_hash and prereg_hash leaked raw exceptions on a non-path
argument — measured across the corpus: OverflowError, TypeError,
FileNotFoundError, 7 raw types per surface. The int case is worse than a wrong
type: os.stat(1) reads a FILE DESCRIPTOR, not a path, so an integer argument
does not merely fail, it inspects stdout.
Widening the except-tuple to catch OverflowError closes exactly one member of
an open set — the next ArithmeticError sibling walks through, and the fd side
effect happens either way because the os call still runs. The floor rejects
before the boundary.
The invariant already existed here: load_bundle implements it verbatim
("bundle path must be a path string, got int (fail-closed)"). It was never
applied to its two siblings.
THREE FIXTURE ERRORS IN MY OWN REGRESSION, all the same class — a number that
spoke about the wrong object:
1. The family test called every surface with ONE argument; verify_prereg takes
two. The resulting "missing 1 required positional argument" counted as a raw
escape: 24 reported failures, none of them about the type floor.
2. Fixed by filling the remaining required args with {} — which made
verify_prereg return "carries no prereg_sha256" BEFORE touching the path.
The test then passed against the PRE-FIX file too, i.e. proved nothing.
3. Only with a claim that actually carries a digest does the corpus reach the
surface. Now: red against the pre-fix file, green against the new one.
2119 tests pass, 162 subtests, ruff clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…03, part) type_confusion_gate classified every non-JSON primary as NON_JSON with the note "covered by tests/test_fuzz_parsers.py". Measured: that file mentions neither evalcard nor prereg. It named a coverage that does not exist — and a named coverage reads exactly like a real one. The note is now RESOLVED: only a test that actually references the surface may be cited, and an entry with no such test says so (deferral_backed: false) instead of citing a file that does not cover it. MY FIRST VERSION OF THE RESOLUTION WAS ITSELF VACUOUS: it accepted `"proofbundle" in source` as a fallback, which matches nearly every test file, so all 26 surfaces came back "backed". A resolution that always finds something is not a resolution — it is the assertion wearing a checker's clothes. Narrowed to require BOTH the function name and the module name. The count did not change (still 0 unbacked), which now means the coverage genuinely exists rather than that the check cannot fail: an invented function name returns [], and prereg.verify_prereg resolves to five real files including the type-floor test written today. HONEST SCOPE: this is the proofbundle half of L1-03. The other half — domain- aware stubs and reachability as a first-class signal in office/governance/berkeley_gate/v4/never_raise_sweep.py — lives in 2bedone, whose push is currently parked on an owner decision. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Fuenf von sechs Linsen der Pflicht-Review-Lane meldeten REJECT. Drei Punkte
sind hier behoben, der Rest ist im Register benannt statt still gelassen.
1. CHANGELOG [Unreleased] behauptete "Nothing under src/ changes, no public
interface gains or loses a field". Gemessen gegen origin/main: acht Dateien,
204 Zeilen unter src/proofbundle/. Der Satz war wahr, als er geschrieben
wurde, und wurde nicht nachgezogen, als der Baum sich bewegte — eine
Aussage, die niemand nachgemessen hat. Sie wird korrigiert statt still
ersetzt, mit der Offenlegung je Aenderungsart, die COMPATIBILITY.md fuer
eine Verschaerfung zuvor akzeptierter Eingaben ausdruecklich verlangt.
Praezedenz dafuer ist 3.2.3 (Finding 15b), das dieselbe Klasse als PATCH
mit ausdruecklicher Offenlegung auslieferte.
2. dsse.py fing als EINZIGE der sechs Flaechen nur BudgetExceeded, waehrend
enforce_structural_budget zwei Geschwister wirft: BudgetExceeded bei
Ueberbreite, BundleFormatError bei Uebertiefe. Heute folgenlos, weil der
Tiefen-Zweig zufaellig genau den Typ wirft, den die Funktion ohnehin
dokumentiert — aber der Kommentar daneben verspricht eine STRUKTURELLE
Eigenschaft, und die haengt dann am Zufall. Die schmale Form stammt aus
Zeile 134, wo sie richtig ist, und wurde auf einen Aufruf mit breiterer
Fehlerflaeche uebertragen. Beide Zweige jetzt gemessen: BundleFormatError.
3. test_path_argument_type_floor.py leitete die Familie ueber
endswith("_path") ab und schloss damit ausgerechnet load_bundle aus, dessen
Parameter schlicht "path" heisst — das Referenzbeispiel aus dem eigenen
Docstring. Zwei Linsen fanden das unabhaengig voneinander. Die Ableitung ist
verbreitert, und ein neuer Test haelt fest, dass das eigene Referenzbeispiel
in der Familie sein muss. Familie 2 -> 3 Mitglieder, 28 -> 42 Subtests.
Die Vorrichtungs-Linse hat MUTATIONSTESTS gefahren und drei der vier neuen
Testdateien als scharf belegt: Riegel entfernt -> Test rot, jeweils mit der
richtigen Meldung. Das ist der Beleg, den ich selbst nicht erbracht hatte.
2120 Tests, 176 Subtests, ruff sauber.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ounds (L2-02, P2)" This reverts commit 2c52596. Owner decision 20260808T1810Z part 4: if no arrangement holds BOTH properties -- cap before the expensive work AND unchanged CLI exit codes -- then 2c52596 is reverted before the PR, as its own revert commit with a reason, not by dropping it from the scope list. 3.7.1 does not become 3.8.0. MEASURED (wf_4d457fb2-1c1, 5 agents, 0 errors, 51 min): none of the three arrangements holds both. p1 raise-instead-of-return: A yes, B no. p2 cheap regex form-check: A no, B no. p3: A no, B yes. MORE IMPORTANT, AND IT CHANGES THE PREMISE: property A does not hold in 2c52596 itself. persample.py:199 calls enforce_structural_budget(opening) one line ABOVE the cap, and _strict_json.py:87-92 walks the list and pushes every element onto a stack -- Omega(n) time AND allocation before the cap is reached. Measured at n=190000: 47.1 ms / 11867 KiB, against 0.087 ms / 2.2 KiB at n=257. A 739x longer list costs 542x time and 5394x memory: linear, not flat. What 2c52596 actually delivers is a constant factor (330 -> 47 ms, 7.0x), not a change of growth class. That Omega(n) path came from 97c929a, the commit immediately before. A and B are mutually exclusive, and that is a lower bound rather than an implementation problem: the old exit code above the cap is 2 exactly when some proof element OR root_b64 would be rejected by b64decode(validate=True), else 1. Deciding "does an invalid element exist" requires reading all n elements in the worst case. A demands O(1). MY OWN EARLIER NUMBER WAS TOO SMALL. I reported ONE divergent input class (invalid base64 above the cap). Measured against the real CLI it is at least TWELVE, including one that no form-check on proof elements can ever see: all elements valid but root_b64 invalid (before 2, after 1). root_b64 is not in the list the cap measures. The CHANGELOG currently states, verbatim, "For every input the outcome is unchanged; only the cost and the detail text differ." That is false for >=12 input classes, and it sits in exactly the paragraph the release gate in RELEASE.md rests on. Precedent already set by this project: docs/release_scope/3.7.1.md kept stash@{0} out for the same reason -- "New behaviour at a public interface is a MINOR, not a PATCH." COMPATIBILITY.md:16 lists "the meaning of exit codes" as one of the four public surfaces. Point 4 (the verdict) is NOT violated: ok stays False either way, nothing flips from fail to pass. Not chosen: p3, which is a revert carrying a new message text -- and that text reads "refused before decoding" after decoding has happened. A false statement to the departing party is worse than no change. L2-02 STAYS OPEN, with the measured numbers and the reason the fix does not fit a patch release. Why the suite never caught the break: no test pins this path. The merkle_path tests hit merkle.verify_inclusion, a different function. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The bullet described 2c52596, which c391117 reverted. Dropping it silently would have been the easier move and the wrong one: the entry carried a claim that was measured FALSE, and readers of 3.7.1 are better served by knowing a change was considered and withdrawn than by finding no trace of it. The false claim was "For every input the outcome is unchanged". Measured against the real CLI, at least twelve input classes above the cap changed exit code from 2 to 1, including one no form-check on proof elements could ever see: root_b64 invalid while every proof element is valid. root_b64 is not in the list the cap measures. This paragraph is the one the release gate in RELEASE.md rests on, which is why a false sentence here mattered more than its size suggests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… beschreibt Die Gegenlesung ueber die Pflicht-Review-Lane hat meine eigene Korrektur von bc3ae70 widerlegt, an drei Stellen. 1. Das Banner nannte "204 lines". Gemessen an bc3ae70 sind es 8 Dateien, 196 Insertions, 3 Deletions. 204 war der Stand bei 239f5aa, also EIN Commit vor dem, der den Satz schrieb — die Zahl war beim Schreiben schon um 9 daneben, und c391117 hat weitere 17 entfernt. Ein Absatz, der sich woertlich mit "a statement nobody re-measured" begruendet, trug selbst eine nicht nachgemessene Zahl. Eine Zaehlung gegen einen wandernden Zweig ist nur an einem benannten Ref wahr; sie nennt jetzt Ref und Kommando. 2. "the release gate in RELEASE.md rests on this paragraph" war zu hoch gegriffen. Gemessen: check_version_and_changelog.py liest NUR Ueberschriften, und die release-scope-Checkbox liest ein Mensch. Es gibt keinen Riegel auf diesem Absatz — was die Sache schlimmer macht, nicht besser, und deshalb steht es jetzt so da. 3. Vier Zahlen standen ohne Quelle, gegen den Maßstab des Repos selbst ("EVERY NUMBER NAMES ITS OBJECT AND ITS SOURCE", check_version_and_changelog.py:23). Die Methode steht jetzt dabei. Die Speicherzahl wurde unabhaengig auf die KiB bestaetigt; die Zeitzahlen sind RAUS, weil zwei Messungen desselben Codes um 28% auseinanderlagen und eine hostabhaengige Zahl in einem Changelog nichts belegt. "Zwoelf Klassen" bleibt als untere Schranke, mit dem Vermerk, dass "input class" hier keine definierte Einheit ist und eine unabhaengige Zaehlung auf 22 kam. Riegel gefahren: check_version_and_changelog exit 0, claims_hygiene PASS. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…e claim being early QITEM-PB-ASSURANCE-SICHTBAR-01 requires all four unpursued checks to stand in the Scorecard section as named causes; Binary-Artifacts (9/10) was missing. Measured 2026-08-10 via the Scorecard API: the deducted point is the checked-in reproduction fixture dist_final/ (wheel+sdist) — the sentence names the measured object. Two claims corrected against measurement: - "Every release is attested" — the three corpus-review pre-releases in the check's five-release window are not; scoped to version releases (v*), whose workflow has carried the attest step since v0.1.0. - "The provenance bundle is now attached as a release asset too" — measured v3.7.0 assets: wheel, sdist, SHA256SUMS. The workflow change takes effect with the NEXT release; published releases were not modified. Re-measured 2026-08-10T07:35:59Z (api.scorecard.dev, commit 6a3011f): 6.5/10, every per-check value identical to 2026-08-07. Owner-GO: GO_OWNER_PB_ASSURANCE_20260807 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Bringt die drei heute gelandeten PRs (#137 expected-origin in der CLI, #138 der gemessene ML-DSA-Grund, #135 Dependabot) auf den 3.7.1-Zweig. Ein Konflikt, in .github/workflows/release.yml, und keine Seite hatte allein recht: - `id: provenance` stammt vom Zweig und wird gebraucht — Zeile 102 liest `steps.provenance.outputs.bundle-path`. Ohne die id ist der Ausdruck leer. - `actions/attest-build-provenance` pinnt main auf 4d101475 (v4.2.2), der Zweig auf 0f67c3f4 (v4.1.1). Die neuere Fassung gewinnt, sie kam aus #135. Aufgeloest als beides: id behalten, Version von main uebernehmen.
Die acht roten Checks von #139 sind acht Laeufer derselben Suite mit EINEM Fehler: FAIL: test_im_echten_checkout_ist_die_ableitung_ein_no_op AssertionError: Lists differ: ['test_relation_statement_rust_parity'] != [] ZWEI ARTEN VON ABWESENHEIT, und die Ableitung hatte eine Regel fuer beide. `test_relation_statement_rust_parity` nennt `tools/pb_verify_rs/target/release/pb_verify_rs`. Diese Datei fehlt auch im VOLLSTAENDIGEN Checkout -- bis jemand `cargo build` laeuft. Ihre Abwesenheit sagt nichts darueber, ob wir in einem sdist sind, und das ist die einzige Frage, die diese Ableitung stellt. Gemessen sind fuenf der neun wurzel-relativen Pfade dieses Moduls Build-Ausgaben unter `target/`. DIE TRENNENDE EIGENSCHAFT IST DIE IGNORE-REGEL DES REPOS, nicht eine Liste von Verzeichnisnamen. Gemessen: tools/pb_verify_rs/target/** IGNORIERT (5 Pfade, alle Build-Ausgaben) tools/pb_verify_rs/crosscheck.py nicht (echte Quelldatei) docs/IN_TOTO_PROFILE.md nicht (der geprunte Blattfall, fuer den die Ableitung existiert) Verzeichnisnamen aufzuzaehlen (`target`, `build`, `dist`, ...) waere wieder Formen sammeln -- genau davor warnt der Kommentar in `modul_ist_repo_kontext` selbst. NUR IM CHECKOUT WIRKSAM: in einem entpackten sdist gibt es kein git, und dort IST Abwesenheit das richtige Signal. Jeder Fehlschlag (git fehlt, kein Repo, exit != 0) faellt fail-safe auf "kein Bauartefakt" zurueck, also auf das bisherige, strengere Verhalten. GEMESSEN, beide Richtungen: Skip-Menge im Checkout ['test_relation_statement_rust_parity'] -> [] sdist-artig (docs fehlt) weiterhin True <- die Ableitung misst noch vollstaendiger Baum False Ruecknahme-Probe ohne den Fix 4 rot, mit ihm 10 gruen volle Suite auf diesem Zweig 2042 passed, 117 skipped, 0 failed EIN EIGENER MESSFEHLER, festgehalten weil er heute zum zweiten Mal passiert ist: mein erster Suitenlauf meldete 31 Fehler, darunter "ein gefaelschter Beleg wurde VERIFIED". Das venv war fuer den Release-Kandidaten gebaut und zeigte per `pip install -e` dorthin -- ich habe #139er Tests gegen Kandidaten-Quelltext gefahren. Mit einer Umgebung, die auf DIESEN Baum zeigt: 0 Fehler. Aufgefallen ist es nur, weil CI und meine Messung sich widersprachen. Dieselbe Wurzel wie die Werkzeugkette, die heute frueh zweieinhalb Stunden lang main gemessen hat: eine Messung, die ihren Gegenstand nicht festhaelt, ist von einer richtigen nicht zu unterscheiden. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
b7n0de
pushed a commit
that referenced
this pull request
Aug 16, 2026
… als der Code
Owner-Entscheid 2026-08-16: "beheb auch gleichzeitig noch die anderen fehler die
aufgefallen sind". Sechs Linsen, ein Gate-Meta-Test mit neun Pflanzungen. Das
Ergebnis in einem Satz: der PRUEFGEGENSTAND hielt jeder Messung stand, die AKTE
UEBER IHN nicht.
## Am Code
ORIGIN-VERGLEICH (B/C aus dem Meta-Test). Zwei eingepflanzte Lockerungen --
`==` -> `.startswith()` und case-insensitive -- wurden von KEINEM der 2030 Tests
gefangen. Der Grund: beide Origin-Tests prueften nur einen VOELLIG FREMDEN Wert,
und dagegen verhaelt sich ein gelockerter Vergleich wie ein exakter. Der
Beinahe-Treffer fehlte, und er ist im Feld der gefaehrliche -- wer einen eigenen
Log betreibt, waehlt dessen Namen selbst.
neu: OriginVergleichIstExakt mit 11 Beinahe-Treffern (Praefix, Suffix,
Gross/Klein, fuehrendes/folgendes Leerzeichen, Zeilenumbruch, Schraegstrich,
Schema, nur-Domain, leer) plus die Gegenrichtung, ohne die ein IMMER-FALSCH-
Vergleich ebenfalls gruen waere.
plus zwei Mutations-Operatoren in scripts/mutation_check.py.
GEMESSEN, beide einzeln gefahren: startswith -> 3 rot, casefold -> 2 rot,
sauber wieder gruen. Ein Korpus ohne Mutant ist eine Behauptung ueber sich selbst.
KONTROLLZEICHEN. `_safe_line()` steht in derselben Datei und wurde in
`_cmd_verify` sechsmal benutzt, in den sieben anderen verify-Funktionen null Mal.
Der `origin` kommt aus der geparsten Beweisdatei. Mitgefixt: `sample-opening` und
`enclave-attestation` (`enclave.py:127` setzt `detail = f"malformed EAT token:
{exc}"`). NICHT mitgefixt und bewusst: `{'OK' if x else 'FAIL'}` ist ein Literal,
`ERROR: {exc}` geht nach stderr -- eine Fundstelle ist kein Defekt.
DAS VERDIKT NENNT DIE ERWARTUNG. Gemessen lieferten drei verschiedene Ursachen --
fremder Origin, falscher log-vkey, verfaelschte Signatur -- BYTE-IDENTISCHES JSON
(gleicher sha256). `inclusion_ok` bleibt in allen dreien True und trennt nichts.
Neu: `expected_origin` im JSON, `None` wenn nicht gefragt (nicht Leerstring --
"nicht gefragt" und "gefragt und leer" sind zwei Lagen).
DER TREFFER-ZWEIG hatte keinen Waechter: der einzige Text-Modus-Test fuhr den
FEHLSCHLAG, alle anderen `--json`. `" (expected)"` war die einzige Verhaltenszeile
des Release ohne Zusicherung.
DER FALSCHE KOMMENTAR IM AUSGELIEFERTEN QUELLTEXT. `cli.py:949` sagte
"expected_origin since 3.6". Gemessen: `git log -S expected_origin --
tlogproof.py` nennt genau einen einfuehrenden Commit, 457b6b8 = v1.3.0. Der
CHANGELOG hatte recht, der Quelltext nicht -- und er wird ausgeliefert.
## An der Auslieferung
DER DOI WURDE VOR DEM OWNER-TOR GEPRAEGT. Gemessen: aktiver Zenodo-Webhook auf
`events=['release']`, das Release entstand ohne `draft:` VOR `publish-pypi`.
Jetzt: Entwurf -> PyPI-Freigabe -> `publish-release` dreht ihn oeffentlich. Bleibt
die Freigabe aus, gibt es keinen DOI. Fail-closed.
2 353 682 BYTES FREMDER 3.6.1-ARTEFAKTE lagen verfolgt in `dist_final/` und
`dist_pkgtest6/` -- am echten GitHub-Archiv von v3.7.0 nachgemessen 19,3 % des
unkomprimierten Quell-Archivs, und ueber den Webhook im zitierbaren Datensatz. Im
sdist waren sie NICHT (die MANIFEST.in-Allowlist hielt). Entfernt + gitignored.
SHA256SUMS trug den `dist/`-Praefix, und RELEASE.md schickt Nutzer genau damit
pruefen: `sha256sum -c SHA256SUMS` meldete fuer beide Zeilen "No such file or
directory / FAILED". Jetzt relativ, Probe gefahren: beide Zeilen OK.
DOKUMENTATION. `--expected-origin` kam in README, SPEC.md, RELEASE.md,
INTEGRATIONS.md und CROSS_IMPLEMENTATION_REPORT.md null Mal vor. Das Release machte
damit `docs/TRUST_ANCHORS.md` unvollstaendig -- die Zeile nennt sich selbst "the
whole trust surface" und war VOR 3.8.0 vollstaendig, weil es keine CLI-Form gab.
Ergaenzt dort und im ausgearbeiteten Beispiel in PUBLIC_TRANSPARENCY_PROFILE.md.
PROSA-VERSIONEN. RELEASE.md und docs/readiness_pack/PROGRESS.md auf 3.8.0. Das zog
eine Manifest-Drift im readiness_pack nach sich (PROGRESS.md ist gepinnt) -- neu
erzeugt, `--check` OK. Der oeffentliche Schluessel wechselte dabei; das ist Bauart,
der Docstring sagt "signed with an EPHEMERAL key generated at [generate time]".
## An der Akte -- und hier lag das meiste
DIE SUITE-ZAHL STAMMTE AUS DER FALSCHEN UMGEBUNG. Die Akte deklarierte
`[pq,pytest,test,dev,anchors]` und berichtete daraus `1 failed, 1970 passed, 116
skipped`. Gemessen sind 113 der 116 Skips `[anchors]`-gegated -- mit installiertem
Extra waeren sie GELAUFEN. Die Zahl kam aus einer Umgebung ohne. Genau dieser
Fehlermodus ist in §4 als RT-08 praeregistriert; die Akte hat ihn benannt und im
selben Lauf begangen.
NACHGEMESSEN mit `[dev,eval,anchors,pq]`: 1 failed, 2085 passed, 9 skipped.
MEINE EIGENE LOCKERUNG §9 WAR WEITER ALS BEHAUPTET. Sie band das Verdikt an
`src/ + tests/ + scripts/` und nannte das "der Code". ZWEI der DREI Orte, die
`check_version_and_changelog` als Versions-Wahrheit erzwingt, lagen ausserhalb --
`pyproject.toml` und `CITATION.cff`. Ausfuehrbar gezeigt an `ed8c3b5`, das IN
diesem Delta liegt: es wechselt den blockierenden Linter und die §9-Messung meldet
sauber. Der Satz "the escape hatch is measurable, so it cannot be argued open" ist
als Formulierung falsch -- die Luke muss nicht aufargumentiert werden, sie war
konstruktiv offen.
Abschnitt 10 angehaengt (§9 bleibt stehen): die Bindemenge umfasst jetzt auch
conformance/ schemas/ formal/ examples/ pyproject.toml CITATION.cff MANIFEST.in
.github/workflows/ -- mit einer REGEL statt einer Liste, und mit der ehrlichen
Grenze, dass die dauerhafte Form die Umkehrung waere (alles bindet ausser einer
begruendeten Ausschlussliste).
ERSTE MESSUNG unter der korrigierten Regel: `release.yml` und `cli.py` HABEN
sich seit f64d35e bewegt. Das Verdikt bindet diesen Digest also nicht mehr --
die richtige Konsequenz meiner eigenen Korrektur.
WEITERE FALSCHE ZAHLEN, je mit der Messung korrigiert:
"27 Commits auf #139" -> reproduziert NUR gegen eine veraltete lokale Ref vom
12.08.; zum Freeze 28, heute 29 mit / 27 ohne Merges
"23 Dateien / 8 Commits" -> Dateien stimmen; die 8 zaehlt Merges mit, ohne sie
sind es 6, und die danebenstehende 13 hatte ausserdem
einen anderen Endpunkt. Drei Unterschiede auf einmal,
keiner genannt.
"all seven reported fields" -> es sind acht; die Aufzaehlung liess ausgerechnet
`witnesses` weg. Seit dieser Runde neun.
"11 Flaechen / 7 Module" -> top-level-only; mit Unterpaketen 12/8. Der Sweep,
der eine handgepflegte Liste als zu eng entlarvt, war
selbst zu eng.
FINDING nannte keinen Digest, obwohl §9 es verlangt -> nachgetragen (f64d35e
bzw. ac0688c fuer die main-Messungen).
DIE KLASSE HINTER ALLEN: eine Zahl ueber eine Population messen und ueber eine
andere berichten. Der CHANGELOG benennt sie selbst -- und begeht sie danach
viermal weiter. Der Klassen-Fix waere, dass jede Zahl ihren Befehl mittraegt, so
wie `pyproject.toml:104-106` es mit dem gepinnten Baum vormacht.
## NICHT behoben, bewusst, mit Owner-Entscheid
(F) Eine 12-Byte-Datei mit dem Wort `adversarial` kippt `pre_tag_audit_gate` von
MISSING auf OK. Der Riegel misst Prosa, nicht Wahrheit; in dieser Bauform ist das
nicht reparierbar. Der Fix waere ein runner-signiertes Receipt ueber die
auditierte SHA -- Tage, nicht Stunden.
(E) Kein Semver-Oberflaechen-Gate: eine Patch-Nummer fuer ein Release, das die CLI
aendert, faellt durch alles.
Owner 2026-08-16: beide als Befund fuehren.
## Gemessen, Abschluss
volle Suite ([dev,eval,anchors,pq]) 1 failed, 2085 passed, 9 skipped
der eine rote ist der BEABSICHTIGTE Audit-Eintrag (TestF7PreTagAudit)
ruff All checks passed
check_version_and_changelog OK
doc_link_check PASS, 91 Links, 0 broken
claims_hygiene PASS, 49 Docs, 0 Verstoesse
readiness_pack_manifest --check OK
type_confusion_gate --strict never_raise_ok=True
test_manifest_gate 2095 gesammelt (Boden 1750)
pre_tag_audit_gate --strict exit 1 <- MUSS rot sein, kein Verdikt
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…e Eigenschaft Der hermetic-cleanroom-Job faehrt die Suite aus dem ENTPACKTEN sdist. Dort gibt es kein git, also gibt `_ist_bauartefakt` fail-safe `False` zurueck -- genau das Verhalten, das der DRITTE Test dieser Klasse ausdruecklich vorhersagt. Die erste Fassung des ersten Tests behauptete `True` unbedingt und fiel deshalb: AssertionError: das Rust-Binary gilt nicht als Bauartefakt -- der Fall kehrt zurueck 1 failed, 1975 passed, 183 skipped Der Test mass damit die UMGEBUNG statt der Eigenschaft. Dieselbe Klasse, gegen die diese ganze Datei steht, eine Ebene hoeher -- und ich hatte die Bedingung im ersten Test schlicht nicht gesetzt, obwohl sie im dritten steht. `_in_git_checkout()` fragt dasselbe, was `_ist_bauartefakt` intern fragt, damit die Bedingung des Tests und die des Codes nicht auseinanderlaufen koennen. Uebersprungen wird NICHT stillschweigend: der Skip sagt, dass die Ignore-Regel hier nicht messbar ist, und nicht messbar ist keine Freigabe. Die Eigenschaft selbst haelt der Cleanroom-Job weiter -- ueber `test_ohne_git_bleibt_das_strengere_alte_verhalten`, nur von der anderen Seite. GEMESSEN: im Checkout 10 passed (scharf), im git-losen `git archive`-Baum uebersprungen statt rot. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
b7n0de
pushed a commit
that referenced
this pull request
Aug 16, 2026
Das Pre-Tag-Tor ist mit --strict der ERSTE Schritt des Release-Workflows und
damit der letzte blockierende Riegel zwischen Tag und PyPI. Eine Linse hat
gemessen, dass es mit dem, wonach es benannt ist, in KEINER Richtung
korreliert. Damit haengt die Frage "kann dieses Release ehrlich geschnitten
werden" daran, ob das Tor durch eine wahre Aussage und NUR durch eine wahre
Aussage erfuellbar ist.
BEIDE Fassungen nebeneinander geladen, gegen denselben Wegwerf-Baum, ein
Record je Vektor. GEGENPROBE ZUERST, weil ein Tor, das immer MISSING sagt, auf
der Falsch-Gruen-Achse perfekt aussaehe: bei leerem Record melden BEIDE MISSING.
A zwoelf Saetze, die einen NICHT gelaufenen Audit behaupten
(mehrere Sprachen; Kommentar, Code-Zaun, Front-Matter, Durchstreichung,
Frage, URL) heute 0/12 richtig · #139 12/12
B vier ehrliche Attestierungen in unserer Hausform
heute 0/4 · #139 0/4
C die kanonische Attestierungszeile beide richtig
D dieselbe Zeile mit der Version eines FRUEHEREN Release
heute akzeptiert (falsch) · #139 abgelehnt
KLASSE A ist der Release-Blocker, und #139 schliesst ihn vollstaendig. Die
heutige Marker/Negations-Paarung ist eine AUFZAEHLUNG: sie listet, wie ein Satz
etwas verneinen kann, und jede nicht gelistete Form liest sich als Behauptung.
Zwoelf von zwoelf kamen durch. #139 dreht die Polaritaet — EINE geschlossene
Vollzeilen-Form, und in einer geschlossenen Form kann keine Verneinung wohnen,
also muss kein Vokabular aufgezaehlt werden.
KLASSE D hatte niemand gemeldet. Bis #139 war der versionsbenannte Ordner der
einzige Anker, also attestierte ein aus einem frueheren Release kopierter Record
das neue, indem er im richtigen Verzeichnis lag.
KLASSE B sieht unveraendert aus und ist KEIN Defekt: in #139 ist Prosa
ausdruecklich praesentational und kann das Verdikt in keiner Richtung bewegen.
Die Kategorie loest sich auf, statt behoben zu werden. Unter dem heutigen Tor
werden dieselben vier Saetze aus dem GEGENTEILIGEN Grund abgelehnt — jeder
enthaelt ein Wort, das die Negationsliste als Verneinung liest. Ein
wahrheitsgemaesser Record in unserem eigenen Stil wird abgelehnt, waehrend
zwoelf unwahre durchkommen.
WAS DAS BELEGT: die bereits gewaehlte Merge-Reihenfolge ist tragend, nicht
kosmetisch. Ohne #139 heisst "das Tor erfuellen", einen Satz zu schreiben, der
ein Markerwort traegt und eine Liste anderer vermeidet — also um eine Pruefung
herumzuschreiben statt eine Tatsache zu attestieren.
WAS ES NICHT BELEGT: dass ein Audit gelaufen ist. #139 nennt seine eigene Grenze
klar — die kanonische Zeile ist provenance-FOERMIG, nicht Provenance.
Die Datei traegt bewusst KEIN Markerwort: die Vektoren wuerden das heutige Tor
kippen, und sie im Release-Record auszuschreiben hiesse, den berichteten Defekt
darin zu reproduzieren. Gegengeprueft: 0 Marker, Tor weiter rc=1, doc_link 91/0,
claims 49/0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
b7n0de
pushed a commit
that referenced
this pull request
Aug 16, 2026
…euert wie entworfen Owner-GO im Chat: "owner go erteilt fuer den merge sorge fuer fortschritt und wir inkludieren alles jetzt best moeglich in das v3.8.0 release". #139 auf main gemergt (518d1ee), main hier zusammengefuehrt. DIE LATTE HAT GETAN, WOFUER SIE GEBAUT WURDE. Gestern trugen P1, P2 und P3 in `test_pre_tag_gate_eigenschaften.py` ein `expectedFailure` mit der Begruendung: "sobald das Tor sie schliesst, meldet unittest UNEXPECTED SUCCESS und zwingt zum Entfernen der Markierung". Genau das ist eingetreten — mit #139s neuem Tor meldete die Suite drei Unexpected success, also drei FEHLSCHLAEGE. Markierungen entfernt, die drei Eigenschaften sind jetzt normale Zusicherungen. Der Absatz bleibt im Docstring stehen, weil er der Beleg ist, dass die Bauform funktioniert: ein `skip` haette geschwiegen, als die Arbeit getan war. CHANGELOG ZUSAMMENGEFUEHRT OHNE VERLUST. Mains `[Unreleased]`-Rubrik trug vier Stichpunkte, die meine Fassung nicht hatte (Strukturbudget auf dem direct-dict-Pfad, zurueckgenommene merkle_path-Kappe, typisierte Fehler auf zwei Pfadargumenten, plus den Banner). GEMESSEN, dass der zugehoerige Code in DIESEM Baum liegt (BundleFormatError in evalcard und prereg vorhanden, json_nodes im Budget) — er wird also mit 3.8.0 ausgeliefert und ist nicht "unveroeffentlicht". Die Rubrik-Ueberschrift entfaellt deshalb, der Wortlaut ist unveraendert uebernommen, mit sichtbarer Herkunftsangabe und die ###-Rubriken am Ende des Abschnitts, damit sie nicht mit den gleichnamigen dieses Release verwechselt werden. STAND JETZT, ehrlich: zwei rote Tests bleiben, beide aus DEMSELBEN Grund und beide richtig. `TestF7PreTagAudit::test_released_version_has_audit_record` und #139s eigener Gegentest `ProsaEntscheidetNicht::test_gegenrichtung_das_echte_repo_besteht_weiterhin` verlangen die kanonische Zeile `pre-tag-adversarial-audit: RUN | version=3.8.0` in der Akte. Sie steht dort NICHT — und sie zu schreiben waere heute eine falsche Attestierung: gelaufen ist eine 3-Linsen-Runde auf dem Origin-Bindungs-Increment (NORMAL 3L/3I), nicht der Tiefenlauf auf dem eingefrorenen Release-Digest. Der Eintrag folgt, wenn der Lauf gelaufen ist, nicht vorher. Batterie: 2232 passed, 9 skipped, 380 subtests.
b7n0de
pushed a commit
that referenced
this pull request
Aug 17, 2026
The mutation gate reported two gaps on 518d1ee and blocked v3.8.0: GAP [relation: cycle detection disabled] SURVIVED (red=1) GAP [relation: verified-flag laxened] SURVIVED (red=1) NEITHER IS A COVERAGE HOLE. Both are the same stale-operator shape, and the cause is a fix that was right. #139 made the ancestor walker apply the gates the direct-edge arm already had (L4-01: the gate was distance-scoped, and the distance is attacker-chosen). It also added a look-ahead so a back-edge is found before descending, making the cycle code independent of sibling order. Both were deliberate defence in depth. The side effect: each operator disables ONE line where the property now rests on TWO. The surviving guard catches the vector, the mutant lives, and the gate cries gap where the defence holds. An operator whose label says "disabled" must actually disable -- otherwise the gate measures the wrong thing in the SAFE direction, and that is how a gate gets ignored. MEASURED, each direction separately, on an extracted tree with the release venv: cycle, line 367 alone -> suite green (look-ahead still catches) cycle, look-ahead alone -> suite green (line 367 still catches) cycle, BOTH -> 4 red, among them test_injected_back_edge_onto_path_is_caught_or_unreachable verified, direct arm alone -> suite green (walker catches one level down) verified, BOTH -> red, tests/test_relation_profile.py:: TestVerifyRelationshipEdges::test_verified_flag_must_be_exactly_true So the killing tests EXIST and always did. They simply cannot kill a half-disabled property. THE FIX is in the operator table, not in the source: an operator may now name several sites (old/new as tuples). Single-string operators are unchanged, so the other 86 keep their exact meaning. Verified with the gate's own machinery on a filtered set: both now report KILLED (red=4 and red=20), 0 gaps. ALSO ADDED, and honestly not required for the above: three truthy-but-not-True vectors (1, "true", ["ja"]) in the distance-invariance generator. The file's own header says every chain test hardcodes "verified": True, so the strictness of `is not True` was untested ACROSS HOP DISTANCES. test_verified_flag_must_be_exactly_true already covers the direct arm; these extend the property to every hop, which is what this file is for. They raise the mutant's red count from 2 to 20 but are not what makes it die. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Was & warum
Dieser Zweig traegt die Vorbereitung fuer 3.7.1 und die auf proofbundle angewandten Ergebnisse des adversarialen Deep-Gate-Laufs: 27 Commits, 38 Dateien, ~2950 eingefuegte Zeilen, entstanden zwischen dem 7. und 12. August.
Er war bisher gepusht, aber ohne Pull Request — und damit ohne einen einzigen CI-Lauf auf seiner Spitze. Das ist der eigentliche Grund, warum 3.7.1 seit dem 12.08. nicht weitergekommen ist: nicht ein rotes Gate, sondern ein fehlender Weg, auf dem die Arbeit ueberhaupt geprueft werden koennte. Dieser PR stellt den Weg her.
Die Fixes mit ihren Gate-Kennungen
97c929a264fcd3f112710239f5aa673baa62c52596→c391117bc3ae70Dazu Identitaets-Achse im Findings-Register (
4bb09c0), relation gates an jedem Hop (ee356c3), Provenance verlaesstdist/und das Publish-Gate wird eine Allowlist (9e523ba), die undeklarierte Versionsstelle (768e299) sowie zehn Doku-Commits.Der Merge von
mainBeim Nachziehen von
main(mit #135, #137 und #138) gab es genau einen Konflikt, in.github/workflows/release.yml, und keine Seite hatte allein recht:id: provenancestammt von diesem Zweig und wird gebraucht — Zeile 102 lieststeps.provenance.outputs.bundle-path.actions/attest-build-provenancesteht aufmain(via ci: bump the actions group with 4 updates #135) auf4d101475(v4.2.2), hier auf0f67c3f4(v4.1.1).Aufgeloest als beides: die
idbleibt, die neuere Action-Version wird uebernommen.Was dieser PR NICHT ist
Kein Release. Kein Tag, kein PyPI-Upload, keine Versionsanhebung. Der Owner-GO fuer 3.7.1 existiert seit dem 07.08. und ist an sieben gemessene Vorbedingungen gebunden; zwei davon (Schnittliste auf dem Release-Stand, CI gruen) sind ohne diesen Merge gar nicht messbar. Der Tag kommt spaeter und auf dem gemergten
main, nachRELEASE.md.Ehrliche Grenze
Der Inhalt der 27 Commits ist in dieser Runde nicht Zeile fuer Zeile nachgelesen worden. Dieser PR macht sie pruefbar, er behauptet nicht, sie seien geprueft. CI laeuft hier zum ersten Mal ueber sie.