Skip to content

feat(proofs): Agda formal verification of the numeric core (issue #1) - #79

Merged
hyperpolymath merged 4 commits into
mainfrom
arena/01a0db37-metamanifold-webui
Sep 26, 2026
Merged

hyperpolymath merged 4 commits into
mainfrom
arena/01a0db37-metamanifold-webui

Conversation

@arena-ai-coding-agent

Copy link
Copy Markdown

What this is

A proof gate for the validated statistics layer, as asked for on #1 and in the
follow-up direction to validate the system with Agda (Lean fallback not needed).

Seven Agda modules, 77 top-level definitions, type-checked by Agda 2.7.0.1
against agda-stdlib 3.0 under --safe --without-K. No postulates, no foreign
code, no proof-irrelevance escape hatch.

Additive only: no existing source file's behaviour changes. The diff is new files
plus appended Justfile recipes, one .gitignore entry, and a new workflow.

Verified locally

Check Result
agda --safe --without-K MetaManifold/All.agda (clean _build) exit 0, empty output — no warnings
proofs/bootstrap.sh (the repo's own entry point) exit 0 → proofs: OK
proofs/tests/axiom-audit.sh 7/7 modules reachable, clean, exit 0
proofs/tests/gate-selftest.sh 10/10 controls behaved correctly, exit 0

Not yet verified: .github/workflows/proofs.yml has not executed on a GitHub
runner. The first run of this PR's checks is its first real test. If it goes red
it will be the bootstrap step (PyPI wheel + stdlib clone + primitive libraries),
not the proofs.

What is proved, and why it is not a test

  • Prelude — Outcome A = value A ⊎ refused Refusal, with outcome-total
    saying there is no third arm. A refactor that forgets a refusal case does not
    type-check. whole-is-one needs no n ≢ 0 hypothesis, because ℚᵘ's
    denominator is suc _ — a zero denominator is unrepresentable by type. That is
    the exact sense in which the refusals are total, and the reason Agda was used.
  • Proportions — relativeAbundance refuses in the order the Julia layer
    checks; proportions sum to exactly 1ℚᵘ; a returned value is c/t.
  • ExactCounts — checkedAdd is exact in both directions: a returned value
    is the true sum, and a refusal happens exactly when the sum leaves the range.
    Proving one direction alone would be worthless, because each alone is satisfied
    by an implementation that is useless. Plus checkedSumOf-is-exact, the fold
    version a table's total rests on.
  • PermutationTest — never-reports-zero holds for every b and B, so
    p = 0 is unprintable by construction rather than by convention. The negative
    control is proved alongside it: naive-estimator-can-report-zero shows the
    estimator this replaces really does return exactly zero.
  • BenjaminiHochberg — the scaling is exactly (M/(j+1))·(n/(d+1)), and a
    bigger family can never make a q-value smaller (a correction that went the
    other way would reward running more tests, silently).
  • DecimalRounding — "correctly rounded to s decimals" as a predicate over
    integers alone, no division inside the specification of rounding. Each printed
    decimal is a type-checked instance; 0.67 for 2/3 is not a number somebody
    typed into two places. Tie handling is stated in both directions
    (IsRounding halves go down, IsRoundHalfUp up) and they are proved to agree
    everywhere else.

A finding, not an assumption

relativeAbundance (suc c) 0 refuses as countExceedsTotal, not
zeroTotal — exact_relative_abundance checks count <= total before
iszero(total). Both orders are defensible; only one is implemented. A test
written from the docstring would have asserted the other.

The gate cannot pass vacuously

proofs/tests/gate-selftest.sh breaks the proofs nine ways on purpose (wrong
known-answer digit, reversed monotonicity, plus-one removed, zero total silently
zero, overflow no longer refused, module dropped from the gate entry, postulate
injected, --safe removed, tie forced the wrong way) and requires each to be
rejected. On its first run it reported 1/10 — because bash -c "$mutator" "$file" binds the path to $0, so no mutation applied. The control reported
"mutation did not apply" rather than passing. That is the point of having it.

Not proved

proofs/residue/ carries every open obligation with an id, a precise statement,
and what would close it — including Checked.refused injectivity, BH
non-negativity, the step-down min envelope, and explicitly: no probability
theory, no IEEE-754 semantics, no proof that numeric_policy.jl implements the
model, nothing about the :ordinary path.
q_i ≥ p_i is false in general and
is deliberately absent.

The Agda-to-Julia bridge is test/fixtures/agda-known-answers.json:
Agda-checked vectors, each annotated with the lemma that fixes it, for the Julia
conformance testset to reproduce. If the two disagree, the Julia layer is wrong.

Docs

  • proofs/PROOF-STATUS.md — full inventory, per-module results, open obligations
  • docs/statistics/formal-verification.md — reviewer-facing: what "validated
    with a proof assistant" does and does not mean here

Refs #1. Does not unblock #2 or #4, which stay gated on the full validation of
#1 and owner approval.

Adds a proof gate for the validated statistics layer: seven Agda modules,
77 top-level definitions, type-checked under --safe --without-K against
agda-stdlib 3.0 with Agda 2.7.0.1.

What is proved, and why it is not a test:

* Prelude -- Outcome A = value A | refused Refusal, with outcome-total saying
  there is no third arm. A refactor that forgets a refusal case does not
  type-check. whole-is-one needs no `n =/= 0` hypothesis because Qu's
  denominator is `suc _`: a zero denominator is unrepresentable by type.
* Proportions -- relativeAbundance refuses in the order the Julia layer checks,
  proportions sum to exactly 1, a returned value is c/t exactly.
* ExactCounts -- checkedAdd is exact in both directions: a returned value IS the
  true sum, and a refusal happens exactly when the sum leaves the range. Plus
  the fold version, checkedSumOf-is-exact, which is what a table's total rests on.
* PermutationTest -- the plus-one estimator never reports zero, for every input,
  alongside a proved negative control showing the estimator it replaces does.
* BenjaminiHochberg -- the q-value scaling is exactly (M/(j+1))*(n/(d+1)), and a
  bigger family can never make a q-value smaller.
* DecimalRounding -- "correctly rounded to s decimals" as a predicate over
  integers alone, with machine-checked known-answer vectors (0.67 for 2/3, etc.)
  and the tie-handling contract stated in both directions.

Finding, not assumption: relativeAbundance (suc c) 0 refuses as
countExceedsTotal, not zeroTotal, because exact_relative_abundance checks
count <= total before iszero(total). A test written from the docstring would
have asserted the other.

The gate cannot pass vacuously:

* proofs/bootstrap.sh pins Agda 2.7.0.1 / stdlib v3.0 and exits non-zero if it
  cannot install them. An absent prover is a failure, never a skip.
* proofs/tests/axiom-audit.sh rejects postulates, FFI, unsound flags, holes,
  missing --safe, and any module not reachable from All.agda.
* proofs/tests/gate-selftest.sh breaks the proofs nine ways on purpose and
  requires each to be rejected. Verified locally: 10/10 controls, exit 0.
* .github/workflows/proofs.yml runs all three plus a check that PROOF-STATUS.md
  matches the tree, and uploads the transcript whether it passed or not.

Everything not proved is in proofs/residue/ with an id, a precise statement, and
what would close it -- including that no probability theory, no IEEE-754
semantics, and no Julia correspondence is claimed.

test/fixtures/agda-known-answers.json carries the Agda-checked vectors into the
Julia conformance testset: if the two disagree, the Julia layer is wrong.

Verified locally: agda --safe --without-K MetaManifold/All.agda -> exit 0 with
empty output; axiom-audit -> clean, 7/7 reachable; gate-selftest -> 10/10;
proofs/bootstrap.sh -> OK. Not yet verified: the workflow has not executed on a
GitHub runner.

Co-authored-by: arena-agent <297053741+arena-agent@users.noreply.github.com>
@coderabbitai

coderabbitai Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Bot user detected.

To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Advanced

Run ID: e5b1d3ab-eb41-46f2-8df8-15de371bf7ab

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@arena-ai-coding-agent

Copy link
Copy Markdown
Author

Do not merge as-is — main was replaced and now has unrelated history

Checked after opening this PR:

Merging would introduce a second root history. That is destructive, so I have
marked this PR draft instead of merging it.

main already has an Agda suite, on a different toolchain

origin/main contains:

proofs/agda/MetaManifold/All.agda
proofs/agda/MetaManifold/Composition/{Tree,Node,Sum}.agda
proofs/agda/MetaManifold/ILR/{SBP,Contrast,Kernel,Invariance,Orthonormal,Comb,Integer}.agda
proofs/agda/MetaManifold/Evidence/{Prelude,Residual,Echo,Warrant,Signed,Decision}.agda
proofs/agda/reject/{Postulate,OrthogonalSelf,KernelWithoutHypothesis,IdentificationWithoutUniqueness}.agda
proofs/agda/metamanifold-proofs.agda-lib
proofs/vectors/evidence_vectors.json

Four concrete collisions, in increasing order of cost:

  1. MetaManifold/All.agda — both sides define module MetaManifold.All.
    Mine imports six statistics modules; main's imports the Composition, ILR and
    Evidence suites. One of the two has to become the union.
  2. Library name — mine is metamanifold (metamanifold.agda-lib), main's is
    metamanifold-proofs. Both have include: . in the same directory, so Agda
    would see two libraries claiming the same tree. Mine must be dropped and my
    modules folded into metamanifold-proofs.
  3. just proofs — main's Justfile already defines that recipe at line 279.
    My appended recipes duplicate the name and would fail to parse.
  4. Toolchain — main pins Agda 2.6.4.3 / agda-stdlib 2.1; this branch was
    developed and verified against Agda 2.7.0.1 / agda-stdlib 3.0. This is the
    real cost. The stdlib-3.0-only APIs used here include
    Data.Rational.Unnormalised (ℚᵘ, mkℚᵘ, _≃_, ↥_/↧_/↧ₙ_),
    toWitness {a? = …}, Data.Integer.Properties.{*-monoˡ-≤-nonNeg, *-monoʳ-≤-nonNeg} with NonNegative instances, and
    Relation.Nullary.Decidable.{toWitness, False}. Under stdlib 2.1 these are
    either absent or differently named. Porting is mechanical but not free, and it
    must be re-verified, not assumed.

Note that main's Evidence.* modules are deliberately stdlib-free
(Agda.Builtin.* only) so they check under any Agda ≥ 2.6.4.3. That is a design
choice worth matching for the statistics modules if the estate wants them to
check without the stdlib — it would also remove the porting problem entirely,
since ℚᵘ is the only thing here that genuinely needs it.

What is still true regardless of the rebase

The mathematics and the gate machinery do not depend on the branch's history:

  • 77 definitions across 6 modules, agda --safe --without-K MetaManifold/All.agda
    → exit 0 with empty output on Agda 2.7.0.1 / stdlib 3.0.
  • proofs/tests/axiom-audit.sh → 7/7 modules reachable, clean.
  • proofs/tests/gate-selftest.sh → 10/10 controls behaved correctly.
  • proofs/bootstrap.sh → proofs: OK.

proofs/tests/gate-selftest.sh (break the proofs nine ways, require each to be
rejected) has no counterpart on main — proofs/agda/reject/ holds modules that
should fail to check, which is the same idea expressed differently, and the two
are complementary rather than duplicative.

Suggested path

Rebase onto 770a614: move the six modules to MetaManifold/Statistics/*, drop
metamanifold.agda-lib, add them to main's All.agda, rename the Justfile
recipes to avoid the collision, and either port to stdlib 2.1 or make them
stdlib-free. Then re-run the gate and re-verify before un-drafting.

@arena-ai-coding-agent
arena-ai-coding-agent Bot marked this pull request as draft September 26, 2026 19:47
…cient

Three defects found by running proofs/bootstrap.sh from an empty proofs/.vendor,
which is the only way to find them:

1. The pinned agda-stdlib tag `v3.0` does not exist. `git clone --branch v3.0`
   fails with "Could not find remote branch v3.0"; the newest tag is v2.4. The
   tree the proofs were developed against was a development snapshot whose own
   library file declares `name: standard-library-3.0`. Pinned by SHA instead
   (2ffa8b7d4e8e818717ad643d184f055a4d1b0447, 2026-09-12) and fetched by
   `git init` + `fetch --depth 1 origin <sha>`, since `--branch` cannot take one.

   The proofs do NOT compile against v2.4: the first failure is
   `Data.Integer.Properties` not exporting `_≡?_`. Recorded as R-TC-1 in
   proofs/residue/toolchain.residue, with the porting targets in the order they
   fail. This matters because `main` pins agda-stdlib 2.1 / Agda 2.6.4.3, so
   landing these modules there requires a port, not a version bump.

2. Agda's library-file location differs between install methods. A fresh PyPI
   venv reads `$XDG_CONFIG_HOME/agda/libraries`; an earlier venv read the data
   directory from `--print-agda-dir`. The script wrote only the latter and failed
   with "Library 'standard-library' not found" while the file it had just written
   was present and correct. It now writes both AND passes `--library-file`
   explicitly on every invocation, so the gate does not depend on ambient config.
   Recorded as R-TC-2 (closed).

3. gate-selftest.sh required `agda` on PATH and died after a successful
   bootstrap, because bootstrap installs into proofs/.vendor/venv. It now
   resolves the binary the same way bootstrap.sh does.

Also: normalise the stdlib `.agda-lib` filename and `name:` field, which differ
between tags (v2.1 ships `standard-library-2.1.agda-lib`).

Re-verified from an empty proofs/.vendor: bootstrap.sh -> "proofs: OK" exit 0
(installs Agda from PyPI, fetches the pinned stdlib, fetches Agda source for the
primitive libraries, type-checks 7 modules, audits clean); gate-selftest.sh ->
10/10 controls, exit 0.

Co-authored-by: arena-agent <297053741+arena-agent@users.noreply.github.com>
@arena-ai-coding-agent

Copy link
Copy Markdown
Author

Correction: the toolchain pin in the first push was wrong, now fixed (b9694a0)

I ran proofs/bootstrap.sh from an empty proofs/.vendor — the only way to
find these — and it exposed three defects. The first contradicts something I
stated earlier:

1. agda-stdlib v3.0 does not exist. git clone --branch v3.0 fails with
Could not find remote branch v3.0; the newest tag is v2.4. The tree these
proofs were developed against was a development snapshot whose own library file
declares name: standard-library-3.0. Now pinned by SHA
(2ffa8b7d4e8e818717ad643d184f055a4d1b0447, 2026-09-12) and fetched with
git init + fetch --depth 1 origin <sha>.

The consequence is the part that matters for this PR: these proofs do not
compile against v2.4
(first failure: Data.Integer.Properties does not export
_≡?_), and main pins agda-stdlib 2.1 / Agda 2.6.4.3. So the rebase
described in my previous comment is not a version bump, it is a port. The failing
APIs, in the order they break, are listed in proofs/residue/toolchain.residue
(R-TC-1).

2. Agda's library-file location differs by install method. A fresh PyPI venv
reads $XDG_CONFIG_HOME/agda/libraries; an earlier venv read the data directory
from --print-agda-dir. The script wrote only the latter and failed with
Library 'standard-library' not found while the file it had just written was
present and correct. It now writes both and passes --library-file
explicitly on every invocation, so the gate no longer depends on ambient config
(R-TC-2, closed).

3. gate-selftest.sh died after a successful bootstrap because it required
agda on PATH while bootstrap installs into proofs/.vendor/venv. It now
resolves the binary the same way bootstrap.sh does.

Re-verified from scratch

From an empty proofs/.vendor, so the bootstrap installed its own toolchain
rather than reusing one:

  • proofs/bootstrap.sh → proofs: OK, exit 0
  • proofs/tests/axiom-audit.sh → 7/7 modules reachable, clean, exit 0
  • proofs/tests/gate-selftest.sh → 10/10 controls behaved correctly, exit 0

The just proofs* recipes are still unexecuted (just is not installed in the
sandbox); the scripts they wrap are.

Still draft, for the reasons in the previous comment: main has unrelated
history and four concrete collisions.

Records what a fresh agent has to know before touching this work: that origin/main
has been replaced with a single-commit history with no common ancestor to this
branch and already carries an Agda suite on a different toolchain; the four
concrete collisions that block a naive merge; the three traps that already cost
time (stdlib has no v3.0 tag and the proofs do not compile against v2.4 or v2.1,
/tmp does not survive between turns, and the self-test's own $0/$1 bug); what is
proved, what is explicitly not, and the Agda syntax traps confirmed the hard way.

Co-authored-by: arena-agent <297053741+arena-agent@users.noreply.github.com>
@hyperpolymath
hyperpolymath marked this pull request as ready for review September 26, 2026 20:42
Signed-off-by: Jonathan D.A. Jewell <6759885+hyperpolymath@users.noreply.github.com>
@hyperpolymath
hyperpolymath merged commit e6726fa into main Sep 26, 2026
7 of 8 checks passed
@hyperpolymath
hyperpolymath deleted the arena/01a0db37-metamanifold-webui branch September 26, 2026 22:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant