diff --git a/.gitignore b/.gitignore index 3409398..8a8741b 100644 --- a/.gitignore +++ b/.gitignore @@ -268,6 +268,11 @@ docs/* !docs/integration/** !docs/formal/ !docs/formal/** +!docs/triage/ +!docs/triage/** +# …but the generated PR-#7 stack patches are derived artefacts: regenerate +# with docs/triage/pr7-split/make-stacks.sh instead of committing them. +docs/triage/pr7-split/patches/*.patch # Test + benchmark run artifacts (not the committed fixtures/baselines) frontend/tests/results/ diff --git a/docs/triage/2026-09-26-notification-backlog.md b/docs/triage/2026-09-26-notification-backlog.md new file mode 100644 index 0000000..e62d837 --- /dev/null +++ b/docs/triage/2026-09-26-notification-backlog.md @@ -0,0 +1,222 @@ + + +# Notification-backlog triage — 2026-09-26 + +Request: clear the [GitHub notifications][notif] backlog — finish chores, +close issues, resolve unmergeable PRs, and split huge PRs into mergeable, +comprehensible pieces. + +Caveat first: the Arena bot token cannot read the personal +`/notifications` page (API: `403 Resource not accessible by +integration`), so this triage reconstructs the backlog from public repo +state instead: all open PRs/issues on the fork +(`hyperpolymath/MetaManifold-WebUI`) and the parent +(`JoshuaJewell/MetaManifold-WebUI`), plus Actions health. If a +notification is not covered below, point at it and it gets triaged next. + +Note: the owner merged fork PR #75 **during this triage** (09:47 UTC); +the report was updated to match. Fork open PRs are now **zero**. + +## TL;DR + +| Item | Verdict | +|------|---------| +| Fork PR #73 (338 files, conflicting) | **Closed by this triage** — wrong-base duplicate of parent #7 | +| Fork PR #75 (71 files, ILR bases) | **Merged by the owner during triage** (09:47 UTC) → fork issue #20 auto-closed. Post-merge verification still owed (below) | +| Parent PRs #8, #10, #11 (300+ files, conflicting) | **Close them** (owner click — bot has no parent write access; commands in `pr7-split/RUNBOOK.md`). Head branches deleted, real deltas already on fork `main` | +| Parent PR #7 (319 files, clean) | **Keep.** Split into 14 verified stacks: `pr7-split/` (`STACKS.md`, `RUNBOOK.md`, `make-stacks.sh`) | +| CI on the fork | **Broken repo-wide since 2026-09-25 22:19 UTC** — every file-based workflow `startup_failure`s with zero jobs, including the #75 merge commit. Settings/platform-side; owner action required (runbook appendix 2) | +| Fork issues (12 open; #20 closed by the #75 merge) | **All remaining stay open** — verified one by one (epic gate, recorded deferrals). Closing any of them now would contradict recorded owner decisions | +| Branch `arena/01a0db67` (no PR) | **Deleted at owner's request** — content preserved in `docs/triage/db67-rescue/` (stale docs branch, not DOI-deleting — see report) | +| Parent issue #12 (owner memo) | Awaits the addressee's eyes; summarised below, not closed | +| Dependabot Updates on fork `main` | **Two consecutive failures** (`36210545146`, `36233872771`); re-run once CI is fixed | + +## Actions taken by this triage + +1. **Closed fork PR #73** with an explanatory comment ([comment][c73]): + same head branch (`feat/stipple-typed-studies-ui` @ `eab8ea0`) is open + upstream as parent #7 where it is CLEAN; the fork-side PR rendered as + 338-file / +65k whole-tree noise after the fork's `git-filter-repo` + rewrite severed the merge base back to the initial commit, and was + CONFLICTING. Branch untouched; nothing lost. +2. Recovered the three **deleted head commits** of parent PRs #8/#10/#11 + by SHA from the fork (`f451f1c0`, `c0be9516`, `b46aee09`) and proved + their content is already on fork `main` (`git cherry` patch-identical: + `6c29b53b`, ancestor, `0ef5eee`). Wrote the close comments + commands + for the owner (bot is `403` on the parent). +3. Built the **parent-#7 split kit** (`docs/triage/pr7-split/`): 14 + disjoint stacks covering all 319 files, verified to reproduce the + branch tree byte-for-byte, with per-stack review guide and runbook. +4. Diagnosed the **CI outage** (evidence below) to a settings/platform + cause with owner fix steps. +5. Audited all open fork issues against merged work: #20 closed by the + #75 merge; none of the rest qualifies for closure (details below). +6. Merged the #75 merge commit into this branch, resolving the one + `.gitignore` clash (kept both sides: `docs/formal/` + `docs/triage/`). + +## The huge PRs, explained + +All four 300+ file PRs are the same disease with different prognoses. +On ~2026-09-24/25 fork `main` was rewritten with `git-filter-repo` +(~260 MiB of dead blobs stripped; HEAD tree byte-identical — see the +parent #7 description). The rewrite severed every merge base back to the +initial commit, so any branch diffed against the "wrong" base renders as +a whole-tree diff: + +- **Parent #7** (319 files, +22,662/−404): branch is `upstream/main` + 3 + commits, so the diff is REAL and the PR is CLEAN. Genuinely large → + split kit provided. +- **Fork #73** (338 files, +65,571/−22): same branch against rewritten + fork `main` → noise + conflicts → closed as duplicate. +- **Parent #8/#10/#11** (~315–345 files, ~+60k): branches cut from + fork `main`, PR'd at upstream `main` (July snapshot) → noise; real + deltas were 2/0/1 commits, all now on fork `main`; branches deleted + → close. + +## Fork PR #75 — merged during triage (verification owed) + +`feat(analysis): PhILR, SBP and balance-dendrogram ILR bases` +(`e2adbb9`, 71 files, +6,918/−78) was merged by the owner at 09:47 UTC +with CI `startup_failure` and CodeFactor FAILURE — i.e. the PR's own +"Closes #20 once CI is green" condition was bypassed, and fork issue +#20 auto-closed. That is the owner's call to make; the owed follow-up: + +1. Once CI runs again, confirm the `CI` + `DOI publication contracts` + lanes pass on the merge commit — the PR author never executed Julia + in their sandbox, so CI is the first real execution of the new tests, + reference implementation and benchmarks. +2. Check whether the CodeFactor FAILURE flagged anything actionable in + the merged code. +3. Re-run Dependabot Updates (failed twice on `main`). + +## CI outage — evidence + +- Since 2026-09-25 22:19 UTC, **100% of file-based workflow runs** end + `startup_failure` with **zero jobs**: `CI`, `DOI publication + contracts`, `Stipple UI contracts` — on `main` pushes, PRs, bot and + Dependabot actors alike (e.g. runs `36210930780`, `36210537997`, + `36196688881`, and the #75 merge runs `36233864893`/`36233864472`). +- The break is **not the workflow files**: `ui.yml` last ran green on + 2026-09-21, is unchanged since, and fails identically now. All + workflows report `active`. Before 22:19 the same files executed jobs + (red, but running). +- Dynamic (non-file) workflows still run: CodeQL succeeds on the same + commits; Dependabot Updates jobs execute (and fail on their own + merits). +- A previous session saw the UI banner: *"Actor is not allowed to + trigger Actions workflows"*; `gh run rerun` is refused + ("workflow file may be broken" — GitHub's stock message for + un-rerunnable startup failures). +- A sibling probe branch (`arena/01a0db23`, "probe whether file-based + workflows can start here") also startup-fails. + +Conclusion: settings- or platform-side block on starting file-based +workflows. Owner steps: read the exact banner on any failed run, +check Settings → Actions → General on both repos, push a no-op commit +from a human account to discriminate actor-block vs global break, and +escalate to GitHub Support if humans fail too. Full steps in +`pr7-split/RUNBOOK.md` appendix 2. **Everything merge-gated waits on +this**: any #7 stack PRs, Dependabot, and post-merge verification of #75. + +## Issues — why the remaining 12 stay open + +| # | Issue | Why it stays open | +|---|-------|-------------------| +| 1 | Validated Julia statistics layer (epic) | 14-item acceptance gate ending in *independent review + explicit owner approval*; the issue text says closing it administratively does not open the symbolic gate | +| 2 | BLOCKED symbolic-engine gate | Blocked by #1 by recorded owner decision; it is the gate, not work | +| 3 | Exact statistics (Fisher/exact-NB/permutation) | Recorded owner *decision* to defer (catalogue item 4, excluded from v1); code shows only refusal/doc mentions | +| 4 | Symbolic engine for formulae | Deferred; gated by #1/#2 | +| 5 | ANCOM-BC / ALDEx2 / Songbird | Deferred; no implementation merged | +| 6 | CladeCumulus phylogenetic integration | Deferred; `clade_cumulus.jl` exists on the PR branch, not merged | +| 7 | Full Evidence Mode | Deferred; PR #74 added only a minimal DOI publication page + launcher | +| 8 | Zenodo integration | PR #74 says explicitly "No issue is being closed by this PR" — needs live sandbox acceptance + baselines first | +| 17 | Multinomial / Dirichlet-Multinomial | Deferred roadmap item, no PR | +| 18 | Occupancy / ZINB / hurdle models | Deferred roadmap item, no PR | +| 19 | Constrained ordinations | Deferred roadmap item, no PR | +| 21 | Advanced zero handling | Deferred roadmap item, no PR (PR #75 shares Agda groundwork with it) | + +Closed during this triage: **#20** (by the #75 merge). Previously +closed issues (#9–#16, e.g. #9 AnalysisConfig v1) stay closed. + +## Flag (resolved): branch `arena/01a0db67-metamanifold-webui` + +**Update: deleted on 2026-09-26 at the owner's request; content preserved +in `docs/triage/db67-rescue/`** (`deleted-branch.patch` recreates the +branch tree byte-for-byte; restore commands in that directory's +`README.md`). Branch head at deletion: `d94ecd4e99dbda819fdc0fe1913d5caaabca414a`. + +Correction to the first triage note: this branch did **not** actively +delete the DOI subsystem. It was three docs commits on the pre-DOI base +`e33e10f`, so diffed against current `main` it merely *appeared* to +delete `src/doi/*` and the ILR work (base-difference noise — a GitHub +3-way merge would not have reverted PR #74/#75). Its genuine content was +a BerryWiki wiki (`docs/wikis/`, 33 pages), `README.adoc`/`EXPLAINME.adoc`, +the autolink spec, and prose updates to `CHANGELOG`/`CONTRIBUTING`/`NOTICE` +— plus a `README.md` deletion and 32 wiki pages missing SPDX headers, +which were the real landmines and the reason for closing it. + +Also already gone: the merged PR #75 branch (`arena/01a0db2e`) was +auto-deleted on merge. Still open and deliberately untouched: the CI +probe branch `arena/01a0db23` (unmerged, no PR — needs its own decision) +and `feat/stipple-typed-studies-ui` (live as parent PR #7). + +## Parent issue #12 (owner memo, 0 comments) + +"Notice to Project Owner: Comprehensive Review & Roadmap Status +(2026-09-25)" — a long plain-English memo addressed to the project +owner explaining the fork's work (fake-number removal, real models, +refuse-to-guess, the four parent PRs). It is FYI, not actionable, and +needs the *addressee's* eyes before closing — left open deliberately. + +## Suggested order of operations for the owner + +1. Fix CI (runbook appendix 2) — unblocks everything. +2. Close parent #8/#10/#11 (one command each, runbook appendix 1). +3. Decide parent #7: Option A merge-as-is (review via `STACKS.md`) or + Option B stacked PRs (`RUNBOOK.md`). +4. ~~Decide `arena/01a0db67`~~ — done: branch deleted, content preserved + in `docs/triage/db67-rescue/`. Remaining: decide the CI probe branch + `arena/01a0db23` (unmerged, no PR). +5. Read parent #12, then close it. +6. Post-merge verification of #75 once CI runs (Julia tests, CodeFactor + findings, Dependabot re-run). + +[notif]: https://github.com/notifications +[c73]: https://github.com/hyperpolymath/MetaManifold-WebUI/pull/73#issuecomment-5845153601 +[p7]: https://github.com/JoshuaJewell/MetaManifold-WebUI/pull/7 + +## Sweep 2 (same day, ~10:30 UTC) — leftovers the bot could still clear + +1. **Kit PR #76: MERGEABLE**, CodeQL + CodeFactor green, semgrep pending. + No CI/DOI checks report (outage). Ready whenever the owner is. +2. **CodeFactor on the #75 merge: 3 × "Complex Method" notices, style + only** — `IlrBasisInputs.tsx:39`, `analysis_config.ts:143`, + `AnalysisConfigEditor.tsx:21` (run `108381688702`). Not correctness + blockers: refactor-or-accept during post-merge verification. +3. **Dependabot: API rerun refused too** ("cannot be rerun" stock + message), so both `main` failures wait on an owner UI re-run or the + next Dependabot cycle after CI is fixed. +4. **Probe branch `arena/01a0db23`: KEEP — do not delete.** True delta is + 14 files, +2,896/−2, all additive, SPDX-clean: the issue-#21 Agda + proofs + `golden.json` fixture, the KYAML pilot (`KYAML.jl`, gate, + docs), and the spent CI probe. Landing notes for whoever opens its + PR: `proofs/agda/README.md` will add/add-conflict with #75's copy + (merge contents by hand); `test/runtests.jl` needs the one-line + kyaml include re-applied next to #75's; `*.agdai` would duplicate + (harmless, tidy it); drop `ci-probe.yml` (probe answered); note its + `EXPLAINME.adoc` differs from the db67 copy archived in + `docs/triage/db67-rescue/`. +5. **Parent #12: §4 walkthrough + §5 Option B are now stale** — they + describe #8/#10/#11 as small reviewable PRs, but those have deleted + branches and render as whole-tree noise. Draft correction for the + owner to paste as a comment (bot is 403 on the parent): + + > Correction (26 Sept triage): PRs #8, #10 and #11 no longer have + > live branches and cannot be reviewed or merged as §4 describes — + > their real changes (2/0/1 commits) are already on the fork's + > `main`, so they should be closed, not reviewed. Please review #7 + > instead (a 14-part review guide now exists on the fork at + > `docs/triage/pr7-split/STACKS.md`). §5 Option B is therefore + > obsolete; options A, C and D stand as written. diff --git a/docs/triage/db67-rescue/README.md b/docs/triage/db67-rescue/README.md new file mode 100644 index 0000000..d1b48f4 --- /dev/null +++ b/docs/triage/db67-rescue/README.md @@ -0,0 +1,54 @@ + + +# Deleted branch `arena/01a0db67` — content preserved here + +On 2026-09-26 the owner asked for this branch to be **closed (deleted)** +during the notification-backlog triage. Before deletion its full content +was preserved in this directory so nothing is lost: + +- `deleted-branch.patch` (260 KB): the branch's true delta + (`e33e10f..d94ecd4`, 41 files, +4,138/−691), verified to recreate the + branch tree byte-for-byte. +- `deleted-branch-files.txt`: the file list (`A`/`M`/`D` per file). + +## What the branch was + +Three docs commits on top of `e33e10f` (fork `main` before the DOI and +ILR merges): a BerryWiki-format wiki (`docs/wikis/`, 33 pages), +`README.adoc` + `EXPLAINME.adoc`, the autolink specification +(`docs/integration/autolink-references.md`), a `.gitignore` negation for +`docs/wikis/`, and prose updates to `CHANGELOG.md`, `CONTRIBUTING.md` +and `NOTICE` (third-party table + acknowledgements). It also deleted +`README.md`, intending the `.adoc` pair to replace it. + +## Correction to the first triage note + +The first version of this triage's report called this branch +"DOI-deleting". That overstated it: diffed against *current* `main` the +branch *appears* to delete `src/doi/*` and the ILR work, but that is +base-difference noise — the branch predates those merges +(`git-filter-repo` era base `e33e10f`), so it never contained them. A +GitHub 3-way merge would **not** have reverted PR #74/#75. The genuine +landmines were the `README.md` deletion and the 32 wiki pages missing +SPDX headers (which would fail `scripts/check-spdx.sh`). + +## Restore (if ever wanted) + +```sh +git fetch https://github.com/hyperpolymath/MetaManifold-WebUI.git \ + arena/01a0dd12-metamanifold-webui +git checkout FETCH_HEAD -- docs/triage/db67-rescue +git checkout -b restore/db67 e33e10f +git apply --index docs/triage/db67-rescue/deleted-branch.patch +git commit -m "restore: db67 docs branch content" +``` + +Before merging a restore: rebase onto current `main`, decide +`README.md` vs `README.adoc` (keep both until decided), and add +`SPDX-License-Identifier: CC-BY-SA-4.0` headers to the wiki pages +(the `_Sidebar.md` file is generated upstream — its header needs to +come from the generator, or the file needs a gate exemption). + +Branch head SHA at deletion: `d94ecd4e99dbda819fdc0fe1913d5caaabca414a`. diff --git a/docs/triage/db67-rescue/deleted-branch-files.txt b/docs/triage/db67-rescue/deleted-branch-files.txt new file mode 100644 index 0000000..af33ded --- /dev/null +++ b/docs/triage/db67-rescue/deleted-branch-files.txt @@ -0,0 +1,41 @@ +M .gitignore +M CHANGELOG.md +M CONTRIBUTING.md +A EXPLAINME.adoc +M NOTICE +A README.adoc +D README.md +A docs/integration/autolink-references.md +A docs/wikis/Deep-Dives--Advanced-Functionality.md +A docs/wikis/Deep-Dives--Compositional-Statistics.md +A docs/wikis/Deep-Dives--Design-Progression.md +A docs/wikis/Deep-Dives--Epistemic-Status.md +A docs/wikis/Deep-Dives--Exact-Arithmetic.md +A docs/wikis/Deep-Dives--Maximum-Likelihood.md +A docs/wikis/Deep-Dives--Type-Theory-Meets-Statistics.md +A docs/wikis/Deep-Dives.md +A docs/wikis/Developers--Architecture-Tour.md +A docs/wikis/Developers--Extending-the-Pipeline.md +A docs/wikis/Developers--REST-API.md +A docs/wikis/Developers--Statistics-Internals.md +A docs/wikis/Developers--Testing-and-Benchmarks.md +A docs/wikis/Developers--Type-System.md +A docs/wikis/Developers.md +A docs/wikis/Home.md +A docs/wikis/Maintainers--Compliance-and-Estate.md +A docs/wikis/Maintainers--Operator-Track.md +A docs/wikis/Maintainers--Releases-and-Distribution.md +A docs/wikis/Maintainers--Steward-Track.md +A docs/wikis/Maintainers.md +A docs/wikis/README.md +A docs/wikis/Status-and-Roadmap.md +A docs/wikis/Users--Analysis-and-Statistics-Today.md +A docs/wikis/Users--Configuration-Reference.md +A docs/wikis/Users--Exploring-Results.md +A docs/wikis/Users--For-Academics.md +A docs/wikis/Users--For-Lab-Professionals.md +A docs/wikis/Users--Install-and-First-Run.md +A docs/wikis/Users--Troubleshooting.md +A docs/wikis/Users--Your-First-Study.md +A docs/wikis/Users.md +A docs/wikis/_Sidebar.md diff --git a/docs/triage/db67-rescue/deleted-branch.patch b/docs/triage/db67-rescue/deleted-branch.patch new file mode 100644 index 0000000..fc1ca8b --- /dev/null +++ b/docs/triage/db67-rescue/deleted-branch.patch @@ -0,0 +1,5095 @@ +diff --git a/.gitignore b/.gitignore +index cdfeec6..89b16f5 100644 +--- a/.gitignore ++++ b/.gitignore +@@ -265,6 +265,8 @@ docs/* + !docs/owner-review-2026-09-25.md + !docs/integration/ + !docs/integration/** ++!docs/wikis/ ++!docs/wikis/** + + # Test + benchmark run artifacts (not the committed fixtures/baselines) + frontend/tests/results/ +diff --git a/CHANGELOG.md b/CHANGELOG.md +index 53edb86..60952a2 100644 +--- a/CHANGELOG.md ++++ b/CHANGELOG.md +@@ -12,6 +12,34 @@ # Changelog + + ## [Unreleased] + ++### Added — the README/EXPLAINME pair, the wiki, and the autolink specification (2026-09-26) ++ ++- **`README.adoc` replaces `README.md`**, per the estate README/EXPLAINME authoring ++ standard (`standards:docs/README-EXPLAINME-STANDARD.adoc`). The README is now the ++ three-layer design history — base R/Python design around raw DADA2, the origin ++ MetaManifold augmentation (JoshuaJewell), and the fork's honesty/typing steps — with ++ diagrammatic progression and shipped/planned markers throughout. The configuration ++ chapters moved to the wiki (they made the README unreadable); the third-party tools ++ table and acknowledgements moved into `NOTICE` (their natural home). ++- **`EXPLAINME.adoc`** (new): the receipts file — claim→implementation→caveat map over ++ every README claim, the dogfooding table, known gaps as CAUTION blocks, and an ++ evidence index. The type-theory/enhanced-statistics deep material deliberately lives ++ in the wiki; EXPLAINME cross-references it rather than re-deriving it. ++- **The GitHub wiki is now the full BerryWiki-format documentation** ++ (`metadatastician/berrywiki` page format: hidden metadata blocks, generated ++ `_Sidebar.md`), sourced from `docs/wikis/` and synced to `MetaManifold-WebUI.wiki.git`: ++ three audience sections (users — with academics and lab-professional tracks; platform ++ maintainers — operator and steward tracks; developers), seven deep dives (design ++ progression, type theory meets statistics, exact arithmetic, maximum likelihood and ++ refusals, compositional statistics and offsets, epistemic status, advanced ++ functionality), and a Status-and-Roadmap board marking everything IN PLACE / PARTIAL / ++ COMING / BLOCKED. ++- **`docs/integration/autolink-references.md`** (new): the complete elaboration of the ++ repository's Settings → Autolink references set (lineage/estate, upstream tools, ++ toolchain, registries), paste-ready and machine-readable, with reserved/omitted cases ++ reasoned. Application in Settings needs Administration permission (one human pass); ++ the file is the source of truth for it. ++ + ### Fixed — the NB test fixture is data a negative binomial describes (2026-09-26) + + - The estimation tests' synthetic table was **under-dispersed** (variance below the mean, +diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md +index 75ad856..72b94bb 100644 +--- a/CONTRIBUTING.md ++++ b/CONTRIBUTING.md +@@ -41,7 +41,7 @@ # Frontend (gates run here) + bun run bench/ # benchmark harness (informational) + bun run check # all three in sequence — must be green + +-# Full application (requires Julia + R/renv per README.md § Prerequisites) ++# Full application (requires Julia + R/renv per README.adoc § Quick start) + ./install.sh # upstream flow + ./start.sh + ``` +diff --git a/EXPLAINME.adoc b/EXPLAINME.adoc +new file mode 100644 +index 0000000..6a72e83 +--- /dev/null ++++ b/EXPLAINME.adoc +@@ -0,0 +1,434 @@ ++// SPDX-License-Identifier: CC-BY-SA-4.0 ++// SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell (hyperpolymath) ++= MetaManifold — EXPLAINME ++:toc: preamble ++:toc-title: Contents ++:icons: font ++:doctype: article ++ ++This is the receipts file: every factual claim in link:README.adoc[README.adoc] ++mapped to the code that implements it, with the honest caveat attached. It is ++written for the sceptical developer or reviewer doing due diligence after the ++README has interested them. The mathematics and type theory behind the ++statistical layer are *not* re-derived here — they run to pages, so they live ++in the link:https://github.com/hyperpolymath/MetaManifold-WebUI/wiki[project wiki] ++and this file points at them. README sells; this file proves; the wiki teaches. ++ ++== Claim-to-implementation map ++ ++=== "MetaManifold is three designs stacked honestly on top of each other." ++ ++[quote, README.adoc] ++____ ++MetaManifold is three designs stacked honestly on top of each other. Each ++layer is still visible in the code; none of them pretends to be the whole ++story. ++____ ++ ++How this is implemented:: ++Layer 1 (raw R/DADA2 + shell/swarm-vsearch) survives as the R bridge in ++`src/pipeline/dada2/dada2_functions.r` and the external-tool stages in ++`src/pipeline/{swarm,vsearch,tools}.jl`. Layer 2 (the orchestration and ++web workbench) is `src/core/config.jl`, `src/core/duckdb_store.jl`, ++`src/server/`, `frontend/`. Layer 3 (the honesty and typing discipline) is ++`src/analysis/{estimation,exact_summaries,numeric_policy,scaling,AnalysisConfig}.jl` ++plus `src/core/epistemic.jl` and the `test/`, `bench/`, CI estate. The ++lineage boundaries are recorded in link:NOTICE[NOTICE] and the full visual ++progression is ++link:https://github.com/hyperpolymath/MetaManifold-WebUI/wiki/Deep-Dives--Design-Progression[Deep-Dives — Design Progression]. ++ ++Caveat:: ++"Stacked honestly" is a claim about presentation as much as code: the ++cutadapt/DADA2/vsearch semantics remain upstream's, and this fork tracks ++application changes, not the science (link:ROADMAP.md[Roadmap], "Out of ++scope"). ++ ++=== "One run takes raw paired-end FASTQs to chimera-filtered, taxonomy-annotated ASV *and* OTU tables" ++ ++[quote, README.adoc] ++____ ++One run takes raw paired-end FASTQs to chimera-filtered, taxonomy-annotated ++ASV **and** OTU tables, with per-stage read accounting and QC (FastQC/MultiQC ++plus DADA2 diagnostics) embedded in the UI. ++____ ++ ++How this is implemented:: ++Typed stage results (`TrimmedReads`, `ASVResult`, `OTUResult`, ++`TaxonomyHits`, `MergedTables`) in `src/core/types.jl`; stage bodies in ++`src/pipeline/`; per-stage counts land in `pipeline_stats.csv`; QC endpoints ++in `src/server/routes/` feed the MultiQC report and DADA2 figures to the UI. ++Stages skip when outputs are fresh (mtime for files, content hash for ++configuration) — `src/pipeline/pipeline freshness logic`. ++ ++Caveat:: ++FastQC and MultiQC were absent from CI for a while (issue #30) so prefilter ++QC was never exercised there; fixed by issue #44 with sha256-pinned installs ++(`config/defaults/tool_versions.yml`). Freshness skipping trusts mtimes for ++files — editing an output behind the engine's back is not detected. ++ ++=== "Analysis configuration is typed and validated (`AnalysisConfig`)" ++ ++[quote, README.adoc] ++____ ++Analysis configuration is typed and validated (`AnalysisConfig`): negative ++binomial GLM, CLR/ILR linear models, logistic models — each with explicit ++constraints, convergence and boundary reporting, or a refusal. Benjamini– ++Hochberg correction is mandatory wherever several tests are reported. ++____ ++ ++How this is implemented:: ++`src/analysis/AnalysisConfig.jl` (immutable struct mirroring the user's ++answers, Nickel schema `config/schemas/analysis_config.ncl`); fits in ++`src/analysis/estimation.jl` called through `Execution.run_analysis` — R ++`MASS::glm.nb`, `stats::lm`, `stats::glm(family=binomial)` with ++identifiability, boundary-estimate and convergence checks that produce named ++unsuccessful states instead of plausible parameters; BH via R `p.adjust`, ++mandatory, with the DANGER banner if anyone tries to disable it. Conditions ++published before implementation in ++link:docs/statistics/method-conditions/parametric-fits.md[method conditions — parametric fits]. ++ ++Caveat:: ++[CAUTION] ++==== ++There has been **no independent statistical review** of this layer. Issue ++#1 requires one before acceptance; it is outstanding. The fits are real and ++tested against R references, but "computed correctly" is not "appropriate for ++your data" — the method conditions say exactly this. ++==== ++ ++=== "Normalisation as honest bookkeeping … never silently swapped" ++ ++[quote, README.adoc] ++____ ++Normalisation as honest bookkeeping: none, rarefaction, or exact ++TSS/CSS/RSS size-factor **offsets** (depth modelled, response unchanged) — ++never silently swapped for relative abundances. ++____ ++ ++How this is implemented:: ++`src/analysis/scaling.jl` implements scaling factors (one positive number per ++sample) and offsets (their logs, after a stated centring) per ++link:docs/statistics/method-conditions/scaling-and-offsets.md[scaling-and-offsets conditions]; ++unknown or aliased methods are refused, lower-case method names compare ++correctly (issue #62), offsets are wired into the fits (issue #62/61), and ++zero-depth samples are healed before any transform (issue #58). Regression ++tests assert the substitutions cannot return. ++ ++Caveat:: ++Offsets are a depth-modelling device, not a compositional solution — the ++conditions document says so explicitly. CSS without `metagenomeSeq` is a ++repo-local implementation of the idea; treat equivalence to the original ++package as unreviewed (covered by the same #1 caution). ++ ++=== "an exact/approximate/rounded numeric policy where higher precision is never sold as exactness" ++ ++[quote, README.adoc] ++____ ++an exact/approximate/rounded numeric policy where higher precision is never ++sold as exactness ++____ ++ ++How this is implemented:: ++`src/analysis/numeric_policy.jl` — an immutable `NumericPolicySpec` with ++modes `:ordinary`, `:exact_counts`, `:high_precision`; only integer and ++rational arithmetic is exact, `BigFloat` is approximate by definition, and a ++float claiming exactness is refused at the boundary. `exact_summaries.jl` ++carries counts and proportions at exact precision (integers beyond 2^53−1 ++included) and labels any float-derived input as approximate. Contracts: ++link:docs/statistics/numeric-contracts.md[numeric contracts]. ++ ++Caveat:: ++None currently known in the policy itself; the scope limit is real though — ++exactness covers descriptive summaries only. Inference is not exact just ++because its inputs are (see ++link:https://github.com/hyperpolymath/MetaManifold-WebUI/wiki/Deep-Dives--Exact-Arithmetic[Deep-Dives — Exact Arithmetic]). ++ ++=== "maximum-likelihood fits (or explicit unsuccessful states) published under method conditions written *before* implementation" ++ ++[quote, README.adoc] ++____ ++Placeholder statistics were removed and guarded ++against by test; maximum-likelihood fits (or explicit unsuccessful states) ++published under method conditions written *before* implementation ++(link:docs/statistics[method catalogue]) ++____ ++ ++How this is implemented:: ++The replaced placeholder derived p-values from `hash(taxon_id)` — a stub ++returned as a result. Its history and the guard against its return are ++documented in ++link:docs/statistics/method-conditions/parametric-fits.md[parametric fits]; ++`test/unit/test_estimation.jl` contains a source-level guard that such ++placeholders do not come back, plus known answers written into fixture data ++and an independent R reference comparison. The catalogue that gates all of ++this is link:docs/statistics/method-catalogue-v1.md[method catalogue v1] ++(owner-approved 2026-09-22). ++ ++Caveat:: ++[CAUTION] ++==== ++Catalogue item 3 (nonparametric tests) and item 4 (exact tests, issue #3) ++are approved in scope but not implemented. Nothing here claims them. ++==== ++ ++=== "0 TypeScript errors under strict + `exactOptionalPropertyTypes`" ++ ++[quote, README.adoc] ++____ ++0 TypeScript errors under strict + ++`exactOptionalPropertyTypes` ++____ ++ ++How this is implemented:: ++`frontend/tsconfig.json` (no-emit gate) with a build sidecar ++(`tsconfig.build.json`); type estate mapped in ++link:docs/types/architecture.md[types architecture]; `bun run typecheck` is ++CI-gated. Third-party coverage is audited in ++link:docs/type-system/category-d-e-closure.md[category D & E closure] — ++including two `FIXME(types)` stubs that declare `unknown`-safe surfaces, never ++`any`. ++ ++Caveat:: ++Two documented exceptions ride alongside: `skipLibCheck: true` (react-router ++6.30.x ships 7 erroneous `.d.ts` entries; retried on react-router 7) and the ++`useAnalysis.alphaFig` `unknown` awaiting a `PlotFigure` narrowing. Both are ++tracked in `docs/compliance/fixme-index.md`. ++ ++=== "pinned toolchains (`mise` + Guix lanes) with sha256-pinned pipeline tools" ++ ++[quote, README.adoc] ++____ ++pinned toolchains (`mise` + Guix lanes) with ++sha256-pinned pipeline tools ++____ ++ ++How this is implemented:: ++`mise.toml` pins julia/bun/node/just to the exact CI versions (R is a ++documented registry-absent exception, restored by `renv.lock`); `guix.scm` + ++`channels.scm` provide the time-machine-pinned peer lane; `install.sh` ++fetches cutadapt/FastQC/MultiQC/vsearch/cd-hit-est/swarm byte-exact against ++the sha256 records in `config/defaults/tool_versions.yml`. Pin ++single-sourcing is enforced by the `coupling-toolchain-pins` test. ++ ++Caveat:: ++The Guix lane carries functional equivalents, not binary identity with the ++download lane (documented in the `guix.scm` header) — use `mise` for exact CI ++parity. Swarm is download-lane-only under Guix. ++ ++=== "Everything at a run's configuration is editable in the browser at every cascade level" ++ ++[quote, README.adoc] ++____ ++Everything at a run's configuration is editable in the browser at every ++cascade level; stale stages are flagged with the exact keys that changed. ++____ ++ ++How this is implemented:: ++The cascade (instance → study → group → run; omitted keys inherit) resolves ++in `src/core/config.jl` and is materialised per run to `run_config.yml` as ++provenance; the REST config routes accept patches at any level and report ++downstream overrides; staleness recomputes stage freshness and the tooltip ++diffs the changed keys. ++ ++Caveat:: ++`run_config.yml` is the only place the fully merged truth appears — editing ++YAML files directly while the server runs can surprise the freshness hash. ++The editors validate whole documents before writing (primers, databases, ++composition) and refuse dangling renames only by *reporting* them, not ++blocking them. ++ ++=== "Functional annotation with dual-classifier consensus (DADA2 bootstrap × vsearch identity)" ++ ++[quote, README.adoc] ++____ ++Functional annotation with dual-classifier consensus (DADA2 bootstrap × ++vsearch identity), contamination curation, manual BLAST override, and an ++append-only FuncDB ledger that survives re-annotation. ++____ ++ ++How this is implemented:: ++`src/annotation/` computes the consensus rank (finest rank where the ++classifiers agree) and a composite confidence score; curation writes ++(contamination flags, manual assignments) are stored separately from the ++derived annotation and re-applied after regeneration; FuncDB entries append to ++a ledger file. ++ ++Caveat:: ++The composite confidence is explicitly "mostly for the sake of curiosity" ++(the README says so) — it is not a calibrated probability. Consensus is ++string equality of labels, which is why both reference formats must come from ++the same release. ++ ++=== Quick start and the gate claim ++ ++[quote, README.adoc] ++____ ++From a clean ++checkout, `just ci` runs every gate that CI runs. ++____ ++ ++How this is implemented:: ++`Justfile` recipes are thin wrappers over the canonical entry points ++(`frontend/package.json` scripts, `scripts/check-*.sh`, the Julia test lanes); ++`.github/workflows/ci.yml` invokes the same commands, so local and CI gates ++are one set. ++ ++Caveat:: ++Fail-loud by design: Julia lanes error without Julia, `test-e2e` errors ++without browsers. "Every gate" means the gated set — benchmarks are ++informational and do not gate. ++ ++== The mathematics, in the wiki ++ ++The type-theoretic and statistical reasoning behind layer 3 is long-form by ++nature (this is closer to a statistics package than a game). It lives in the ++wiki, cross-linked both ways. Start from the ++link:https://github.com/hyperpolymath/MetaManifold-WebUI/wiki/Deep-Dives[deep dives hub]. ++ ++[cols="1,2,2", options="header"] ++|=== ++| README claim area | Type-theory thread | Statistics thread (wiki page) ++ ++| Typed, validated configuration; refusals as results ++| Refinement-style validation of `AnalysisConfig`; third-party type closure (category D/E) ++| ML estimation with identifiability, convergence, boundary states — link:https://github.com/hyperpolymath/MetaManifold-WebUI/wiki/Deep-Dives--Maximum-Likelihood[Deep-Dives — Maximum Likelihood] ++ ++| Exact counts and rationals ++| Mode-indexed numeric policy; exact/approximate/rounded as distinct types of claim ++| Exact descriptive summaries; what exact inference cannot be — link:https://github.com/hyperpolymath/MetaManifold-WebUI/wiki/Deep-Dives--Exact-Arithmetic[Deep-Dives — Exact Arithmetic] ++ ++| Offsets vs transforms ++| `Offset` is not `Transform` — the type distinction is the safety property ++| Size factors, TSS/CSS/RSS, compositional bias — link:https://github.com/hyperpolymath/MetaManifold-WebUI/wiki/Deep-Dives--Compositional-Statistics[Deep-Dives — Compositional Statistics] ++ ++| Epistemic receipts; DANGER banner ++| Σ-types, identity types, factive modalities and warrants without soundness (`src/core/epistemic.jl`, shadows of the Agda echo/epistemic/residual-evidence types) ++| Unsuccessful states as first-class statistical output — link:https://github.com/hyperpolymath/MetaManifold-WebUI/wiki/Deep-Dives--Epistemic-Status[Deep-Dives — Epistemic Status] ++ ++| Advanced functionality (planned) ++| Symbolic engine (#2) — blocked on the numeric layer's review ++| Exact tests, multinomial/DM, occupancy, constrained ordination, PhILR/SBP, zero handling — link:https://github.com/hyperpolymath/MetaManifold-WebUI/wiki/Deep-Dives--Advanced-Functionality[Deep-Dives — Advanced Functionality] ++|=== ++ ++Where the two threads meet — the central argument that a statistical result ++is only as good as the *type of claim* its numbers can carry — is developed in ++link:https://github.com/hyperpolymath/MetaManifold-WebUI/wiki/Deep-Dives--Type-Theory-Meets-Statistics[Deep-Dives — Type theory meets statistics]. ++ ++== Dogfooded across the account ++ ++[cols="1,2,2", options="header"] ++|=== ++| Technology / pattern | Used here | Also used in ++ ++| README + EXPLAINME authoring standard (the pair you are reading) ++| Root `README.adoc` + `EXPLAINME.adoc`, claim→implementation map and all ++| link:https://github.com/hyperpolymath/standards[standards] (the standard itself), link:https://github.com/hyperpolymath/rsr-template-repo[rsr-template-repo] ++ ++| BerryWiki page format (hidden metadata, generated sidebar, plain-Markdown survival) ++| The whole link:https://github.com/hyperpolymath/MetaManifold-WebUI/wiki[project wiki], sourced from `docs/wikis/` ++| link:https://github.com/metadatastician/berrywiki[berrywiki]'s own wiki — the format's first dogfood ++ ++| Justfile command surface (Makefiles banned estate-wide) ++| `Justfile` (43 thin-wrapper recipes incl. pin codegen + drift gate) ++| link:https://github.com/hyperpolymath/standards[standards], link:https://github.com/hyperpolymath/rsr-template-repo[rsr-template-repo] ++ ++| Pinned dual-lane toolchain (mise exact pins + Guix time-machine) ++| `mise.toml`, `guix.scm`, `channels.scm`, sha256-pinned `install.sh` ++| rsr-template-repo doctrine; the estate's reproducibility practice ++ ++| Publish method conditions *before* implementation, then hold code to them ++| link:docs/statistics/method-conditions/[method conditions] + `test/unit/test_estimation.jl` ++| The estate statistics track (link:https://github.com/hyperpolymath/statistikles[statistikles]) adopts the same gate ++|=== ++ ++== Known gaps ++ ++[CAUTION] ++==== ++*Independent statistical review (issue #1):* the estimation and exact-summary ++layers are implemented and tested, but issue #1's acceptance criteria include ++an independent review that has not happened. Read every inference result as ++"computed as documented", not "reviewed". ++==== ++ ++[CAUTION] ++==== ++*Exact statistical tests (issue #3) and the milestone-3 deferred suite ++(issues #17–21: multinomial/DM, occupancy, constrained ordinations, PhILR/SBP, ++advanced zero handling):* approved in scope, specified in ++link:docs/issues/milestone3/[docs/issues/milestone3], **not implemented**. ++The README marks them COMING and this file repeats that. ++==== ++ ++[CAUTION] ++==== ++*Symbolic engine (issue #2):* deliberately **BLOCKED** until the numeric ++statistics layer passes real-data validation. Do not read the AnalysisConfig ++formula strings as a symbolic algebra — they are parsed and fitted, not ++manipulated. ++==== ++ ++[CAUTION] ++==== ++*Stipple/Vue migration and standalone releases:* agreed requirements, partial ++delivery (the first read-only UI slice). The legacy React app is the default; ++no standalone offline archive has been built yet. Source of truth: ++link:docs/migration/STATUS.md[migration status]. ++==== ++ ++[CAUTION] ++==== ++*Deployment posture:* local single-user operation only. The current server is ++not multiuser-ready and must not be exposed as if it were (a standing line in ++link:docs/migration/STATUS.md[migration status]). ++==== ++ ++[CAUTION] ++==== ++*README screenshots:* the UI captures referenced by the origin README were ++never committed to this repository; the design progression in `README.adoc` ++is therefore diagrammatic. Screenshots will land on the wiki pages when ++captured — the pages say which ones lack them. ++==== ++ ++Type-estate gaps (`skipLibCheck` exception, `alphaFig` `unknown`, the two ++`FIXME(types)` stubs) are documented with their exit criteria in ++`docs/compliance/fixme-index.md` and are not repeated as cautions here. ++ ++== Evidence index ++ ++[cols="2,3", options="header"] ++|=== ++| Path | Proves ++ ++| `test/unit/test_estimation.jl` ++| Known-answer fits, independent R reference comparison (coefficients and BH against `p.adjust`), negative controls for every refusal path, and the guard that placeholder statistics never return. ++ ++| `test/unit/test_exact_summaries.jl` ++| Hand-derived exact answers, an independent `fractions.Fraction` cross-check, negative controls (float claiming exactness, negative counts, budget overrun) and the "pipeline does not call this module by default" property. ++ ++| `test/unit/` numeric boundary suites ++| Boundary values (e.g. counts beyond 2^53−1) are measured, not assumed (issue #52's lesson). ++ ++| `src/analysis/scaling.jl` + its tests ++| TSS/CSS/RSS produce offsets; aliases and silent substitutions are refused (issues #16, #61–62). ++ ++| `docs/statistics/method-conditions/*.md` ++| The conditions documents that predate their implementations — the receipts that code is held to documents, not the reverse. ++ ++| `frontend/tests/` + `frontend/tsconfig.json` ++| The strict type gate (0 errors) and the behavioural suite behind `bun run check`. ++ ++| `test/` coupling pins test ++| Toolchain pins are single-sourced (`.bun-version` generated from `mise.toml`; overlaps drift-checked). ++ ++| `bench/*/baseline.json` ++| Recorded performance baselines the informational bench lane compares against. ++|=== ++ ++== Licence ++ ++This document is licensed under CC BY-SA 4.0. Code receipts point at ++AGPL-3.0-only / MPL-2.0 sources per link:NOTICE[NOTICE]. ++ ++SPDX-License-Identifier: CC-BY-SA-4.0 +diff --git a/NOTICE b/NOTICE +index 86968b7..03a2e03 100644 +--- a/NOTICE ++++ b/NOTICE +@@ -29,11 +29,47 @@ unambiguous; it is not a claim of copyright ownership either way. + ## Third-party components + + MetaManifold orchestrates, but does not vendor, third-party command-line +-tools (cutadapt, DADA2, SWARM, vsearch, cd-hit-est, R/vegan). Each tool +-retains its own licence; see `README.md § Third-party tools` for the +-runtime dependency list and upstream references. Frontend npm dependencies +-are declared in `frontend/package.json` / `frontend/bun.lock` under their +-own licences. ++tools. Each is fetched from its upstream source by `install.sh` and is ++subject to its own licence; no third-party binaries are included in this ++repository. ++ ++| Tool | License | Source | ++|---|---|---| ++| [cutadapt](https://github.com/marcelm/cutadapt) | MIT | PyPI | ++| [FastQC](https://github.com/s-andrews/FastQC) | GPL v3 | Babraham Bioinformatics | ++| [MultiQC](https://github.com/MultiQC/MultiQC) | GPL v3 | PyPI | ++| [DADA2](https://benjjneb.github.io/dada2/) | LGPL v3 | Bioconductor | ++| [swarm](https://github.com/frederic-mahe/swarm) | GPL v3 | GitHub Releases | ++| [vsearch](https://github.com/torognes/vsearch) | GPL v3 | GitHub Releases | ++| [cd-hit](https://github.com/weizhongli/cdhit) | GPL v2+ | GitHub Releases / apt | ++ ++R/vegan and the R runtime are system components under their own licences ++(GPL family). Frontend npm dependencies are declared in ++`frontend/package.json` / `frontend/bun.lock` under their own licences. ++ ++## Acknowledgements and lineage ++ ++This pipeline draws on the following prior work: ++ ++- **Frédéric Mahé**: [Fred's metabarcoding pipeline](https://github.com/frederic-mahe/swarm/wiki/Fred's-metabarcoding-pipeline) ++ informed the overall workflow architecture, namely the sequencing of ++ primer trimming, `swarm.jl`, vsearch-based taxonomy assignment, and the ++ final table merge/filter stages. ++- **Benjamin J. Callahan _et al._**: [DADA2 tutorial](https://benjjneb.github.io/dada2/tutorial.html), ++ used under [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/), on ++ which `dada2.jl` and its modules are based. ++ ++The following colleagues at the **Department of Parasitology, Charles ++University** (Faculty of Science, BIOCEV, Vestec, Czech Republic) ++contributed to this work: ++ ++- **Mgr. Jiří Novák** (supervisor): scripts from which several modules and ++ configurations were adapted. ++- **doc. Mgr. Vladimír Hampl**: provided laboratory access and resources. ++- **Mgr. Paulína Pristašová**: <3. ++ ++Copyright © 2026 Joshua Benjamin Jewell (origin design) and the ++hyperpolymath fork authors. + + ## Notices required by AGPL-3.0 + +diff --git a/README.adoc b/README.adoc +new file mode 100644 +index 0000000..a8fb05b +--- /dev/null ++++ b/README.adoc +@@ -0,0 +1,268 @@ ++// SPDX-License-Identifier: CC-BY-SA-4.0 ++// SPDX-FileCopyrightText: 2026 Joshua Benjamin Jewell; 2026 Jonathan D.A. Jewell (hyperpolymath) ++= MetaManifold ++:toc: preamble ++:toc-title: Contents ++:icons: font ++:doctype: article ++ ++// ── Licensing ─────────────────────────────────────────────────────────────────────── ++image:https://img.shields.io/badge/Code-AGPL--3.0-blue.svg?logo=gnu[Code licence: AGPL-3.0,link="LICENSE"] ++image:https://img.shields.io/badge/Docs-CC--BY--SA--4.0-blue.svg?logo=creativecommons[Docs licence: CC-BY-SA-4.0,link="https://creativecommons.org/licenses/by-sa/4.0/"] ++// ── Toolchain ─────────────────────────────────────────────────────────────────────── ++image:https://img.shields.io/badge/Julia-1.12.5-9558B2?logo=julia[Julia 1.12.5,link="https://julialang.org"] ++image:https://img.shields.io/badge/R-%E2%89%A54.0-276DC3?logo=r[R at least 4.0,link="https://www.r-project.org"] ++image:https://img.shields.io/badge/Bun-1.3.10-F9F1E1?logo=bun[Bun 1.3.10,link="https://bun.sh"] ++// ── Continuous integration ────────────────────────────────────────────────────────── ++image:https://github.com/hyperpolymath/MetaManifold-WebUI/actions/workflows/ci.yml/badge.svg[CI,link="https://github.com/hyperpolymath/MetaManifold-WebUI/actions/workflows/ci.yml"] ++ ++Amplicon metabarcoding from raw paired-end Illumina reads to filtered, ++taxonomy-annotated ASV/OTU tables — orchestrated by one Julia engine, ++driven from one browser workbench, and increasingly explicit about what its ++statistics can and cannot claim. ++ ++For researchers who need reproducible microbiome analysis (clinical ++parasitology, environmental eDNA, teaching datasets) and for the ++professionals who run those analyses as a service. ++ ++[NOTE] ++.OpenSSF badges are planned, not claimed ++==== ++The estate's README standard asks for OpenSSF Best Practices and Scorecard ++badges on public repositories. MetaManifold is not enrolled yet; the badges ++will appear here on the day enrolment does, not before. ++==== ++ ++== The design in three layers ++ ++MetaManifold is three designs stacked honestly on top of each other. Each ++layer is still visible in the code; none of them pretends to be the whole ++story. (The full lineage, with file-level receipts, is ++link:https://github.com/hyperpolymath/MetaManifold-WebUI/wiki/Deep-Dives--Design-Progression[in the wiki].) ++ ++=== Layer 1 — the base design: R and Python around DADA2 ++ ++The scientific core is the established open-source pipeline as it is used ++*raw*, without us: DADA2 in R for exact sequence variants, and the ++swarm/vsearch shell tradition (Fred's metabarcoding pipeline) for OTUs. ++Before MetaManifold, a lab ran these as separate scripts and stitched the ++tables by hand. ++ ++[source] ++---- ++ R scripts (DADA2 tutorial lineage) Shell / Python (swarm pipeline lineage) ++ ────────────────────────────────── ──────────────────────────────────────── ++ filterAndTrim() cutadapt — primer trimming ++ learnErrors() vsearch --fastq_mergepairs — merge ++ dada() ── ASVs ── swarm -d 1 ── OTU clusters ++ mergePairs() vsearch --uchime_denovo — chimera filter ++ makeSequenceTable() vsearch --usearch_global — taxonomy ++ removeBimeraDenovo() ++ assignTaxonomy() ++ │ │ ++ └──────── counts + taxonomy ─────────────┘ ++ spreadsheets, ad-hoc glue, no shared provenance ++---- ++ ++*In place since the beginning:* DADA2 denoising and taxonomy (R, pinned by ++`renv.lock`), swarm clustering, vsearch alignment. *Never ours to change:* the ++science of these tools. This fork tracks application changes; it does not fork ++the science. ++ ++=== Layer 2 — the MetaManifold augmentation (JoshuaJewell's design) ++ ++The origin design wraps those raw lanes in one configurable Julia ++orchestrator and one web UI. Both lanes run side by side per run; their ++tables merge into a per-run DuckDB store the browser can query. ++ ++[source] ++---- ++ data/{study}/[{group}/]{run}/*.fastq.gz ++ │ ++ cutadapt · primer trimming ++ │ ++ ├─── DADA2 ASV lane ────────────────┬─── SWARM OTU lane ──────────────┐ ++ │ filter & trim → learn errors → │ merge pairs → dereplicate → │ ++ │ denoise → merge pairs → │ cluster (swarm) → │ ++ │ length filter → chimera cull → │ chimera filter → │ ++ │ taxonomy assign* → cd-hit-est* │ vsearch taxonomy │ ++ │ → vsearch* │ │ ++ └───────────────┬───────────────────┴───────────────┬────────────────┘ ++ │ │ ++ merge_taxa; join tables; apply filters │ ++ │ │ ++ └────────► DuckDB results store ◄───┘ ++ │ ++ Oxygen.jl REST + SSE ◄───────┴───────► React SPA workbench ++ (config cascade, runs & jobs, QC, results explorer, ++ annotation & curation, composition views) ++---- ++ ++*In place (layer 2):* the parallel ASV/OTU pipelines; the configuration ++cascade (instance → study → group → run, merged to `run_config.yml` as ++provenance); primer/database/composition editors in the UI; the results ++explorer with saved filter presets and Excel export; the functional ++annotation layer with dual-classifier consensus, contamination tagging and a ++FuncDB ledger; live job progress over server-sent events; remote (SSH) ++offload of the memory-heavy taxonomy step. *Coming at this layer:* the ++Julia-authored Stipple UI migration and standalone offline release archives ++(link:docs/migration/STATUS.md[migration status]). ++ ++=== Layer 3 — the steps this fork adds (hyperpolymath, 2026-09) ++ ++The fork's contribution is a discipline of honesty and typing around the ++science layer 2 runs: statistics that compute real estimates or refuse; ++numbers that never claim more precision than they carry; a typed frontend ++estate; and an engineering harness that keeps every claim checkable. ++ ++[source] ++---- ++ JoshuaJewell's MetaManifold (layer 2) ++ │ ++ ├── typed stages & a strict TypeScript estate ──── src/core/types.jl, ++ │ frontend/src/types/ ++ ├── tests, benchmarks, CI gates ────────────────── test/, frontend/tests/, ++ │ bench/, .github/workflows/ ++ ├── AnalysisConfig v1 ──────────────────────────── src/analysis/AnalysisConfig.jl ++ │ (NB-GLM, CLR/ILR + LM, logistic; BH mandatory; DANGER banner) ++ ├── real estimation or explicit refusal ────────── src/analysis/estimation.jl ++ │ (MASS::glm.nb / lm / glm fits with identifiability, ++ │ convergence and boundary reporting) ++ ├── exact counts & rationals ───────────────────── src/analysis/numeric_policy.jl, ++ │ src/analysis/exact_summaries.jl ++ ├── exact TSS/CSS/RSS offsets, no substitutions ── src/analysis/scaling.jl ++ └── epistemic receipts & refusals as results ───── src/core/epistemic.jl ++ ▼ ++ where this is heading (tracked, not promised as present) ++ ├── exact tests (Fisher, exact NB, permutation) ──── issue #3 [COMING] ++ ├── compositional methods (ANCOM-BC, ALDEx2, …) ──── issue #5 [COMING] ++ ├── occupancy · ordination · PhILR/SBP · zeros ───── issues #17–21 [COMING] ++ ├── CladeCumulus & Full Evidence Mode ────────────── issues #6–7 [COMING, scaffolded] ++ ├── Zenodo DOI minting ───────────────────────────── issue #8 [COMING] ++ └── symbolic formula engine ──────────────────────── issue #2 [BLOCKED ++ until the numeric layer has passed independent review] ++---- ++ ++*In place (layer 3):* placeholder statistics were removed and guarded ++against by test; maximum-likelihood fits (or explicit unsuccessful states) ++published under method conditions written *before* implementation ++(link:docs/statistics[method catalogue]); an exact/approximate/rounded ++numeric policy where higher precision is never sold as exactness; exact ++TSS/CSS/RSS offsets with refusals instead of silent substitutions; zero-depth ++samples healed before transforms; 0 TypeScript errors under strict + ++`exactOptionalPropertyTypes`; pinned toolchains (`mise` + Guix lanes) with ++sha256-pinned pipeline tools. *Coming and blocked items are listed as exactly ++that*, on the ++link:https://github.com/hyperpolymath/MetaManifold-WebUI/wiki/Status-and-Roadmap[Status and Roadmap] ++page of the wiki. ++ ++The layers nest — that is the whole picture: ++ ++[source] ++---- ++┌─ Layer 3 · verified statistics, typed estate, engineering gates ─────────────────┐ ++│ ┌─ Layer 2 · Julia orchestrator + WebUI: both lanes, DuckDB, config cascade ──┐ │ ++│ │ ┌─ Layer 1 · base design: raw DADA2 (R) + swarm/vsearch (shell) pipelines ─┐│ │ ++│ │ │ raw FASTQs → denoise or cluster → taxonomy → count tables ││ │ ++│ │ └──────────────────────────────────────────────────────────────────────────┘│ │ ++│ └─────────────────────────────────────────────────────────────────────────────┘ │ ++└─────────────────────────────────────────────────────────────────────────────────┘ ++---- ++ ++== What it does today ++ ++* One run takes raw paired-end FASTQs to chimera-filtered, taxonomy-annotated ++ ASV **and** OTU tables, with per-stage read accounting and QC (FastQC/MultiQC ++ plus DADA2 diagnostics) embedded in the UI. ++* Analysis on request, per run and across runs: alpha diversity with ++ significance testing that reports its status rather than degrading quietly, ++ taxonomic composition bars, organism composition categories (protozoa, ++ helminths, fungi, host, …), taxon overlap (Euler/UpSet), NMDS and PERMANOVA ++ via the locked R runtime, pipeline-stage read summaries. ++* Analysis configuration is typed and validated (`AnalysisConfig`): negative ++ binomial GLM, CLR/ILR linear models, logistic models — each with explicit ++ constraints, convergence and boundary reporting, or a refusal. Benjamini– ++ Hochberg correction is mandatory wherever several tests are reported. ++* Normalisation as honest bookkeeping: none, rarefaction, or exact ++ TSS/CSS/RSS size-factor **offsets** (depth modelled, response unchanged) — ++ never silently swapped for relative abundances. ++* Functional annotation with dual-classifier consensus (DADA2 bootstrap × ++ vsearch identity), contamination curation, manual BLAST override, and an ++ append-only FuncDB ledger that survives re-annotation. ++* Everything at a run's configuration is editable in the browser at every ++ cascade level; stale stages are flagged with the exact keys that changed. ++ ++[NOTE] ++.Planned — tracked, not claimed ++==== ++Exact statistical tests (issue #3) · compositional methods such as ANCOM-BC ++and ALDEx2 (issue #5) · occupancy and zero-inflated models, constrained ++ordinations, PhILR/SBP ILR bases, advanced zero handling (issues #17–21) · ++CladeCumulus phylogenetic explorer and Full Evidence Mode (issues #6–7) · ++Zenodo DOI minting (issue #8) · a symbolic formula engine (issue #2, formally ++**blocked** pending independent statistical validation of the numeric layer). ++The wiki's link:https://github.com/hyperpolymath/MetaManifold-WebUI/wiki/Status-and-Roadmap[Status and Roadmap] ++page carries the complete board. ++==== ++ ++== Quick start ++ ++[source,bash] ++---- ++git clone https://github.com/hyperpolymath/MetaManifold-WebUI.git ++cd MetaManifold-WebUI ++bash install.sh # Julia + R deps, sha256-pinned pipeline tools ++bash start.sh # builds the frontend on first run, serves on :8080 ++---- ++ ++Put paired-end `.fastq.gz` files under `data/MyProject/run_A/` (Illumina ++naming), open `http://localhost:8080`, create a study and launch a run. ++Bundled sample data: `data/MiSeq_SOP/` (the mothur MiSeq SOP set). ++ ++Toolchain of record: `mise.toml` (Julia 1.12.5, Bun 1.3.10, Node 20.20.2, ++just 1.43.1) with a Guix peer lane (`guix.scm` + `channels.scm`); R is a ++documented system exception restored byte-pinned by `renv.lock`. From a clean ++checkout, `just ci` runs every gate that CI runs. ++ ++Wondering how this works? See link:EXPLAINME.adoc[]. ++ ++== Repository layout ++ ++[cols="1,3", options="header"] ++|=== ++| Path | What lives there ++ ++| `src/pipeline/` | Stage implementations: cutadapt, DADA2 (R bridge), swarm, vsearch, cd-hit-est, merge_taxa — typed results, freshness-based skipping. ++| `src/analysis/` | Diversity, estimation, exact summaries, numeric policy, scaling/offsets, AnalysisConfig, chart builders. ++| `src/core/` | Config cascade, DuckDB store, provenance, epistemic statuses, R runtime bridge, validation. ++| `src/server/` | Oxygen.jl REST API + SSE job stream (`routes/`). ++| `frontend/` | TypeScript + React + Vite SPA (strict-typed estate; tests and benches alongside). ++| `ui/` | The opt-in Stipple/Vue UI migration slice (legacy React app remains default). ++| `config/` | Defaults, schemas, primers, databases, composition filters, CI fixtures. Per-study/run overrides live under `data/`, `projects/`. ++| `docs/` | Statistics method catalogue and conditions, type-system notes, compliance and migration records. ++| `docs/wikis/` | Source of the GitHub wiki (BerryWiki page format); synced to `MetaManifold-WebUI.wiki.git`. ++| `test/`, `bench/`, `frontend/tests/` | Julia tests, benchmark lanes with baselines, frontend unit/integration/e2e lanes. ++| `Justfile`, `install.sh`, `start.sh` | Every task is a `just` recipe; the two shell entry points above. ++|=== ++ ++== Documentation ++ ++* link:EXPLAINME.adoc[EXPLAINME] — receipts: every claim here mapped to code, with caveats. ++* link:https://github.com/hyperpolymath/MetaManifold-WebUI/wiki[The project wiki] — the long-form documentation, in three audiences: *users* (with separate tracks for academics and lab professionals), *platform maintainers* (operator and repo-steward tracks), and *developers*, plus the deep dives this README deliberately does not attempt. ++* link:docs/statistics/method-catalogue-v1.md[method catalogue] and link:docs/statistics/method-conditions/[method conditions] — what each statistical method accepts, computes, and refuses. ++* link:docs/reproducibility.md[Reproducibility] — toolchain source of truth; link:docs/compliance/[compliance records] — RSR and standards alignment. ++* link:CONTRIBUTING.md[Contributing] · link:SECURITY.md[Security] · link:ROADMAP.md[Roadmap] · link:CHANGELOG.md[Changelog] · link:CITATION.cff[Citing this work] ++ ++== Licence ++ ++Source code is licensed under the GNU Affero General Public License v3.0 — ++see link:LICENSE[license]; upstream-authored files carry `AGPL-3.0-only` ++headers and fork-authored files `MPL-2.0`/`CC-BY-SA-4.0` per the estate ++licence policy, all summarised in link:NOTICE[notice]. ++This documentation (README.adoc) is licensed under CC BY-SA 4.0. ++Pipeline tools are orchestrated, not vendored; each retains its own licence — ++see the third-party note in link:NOTICE[notice]. ++ ++Copyright © 2026 Joshua Benjamin Jewell (origin design) and the ++hyperpolymath fork authors. Full attribution, including the DADA2 tutorial ++(CC BY 4.0) and Fred's metabarcoding pipeline lineages, is in link:NOTICE[notice]. +diff --git a/README.md b/README.md +deleted file mode 100644 +index f76644d..0000000 +--- a/README.md ++++ /dev/null +@@ -1,685 +0,0 @@ +- +-# MetaManifold +- +-[![License: AGPL-3.0](https://img.shields.io/badge/License-AGPL--3.0-blue.svg)](LICENSE) +-[![Julia 1.12.5](https://img.shields.io/badge/Julia-1.12.5-9558B2?logo=julia)](https://julialang.org) +-[![R ≥ 4.0](https://img.shields.io/badge/R-%E2%89%A54.0-276DC3?logo=r)](https://www.r-project.org) +-[![CI](https://github.com/hyperpolymath/MetaManifold-WebUI/actions/workflows/ci.yml/badge.svg)](https://github.com/hyperpolymath/MetaManifold-WebUI/actions/workflows/ci.yml) +- +-MetaManifold wraps standard amplicon sequencing workflows into a single configurable Julia orchestrator: from raw paired-end Next Generation Sequencing reads through denoising, taxonomy assignment, taxonomic filtering, and functional annotation, with interactive configuration and analysis in the browser. +- +-## Overview +- +-MetaManifold consists of a Julia backend (pipeline engine + REST API) and a TypeScript/React frontend. The pipeline runs FastQC, MultiQC, cutadapt, DADA2, SWARM, vsearch, and cd-hit-est under the hood; results are stored in per-run DuckDB databases and served to the frontend as interactive Plotly charts and filterable tables. Pipeline configuration is editable directly in the web UI at every cascade level (see [Configuration](#configuration)), and a functional-annotation layer supports manual curation. +- +-**Pipeline stages** +- +-``` +-Raw FASTQs (data/{study}/[{group}/]{run}/*.fastq.gz) +- │ +- cutadapt, primer trimming +- │ +- ├────────────────────────────┐ +- │ │ +- DADA2*, ASV; SWARM*, OTU; +- filter & trim merge pairs +- learn error rates dereplicate +- denoise + merge chimera filter +- length filter cluster OTUs +- chimera removal │ +- taxonomy assign* │ +- │ │ +- cd-hit-est*, demultiplex │ +- │ │ +- vsearch* vsearch, global alignment +- │ │ +- ├────────────────────────────┘ +- │ +- merge_taxa; +- join ASV tables* +- join counts-taxonomy +- apply filters +- │ +- DuckDB results store +-``` +-*optional +- +-**Analysis** +- +-Once a run completes, analysis is performed on request through the web UI, both per run and across runs within a study: +- +-- Alpha diversity (richness, Shannon, Simpson), per sample or as cross-run comparison boxplots with significance testing +-- Taxonomic composition bar charts at a chosen rank, relative or absolute +-- Organism-composition charts that classify ASVs/OTUs into biological categories (a "Composition" view, e.g. protozoa, helminths, fungi, host) +-- Taxon overlap across runs as proportional Euler or UpSet plots +-- Pipeline stage read-count summaries +-- NMDS ordination (Bray-Curtis, via R/vegan) +-- PERMANOVA (via R/vegan) +- +-Counts may be normalised before analysis (none, rarefaction to a fixed or auto-resolved depth, or relative sum scaling), and contamination-flagged taxa may be included or excluded. All analysis charts are returned as Plotly JSON and rendered interactively in the browser. +- +-## Prerequisites +- +-**One-command toolchain (recommended — the repo is standalone):** the pinned +-dev toolchain lives in `mise.toml` (julia 1.12.5, bun 1.3.10, node 20.20.2, +-just 1.43.1 — exact CI pins) with `guix.scm`/`channels.scm` as the Guix +-peer lane and `.envrc` for direnv auto-activation: +- +-```bash +-curl https://mise.run | sh && just bootstrap # or: guix time-machine -C channels.scm -- shell -D -f guix.scm +-just ci # the proof: all gates green +-just setup-full # full first-run: + Julia deps + sha256-pinned pipeline tools +-just start # launches the server (estate launcher, :8080) +-``` +- +-Then every task is a `just` recipe (`just` lists them). R remains a system +-install (not in the mise registry — documented exception in +-`docs/reproducibility.md`, which is the toolchain source of truth). +-Pipeline tools (cutadapt, MultiQC, FastQC, cd-hit-est, vsearch, swarm) are +-fetched byte-exact by `install.sh` against the sha256-pinned records in +-`config/defaults/tool_versions.yml`; the Guix shell also carries functional +-equivalents for development. +- +-- **Julia** >= 1.0, pinned 1.12.5 via `mise.toml` (installed automatically by `install.sh` if missing) +-- **Julia** 1.12.5 exactly, pinned via `mise.toml` (installed automatically by `install.sh` if missing) +- - Ubuntu/Debian: `sudo apt install r-base` +- - macOS: `brew install r` or [CRAN package](https://cran.r-project.org/bin/macosx/) +-- **Bun** >= 1.3.10 for building the frontend (pinned in `.bun-version`; CI reads the same version from `config/defaults/tool_versions.yml`). `bun install` in `frontend/` pulls all JS dependencies, including `react-chart-editor` and `react-plotly.js`. The chart editor is fed `plotly.js-dist-min` rather than full `plotly.js` to keep the bundle size manageable. +- +-## Installation +- +-```bash +-git clone https://github.com/JoshuaJewell/MetaManifold-WebUI.git +-cd MetaManifold-WebUI +-bash install.sh +-``` +- +-`install.sh` will check for Julia and R, install Julia and R dependencies, locate or download each external tool (cutadapt, FastQC, MultiQC, vsearch, cd-hit-est). Tool paths can be configured manually in `config/tools.yml`, including by SSH if you wish to use a server-hosted binary. +- +-To update: +-```bash +-bash install.sh --update +-``` +- +-### Reproducing the R environment +- +-The R-side dependencies (DADA2, vegan, and their transitive packages) are +-pinned with [`renv`](https://rstudio.github.io/renv/). The lockfile lives at +-`renv.lock` and a project-local library is created on first activation. +- +-```bash +-Rscript -e 'if (!requireNamespace("renv", quietly=TRUE)) install.packages("renv", repos="https://cloud.r-project.org"); renv::restore(prompt = FALSE)' +-``` +- +-Subsequent invocations of `Rscript` or `R` from the repository root will pick +-up the project library automatically via the committed `.Rprofile`. To add or +-upgrade a package, install it inside the project (`renv::install(...)`) and +-record the change with `renv::snapshot()`. +- +-### Tool paths +- +-`install.sh` generates `config/tools.yml`. You can edit it manually, for example, to point a tool at a remote server: +- +-```yaml +-vsearch: +- path: "user@bioserver:/home/user/software/vsearch" +-``` +- +-When a remote SSH path is set, the pipeline routes that tool's invocations through +-`ssh`. `config/defaults/tools.yml` contains the full config format. +- +-## Quick start +- +-### 1. Place paired-end FASTQs under data/ +-Put `.fastq.gz` files into `data/MyProject/run_A`. +- +-### 2. Start the server (builds frontend on first run) +-`bash start.sh` +- +-### 3. Open http://localhost:8080 +- +-The web UI lets you create studies, configure pipeline parameters, launch runs, and explore results interactively. All state lives in the filesystem under `data/` and `projects/`. +- +-| Environment variable | Default | Description | +-| ------------------------- | ----------------- | ------------- | +-| `JULIA_METAMANIFOLD_PORT` | `8080` | Server port | +-| `JULIA_METAMANIFOLD_ROOT` | working directory | Project root | +-| `JULIA_THREADS` | `8` | Julia threads | +- +-Or run the Julia server directly: +- +-```bash +-julia --project=. src/server/server.jl +-``` +- +-## Web interface +- +-The browser interface is the primary way to drive MetaManifold. Beyond creating studies and launching runs, it offers the following: +- +-### Editing configuration +- +-Every pipeline setting can be edited in the UI without touching a YAML file. Configuration is presented as collapsible accordion sections (study design, primer trimming, DADA2 denoising and taxonomy, OTU clustering, annotation, analysis) and can be set at any level of the cascade: the instance-wide defaults, a study, a group, or an individual run. Edits at a finer level override coarser ones (see [Configuration](#configuration) for the cascade rules). When a setting changes, the affected pipeline stages are flagged as stale, so it is clear which outputs a re-run would regenerate; a tooltip lists exactly which keys changed and at which level. +- +-### Pipeline runs and outputs +- +-A run's page is the working surface for that run. It is where the run-level configuration above is edited, where the full pipeline or any individual stage (including the DADA2 substages) is launched, and where each stage's status is shown. Long-running jobs report progress live through a server-sent event stream, and the jobs panel lets you watch or cancel them. As stages complete, their outputs become available through the views that follow: quality reports, the results table, functional annotation, and composition. +- +-

+- A run page showing per-stage status, launch controls, and run-level configuration +-

+- +-

+- Live pipeline progress: the jobs panel and the server-sent event stream +-

+- +-### QC +- +-Raw-read QC (FastQC aggregated by MultiQC) and the DADA2 quality, denoising, merging, and taxonomy diagnostics are embedded in the UI, each with the relevant per-stage configuration alongside and a re-run control. +- +-

+- QC view embedding the MultiQC report and DADA2 quality diagnostics +-

+- +-### Results explorer +- +-Each run's merged taxonomy-and-count table, and any derived tables, can be browsed interactively. The table supports per-column filtering (text search, numeric range, include/exclude lists) and a global text filter, column sorting, configurable pagination, and column visibility toggles including taxonomy-source presets (VSEARCH-only, DADA2-only, or all) and a switch for the per-sample count columns. Sequences carry BLAST links, and OTU rows can be expanded to their constituent sequences. Frequently used filters can be saved as named presets and reapplied; filtered tables can be saved back into the run or exported to Excel (`.xlsx`). +- +-

+- Results explorer table with per-column filters and taxonomy-source column presets +-

+- +-### Annotation +- +-The annotation view applies a functional database (`funcdb`) to a run's merged table, independently for the VSEARCH and DADA2 taxonomies. For each sequence it attaches functional metadata (function, associated organism and material, environment, pathogen status, and free-text notes) down to a configurable maximum rank. It also derives a 'consensus rank' (the finest rank at which the two classifiers agree) and a composite confidence score combining DADA2's bootstrap support at this rank and VSEARCH's percent identity (mostly for the sake of curiosity). +- +-Curation is supported directly in the view: +- +-- **Contamination tagging:** mark a taxon as contamination (yes / no / unassigned); the flag applies to all rows sharing that rank and taxon, with a live summary of affected reads. +-- **Manual BLAST assignment:** override the assignment for an individual sequence inline. +-- **FuncDB ledger:** add a new functional entry for a taxon, prefilled from the selected row. Entries are written to an append-only ledger and become available to subsequent annotation runs. User edits (contamination flags, manual assignments) are preserved when annotations are regenerated. +- +-

+- Annotation view showing consensus rank, confidence score, and contamination tagging controls +-

+- +-### Composition +- +-The composition view classifies each ASV/OTU into a biological category and renders per-sample or pooled stacked bar charts. Category sets live in `config/composition.yml` (the bundled `default` set covers protozoa, helminths, fungi, host, plants, and invertebrates); each category references a named taxonomic filter from the `filters:` library in that same file. Both are editable from the Compositions page under SYSTEM in the sidebar. A category summary precedes the chart, and a quality filter can cap the number of unresolved taxonomic placeholders admitted. +- +-

+- Composition view with a stacked organism-category bar chart and category summary +-

+- +-## Configuration +- +-Pipeline settings use a cascade: each level overrides the one above it, and any key you omit is inherited from the nearest ancestor. The fully merged result is written to `run_config.yml` at runtime; that is the single place to see exactly what was used for a run. +- +-Settings can be edited in the web UI (per-study, per-group, or per-run) or as YAML files directly. +- +-| File | Purpose | +-|------|---------| +-| `config/defaults/` | Canonical defaults for every setting; do not edit | +-| `config/composition.yml` | Composition library: named taxonomic filters and the category sets that reference them | +-| `config/presets/` | Saved table-view filter presets, written from the Tables view | +-| `config/databases.yml` | Database URIs and optional local paths. Editable from the Databases page under SYSTEM in the sidebar | +-| `config/primers.yml` | Primer sequences and pair definitions. Editable from the Primers page under SYSTEM in the sidebar | +-| `config/tools.yml` | Tool binary paths (cutadapt, FastQC, MultiQC, vsearch, cd-hit-est) | +-| `config/pipeline.yml` | Machine-level overrides (lowest user-editable precedence) | +-| `data/{name}/pipeline.yml` | Study-level overrides | +-| `data/{name}/{group}/pipeline.yml` | Group-level overrides (intermediate directories) | +-| `data/{name}/{run}/pipeline.yml` | Run-level overrides (highest precedence) | +-| `projects/{name}/{run}/run_config.yml` | Generated merged config (provenance); do not edit | +- +-Each `pipeline.yml` stub is created with a comment block explaining that level's role. Write only the keys you want to change; omit the rest. +- +-### Configuring databases (`config/databases.yml`) +- +-This is the single place to manage DB URIs shared across all projects. +- +-```yaml +-databases: +- dir: "./databases" +- pr2: +- dada2: +- uri: "https://..." # DADA2-format FASTA (downloaded on first use) +- local: ~ # set to a local path to skip download +- vsearch: +- uri: "https://..." # vsearch-format FASTA +- local: ~ +-``` +- +-Edit this on the Databases page under SYSTEM in the sidebar, or in the YAML directly. The page edits the shared cache directory (`dir`) and, per database, the dada2 and vsearch source URIs, a `local:` override for a file already on disk, `remote_path` (dada2 only) for a file already present on the remote taxonomy host, the ordered taxonomy `levels`, the `vsearch_format` parser selector, and the taxonomy `corrections`. Adding and removing a database is supported, not just retuning PR2. `vsearch_format` offers exactly `pr2` and `generic`: only the literal `pr2` selects pipe-separated parsing, and anything else is parsed generically. +- +-Removing or renaming a database, or changing its `levels`, is allowed, but the save reports which studies it affects. The warning resolves the real config cascade, so it names the studies that inherit the database without naming it, not merely those that mention it explicitly. +- +-Both formats of one database should come from the same reference release: the dual-classifier consensus compares DADA2 and VSEARCH labels for string equality, so references drawn from different releases score genuine agreements as disagreements. The editor warns on a version-token mismatch between the two URIs, but this is a filename heuristic and cannot warn for a database whose URIs carry no version. +- +-### Defining primer pairs `primers.yml` +- +-Maps primer names to sequences and defines which forward/reverse sequences constitute a pair: +- +-```yaml +-Forward: +- PrimerF: "CCAGCASCYGCGGTAATTCC" +- +-Reverse: +- Primer1R: "ACTTTCGTTCTTGATYRA" +- Primer2R: "DCTKTCGTYCTTGATYRA" +- +-Pairs: +- - PrimerPair1: +- - PrimerF +- - Primer1R +- - PrimerPair2: +- - PrimerF +- - Primer2R +-``` +- +-Store all primer pairs in here and reference whichever combinations you need per project. Shared primers across pairs (same forward primer in two pairs) are automatically deduplicated in the `cutadapt` invocation since otherwise it complains a bit. If you need duplicates, you must create the same sequence under a different name. +- +-Edit this on the Primers page under SYSTEM in the sidebar, or in the YAML directly. The page adds and removes primers and composes pairs from them, validating each sequence against the IUPAC base set as you type. The whole document is validated before it lands on disk, so a pair naming a primer that does not exist is rejected and the file is left untouched. +- +-Pair names are referenced by `cutadapt.primer_pairs` in `pipeline.yml`. Removing or renaming a pair that a study still references is permitted, but the save reports which studies, groups, or runs named it, so the dangling reference is never silent. Renaming a primer carries its pairs with it automatically. +- +-### Configuring cutadapt (`cutadapt:` in `pipeline.yml`) +- +-Selects which primer pairs to apply and controls trimming behaviour. +- +-```yaml +-cutadapt: +- # Names must match keys in the Pairs section of config/primers.yml. +- primer_pairs: +- - PrimerPair1 +- - PrimerPair2 +- min_length: 200 # discard reads shorter than this after trimming (-m) +- discard_untrimmed: true # drop reads where no adapter was found (--discard-untrimmed) +- cores: 0 # parallel cores; 0 = auto-detect (-j) +- quality_cutoff: ~ # 3' quality trimming cutoff, null to disable (-q) +- error_rate: ~ # max adapter mismatch rate, null = cutadapt default (-e) +- overlap: ~ # min adapter overlap length, null = cutadapt default (-O) +- optional_args: "" # additional flags passed verbatim to cutadapt +-``` +- +-### Configuring DADA2 (`dada2:` in `pipeline.yml`) +- +-```yaml +-dada2: +- file_patterns: +- mode: "paired" # paired | forward | reverse +- +- # Filter and trim; DADA2's filterAndTrim(): +- filter_trim: +- trunc_q: 2 +- trunc_len: [220, 220] # [forward, reverse]; first value used for single-end mode +- max_ee: [3, 3] # maximum expected errors in F and R reads +- min_len: 175 +- max_n: 0 +- match_ids: true +- rm_phix: true +- +- # Denoising; learnErrors() and dada(): +- dada: +- seed: 123 +- nbases: 200000000 +- max_consist: 15 +- pool_method: "pseudo" # none | pseudo | true +- +- # Merging; mergePairs(), paired mode only: +- merge: +- min_overlap: 20 +- max_mismatch: 0 +- trim_overhang: true +- +- # ASV length filtering and chimera removal: +- asv: +- band_size_min: 200 # null to skip length filtering +- band_size_max: 430 +- denovo_method: "consensus" # consensus | pooled | per-sample +- +- # Taxonomy; assignTaxonomy() against the configured database: +- taxonomy: +- database: pr2 # key into config/databases.yml +- multithread: 4 # threads for assignTaxonomy(); higher values increase memory use +- min_boot: 0 # minimum bootstrap confidence to retain (0-100) +- # Taxonomy rank names are read from databases.yml (the `levels:` key under +- # each database entry). Do not set them here. +- +- # Optional: offload the memory-intensive assignTaxonomy() step to a remote +- # server via SSH. Omit or set host to null to run locally. +- # DISCLAIMER: You are solely responsible for ensuring you have authorisation +- # to use the configured host. See config/defaults/pipeline.yml for the full disclaimer. +- remote: +- host: ~ # user@hostname +- identity_file: ~ # path to SSH private key; null to use password auth +- rscript: "Rscript" # path to Rscript on the server +- staging_dir: "/absolute/path/on/server" +- # To avoid transferring the database each run, set dada2.remote_path under +- # the relevant database entry in config/databases.yml instead. +- +- # Output filename prefixes (all written to dada2/Tables/): +- output: +- seq_table_prefix: "seqtab_nochim" +- fasta_prefix: "asvs" +- taxa_prefix: "taxonomy" +-``` +- +-**Outputs written to `projects/{name}/{run}/dada2/Tables/`:** +- +-| File | Contents | +-| ---------------------------| --------------------------------------------------------| +-| `seqtab_nochim.csv` | Chimera-free ASV count table (samples x ASVs) | +-| `asvs.fasta` / `asvs.csv` | ASV sequences with short identifiers (seq1, seq2, ...) | +-| `taxonomy.csv` | Taxonomy assignments per ASV | +-| `taxonomy_bootstraps.csv` | Bootstrap confidence values per rank | +-| `taxonomy_combined.csv` | Taxonomy ├ bootstrap columns combined | +-| `tax_counts.csv` | Taxonomy ├ per-sample counts | +-| `asv_counts.csv` | ASV sequences ├ per-sample counts (no taxonomy) | +-| `pipeline_stats.csv` | Read counts retained at each pipeline stage | +- +-### Configuring vsearch (`vsearch:` in `pipeline.yml`) +- +-Controls the alignment thresholds used when assigning taxonomy against the reference database. Per run, this provides the same configuration for both ASV and OTU pipeline if they are running parallel. +- +-```yaml +-vsearch: +- identity: 0.75 # minimum sequence identity (--id) +- query_cov: 0.8 # minimum fraction of query covered (--query_cov) +- maxaccepts: ~ # stop after this many hits per query, null = vsearch default +- maxrejects: ~ # max rejected candidates, null = vsearch default +- strand: ~ # "plus" or "both"; null = vsearch default +- optional_args: "" # additional flags passed verbatim to vsearch +-``` +- +-### Configuring cd-hit-est (`cdhit:` in `pipeline.yml`) +- +-Optional clustering step that collapses near-identical ASVs before vsearch taxonomy assignment. Used here for when using primers in multiplex, to reduce inflation from same sequences from different primers appearing different. +- +-```yaml +-cdhit: +- identity: 1 # sequence identity threshold (-c) +- threads: 0 # worker threads; 0 = all available (-T) +- optional_args: "" # additional flags passed verbatim to cd-hit-est +-``` +- +-### Configuring swarm (`swarm:` in `pipeline.yml`) +- +-OTU clustering pipeline run in parallel with DADA2. Produces an OTU count table and FASTA which are carried through vsearch taxonomy assignment and `merge_taxa` alongside the ASV outputs. +- +-```yaml +-swarm: +- differences: 1 # -d: max differences between sequences in the same cluster +- threads: 0 # -t: worker threads; 0 = all available +- chimera_check: true # run vsearch --uchime_denovo before clustering +- min_abundance: 2 # --minsize: discard singleton dereps before clustering +- fastq_minovlen: 20 # min overlap for paired-end merging +- identity: 0.97 # --id: threshold for mapping reads back to OTU seeds +- optional_args: "" # additional flags passed verbatim to swarm +-``` +- +-### Configuring merge_taxa (`merge_taxa:` in `pipeline.yml`) +- +-Controls which filter configs are applied when merging taxonomy and count tables. `merged.csv` (unfiltered) is always written; each entry in `filters` produces an additional filtered CSV. +- +-```yaml +-merge_taxa: +- filters: +- - "protist_filter.yml" # -> merged/protist_filter.csv +-``` +- +-Each entry names a filter in the `filters:` library of `config/composition.yml`. Remove all entries (or set `filters: []`) to produce only the unfiltered `merged.csv`. +- +-### Configuring analysis (`analysis:` in `pipeline.yml`) +- +-Controls the defaults applied to the analysis charts (alpha diversity, taxa bar, NMDS, etc.). Per-chart choices such as the taxonomic rank, and relative/absolute abundance are selected interactively in the UI and are not config keys. +- +-```yaml +-analysis: +- exclude_categories: # composition categories to drop from figures; [] to keep all +- - {set: contamination, category: Contaminant, apply_to: [diversity, taxa, venn]} +- # apply_to surfaces: diversity | taxa | composition | venn +- # (omit apply_to to act on every surface) +- normalisation: none # none | rarefaction | rss (relative sum scaling) +- normalisation_depth: 0 # rarefaction depth; 0 = auto (min positive library size) +- alpha: +- show_points: true # overlay individual sample points on boxplots +- annotate_significance: false # annotate pairwise significance on grouped alpha +- pairwise_brackets: false # draw significance brackets between groups +- paired_samples: false # treat samples as paired in the significance test +- significance_test: "kruskal_wallis" # test used for group comparison +- nmds: +- max_stress: 0.2 # warn if NMDS stress exceeds this value +-``` +- +-### Configuring annotation (`annotation:` in `pipeline.yml`) +- +-Controls the functional-annotation layer applied in the Annotation view. +- +-```yaml +-annotation: +- max_rank: "species" # finest rank to which functional metadata is attached +-``` +- +-### Configuring taxonomic filtering (`filters:` in `config/composition.yml`) +- +-Each named filter in the `filters:` library of `config/composition.yml` defines one biological group to extract from the merged table. A category set references these filters by name, and the same filters back the `merge_taxa.filters` stage, which produces one additional CSV per entry. Edit them on the Compositions page under SYSTEM in the sidebar, or in the YAML directly. +- +-Saved table-view presets are a separate concern and live in `config/presets/`; the Tables view reads and writes them. +- +-#### Database-specific filters +- +-Each filter carries a `databases:` key so that it is only applied when the active database matches. The following filters ship in the library: +- +-| Category | PR2 match | +-|----------|-----------| +-| `bacteria_archaea` | `Domain` = Bacteria\|Archaea | +-| `environmental_protozoa` | `Subdivision` = Cercozoa\|Gyrista\|Ciliophora\|Chrompodellids | +-| `fungi` | `Subdivision` = Fungi | +-| `helminths` | `Class` = Nematoda (excl. *Miculenchus*) | +-| `parasitic_protozoa` | `Subdivision` = Apicomplexa\|Parabasalia\|Fornicata\|Bigyra | +-| `plants_invertebrates` | Exclusion-based (PR2 ranks) | +-| `protist` | Exclusion-based (PR2 ranks) | +-| `vertebrates` | `Class` = Craniata | +- +-Example: +- +-```yaml +-# fungi.pr2.yml +-databases: [pr2] +- +-filters: +- - column: Subdivision +- pattern: Fungi +- action: keep # keep rows matching the pattern (default action is exclude) +- +-remove_empty: +- - Subdivision +-``` +- +-#### Filter file format +- +-```yaml +-databases: [pr2] # omit to apply regardless of active database +- +-mappings: # optional column remapping applied before filters +- - source_column: Division +- target_column: Supergroup +- values: { Rhizaria: Rhizaria, Alveolata: Alveolata } +- +-filters: +- - column: Domain +- pattern: "Bacteria|Archaea" +- regex: true # false (default) = substring match +- action: exclude # exclude (default) | keep +- +-remove_empty: # remove rows where this column is blank or "NA" +- - Domain +-``` +- +-## Deployment +- +-### Local (single machine) +- +-```bash +-bash start.sh +-``` +- +-Open `http://localhost:8080`. The backend serves the frontend automatically. +- +-## Engineering gates +- +-The fork maintains an engineering estate around the application. From a +-clean checkout (`frontend/`): +- +-| Gate | Command | Authority | +-|---|---|---| +-| Strict typecheck | `bun run typecheck` | 0 errors, gated | +-| Unit + integration tests | `bun test` | gated (no DOM lane) | +-| Benchmarks | `bun run bench/` | informational, no gate | +-| Everything above | `bun run check` | combined pre-push gate | +-| Licence headers | `scripts/check-spdx.sh` | gated | +-| Formatting | `scripts/check-format.sh` | gated | +-| Lint (tsc semantics + shell) | `scripts/check-lint.sh` | gated | +- +-CI runs the same gates (see `.github/workflows/ci.yml`: repo-hygiene job, +-then the pinned Julia and frontend jobs). Contributor setup, commit and +-branch conventions: `CONTRIBUTING.md`. Frontend reproducibility: +-`docs/reproducibility.md`. Type estate map: `docs/types/architecture.md`. +-Test inventory and metrics: `docs/testing/coverage.md`. +- +-## Input data +- +-Place paired-end FASTQ files under `data/{project_name}/` following Illumina naming: +- +-``` +-data/MyProject/SampleName_*_L001_R1_001.fastq.gz +-data/MyProject/SampleName_*_L001_R2_001.fastq.gz +-``` +- +-For multi-run projects, nest runs in subdirectories. The server detects any directory containing `.fastq.gz` files as a leaf run and creates a matching project directory under `projects/{project_name}/`. +- +-## Output structure +- +-All outputs for a given run live under `projects/{project_name}/{run}/`: +- +-``` +-projects/{project_name}/{run}/ +-├── cutadapt/ # Trimmed FASTQ pairs and logs +-│ └── logs/ +-├── QC/ +-│ ├── fastqc/ # Per-file FastQC HTML reports +-│ ├── multiqc_report.html # MultiQC summary across all samples +-│ └── logs/ +-├── dada2/ +-│ ├── Tables/ +-│ │ ├── seqtab_nochim.csv # ASV count table +-│ │ ├── asvs.fasta # ASV sequences +-│ │ ├── asvs.csv # ASV sequence index +-│ │ ├── taxonomy.csv # Taxonomy assignments +-│ │ ├── taxonomy_bootstraps.csv +-│ │ ├── taxonomy_combined.csv +-│ │ ├── tax_counts.csv # Taxonomy + per-sample counts +-│ │ ├── asv_counts.csv # Sequences + per-sample counts +-│ │ └── pipeline_stats.csv +-│ ├── Figures/ # Quality profile and error rate PDFs +-│ ├── Checkpoints/ # RData checkpoints for stage resumption +-│ └── Logs/ # Per-stage R logs +-├── cdhit/ +-│ ├── asvs.fasta # Clustered ASV sequences +-│ └── asvs.fasta.clstr # Cluster membership file +-├── swarm/ +-│ ├── otus.fasta # OTU representative sequences +-│ ├── otus.count_table.csv # OTU count table (samples x OTUs) +-│ └── logs/ +-├── vsearch/ +-│ ├── taxonomy.tsv # Top-hit taxonomy assignments (ASV or OTU) +-│ └── logs/ +-└── merged/ +- ├── merged.csv # Merged taxonomy + counts (all taxa) +- ├── protist_filter.csv # Filtered subset (one per merge_taxa.filters entry) +- └── results.duckdb # DuckDB database for API queries +-``` +- +-## REST API +- +-The server exposes a REST API under `/api/v1/`. Key endpoint groups: +- +-| Group | Endpoints | Description | +-| ----------------| ---------------------------------------------------------------------------------------------------------------| ------------------------------------------------------------------------------------------| +-| Studies | `GET/POST/DELETE /studies`, `POST .../rename` | List, create, rename, delete studies | +-| Groups | `POST/DELETE /studies/{study}/groups`, `POST .../rename` | Create, rename, delete groups | +-| Runs | `GET/POST/DELETE /studies/{study}/runs`, `POST .../rename` | List, create, rename, delete runs | +-| Config | `GET/PATCH/DELETE .../config`, `GET .../config/overrides` | Read and edit config at any cascade level; list downstream overrides | +-| Primers | `GET /primers`, `GET /primers/document`, `PUT /primers` | List pair names; read and replace the whole primers document (validated before writing) | +-| Pipeline | `POST .../pipeline`, `POST .../stages/{stage}` | Launch full-study, single-run, or individual-stage jobs | +-| Jobs | `GET/DELETE /jobs`, `GET /jobs/{id}/logs` | Monitor and cancel running pipeline jobs | +-| Events | `GET /events` | Server-sent event stream of real-time job and stage updates | +-| Results | `GET/POST/DELETE .../results/tables/...` | List, query, filter, save, export (`.xlsx`), and delete tables; OTU member drill-down | +-| QC | `GET .../results/qc`, `GET .../results/dada2` | MultiQC report metadata and DADA2 figures, logs, stats | +-| Analysis | `POST .../analysis/{alpha,taxa-bar,venn}`, `GET .../analysis/{pipeline-stats,ranks}` | Per-run charts and rank discovery | +-| Cross-run | `POST /studies/{study}/analysis/{alpha,taxa-bar,nmds,permanova,venn}` | Comparison, NMDS, PERMANOVA, taxon overlap across runs | +-| Composition | `GET /category-sets`, `POST .../composition/{source}/{build,query,distinct,analysis}` | Category-set listing and organism-composition tables and charts | +-| Annotation | `GET/POST .../annotations/{source}/...`, `POST /funcdb/entries`, `PATCH .../{contamination,blast-assignment}` | Generate and query annotations, curate contamination and assignments, add FuncDB entries | +-| Filter presets | `GET/POST/DELETE /filter-presets` | Save, list, delete reusable table filters | +-| Databases | `GET /databases`, `GET /databases/document`, `PUT /databases`, `POST /databases/{key}/download` | List and download taxonomy databases; read and replace the whole databases document (validated before writing, returns advisory warnings) | +-| System | `POST /init`, `GET /capabilities` | Initialise project directories; report server capabilities (e.g. R availability) | +- +-All responses are JSON. Analysis endpoints return Plotly chart specifications. +- +-## Architecture +- +-``` +-frontend/ TypeScript + React + Vite (SPA) +-src/ +- core/ Types, config cascade, validation, DuckDB store, logging +- pipeline/ Pipeline stages (cutadapt, dada2, swarm, vsearch, cd-hit-est, merge_taxa) +- annotation/ Functional database (funcdb): dual-classifier consensus and curation +- analysis/ Diversity metrics + Plotly chart builders +- server/ Oxygen.jl HTTP server +- routes/ REST API route handlers +-config/ Default configs, filters, CI fixtures +-data/ Input FASTQs (user-managed) +-projects/ Pipeline outputs (generated) +-``` +- +-Each pipeline stage returns a typed result (`TrimmedReads`, `ASVResult`, `OTUResult`, `TaxonomyHits`, `MergedTables`) and skips automatically if outputs are already up to date (mtime-based for files, content-hash-based for configuration). Rerunning after a config change only re-executes the minimum necessary stages. +- +-## Third-party tools +- +-This project orchestrates the following tools. Each is fetched from its upstream source by `install.sh` and is subject to its own licence; no third-party binaries are included in this repository. +- +-| Tool | License | Source | +-| -------------------------------------------------| ---------| -------------------------| +-| [cutadapt](https://github.com/marcelm/cutadapt) | MIT | PyPI | +-| [FastQC](https://github.com/s-andrews/FastQC) | GPL v3 | Babraham Bioinformatics | +-| [MultiQC](https://github.com/MultiQC/MultiQC) | GPL v3 | PyPI | +-| [DADA2](https://benjjneb.github.io/dada2/) | LGPL v3 | Bioconductor | +-| [swarm](https://github.com/frederic-mahe/swarm) | GPL v3 | GitHub Releases | +-| [vsearch](https://github.com/torognes/vsearch) | GPL v3 | GitHub Releases | +-| [cd-hit](https://github.com/weizhongli/cdhit) | GPL v2+ | GitHub Releases / apt | +- +-## Acknowledgements +- +-This pipeline draws on the following prior work: +- +-- **Frédéric Mahé**: [Fred's metabarcoding pipeline](https://github.com/frederic-mahe/swarm/wiki/Fred's-metabarcoding-pipeline) informed the overall workflow architecture, namely the sequencing of primer trimming, `swarm.jl`, vsearch-based taxonomy assignment, and the final table merge/filter stages. +-- **Benjamin J. Callahan _et al._**: [DADA2 tutorial](https://benjjneb.github.io/dada2/tutorial.html), used under [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/), on which `dada2.jl` and its modules are based. +- +-The following colleagues at the **Department of Parasitology, Charles University** (Faculty of Science, BIOCEV, Vestec, Czech Republic) contributed to this work: +- +-- **Mgr. Jiří Novák** (supervisor): scripts from which several modules and configurations were adapted. +-- **doc. Mgr. Vladimír Hampl**: provided laboratory access and resources. +-- **Mgr. Paulína Pristašová**: <3. +- +-## Licence +- +-Copyright © 2026 Joshua Benjamin Jewell. +- +-Source code is licensed under the [GNU Affero General Public License v3.0](LICENSE). +- +-This documentation (README.md) is licensed under [CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/). +- +-File-level identifier annotations and the fork/upstream licence split are +-summarised in [`NOTICE`](NOTICE); canonical texts live in +-[`LICENSES/`](LICENSES/). Security reporting: [`SECURITY.md`](SECURITY.md). +diff --git a/docs/integration/autolink-references.md b/docs/integration/autolink-references.md +new file mode 100644 +index 0000000..4ad15b0 +--- /dev/null ++++ b/docs/integration/autolink-references.md +@@ -0,0 +1,121 @@ ++ ++ ++# Autolink references — complete specification for repository settings ++ ++**Status: specified 2026-09-26; pending application in Settings → Autolink references.** ++The Arena GitHub App token carries contents/issues permissions but not ++Administration, which the autolinks REST API (`GET/POST /repos/{owner}/{repo}/autolinks`) ++requires. This document is therefore the authoritative, paste-ready elaboration of ++the repository's autolink reference set. Applying it needs one human pass in the ++GitHub UI (or a token with `administration:write`, after which `gh api` can apply ++the table verbatim). ++ ++## What an autolink reference does ++ ++GitHub's **Autolink references** (Settings → Autolink references) let a bare token ++such as `SWARM-413` render as a link everywhere the tracker syntax `#413` cannot: ++commit messages, README/EXPLAINME/wiki prose, issue and PR bodies across forks, ++release notes, and any cross-repository mention. Each entry pairs an alphanumeric ++**prefix** with a **URL template** containing ``; GitHub appends the digits ++that follow the prefix to the template. ++ ++Native `#123` references (this repository's own issues and pull requests) already ++autolink and are deliberately **not** duplicated here. ++ ++## Design rules used below ++ ++1. **Every prefix names the tracker it resolves to** — no generic `GH-` or `REF-` ++ tokens. A reader of `VSEARCH-2203` knows which project's issue to open. ++2. **One tracker per prefix, issues-number space** — GitHub issue and PR numbers ++ share a number space, so `/issues/` redirects to the pull request when the ++ number is a PR. One entry covers both. ++3. **Coverage follows real references made by this repository** (from `README`, ++ `docs/`, `NOTICE`, `CONTRIBUTING`, commit history): the upstream fork, the ++ hyperpolymath estate, the seven orchestrated bioinformatics tools, the ++ runtime/toolchain dependencies, and the planned publication registry. ++4. **Uppercase prefixes** by convention, matching the `JIRA-123` idiom GitHub ++ documents. ++ ++## The reference set (paste-ready) ++ ++Apply in Settings → Autolink references, in this order. The UI takes two fields ++per row: **Link prefix** and **URL**. ++ ++### Tier 1 — lineage and estate (apply first) ++ ++| Link prefix | URL | Elaboration | ++|---|---|---| ++| `JJ-` | `https://github.com/JoshuaJewell/MetaManifold-WebUI/issues/` | Upstream (origin) repository — issues *and* pull requests. The owner-review correspondence ("PR #7, #8, #10, #11") means upstream PR numbers in particular. | ++| `MMW-` | `https://github.com/hyperpolymath/MetaManifold-WebUI/issues/` | This fork, explicit form. Useful in the wiki and in cross-repo contexts (standards, upstream, docs mirrors) where a bare `#n` resolves against the wrong tracker. | ++| `STD-` | `https://github.com/hyperpolymath/standards/issues/` | Organization-wide standards and specifications (README/EXPLAINME standard, A2ML, RSR canon, licence policy). Referenced by `docs/compliance/*`. | ++| `RSR-` | `https://github.com/hyperpolymath/rsr-template-repo/issues/` | The Rhodium Standard Repository template this repo is aligned against (`docs/compliance/rsr-alignment.md`). | ++| `STAT-` | `https://github.com/hyperpolymath/statistikles/issues/` | The estate's neurosymbolic statistics assistant — the intended consumer-side companion of this repo's method catalogue. | ++| `LITHO-` | `https://github.com/hyperpolymath/Lithoglyph/issues/` | LithoglyphDB — named in `ROADMAP.md` as owner of storage/journal/provenance internals. | ++| `GNPL-` | `https://github.com/hyperpolymath/GNPL/issues/` | GNPL (Lithoglyph's narration/projection language) — named in `ROADMAP.md` as provenance-adjacent owner. | ++| `BERRY-` | `https://github.com/metadatastician/berrywiki/issues/` | BerryWiki — the GitHub-wiki page format this repository's wiki is authored in (see `docs/wikis/`). | ++ ++### Tier 2 — orchestrated and acknowledged upstream tools ++ ++These are the tools the pipeline shells out to (and the two lineage pipelines ++credited in `NOTICE`/Acknowledgements). Bug triage in this repo routinely lands ++in an upstream tracker; autolinks make that one gesture. ++ ++| Link prefix | URL | Elaboration | ++|---|---|---| ++| `DADA2-` | `https://github.com/benjjneb/dada2/issues/` | ASV inference engine (R/Bioconductor); the tutorial lineage of the base design. | ++| `SWARM-` | `https://github.com/frederic-mahe/swarm/issues/` | OTU clustering; also the home of Fred's metabarcoding pipeline (layer-1 lineage). | ++| `VSEARCH-` | `https://github.com/torognes/vsearch/issues/` | Pair merging, chimera checks, taxonomy-by-alignment for both lanes. | ++| `CUTADAPT-` | `https://github.com/marcelm/cutadapt/issues/` | Primer/adapter trimming, stage one of every run. | ++| `FASTQC-` | `https://github.com/s-andrews/FastQC/issues/` | Raw-read QC (pre-filter lane). | ++| `MULTIQC-` | `https://github.com/MultiQC/MultiQC/issues/` | QC aggregation report embedded in the run view. | ++| `CDHIT-` | `https://github.com/weizhongli/cdhit/issues/` | cd-hit-est optional ASV clustering (multiplex inflation control). | ++| `MOTHUR-` | `https://github.com/mothur/mothur/issues/` | Source of the bundled MiSeq SOP sample data (`data/MiSeq_SOP/`). | ++ ++### Tier 3 — runtime, frontend and toolchain dependencies ++ ++| Link prefix | URL | Elaboration | ++|---|---|---| ++| `JULIA-` | `https://github.com/JuliaLang/julia/issues/` | Orchestrator language (pinned 1.12.5 via `mise.toml`). | ++| `RENV-` | `https://github.com/rstudio/renv/issues/` | R environment pinning (`renv.lock`); the reproducibility story for the R half. | ++| `BUN-` | `https://github.com/oven-sh/bun/issues/` | Frontend toolchain (pinned in `.bun-version`). | ++| `VITE-` | `https://github.com/vitejs/vite/issues/` | Frontend build. | ++| `PLOTLY-` | `https://github.com/plotly/plotly.js/issues/` | All analysis charts are Plotly JSON; `plotly.js-dist-min` has no published types (see `docs/types/architecture.md`). | ++| `REACT-` | `https://github.com/facebook/react/issues/` | SPA framework (18.x). | ++| `DUCKDB-` | `https://github.com/duckdb/duckdb/issues/` | Per-run results store. | ++| `STIPPLE-` | `https://github.com/GenieFramework/Stipple.jl/issues/` | Target framework of the Julia-authored UI migration (`docs/migration/STATUS.md`). | ++| `GENIE-` | `https://github.com/GenieFramework/Genie.jl/issues/` | Stipple's HTTP substrate; the migration UI's server lane. | ++ ++### Tier 4 — publication and registries ++ ++| Link prefix | URL | Elaboration | ++|---|---|---| ++| `ZENODO-` | `https://zenodo.org/records/` | Zenodo record IDs are flat integers, so `ZENODO-14012345` resolves to the record (and its DOI landing page). Prepared for the DOI-minting work of issue #8. | ++ ++## Reserved and deliberately omitted ++ ++| Candidate | Why it is out (for now) | ++|---|---| ++| `CVE-…` / `OSV-…` | CVE ids are two-part (`CVE-2026-12345`); GitHub autolinks match a single numeric run after the prefix. A `CVE-` prefix would mis-render `CVE-2026-12345` as `CVE-2026` + junk. OSV ids are alphanumeric. Link security advisories explicitly instead. | ++| `ADR-…` | The standards repo's decision records have slug-bearing filenames (`ADR-003-workflow-pin-staleness-window.adoc`); a ``-only template cannot reconstruct the slug. Reference them as `standards:docs/decisions/…` paths. | ++| `PR2-…` | The PR2 reference database versions are release tags with dots, not integers. Cite the release URL (see `config/databases.yml`). | ++| `JIRA-`-style generic tokens | Rejected by design rule 1. | ++ ++## Verification checklist (after application) ++ ++- [ ] `JJ-1` resolves to `https://github.com/JoshuaJewell/MetaManifold-WebUI/issues/1`. ++- [ ] `MMW-16` resolves to this fork's issue #16 (TSS/CSS/RSS offsets). ++- [ ] `SWARM-413` resolves to the swarm tracker. ++- [ ] `ZENODO-3265123` resolves to `https://zenodo.org/records/3265123`. ++- [ ] A commit message containing `DADA2-1948` renders linked in the commit list. ++- [ ] The wiki page `Developers--REST-API` (in `docs/wikis/`) renders its `JULIA-…`/`DUCKDB-…` references linked. ++ ++## Maintenance ++ ++Adding a tracker = one table row here + one Settings entry. Keep the tiers: they ++are the review order. If Administration access is later granted to the estate's ++automation, this file is the source of truth an `apply-autolinks` script should ++read — the table above is deliberately machine-parseable (three columns, no ++spans). +diff --git a/docs/wikis/Deep-Dives--Advanced-Functionality.md b/docs/wikis/Deep-Dives--Advanced-Functionality.md +new file mode 100644 +index 0000000..ada5d60 +--- /dev/null ++++ b/docs/wikis/Deep-Dives--Advanced-Functionality.md +@@ -0,0 +1,152 @@ ++ ++ ++# Advanced functionality — the coming suite ++ ++**Status: COMING throughout (one item BLOCKED by design).** The advanced ++statistical and exploration suite, in its approved order, with the design ++reasoning and the risks each item carries. This is the "what is coming" ++half of the honesty contract — written down in depth so nobody has to guess. ++Specifications of record: `docs/issues/milestone3/` (six numbered specs), ++`docs/milestones/02-deferred-issues.md`; issues #3, #5–8, #17–21. ++ ++## The gate that governs all of it ++ ++Every method below enters through the same door as the current layer: ++**publish its supported/unsupported conditions first, implement second.** ++The conditions template (already in force) demands: response types and zero ++handling; study-design support (pairing, blocking, repeated measures, ++covariates, depth, compositionality); overdispersion and depth policy ++named, not defaulted; uncertainty (interval method, effect size, BH — ++mandatory); diagnostics including "the assumption most likely to be wrong"; ++computational limits and the states at them. ++ ++The approved dependency order (the queue, honestly): ++ ++## 1. Exact statistical tests — issue #3 (the flagship) ++ ++*Why:* asymptotics fail at small n and sparse features — rare biosphere, ++low-biomass, clinical cohorts. The motivating case from the issue: pathogen ++detection where 2/3 cases vs 0/20 controls *should* be significant. ++ ++*Scope:* Fisher's exact (2×2 presence/absence vs group); exact negative ++binomial test (the edgeR `exactTest` shape: exact tail mass under the NB ++model); permutation-based exact p-values for NB-GLM (parametric bootstrap); ++exact CLR/ILR inference by permuting Aitchison distances. AnalysisConfig ++methods `exact_fisher`, `exact_nb`, `permutation_nb`. ++ ++*Design notes and risks (why it is not trivial):* Fisher on large tables ++needs the network algorithm (O(n!) is not a strategy); permutation storage ++(B=10⁴ resamples × 10⁴ taxa) is ~GB — resample-wise batching required; ++**exact is not assumption-free** — exchangeability under the null is still ++required and the context help must say so (exact tests can be uselessly ++conservative); edgeR/BiocParallel enter `renv.lock` with all the fragility ++that implies. Acceptance: p-values match R's `fisher.test` on known tables; ++benchmarks show <10% regression on existing methods; estimated runtime shown ++before launch (it will sit behind the Advanced expander). ++ ++*Note the naming:* "exact test" means the tail mass is enumerated rather ++than asymptotically approximated — see ++[Exact Arithmetic](Deep-Dives--Exact-Arithmetic) for why that is still not ++"exact numbers". ++ ++## 2. Multinomial and Dirichlet-Multinomial — issue #17 ++ ++*Why:* bridges count models and compositional geometry — effect sizes that ++are log-contrasts, compositionally coherent, **no pseudocount**. The ++Songbird-shaped answer to "p-values on a simplex are awkward". ++ ++*Risks recorded in the spec:* high-dimensional non-convex optimisation (DM ++especially); 10–100× runtime; reference-taxon choice changes the story ++(unstable bases must surface); determinism is hard if any learned component ++sneaks in (the layer's determinism rule forbids that). ++ ++## 3. Occupancy models — issue #18 ++ ++*Why:* presence/absence with imperfect detection — distinguishing true ++absence from "not detected", the zero-inflation/hurdle family. Reduces false ++negatives for rare taxa. ++ ++*Risks:* non-identifiability when detection and occupancy share parameters ++(single-visit adaptation is genuinely controversial); ZINB vs NB ++overfitting — the diagnostics must make the choice visible, not default. ++ ++## 4. Constrained ordinations — issue #19 ++ ++*Why:* RDA/CCA/CAP/dbRDA with permutation tests and variance partitioning — ++beta diversity *explained by covariates*, the community-level complement to ++per-taxon models. ++ ++*Risks:* RDA with Bray–Curtis is a misuse magnet (distance-based form is ++the right tool — the help must say which); permutation p-values are only as ++good as the exchangeability structure; 999-permutation cost at scale; ++biplot UI is its own project. ++ ++## 5. PhILR / SBP ILR bases — issue #20 ++ ++*Why:* biologically meaningful ILR balances — phylogenetic ILR (balances as ++clades), sequential binary partitions (hypothesis-driven), balance ++dendrograms for display. This is where the ilr's arbitrary basis becomes a ++scientific choice. ++ ++*Risks:* SBP flexibility is a p-hacking surface (pre-registration guidance ++belongs in the UI copy); phylogeny accuracy bounds the method; O(n²) memory ++(~800 MB at 10k taxa) — budgeted in the conditions. ++ ++## 6. Advanced zero handling and dispersion — issue #21 ++ ++*Why:* zeros distort transforms ([Compositional ++Statistics](Deep-Dives--Compositional-Statistics)); glmGamPoi gives faster, ++stabler dispersion for big count matrices; Bayesian multiplicative ++replacement propagates zero-replacement uncertainty instead of laundering ++it through a pseudocount. ++ ++*Risks:* replacement choices are conclusions, not settings — the receipt ++must carry which was used; dependency weight (glmGamPoi, zCompositions). ++ ++## 7. The compositional grand family — issue #5 ++ ++ANCOM-BC, ALDEx2, Songbird-class methods — log-contrast inference with bias ++corrections. Landed behind and dependent on the above. ++ ++## The exploration layer — issues #6–8 ++ ++- **CladeCumulus (#6)** — cumulative cladistic explorer: tree with ++ cumulative frequencies, epistemic colour coding ++ ([Epistemic Status](Deep-Dives--Epistemic-Status)), cloud sizing by ++ residual count, drag-and-drop re-partitioning with live ++ `present_in_every_admissible_world` validation. Scaffold in place; the UI ++ rides its branch. ++- **Full Evidence Mode (#7)** — the epistemic editor and fibre visualiser ++ (the `avec_fibre` story made visible). ++- **Zenodo DOI minting (#8)** — automated citable release of a study's ++ artefacts; a recoverable publication-and-citations slice is in draft ++ (PR #74). ++ ++## The blocked one — the symbolic engine, issue #2 ++ ++Formula manipulation, contrast derivation, provenance-carrying algebra — ++**BLOCKED**, deliberately and formally: a symbolic layer over unvalidated ++numeric machinery would launder approximate results through exact-looking ++notation. The unblock condition is named (validation of the numeric layer on ++real biological data — issue #1's review is on that path). This is not ++neglect; it is the same refusal discipline the runtime enforces, applied to ++project management. When it comes, it is where identity types earn their ++keep — proof-carrying manipulation, not string rewriting ++([Type Theory Meets Statistics](Deep-Dives--Type-Theory-Meets-Statistics)). ++ ++## What all of this does to the docs ++ ++Each landing updates, in order: its method-conditions document (first), ++the catalogue, [Analysis and Statistics Today](Users--Analysis-and-Statistics-Today), ++[Status and Roadmap](Status-and-Roadmap), and the EXPLAINME gaps list. The ++README's "Planned" note shrinks as items cross to IN PLACE — and only then. +diff --git a/docs/wikis/Deep-Dives--Compositional-Statistics.md b/docs/wikis/Deep-Dives--Compositional-Statistics.md +new file mode 100644 +index 0000000..bdf9fdc +--- /dev/null ++++ b/docs/wikis/Deep-Dives--Compositional-Statistics.md +@@ -0,0 +1,122 @@ ++ ++ ++# Compositional statistics and offsets ++ ++**Status: IN PLACE (offsets); COMING (the compositional method suite).** ++Why sequencing depth is *modelled* here rather than normalised away, and ++what a genuinely compositional answer would add. Conditions of record: ++`docs/statistics/method-conditions/scaling-and-offsets.md`; code: ++`src/analysis/scaling.jl`. ++ ++## The problem in one paragraph ++ ++Sequencing produces counts with an arbitrary library size: the same ++underlying community sequenced deeper yields larger numbers everywhere. The ++classical reflex — divide by library size, analyse proportions — quietly ++changes the *type* of the response (counts → compositions) and destroys the ++mean–variance relationship count models depend on. Worse, compositions live ++on a simplex where "more of A" forces "less of B" — an artefact of closure, ++not biology. Everything in this area is about which of these distortions you ++are willing to pay for. ++ ++## The vocabulary the code enforces ++ ++| Term | Shape | What it touches | May it change the response? | ++|---|---|---|---| ++| **Scaling factor** | one positive number per sample | nothing yet — it is a number | n/a | ++| **Offset** | log of the factor (after a stated centring) | the linear predictor of a count model | **no** — counts stay counts | ++| **Transform** | new table (CLR, ILR, relative) | the response itself | yes — and then a *different model family* applies | ++ ++The type distinction is the safety property: an `Offset` cannot be used ++where a `Transform` is required or vice versa, and `nb_glm` consumes counts +++ offsets while `clr_lm`/`ilr_lm` consume transforms. This is the "no silent ++substitution" rule at the type level. ++ ++## The three offsets: TSS, CSS, RSS ++ ++- **TSS (total sum scaling) factors** — library size (its log, centred). ++ The honest baseline: model depth as exposure. (The name is historically ++ overloaded with "convert to relative abundance" — which is the *transform* ++ of the same name. Here TSS produces **factors**, not proportions. That ++ disambiguation is half the point of the conditions document.) ++- **CSS (cumulative sum scaling) factors** — the `metagenomeSeq` idea: ++ normalise by the cumulative sum up to a percentile of the count ++ distribution, robust to a heavy tail of features. Implemented locally; ++ treat package-equivalence as unreviewed. ++- **RSS (relative sum scaling) factors** — sum-scaling relative to a ++ reference sum. Distinct from TSS factors in centring; again a factor, not ++ a transform. ++ ++These are **not** "normalisation" in the chart-facing sense — the ++`analysis.normalisation` config key (none/rarefaction/relative for display) ++is a separate concern ([Configuration Reference](Users--Configuration-Reference)). ++ ++## What this replaced — the three silent substitutions ++ ++The conditions document opens with the historical failures, because they are ++the justification for the whole design: ++ ++1. `TSS` divided counts by library size and returned **proportions** — the ++ exact compositional transform the design was written to avoid — for every ++ method except `nb_glm`. ++2. `CSS` and `RSS` were **aliased to `relative` outright** (with a `@warn`): ++ under `nb_glm` a proportion table met a count model. ++3. `size_factors` was *named* "DESeq2 median-of-ratios" and *implemented* as ++ library size over its own geometric mean — TSS wearing another method's ++ name. ++ ++Each is now refused at the boundary; regressions are test-locked; method ++names compare case-insensitively (#62) so `tss`/`TSS` cannot redden a ++compliant run (#46's echo of the capital-letter incident). The standing ++rule when document and code disagree: **the document is right and the file ++is the bug.** ++ ++## What offsets deliberately do not solve ++ ++Offsets handle *depth*. They do not handle *compositionality* — the simplex ++constraint remains in the data whether or not the model sees exposure. For ++differential abundance that respects the geometry, you need the compositional ++suite, which is **COMING**: ++ ++- **CLR/ILR + linear models** — **PARTIAL/IN PLACE** as `clr_lm`/`ilr_lm` ++ (transforms + Gaussian models, with the zero-replacement question ++ currently handled by the configured zero policy; improvements below). ++- **Multinomial / Dirichlet-Multinomial regression** (#17) — effect sizes ++ living on the simplex; no pseudocount; Songbird-like. Hard: high-dimensional ++ non-convex optimisation, reference-taxon choice, determinism. ++- **ANCOM-BC, ALDEx2, Songbird-class methods** (#5) — log-contrast families ++ with their own bias corrections. ++- **PhILR / SBP ILR bases** (#20) — balances as clades (phylogenetic ILR), ++ hypothesis-driven sequential binary partitions, balance dendrograms; the ++ *basis* of the ilr is where biology enters. ++- **Advanced zero handling** (#21) — zeros are the compositionalist's ++ nightmare: sampling zeros vs structural zeros. Bayesian multiplicative ++ replacement and glmGamPoi dispersion are queued ahead of trusting ILR ++ fully. ++- **Occupancy models** (#18) — "absent" vs "undetected" is a latent-variable ++ question at the count/zero boundary. ++ ++Until those land, the honest position (which the software enforces) is: ++report offsets-based `nb_glm` effects as *depth-modelled count effects*, and ++`clr_lm`/`ilr_lm` effects as *log-contrast effects under the stated zero ++policy* — and never call either "compositional differential abundance ++analysis". ++ ++## Reading list (for the statistically curious) ++ ++Aitchison's log-contrast geometry is the ground (compositions as log-ratios); ++Gloor et al. on why relative abundance misleads; the `metagenomeSeq` CSS ++paper for the cumulative-sum idea; Nearing et. al. (the Songbird/DIAMOND ++line) for multinomial regression as the compositional alternative to ++p-value tables. The repository's method-conditions documents cite what each ++implementation actually follows. +diff --git a/docs/wikis/Deep-Dives--Design-Progression.md b/docs/wikis/Deep-Dives--Design-Progression.md +new file mode 100644 +index 0000000..ee86ca8 +--- /dev/null ++++ b/docs/wikis/Deep-Dives--Design-Progression.md +@@ -0,0 +1,119 @@ ++ ++ ++# Design progression ++ ++**The three-layer lineage in full**, with the receipts. The README shows the ++same progression in brief; this page is the extended cut. ++ ++## Layer 1 — the base design: R and Python around DADA2 ++ ++Two established open-source traditions, used *raw*, before any MetaManifold ++code exists: ++ ++**The DADA2 tutorial lineage (R).** Callahan et al.'s exact-sample variant ++pipeline as it is run from scripts: `filterAndTrim` (quality trim + trunc), ++`learnErrors` (error rate learning), `dada` (the OME-picking-free denoiser), ++`mergePairs`, `makeSequenceTable`, `removeBimeraDenovo`, `assignTaxonomy` ++against a reference FASTA. In this repository that R lineage still exists, ++almost verbatim, as `src/pipeline/dada2/dada2_functions.r` — the Julia side ++wraps it, it does not reimplement it. ++ ++**Fred's metabarcoding pipeline lineage (shell/Python + swarm/vsearch).** ++Mahé's swarm pipeline: cutadapt primer trimming, vsearch pair merging, ++dereplication, `swarm -d 1` single-linkage clustering, `--uchime_denovo` ++chimera culling, vsearch global-alignment taxonomy, then table merge and ++taxonomic filtering. Again: `src/pipeline/swarm.jl` and the vsearch stage ++are orchestrators around the same tools. ++ ++**Acknowledgements are load-bearing here.** The DADA2 tutorial is CC BY 4.0; ++Fred's pipeline is credited in `NOTICE`. Layer 1 is *upstream science* — the ++fork's own standing rule is that it tracks application changes and does not ++fork the science. ++ ++What layer 1 lacks: shared provenance (tables stitched by hand), single-lane ++operation (you chose ASVs *or* OTUs per script), and any contract about what ++the numbers claimed. ++ ++## Layer 2 — JoshuaJewell's MetaManifold augmentation ++ ++The origin design (Joshua Benjamin Jewell, Department of Parasitology, ++Charles University) makes three moves: ++ ++1. **Orchestration.** A Julia engine runs *both* lanes per run — ASV and OTU ++ from one launch — with cutadapt in front, cd-hit-est optional for ++ multiplex inflation, `merge_taxa` behind. Each stage is a typed result ++ (`TrimmedReads` → `ASVResult`/`OTUResult` → `TaxonomyHits` → ++ `MergedTables`) with mtime/hash freshness so re-runs are minimal. ++2. **State.** Per-run DuckDB (`merged/results.duckdb`) as the queryable ++ results store; `projects/{study}/{run}/` as the complete output tree; ++ `run_config.yml` as merged-configuration provenance. ++3. **A workbench.** Oxygen.jl REST + SSE and a React SPA: the config cascade ++ editable at every level (instance → study → group → run), run and job ++ control with live progress, embedded QC (FastQC/MultiQC + DADA2 ++ diagnostics), the results explorer (filters, presets, OTU drill-down, ++ xlsx), functional annotation (dual-classifier consensus, contamination ++ curation, FuncDB ledger), composition views (biological categories). ++ ++The pipeline diagram in the README's layer-2 section *is* this design. It is ++still the shipping architecture. ++ ++What layer 2 left open: statistics that were wired to placeholders; numbers ++whose precision claims were unchecked; a large frontend with no type estate; ++and no engineering harness to keep either honest. ++ ++## Layer 3 — the hyperpolymath steps (2026-09) ++ ++A discipline of honesty and typing wrapped around layer 2, delivered as a ++patch series (Milestone 2 → AnalysisConfig v1 → the statistics layer): ++ ++``` ++ discipline what it forbids where ++ ───────────────────── ───────────────────────────────────── ───────────────────────── ++ typed estate `any`, `@ts-ignore`, ambient sprawl frontend/src/types, tsconfig ++ engineering gates untestable claims, unpinned tools test/, bench/, ci, Justfile ++ numeric policy precision-laundering (float as exact) numeric_policy.jl ++ exact summaries inference dressed as description exact_summaries.jl ++ real estimation placeholder-as-result (hash-based p) estimation.jl + guard test ++ offsets silent substitution (TSS→relative …) scaling.jl ++ epistemic receipts bare numbers without warrants epistemic.jl ++ refusal-first guessing when a method cannot run everywhere above ++``` ++ ++The progression is not "rewrite in another stack" — each step *tightens a ++claim* the earlier layers were making loosely. Layer 2 said "here is a ++p-value"; layer 3 asks "what type of claim is that p-value, and what ++warrants it?" — and where the answer is "none", the software now returns a ++named refusal instead. ++ ++**In place vs coming (layer 3):** the table above is shipped (guard tests ++included). The forward path — exact tests, the compositional/occupancy/ ++ordination suite, PhILR/SBP, advanced zeros, CladeCumulus, Evidence Mode, ++Zenodo, and the formally blocked symbolic engine — is ++[Advanced Functionality](Deep-Dives--Advanced-Functionality) and the ++[Status and Roadmap](Status-and-Roadmap) board. ++ ++## The nested picture ++ ++``` ++┌─ Layer 3 · verified statistics, typed estate, engineering gates ─────────────┐ ++│ ┌─ Layer 2 · Julia orchestrator + WebUI: both lanes, DuckDB, config cascade ┐│ ++│ │ ┌─ Layer 1 · base design: raw DADA2 (R) + swarm/vsearch (shell) pipelines ┐│ ++│ │ │ raw FASTQs → denoise or cluster → taxonomy → count tables ││ ++│ │ └─────────────────────────────────────────────────────────────────────────┘│ ++│ └───────────────────────────────────────────────────────────────────────────┘│ ++└───────────────────────────────────────────────────────────────────────────────┘ ++``` ++ ++Each layer's code is still legible in the tree — that is what "stacked ++honestly" means: you can read layer 1's R, layer 2's orchestrator, and ++layer 3's contracts without a archaeology dig. +diff --git a/docs/wikis/Deep-Dives--Epistemic-Status.md b/docs/wikis/Deep-Dives--Epistemic-Status.md +new file mode 100644 +index 0000000..f9f39fe +--- /dev/null ++++ b/docs/wikis/Deep-Dives--Epistemic-Status.md +@@ -0,0 +1,91 @@ ++ ++ ++# Epistemic status and receipts ++ ++**Status: PARTIAL — the epistemic core is IN PLACE (`src/core/epistemic.jl`); ++the Evidence Mode UI is COMING (#7).** What "epistemic receipts" are, the ++type theory they shadow, and why a warrant is not a proof. ++ ++## The problem this layer addresses ++ ++A pipeline that outputs numbers asks its users to believe things: that the ++fit converged, that the method's preconditions held, that the number is not ++a stub. Most pipelines leave those beliefs implicit — they hand over bare ++numbers and let the README do the warranting. When the README is wrong (or ++the reader does not read it), nothing in the output says so. ++ ++The epistemic layer makes the warrant **travel with the result**. Every ++carried value can wear its status — how it is admissible — as a first-class ++part of its representation. ++ ++## The Agda lineages being shadowed ++ ++`src/core/epistemic.jl` states its provenance openly: finite, executable ++shadows of three formal type families (developed in the estate's Agda work — ++echo-types, epistemic-types, residual-evidence-types). The mapping: ++ ++| Formal notion (Agda) | Shape | Julia shadow | Statistical reading | ++|---|---|---|---| ++| **Echo** | `Echo f y := Σ (x:A), (f x ≡ y)` — a value plus evidence of how it arose (total space of a fibration; `Σ B (Echo f) ≃ A`) | `EchoFiber`, `avec_fibre` / `sans_fibre` | a result with its derivation (method, policy, trace) — vs a bare result | ++| **Identity type** | `f x ≡ y` — evidence of sameness, inspectable | equality evidence in validation | "this output matches the config that promised it" (freshness hashes, `run_config.yml`) | ++| **Modality `E κ A`** / **FactiveModality** | framed belief; `reflect` | `Modality`, `FactiveModality` | a result admitted under a stated frame (the model, the policy mode) | ++| **Warrant (without soundness)** | grounds for belief, *no* soundness proof attached | `Warrant`, `SoundWarrant` (the latter opt-in) | convergence checks, BH decisions, bootstrap floors — reasons to believe, not truth | ++| **Residual-evidence triad** | `Candidate`, `Holds`, `Identified` (actual-world-sound) | `Candidate`, `Case`, `Holds`, `Identified` | candidate explanation → survives checks → identified in the actual world | ++| **Admissible worlds** | `present_in_every_admissible_world` (and the `some` / `absent` variants) | same names, exported | robustness across genuinely open analysis choices | ++ ++**Warrants without soundness** is the load-bearing choice. A convergence ++check is not a proof that the estimate is *true*; it is grounds for admitting ++the number at all. The formalism refuses to smuggle soundness in — the ++statistics agrees (an ML estimate under a wrong model is still wrong), and ++so does the user-facing copy ("computed as documented, not reviewed"). ++ ++## The wire and the UI ++ ++- **`avec_fibre` column** ("with fibre") — the Echo idea on the wire: a ++ displayed result can carry the fibre of its derivation. `sans_fibre` ++ exists for the bare form, so the distinction is representable both ways. ++- **Status vocabulary** (`EPISTEMIC_STATUSES`) — includes ++ `present_in_every_admissible_world`-family claims; CladeCumulus (#6) will ++ colour cladistic clouds by exactly these. ++- **The DANGER banner** — the loud edge of the same discipline: a ++ configuration that would disable BH or set an out-of-policy advanced flag ++ is a *visible type* (banner + structured log), not a silent option. ++- **Refusals** ([Maximum Likelihood](Deep-Dives--Maximum-Likelihood)) are ++ the degenerate receipt: the warrant failed, and the system says which one. ++ ++## Validation semantics ++ ++`present_in_every_admissible_world` (exported and used in validation) is a ++modal quantification: a claim is admitted only if it survives every ++admissible reading of the data's genuine ambiguities (zero policy, ++normalisation story, grouping). Its siblings — ++`present_in_some_admissible_world`, `absent_in_every_admissible_world` — ++give the lattice that CladeCumulus and Evidence Mode will render. This is ++Kripke-flavoured: truth relative to worlds, with the "actual world" entering ++through the residual-evidence `Holds → Identified` path. ++ ++## What is here and what is coming ++ ++**IN PLACE:** the module (types, statuses, the `avec_fibre` column, the ++validation predicates); the DANGER banner; unsuccessful states as receipts ++across the analysis layer; `clade_cumulus.jl` data structures and live ++`present_in_every_admissible_world` validation logic (scaffold). ++ ++**COMING:** Full Evidence Mode (#7) — the epistemic editor and fibre ++visualiser; CladeCumulus's UI (#6) — cumulative frequencies over a cladistic ++tree, epistemic colour coding, cloud sizing by residual count, ++drag-and-drop re-partitioning with live admissible-world validation. ++ ++**Not claimed:** soundness. Nowhere does the system claim a warrant is true. ++`SoundWarrant` exists as a distinct type precisely so that soundness is an ++extra, stated property — never a default. +diff --git a/docs/wikis/Deep-Dives--Exact-Arithmetic.md b/docs/wikis/Deep-Dives--Exact-Arithmetic.md +new file mode 100644 +index 0000000..3543f40 +--- /dev/null ++++ b/docs/wikis/Deep-Dives--Exact-Arithmetic.md +@@ -0,0 +1,111 @@ ++ ++ ++# Exact arithmetic ++ ++**Status: IN PLACE** (descriptive layer). Where exactness lives, what it ++buys, and the line it must not cross. Contracts: the repository's ++`docs/statistics/numeric-contracts.md`; code: `src/analysis/numeric_policy.jl`, ++`src/analysis/exact_summaries.jl`. ++ ++## The three kinds of number ++ ++Most statistics are not exact, and no policy makes them so. What a policy ++can do is stop three different claims from wearing one representation: ++ ++| Kind | What it means | May be called a fact? | ++|---|---|---| ++| **exact** | an integer count, or a rational built from counts; no rounding in its history | yes — as a count or proportion | ++| **approximate** | any floating value, **including arbitrary precision** | no — it is a number | ++| **rounded** | an approximation cut to N digits for display | no — it is a rendering | ++ ++The sentence the whole layer turns on: **higher precision is not exactness.** ++`BigFloat` at 4096 bits rounds; it merely rounds further away. The only ++exact arithmetic here is integer and rational. ++ ++## The policy as a type ++ ++`numeric_policy(; mode, precision_bits, max_denominator_bits, round_digits)` ++builds an immutable `NumericPolicySpec`, validated at construction: ++ ++| Field | Modes / limits | Default | ++|---|---|---| ++| `mode` | `:ordinary`, `:exact_counts`, `:high_precision` | `:ordinary` | ++| `precision_bits` | 2 … 1 000 000 | 53 (Float64's significand) | ++| `max_denominator_bits` | ≥ 32 | 4096 | ++| `round_digits` | ≥ 0 | 6 | ++ ++`:ordinary` is Float64 throughout and is the only mode under which ++previously saved analyses are unchanged (the backward-compatibility line). ++The load-bearing design: **nothing switches mode on its own.** A caller that ++needs `:exact_counts` and is handed `:ordinary` *fails through `assert_mode`* ++— it does not receive a Float64 that looks like the exact answer. That is ++the "no coercion `Approximate → Exact`" rule from ++[Type Theory Meets Statistics](Deep-Dives--Type-Theory-Meets-Statistics) ++made executable. ++ ++## What exactness buys in practice ++ ++`exact_summaries.jl` (catalogue item 1) computes counts and proportions at ++exact precision: ++ ++- **Counts as integers without a ceiling.** Read counts beyond 2⁵³−1 — where ++ every float silently loses integers — stay exact. (This was not ++ hypothetical: the boundary audit in issue #52 *measured* where Float64 ++ stops carrying consecutive integers.) ++- **Proportions as rationals of counts.** 2/3 stays 2/3. A proportion ++ supplied as text (`"2/3"`, `"4/6"`) is exact; one supplied as a float is ++ accepted **only as an approximation and labelled as one** wherever shown. ++- **Refusals as facts.** A float claiming exactness is refused; a negative ++ count is refused; a total-zero proportion is refused with a name, not ++ rendered as 0. ++ ++Independent reference: the suite compares value-by-value against Python's ++`fractions.Fraction` (skipping loudly by name if `python3` is absent) and ++plants hand-derived known answers. ++ ++## Two defects the conditions document caught ++ ++Worth recording because they show why conditions precede implementation: ++ ++1. `to_display` printed a rendering labelled *6dp* — and then printed eighty ++ digits after the point. That is the rounded/exact boundary leaking; the ++ document ruled, the code changed. ++2. Exact rationals rendered as Julia's `2//3` — implementation syntax leaking ++ into a human's display. Again: document right, file wrong. ++ ++## The line exactness must not cross ++ ++**No inference is exact just because its inputs are.** The descriptive layer ++claims facts; a p-value, an interval, or an MLE is approximate by ++construction (a tail mass or an optimiser output), and the policy does not ++pretend otherwise. What does *not* exist yet, and is easy to overread into ++this layer: ++ ++- **Exact statistical tests** (Fisher's exact, exact NB, permutation ++ PERMANOVA) — issue #3, **COMING**. "Exact tests" there will mean the tail ++ mass is enumerated/permuted rather than asymptotically approximated — a ++ claim about the *reference distribution*, still living in the ++ approximate/reporting column for display purposes. The naming collision is ++ deliberate and worth meditating on: "exact test" ≠ "exact number". ++- Small-n validity: today's honest default for tiny samples is the exact ++ descriptive summary plus "no valid inferential test computed" — the ++ catalogue endorses exactly this default. ++ ++## Why rationals and not decimals ++ ++Decimals (fixed-point, arbitrary or not) are exact only for dyadic-friendly ++fractions and become a new rounding story otherwise; rationals of integers ++are the free field over ℤ and carry no rounding history at all. The ++denominator budget (`max_denominator_bits`) exists so that accumulated ++products cannot grow unbounded — at the budget the policy *refuses or ++degrades explicitly*, never silently simplifies. +diff --git a/docs/wikis/Deep-Dives--Maximum-Likelihood.md b/docs/wikis/Deep-Dives--Maximum-Likelihood.md +new file mode 100644 +index 0000000..7f3d91e +--- /dev/null ++++ b/docs/wikis/Deep-Dives--Maximum-Likelihood.md +@@ -0,0 +1,113 @@ ++ ++ ++# Maximum likelihood and refusals ++ ++**Status: PARTIAL — implemented and reference-tested; independent review ++(#1) outstanding.** How estimation works here, why the return type is a sum, ++and what each unsuccessful state means. Conditions of record: ++`docs/statistics/method-conditions/parametric-fits.md`; code: ++`src/analysis/estimation.jl`. ++ ++## What a fit is, here ++ ++Per-feature estimation of a **declared** model — one row per feature, the ++design named by the configuration's formula. Four methods, and for each one ++exactly one response type: ++ ++| Method | Response | Model | Fit in | ++|---|---|---|---| ++| `nb_glm` | raw counts (non-negative) | negative binomial GLM, log link | R `MASS::glm.nb` | ++| `clr_lm` | CLR-transformed values | Gaussian linear model | R `stats::lm` | ++| `ilr_lm` | ILR balances | Gaussian linear model | R `stats::lm` | ++| `logistic` | 0/1 presence | binomial GLM, logit link | R `stats::glm(family=binomial)` | ++ ++Two structural rules: ++ ++- **No response mixing.** Counts go to `nb_glm`; transforms go to their ++ linear models; presence to logistic. A proportion table where a count ++ model expects counts is refused at the boundary (the historical ++ alias-and-warn path is gone). ++- **Determinism.** No resampling, no permutation, no random start. The ++ provenance records the caller's seed *as "not a parameter of any number ++ here"* — an honesty line that exists because seeds are easy to imply and ++ hard to disclaim later. ++ ++Maximum likelihood is the estimator because the models justify it: for an ++NB-GLM with log link the likelihood is (given dispersion) a GLM likelihood ++solved by IRLS (`MASS::glm.nb` jointly estimates dispersion by profile ++likelihood); for the Gaussian and binomial cases the same machinery is exact ++for the canonical links. What ML *buys* is the classical asymptotic ++apparatus — standard errors, z/Wald tests, intervals — and what it *costs* ++is that every one of those is an asymptotic claim that must survive its ++preconditions. Hence the next section. ++ ++## The return type is a sum, not a maybe ++ ++A fit returns **estimates or a named unsuccessful state** — never a number ++dressed for the occasion. The states, and what each means: ++ ++| State | Meaning | Typical cause | ++|---|---|---| ++| **Non-convergence** | IRLS/profile iteration hit its cap without settling | separation, extreme overdispersion, degenerate design | ++| **Non-identifiable** | the design matrix does not determine the parameters | collinear covariates, empty cells in a factor | ++| **Boundary estimate** | a parameter sits at a limit (e.g. fitted dispersion) | true boundary or model misspecification — reported either way | ++| **Precondition failure** | the declared response/design contract is violated | wrong response type, negative counts, unmodelled pairing | ++| **Resource limit** | time/memory bounds exceeded | huge feature counts × expensive fits — bounded by contract | ++| **Not implemented** | the method or its R package is absent from `renv.lock` | deliberately uninstalled dependency — refused, never faked | ++ ++The states are *results*. Users see them in tables where a p-value would be; ++the UI renders them as such; the paper-language translation is "not tested" ++(see [For Academics](Users--For-Academics)). ++ ++## What the tests hold the code to ++ ++The suite's idiom is the reason to trust any of this (all in ++`test/unit/test_estimation.jl`): ++ ++1. **Known answers written into the data** — fixtures encode a planted ++ effect (counts engineered to give a known coefficient direction and rough ++ magnitude); the test checks the fit against the planting, not a snapshot. ++2. **Independent reference** — the same data fitted directly in R outside ++ the pipeline, compared coefficient by coefficient, and BH compared ++ against `p.adjust`. The pipeline and the reference share R's ++ implementations, so this checks the *wiring* (design matrix, offsets, ++ response construction), which is where pipeline bugs actually live. ++3. **Negative controls for every refusal path** — each state in the table ++ above has a test that provokes it. ++4. **The placeholder guard** — see below. ++ ++## The history this replaced (why the guard exists) ++ ++Before 2026-09-25, `run_analysis` returned per-feature "statistics" ++computed as `0.01 + (hash(taxon_id) % 100) / 1000.0` — deterministic ++nonsense derived from each taxon's *name*, labelled as results. The lesson ++the conditions document states flatly: **a stub that is returned as a result ++is a wrong answer.** The placeholder is deleted and a source-level guard ++test now fails if anything of that shape returns. If you are tempted to stub ++a statistic to unblock a UI: return an unsuccessful state instead. The UI ++already knows how. ++ ++## Honest limits (the #1 caution, expanded) ++ ++- "Computed as documented" ≠ "appropriate for your data." The conditions ++ document says nothing about whether an NB-GLM suits your study design — ++ that remains the analyst's judgement (and the document names the ++ assumption most likely to be wrong for each method). ++- The asymptotics behind Wald intervals and z-tests degrade at small n and ++ sparse features — the regime the **exact tests** (#3, COMING) exist for. ++ Until they land, small-n claims should lean on ++ [Exact Arithmetic](Deep-Dives--Exact-Arithmetic)'s descriptive layer. ++- CSS (in `scaling.jl`) is a repo-local implementation of a published idea; ++ equivalence to `metagenomeSeq` is unreviewed — same review umbrella. ++- Multiple testing is BH everywhere several tests are reported. This is ++ mandatory, enforced, and deliberately not configurable downward. +diff --git a/docs/wikis/Deep-Dives--Type-Theory-Meets-Statistics.md b/docs/wikis/Deep-Dives--Type-Theory-Meets-Statistics.md +new file mode 100644 +index 0000000..b081e2f +--- /dev/null ++++ b/docs/wikis/Deep-Dives--Type-Theory-Meets-Statistics.md +@@ -0,0 +1,118 @@ ++ ++ ++# Type theory meets statistics ++ ++**Status: IN PLACE as discipline and code shape.** This is the central ++argument of layer 3: most statistical misconduct in pipelines is a *type ++error* — a claim of one kind wearing the representation of another. The ++fix is to make kinds of claim into kinds of value. ++ ++## The identification ++ ++Start from three observations a statistician and a type theorist would state ++differently but agree on: ++ ++1. **A p-value and a count are not the same kind of thing.** A count is a ++ fact about a sample (finite, exactable). A p-value is the tail mass of a ++ reference distribution *under a model* — approximate by construction, ++ meaningful only where the model's preconditions hold. ++2. **"The fit did not converge" is a result.** It is information about the ++ data/design combination. Pipelines that render it as `NA`, `0`, or a ++ plausible-looking estimate have converted information into misinformation ++ — a coercion between kinds that should not exist. ++3. **Precision is not exactness.** A 4096-bit float is still an ++ approximation of a real number. Treating `BigFloat` output as "exact" is ++ a category error, not a rounding improvement. ++ ++In type-theoretic terms: these are **distinct types of claim**, and the ++representations must not be silently coercible. The numeric policy ++([Exact Arithmetic](Deep-Dives--Exact-Arithmetic)) makes (3) a runtime ++discipline; the unsuccessful-state machinery ++([Maximum Likelihood](Deep-Dives--Maximum-Likelihood)) makes (2) a return ++type; the offset/transform distinction ++([Compositional Statistics](Deep-Dives--Compositional-Statistics)) makes ++(1)-adjacent confusions unrepresentable. ++ ++## The vocabulary, and where it comes from ++ ++The estate's epistemic layer is grounded in formal type theory — the Agda ++lineages that `src/core/epistemic.jl` shadows executably (see ++[Epistemic Status](Deep-Dives--Epistemic-Status) for the full mapping). ++The terms that matter for statistics: ++ ++- **Σ-types (dependent pairs).** `EchoFiber` — Σ(x:A), (f x ≡ y) — is a ++ value *together with evidence of how it arose*. A statistical result here ++ is shaped the same way: a number plus its computation trace (method, ++ policy mode, refusals avoided). The `avec_fibre` column ("with fibre") is ++ this idea on the wire: a result carries the fibre of its derivation. ++- **Identity types.** `f x ≡ y` is not boolean equality; it is *evidence ++ that two things are the same*, which can be inspected. "This table equals ++ what the config promised" is checked as evidence (freshness hashes, ++ `run_config.yml`), not trusted. ++- **Modalities and warrants (without soundness).** A `Warrant` is a ++ *reason to believe* — κ-modality framed — explicitly **without** a ++ soundness proof. This is exactly the epistemology of a p-value or a ++ convergence check: grounds for a claim, not the claim's truth. The ++ type theory refuses to let software pretend otherwise, and so does the ++ statistics: BH-adjusted p-values are warrants for ranking features, not ++ certificates of biological difference. ++- **Admissible worlds.** `present_in_every_admissible_world` quantifies over ++ candidate interpretations — a modal notion. Statistically this is ++ robustness: a feature that "holds in every admissible world" survives the ++ analysis choices that are genuinely open (zero policy, normalisation ++ story). CladeCumulus colours by it, deliberately. ++ ++## Where TypeScript carries the same load ++ ++The frontend estate is not where Σ-types live, but it enforces the same ++"no silent coercion" rule at the wire boundary: ++ ++- `unknown` + narrowing instead of `any` — a value from outside the type ++ boundary must *earn* its type, with evidence (runtime checks), exactly ++ like a warrant. ++- `exactOptionalPropertyTypes` — "absent" and "present-but-undefined" are ++ distinct, the same distinction that makes a refusal different from an ++ empty result. This one option choice is why #31's rule ("unknown, never ++ silently not-significant") is implementable in the UI without special ++ cases. ++- Category D/E closure — third-party types are audited like dependencies; ++ the two `FIXME(types)` stubs are *declared* coercions (documented debt), ++ not discovered ones. ++ ++## The payoff, concretely ++ ++| Statistical honesty rule | Type-theoretic shape | Enforced by | ++|---|---|---| ++| Never a placeholder as a result | no coercion `Stub → Estimate` | guard test (estimation suite) | ++| Refusal is a result | sum type `Estimate ⊎ UnsuccessfulState` | `estimation.jl` returns, UI renders states | ++| Floats never claim exactness | no coercion `Approximate → Exact` | `numeric_policy.jl` (`assert_mode`), boundary tests | ++| Offsets ≠ transforms | distinct types, no implicit map | `scaling.jl` + conditions doc | ++| Significance never silently degrades | `Unknown ⊎ NotSignificant ⊎ Significant` | analysis status plumbing | ++| BH cannot be quietly disabled | the dangerous configuration is a *loud* type (DANGER banner + log) | AnalysisConfig validation | ++ ++None of this requires a dependently-typed host language — Julia and ++TypeScript reach the same discipline by *convention made mechanical* ++(validation at boundaries, sum-typed returns, guard tests). The Agda ++lineages matter as the conceptual ground: the shadows in `epistemic.jl` are ++executable glossary entries, keeping the vocabulary honest. ++ ++## Where it goes next ++ ++The advanced suite ([Advanced Functionality](Deep-Dives--Advanced-Functionality)) ++pushes the same identification further: exact tests are "results whose tail ++mass is computed, not approximated"; multinomial regression is "effects ++that live on the simplex, typed as such"; occupancy models are "absence as ++a latent variable, not a zero". The symbolic engine (#2, blocked) is where ++the identity types become load-bearing — proof-carrying formula ++manipulation rather than string rewriting — which is precisely why it waits ++for the numeric layer to earn its review (#1). +diff --git a/docs/wikis/Deep-Dives.md b/docs/wikis/Deep-Dives.md +new file mode 100644 +index 0000000..518b37b +--- /dev/null ++++ b/docs/wikis/Deep-Dives.md +@@ -0,0 +1,56 @@ ++ ++ ++# Deep dives ++ ++This is where the mathematics and the engineering receipts live — the ++material that would otherwise make the README long and the EXPLAINME ++unreadable. The repository's ++[EXPLAINME.adoc](https://github.com/hyperpolymath/MetaManifold-WebUI/blob/main/EXPLAINME.adoc) ++maps README claims to code and points here for the reasoning; this section is ++the reasoning. ++ ++Written for readers comfortable with statistics at graduate level and ++curious about the type-theoretic discipline around it. Academic and lab ++readers alike: the user-facing contracts are ++[Analysis and Statistics Today](Users--Analysis-and-Statistics-Today); this ++section is *why those contracts are shaped that way*. ++ ++## The seven dives ++ ++| Dive | Question it answers | Status of its subject | ++|---|---|---| ++| [Design Progression](Deep-Dives--Design-Progression) | How did we get from raw R/Python scripts to this, and what did each layer add? | historical + IN PLACE | ++| [Type Theory Meets Statistics](Deep-Dives--Type-Theory-Meets-Statistics) | What do types have to do with honest statistics? | IN PLACE (the discipline) | ++| [Exact Arithmetic](Deep-Dives--Exact-Arithmetic) | When is a number exact, and what does that buy? | IN PLACE (descriptive layer) | ++| [Maximum Likelihood](Deep-Dives--Maximum-Likelihood) | How does estimation work here, and what are "unsuccessful states"? | PARTIAL (review pending) | ++| [Compositional Statistics](Deep-Dives--Compositional-Statistics) | Why offsets instead of normalising counts away? | IN PLACE (offsets); COMING (compositional methods) | ++| [Epistemic Status](Deep-Dives--Epistemic-Status) | What are "epistemic receipts", Σ-types, and warrants without soundness? | IN PLACE (core); COMING (Evidence Mode UI) | ++| [Advanced Functionality](Deep-Dives--Advanced-Functionality) | What is the exact/compositional/occupancy/ordination suite, and why is it late? | COMING / BLOCKED | ++ ++## The thread through all seven ++ ++A statistical result is only as good as the **type of claim** its numbers can ++carry. Most of the engineering here is making that notion executable: ++ ++- exact / approximate / rounded are *different types of claim* and must not ++ wear the same representation ([Exact Arithmetic](Deep-Dives--Exact-Arithmetic)); ++- a fit is either warranted (converged, identified, within bounds) or it is ++ an *unsuccessful state* — there is no third outcome called "a number" ++ ([Maximum Likelihood](Deep-Dives--Maximum-Likelihood)); ++- depth is a modelled quantity (an offset), not a dial to normalise away ++ ([Compositional Statistics](Deep-Dives--Compositional-Statistics)); ++- and every carried result travels with its *warrant* — the reason it is ++ admissible — without anyone claiming the warrant is the truth ++ ([Epistemic Status](Deep-Dives--Epistemic-Status)). ++ ++That thread is why the deep dives are one section and not seven essays. +diff --git a/docs/wikis/Developers--Architecture-Tour.md b/docs/wikis/Developers--Architecture-Tour.md +new file mode 100644 +index 0000000..f23ba7e +--- /dev/null ++++ b/docs/wikis/Developers--Architecture-Tour.md +@@ -0,0 +1,105 @@ ++ ++ ++# Architecture tour ++ ++**Status: IN PLACE.** A walk from an HTTP request to a count table and back ++to a chart. Companion: `EXPLAINME.adoc` (claim receipts) and ++`docs/types/architecture.md` (frontend type placement). ++ ++## The shape ++ ++``` ++frontend/ TypeScript + React + Vite (SPA) ++ src/api/ REST client + canonical wire types (src/api/types.ts) ++ src/types/ the domain type estate (API boundary, plotly vocabulary) ++ src/views/ Run view, results explorer, annotation, composition, ... ++ tests/ unit / integration / e2e-lane / fixtures ++src/ ++ core/ types.jl, config.jl (cascade), duckdb_store.jl, project.jl, ++ databases.jl, primers_library.jl, composition_library.jl, ++ validate.jl, r_runtime.jl, epistemic.jl, log.jl ++ pipeline/ tools.jl (cutadapt/QC wrappers), dada2/ (R bridge), ++ swarm.jl, merge_taxa.jl — typed stage results ++ annotation/ funcdb: dual-classifier consensus, curation, ledger ++ analysis/ diversity.jl, analysis.jl, estimation.jl, exact_summaries.jl, ++ numeric_policy.jl, scaling.jl, AnalysisConfig.jl, Execution.jl, ++ clade_cumulus.jl (scaffold) ++ server/ Oxygen.jl HTTP server + routes/ (REST + SSE) ++config/ defaults/, schemas/ (analysis_config.schema.json + .ncl), ++ composition.yml, databases.yml, primers.yml, tool_versions.yml ++data/, projects/ input FASTQs (user) / outputs (generated) ++ui/ the opt-in Stipple/Vue migration slice (PR #73 in flight) ++``` ++ ++## The typed stage spine ++ ++Each pipeline stage returns a typed result — `TrimmedReads`, `ASVResult`, ++`OTUResult`, `TaxonomyHits`, `MergedTables` (`src/core/types.jl`) — and ++skips itself when outputs are current: **mtime freshness for files, content ++hash for configuration**. Rerunning after a config change re-executes the ++minimum set; the UI flags exactly the stages a change would regenerate and ++names the changed keys at their cascade level. This freshness contract is ++the quiet backbone: everything user-facing ("why is this stale?") reduces to ++it. ++ ++## A request's life (analysis chart) ++ ++``` ++browser (Run view) ++ → POST /api/v1/.../analysis/{alpha|taxa-bar|venn|nmds|permanova} ++ → server/routes (authz-free single-user layer; validation in core/validate.jl) ++ → analysis/Execution.run_analysis ++ → numeric_policy mode assertion (exact work needs exact mode) ++ → diversity/estimation/scaling as the AnalysisConfig directs ++ → R bridge (r_runtime.jl) where the method is R-backed ++ → refusal (named unsuccessful state) where it cannot run ++ → Plotly chart JSON back to the browser → plotly.js-dist-min render ++``` ++ ++Chart JSON, not HTML fragments: the frontend is a thin renderer over ++server-built Plotly specifications (with the chart-editor seam for ++cosmetics — one of the Stipple migration's known parity risks). ++ ++## The R boundary ++ ++R (DADA2, vegan, MASS, stats) enters through `src/core/r_runtime.jl` and the ++DADA2 module's `dada2_functions.r`. Design rules: the R session is a ++capability (its absence is reported, never faked — `GET /api/v1/capabilities`); ++R's `NA` is read as missing, not as a value (a real bug once killed every ++parametric fit — #66); package availability is decided by `renv.lock`, and ++"package absent" yields "Not Implemented", never a guess. ++ ++## Where state lives ++ ++| State | Home | Notes | ++|---|---|---| ++| Inputs | `data/{study}/[{group}/]{run}/*.fastq.gz` | user-owned; leaves = runs | ++| Outputs | `projects/{study}/{run}/…` | generated; `run_config.yml` is provenance | ++| Tables the UI queries | `merged/results.duckdb` (per run) | via `core/duckdb_store.jl` | ++| Curation (contamination, BLAST overrides) | separate from derived annotation | survives re-annotation | ++| FuncDB ledger | append-only | survives re-annotation | ++| Config cascade | `config/`, `data/**/pipeline.yml`, `projects/**/run_config.yml` | merged truth at run_config.yml | ++| Jobs | in-memory + logs | cancel via `DELETE /jobs/{id}` | ++ ++## Known structural seams (read before moving things) ++ ++- The **frontend wire types** (`src/api/types.ts`, upstream file) are the API ++ boundary of record; `src/types/api/` fakes nothing — endpoint→source map in ++ `src/types/api/index.ts`. ++- **Plotly-chain modules** are import-blocked in the DOM-less bun test lane; ++ the DOM lane decision (playwright vs harness) is pending before the e2e set ++ grows (`frontend/tests/unit/plotly-chain.todo.test.ts`). ++- The **Stipple slice** (`ui/`) is intentionally isolated (own Julia ++ environment) — do not share the HTTP-2 environment with the Oxygen HTTP-1 ++ server (a migration-record line). ++- **CladeCumulus** (`src/analysis/clade_cumulus.jl`) is scaffolded ++ structures only; the real implementation rides its own branch/issue (#6). +diff --git a/docs/wikis/Developers--Extending-the-Pipeline.md b/docs/wikis/Developers--Extending-the-Pipeline.md +new file mode 100644 +index 0000000..43aea2b +--- /dev/null ++++ b/docs/wikis/Developers--Extending-the-Pipeline.md +@@ -0,0 +1,95 @@ ++ ++ ++# Extending the pipeline ++ ++**Status: IN PLACE** (the extension points; the queue of things to add with ++them is [Status and Roadmap](Status-and-Roadmap)). ++ ++## Adding a pipeline stage ++ ++1. New stage module under `src/pipeline/` returning a typed result from ++ `src/core/types.jl` (or a new type added there, with its consumers). ++2. **Freshness contract first:** declare what outputs the stage writes and ++ which config keys invalidate them. The mtime/hash skip logic and the UI's ++ stale-flagging both read this; a stage without it will re-run forever or ++ never. ++3. Wire into `Execution`/the run orchestrator at the right point in the ++ graph (the two lanes and their joins are drawn in ++ [Design Progression](Deep-Dives--Design-Progression)). ++4. Config block: defaults in `config/defaults/pipeline.yml` (canonical, do ++ not edit casually) + cascade merge in `core/config.jl` semantics + UI ++ accordion section. ++5. Tool wrapper (if it shells out): `src/pipeline/tools.jl` pattern — path ++ resolution from `config/tools.yml`, optional SSH routing, and a pinned ++ record in `config/defaults/tool_versions.yml` (version + URL + sha256) so ++ `install.sh` fetches it byte-exactly. An unpinned tool cannot merge. ++6. Tests: fixture under `test/fixtures/`, unit test with known outputs, ++ negative controls for each refusal. ++ ++## Adding a statistical method — the gated path ++ ++This is the extension the project cares most about, so it is the most ++gated: ++ ++1. **Conditions document** in `docs/statistics/method-conditions/` — ++ response types accepted (and what happens to zeros and all-zero samples), ++ study designs supported (pairing, blocking, repeated measures, ++ covariates, depth, compositionality), overdispersion/depth policy, ++ uncertainty (interval method, effect size, BH mandatory), diagnostics ++ and the assumption most likely to be wrong, computational limits and ++ what happens at them (`ResourceLimitError`-style states). Published and ++ reviewed **before** code. ++2. **Catalogue registration** in `docs/statistics/method-catalogue-v1.md` ++ (or its successor) — scope approval is a separate step from behaviour. ++3. Implementation that can only refuse or do what the document says. If the ++ document and the code disagree, **the document is right and the file is ++ the bug** (the standing rule; it has already found real defects). ++4. Tests in the estimation suite's idiom: known answers written into the ++ fixture data (not read back out of the fit), an independent reference ++ (R or Python) compared value-by-value, negative controls for every ++ refusal path, and a guard that stubs cannot return. ++5. AnalysisConfig surface: new method token + validation + UI copy that ++ surfaces refusals as first-class results. ++ ++**Queue discipline:** items land in the approved dependency order ++([Analysis and Statistics Today](Users--Analysis-and-Statistics-Today) ++"coming" list). Jumping the queue requires lifting a deferral in writing (as ++happened for TSS/CSS/RSS offsets on 2026-09-25 — see the header of ++`method-conditions/scaling-and-offsets.md` for how to do it honestly). ++ ++## Adding a composition filter or category set ++ ++Edit `config/composition.yml`'s `filters:` library (or add a filter file with ++the same shape: `databases:`, `mappings:`, `filters:`, `remove_empty:`) and ++reference it from a category set or `merge_taxa.filters`. Filters are ++database-scoped (`databases: [pr2]` etc.) and both editable from the ++Compositions UI and validatable as YAML. Remember the consensus constraint: ++patterns are matched against labels from a specific reference release. ++ ++## Adding UI surface ++ ++React lane (default): views in `frontend/src/views/`, wire types first ++([Type System](Developers--Type-System)). Charts consume server-built Plotly ++JSON — build specs server-side (`analysis/` chart builders) unless the ++interaction is purely cosmetic (then the chart-editor seam). ++ ++Stipple lane (`ui/`, migration): small JS adapters allowed; **no new ++application TypeScript**; parity risks (chart editor, custom table, ++Euler/UpSet) are named in `docs/migration/STATUS.md` — do not start them ++without reading that. ++ ++## Adding an API endpoint ++ ++Follow the grouped layout in `src/server/routes/`; JSON in/out; SSE only for ++event streams; update the endpoint→SOURCE map in `frontend/src/types/api/` ++in the same change. [REST API](Developers--REST-API). +diff --git a/docs/wikis/Developers--REST-API.md b/docs/wikis/Developers--REST-API.md +new file mode 100644 +index 0000000..aef5fae +--- /dev/null ++++ b/docs/wikis/Developers--REST-API.md +@@ -0,0 +1,70 @@ ++ ++ ++# REST API ++ ++**Status: IN PLACE** (surface summary; route handlers in `src/server/routes/` ++are the truth). All responses are JSON; analysis endpoints return Plotly ++chart specifications. Base: `/api/v1/`. Single-user, **no authn** — see ++[Operator Track](Maintainers--Operator-Track) before binding anywhere but ++localhost. ++ ++## Endpoint groups ++ ++| Group | Endpoints | Description | ++|---|---|---| ++| Studies | `GET/POST/DELETE /studies`, `POST .../rename` | list, create, rename, delete studies | ++| Groups | `POST/DELETE /studies/{study}/groups`, `POST .../rename` | group management (pooled-run navigation) | ++| Runs | `GET/POST/DELETE /studies/{study}/runs`, `POST .../rename` | run management; leaves on disk are discovered | ++| Config | `GET/PATCH/DELETE .../config`, `GET .../config/overrides` | read/edit any cascade level; list downstream overrides | ++| Primers | `GET /primers`, `GET /primers/document`, `PUT /primers` | pair names; whole-document read/replace (validated before write) | ++| Pipeline | `POST .../pipeline`, `POST .../stages/{stage}` | launch full study, single run, or one stage (DADA2 substages included) | ++| Jobs | `GET/DELETE /jobs`, `GET /jobs/{id}/logs` | monitor, cancel, fetch logs | ++| Events | `GET /events` | **SSE** stream: real-time job and stage updates | ++| Results | `GET/POST/DELETE .../results/tables/...` | list, query, filter, save, export (`.xlsx`), delete; OTU member drill-down | ++| QC | `GET .../results/qc`, `GET .../results/dada2` | MultiQC metadata; DADA2 figures, logs, stats | ++| Analysis | `POST .../analysis/{alpha,taxa-bar,venn}`, `GET .../analysis/{pipeline-stats,ranks}` | per-run charts; rank discovery | ++| Cross-run | `POST /studies/{study}/analysis/{alpha,taxa-bar,nmds,permanova,venn}` | comparisons, NMDS, PERMANOVA, overlap across runs | ++| Composition | `GET /category-sets`, `POST /composition/{source}/{build,query,distinct,analysis}` | category sets; organism-composition tables and charts | ++| Annotation | `GET/POST .../annotations/{source}/...`, `POST /funcdb/entries`, `PATCH .../{contamination,blast-assignment}` | generate/query annotations; curate; FuncDB ledger | ++| Filter presets | `GET/POST/DELETE /filter-presets` | reusable table filters | ++| Databases | `GET /databases`, `GET /databases/document`, `PUT /databases`, `POST /databases/{key}/download` | whole-document read/replace (validated; returns advisory warnings), explicit downloads | ++| System | `POST /init`, `GET /capabilities` | project directory initialisation; server capabilities (e.g. R availability) | ++ ++## Contracts worth knowing before a client change ++ ++- **Plotly JSON is the chart wire format** — server-built specs; the client ++ renders (`plotly.js-dist-min`) and may adjust cosmetics via the ++ chart-editor seam. A new chart means a server-side builder. ++- **Validation is server-side and total** (`core/validate.jl`): primers and ++ databases documents are validated whole before any write — a rejected ++ document leaves the file untouched. Same for config patches (bad cascade ++ keys refuse; dangling renames are *reported*, not blocked). ++- **Refusals travel as structured unsuccessful states** on analysis ++ endpoints — treat them as first-class responses in every client ++ (including the Stipple one), never as empty successes (#31's rule ++ generalised). ++- **SSE (`/events`)** is the only push channel; reconnect snapshots and ++ bounded subscriptions are on the Stipple migration's checklist ++ (`docs/migration/STATUS.md`) — design clients to re-sync, not to assume a ++ gapless stream. ++- **Capabilities** (`/capabilities`) is how a client learns that R is absent ++ before offering R-backed methods — check it, do not infer from errors. ++ ++## Adoc-era notes ++ ++- The endpoint→SOURCE map in `frontend/src/types/api/index.ts` must move in ++ the same change as any route (the type estate's rule). ++- **COMING:** the API is versioned at `/api/v1/` for the React client; the ++ Stipple client consumes the same surface through a validated backend ++ adapter (`ui/`), and recoverable Zenodo publication endpoints ride issue ++ #8 (draft PR #74). +diff --git a/docs/wikis/Developers--Statistics-Internals.md b/docs/wikis/Developers--Statistics-Internals.md +new file mode 100644 +index 0000000..71d037c +--- /dev/null ++++ b/docs/wikis/Developers--Statistics-Internals.md +@@ -0,0 +1,91 @@ ++ ++ ++# Statistics internals ++ ++**Status: PARTIAL — implemented with strong tests; independent review (#1) ++outstanding.** The module map and the contracts; the mathematics is in ++[Deep Dives](Deep-Dives). ++ ++## The modules and what each may claim ++ ++| Module | Owns | May claim | May never claim | ++|---|---|---|---| ++| `numeric_policy.jl` | `NumericPolicySpec` (modes `:ordinary`, `:exact_counts`, `:high_precision`) | exactness of integers/rationals; bounded denominators; display rounding | that high precision is exactness | ++| `exact_summaries.jl` | counts + proportions at exact precision | exact descriptive facts (incl. counts > 2⁵³−1) | any inference | ++| `estimation.jl` | per-feature ML fits (`nb_glm`, `clr_lm`, `ilr_lm`, `logistic`) | estimates **or** named unsuccessful states | a number when a fit fails | ++| `scaling.jl` | size factors + offsets (TSS/CSS/RSS) | depth modelling without touching the response | that offsets solve compositionality | ++| `AnalysisConfig.jl` + `Execution.jl` | immutable config; the run path | what the user actually asked for | silent substitution of a nearby method | ++| `diversity.jl` / `analysis.jl` | indices, charts, NMDS/PERMANOVA | descriptive + vegan-backed ordination | significance without status | ++| `epistemic.jl` | statuses, receipts, DANGER banner | *how* a result is warranted | soundness of a warrant (deliberately) | ++ ++The contracts the code is held to (published **before** implementation, and ++the tests check code against documents, not the reverse): ++ ++- `docs/statistics/numeric-contracts.md` — exact/approximate/rounded; what ++ may cross each boundary; what is refused (`assert_mode` fails rather than ++ handing a Float64 that looks exact). ++- `docs/statistics/method-conditions/exact-descriptive-summaries.md` — ++ catalogue item 1. Surfaced two real defects under the layer (`to_display` ++ labelling a rendering *6dp* then printing 80 digits; `2//3` syntax leaking ++ into user text) — the conditions doc caught both. ++- `docs/statistics/method-conditions/parametric-fits.md` — catalogue item 2: ++ the four methods' response types, refusals, determinism (no resampling, no ++ random start; the recorded seed is honestly labelled "not a parameter of ++ any number here"). ++- `docs/statistics/method-conditions/scaling-and-offsets.md` — factors vs ++ offsets vs **transforms**; the three historical substitutions it forbids ++ (TSS→proportions; CSS/RSS→`relative`; `size_factors`→TSS wearing DESeq2's ++ name) and why "the document is right and the file is the bug". ++ ++## The run path ++ ++`Execution.run_analysis` is the only door the server uses: parse/validate ++`AnalysisConfig` → assert numeric mode → dispatch to the modules → compose ++Plotly specs. Ordering rules that bite newcomers: ++ ++1. Zero-depth samples are **healed before** any transform (#58) — transforms ++ are not zero-safe and are not claimed to be. ++2. Offsets attach at the fit (log of the size factor, stated centring); ++ the response stays counts. ++3. BH (`p.adjust`, R) is attached wherever several tests are reported — the ++ DANGER banner path exists to make disabling it loud. ++4. A refusal short-circuits *with a name* (non-convergence, non-identifiable, ++ boundary estimate, precondition failure, resource limit). Downstream chart ++ builders receive the state and render it as such. ++ ++## Where the R lives ++ ++`MASS::glm.nb` (NB-GLM), `stats::lm` (CLR/ILR), `stats::glm(family=binomial)` ++(logistic), `p.adjust` (BH), vegan (`vegdist`/`adonis`-family for NMDS/ ++PERMANOVA) — reached through `src/core/r_runtime.jl`. R's `NA` is read as ++missing (#66's lesson). The independent reference tests fit the same fixtures ++in R and compare coefficient-by-coefficient — that suite is the reason to ++trust the bridge at all. ++ ++## Anti-patterns this layer has already survived (do not reintroduce) ++ ++- **Placeholder-as-result:** p-values once derived from `hash(taxon_id)`. ++ Removed; a source-level guard test now fails if anything similar returns. ++ If you are tempted to stub a statistic to unblock a UI, return an ++ unsuccessful state — the UI already knows how to render one. ++- **Silent method substitution:** accepting `CSS` and computing `relative`. ++ Method names compare in lower case (#62) and unknown/aliased methods refuse. ++- **Precision-laundering:** printing a rounded display labelled as exact. ++ `to_display` is contract-tested. ++ ++## Extending it ++ ++Adding a method = conditions document first (response types, design support, ++zeros, overdispersion/depth, uncertainty + multiple testing, diagnostics, ++computational limits), then code, then tests with known answers + negative ++controls for every refusal path. See ++[Extending the Pipeline](Developers--Extending-the-Pipeline). +diff --git a/docs/wikis/Developers--Testing-and-Benchmarks.md b/docs/wikis/Developers--Testing-and-Benchmarks.md +new file mode 100644 +index 0000000..755a1a2 +--- /dev/null ++++ b/docs/wikis/Developers--Testing-and-Benchmarks.md +@@ -0,0 +1,87 @@ ++ ++ ++# Testing and benchmarks ++ ++**Status: IN PLACE** (with two deliberate COMING gates named below). Test ++inventory and metrics: `docs/testing/coverage.md`; infrastructure: ++`docs/testing/infrastructure.md`; facet taxonomy: `docs/testing/taxonomy-facets.md`. ++ ++## The lanes ++ ++| Lane | Command | Gate? | Notes | ++|---|---|---|---| ++| Frontend unit + integration | `bun test` (in `frontend/`) | yes | DOM-less lane; type-level assertions included | ++| Frontend combined | `bun run check` | yes (pre-push) | typecheck + tests + bench invocation | ++| Julia unit/integration | `just` test recipes | yes | fixtures under `test/fixtures/` | ++| e2e | `just test-e2e` | opt-in | fails loudly without browsers (by design) | ++| Benchmarks | `bun run bench/`, `bench/` | **informational** | compares vs `bench/*/baseline.json`; no gate | ++| Repo hygiene | `just ci` | yes | SPDX, format, lint, pins drift | ++ ++Everything CI runs is `just ci`. If a lane is missing its tool it fails ++loudly rather than skipping — silence would be a lie about coverage. ++ ++## The four test idioms that matter most ++ ++1. **Known answers written into the data.** Estimation fixtures embed ++ outcomes in the counts; the test reads the fit against the planted truth ++ — never a snapshot read-back of whatever the fit produced. ++2. **Independent reference comparison.** The same fixtures fitted directly ++ in R (coefficients, then BH against `p.adjust`), or exact summaries ++ compared against Python's `fractions.Fraction` (skipping loudly by name ++ when `python3` is absent). ++3. **Negative controls for every refusal path.** Each unsuccessful state has ++ a test that *provokes* it. A refusal path without a test is a rumour. ++4. **Guards against regression to dishonesty.** Source-level guard: the ++ placeholder-statistics pattern cannot return (estimation suite); the ++ "pipeline does not call `exact_summaries` unless selected" property; ++ the silent-substitution regressions (TSS/CSS/RSS aliases) cannot return ++ (scaling suite). ++ ++Boundary values are **measured, not assumed** — the numeric-boundary suite ++(#52's lesson) asserts where Float64/int behaviour actually breaks (counts ++beyond 2⁵³−1, denominator budgets, rounding). ++ ++## Type-level assertions ++ ++20+ assertions pin the frontend domain model (`frontend/tests/`). They are ++tests of *shape* — the compiler-adjacent half of the honesty contract. The ++DOM-less lane's blind spot is bounded and known: Plotly-chain modules ++(`PlotlyChart`, `ChartCustomiser`, `ChartEditorInner`, `AnnotationPanel`, ++`RunView`) are import-blocked there; `TODO(tests/e2e-lane)` marks the spot. ++ ++## Benchmarks ++ ++`bench/` carries baselines for: table loading, duckdb aggregation, tree ++rendering, PERMANOVA/NMDS, epistemic parsing, analysis config — plus a ++comprehensive lane and a layer-1 mock-recovery harness (fetch/evaluate/ ++report over `datasets.yml`). Comparison output is informational; promotion ++to a non-gating regression **alert** waits for a multi-week stability window ++(deliberate, per the roadmap). ++ ++**COMING (gates, deliberately deferred):** the **DOM test lane** decision ++(playwright vs a DOM harness — queued at the e2e-lane TODO) and the ++**coverage gate** (lands against a recorded baseline in ++`docs/testing/coverage.md`, never an arbitrary percentage). ++ ++## Writing a test for this repo (cheat sheet) ++ ++- Julia statistical code → idioms 1–4 above; if you cannot plant a known ++ answer, the method is not ready for a test, which usually means its ++ conditions document is not ready either. ++- Frontend → DOM-less lane first; if the module touches Plotly chains, it ++ joins the import-blocked list *explicitly* (with a TODO) rather than ++ quietly skipping. ++- Fixtures stay small (`test/fixtures/`, `frontend/tests/fixtures/`) — the ++ MiSeq SOP sample data is for runs, not unit tests. ++- CI names failing assertions since #65 — write assertions you would want ++ named in red. +diff --git a/docs/wikis/Developers--Type-System.md b/docs/wikis/Developers--Type-System.md +new file mode 100644 +index 0000000..752196a +--- /dev/null ++++ b/docs/wikis/Developers--Type-System.md +@@ -0,0 +1,93 @@ ++ ++ ++# Type system ++ ++**Status: PARTIAL — strict foundation IN PLACE; two documented exceptions.** ++This page is the placement map and the rules; the reasoning that connects ++types to *statistical claims* is ++[Deep Dives — Type Theory Meets Statistics](Deep-Dives--Type-Theory-Meets-Statistics). ++ ++## The two type estates ++ ++### Frontend (TypeScript) — strict by gate ++ ++`frontend/tsconfig.json` (no-emit gate) + `tsconfig.build.json` sidecar; ++`bun run typecheck` is CI-gated at **0 errors** under `strict` + ++`exactOptionalPropertyTypes` + `verbatimModuleSyntax`. Placement hierarchy ++(mapped in `docs/types/architecture.md`): ++ ++``` ++src/api/types.ts canonical REST wire types (upstream file, boundary of record) ++src/types/api/ DOMAIN A — API boundary leaves (tables, cosmetics, SOURCE map) ++src/types/plotly.ts DOMAIN 0 — hand-written plotly.js subset vocabulary ++src/types/declarations.d.ts the ONLY home of ambient module declarations ++``` ++ ++Rules of the estate (enforced in review and mostly in compiler): ++ ++- **Never `any`. Never `@ts-ignore`.** Use `unknown` and narrow. Type-only ++ imports must be `import type` (machine-checked via `verbatimModuleSyntax`). ++- The two third-party libraries with no published types ++ (`plotly.js-dist-min`, `react-chart-editor`) carry `FIXME(types)` ++ **unknown-safe** stubs in `declarations.d.ts` — replacing upstream's own ++ silent `treat its exports as any` — each with a tracking pointer in ++ `docs/type-system/category-d-e-closure.md`. ++- Category D/E closure (third-party and framework type coverage) is audited ++ and closed or explicitly parked with evidence — see that document. ++ ++**The two documented exceptions (PARTIAL's meaning):** ++ ++1. `skipLibCheck: true` — react-router 6.30.x bundles 7 erroneous `.d.ts` ++ entries, unfixable in-repo; retried on react-router 7. Tracked with exit ++ criteria in `docs/compliance/fixme-index.md`. ++2. `useAnalysis.alphaFig` is `unknown` — narrowing to `PlotFigure` requires ++ threading the chart response type through analysis state; listed in ++ `docs/types/architecture.md § Known gaps`. ++ ++20+ type-level assertions pin the domain model (the strict-TS foundation's ++165→0 error journey is logged in `docs/type-system/strict-mode-foundation.md`). ++ ++### Julia — the typed stage spine ++ ++`src/core/types.jl` defines the pipeline's result types (`TrimmedReads`, ++`ASVResult`, `OTUResult`, `TaxonomyHits`, `MergedTables`); `AnalysisConfig` ++is an immutable struct whose construction validates. Julia's dispatch and ++the validation layer together do the work dependent types would do in Idris: ++the *shape* of a result is checked at boundaries (config construction, R ++returns, table loads), not assumed. ++ ++### The epistemic shadow types ++ ++`src/core/epistemic.jl` is where the type-theoretic material becomes ++executable: finite Julia shadows of the Agda lineages (echo-types: ++`EchoFiber` as a Σ-type Σ(x:A), (f x ≡ y); epistemic-types: `Modality`, ++`FactiveModality`, `Warrant` *without soundness*; residual-evidence-types: ++`Candidate`, `Holds`, `Identified`). The wire column `avec_fibre`, statuses ++like `present_in_every_admissible_world`, and the DANGER-banner discipline ++all hang off this module. Deep end: ++[Deep Dives — Epistemic Status](Deep-Dives--Epistemic-Status). ++ ++## Adding a type — the checklist ++ ++**Frontend:** leaf types go in `src/types/api/`; wire types extend ++`src/api/types.ts` *with the endpoint→SOURCE map updated*; plotly vocabulary ++extends `src/types/plotly.ts`. Ambient declarations go nowhere except ++`declarations.d.ts` (with `FIXME(types)` header if third-party). Grep for ++`any` before pushing — `bun run check` will not catch a creative escape. ++ ++**Julia:** stage results extend `src/core/types.jl` and the stage's ++freshness contract; config keys extend the schema (`config/schemas/`) and ++`core/config.jl` merge semantics together or the cascade will lie. ++ ++**Both:** if the type encodes a *statistical claim* (exactness, refusal, ++status), it needs a line in the relevant method-conditions document first — ++types are how the honesty rules stop being conventions. +diff --git a/docs/wikis/Developers.md b/docs/wikis/Developers.md +new file mode 100644 +index 0000000..71c5be5 +--- /dev/null ++++ b/docs/wikis/Developers.md +@@ -0,0 +1,65 @@ ++ ++ ++# Developers ++ ++This section is for people changing the code: where things live, why the ++types are the way they are, how the statistics layer is wired, and how to ++add to the pipeline without breaking the honesty contract. ++ ++Read this section with the repository open. Every claim here names its file; ++the deep *why* is in [Deep Dives](Deep-Dives). ++ ++## Learning path ++ ++1. [Architecture Tour](Developers--Architecture-Tour) — the moving parts and ++ the typed stage spine. ++2. [Type System](Developers--Type-System) — the frontend type estate, the ++ Julia type spine, and the rules for adding either. ++3. [Statistics Internals](Developers--Statistics-Internals) — numeric ++ policy, estimation, exact summaries, scaling/offsets, AnalysisConfig, and ++ the epistemic layer. ++4. [Extending the Pipeline](Developers--Extending-the-Pipeline) — adding a ++ stage, a method, a filter; the gate that keeps refusals honest. ++5. [Testing and Benchmarks](Developers--Testing-and-Benchmarks) — lanes, ++ fixtures, baselines, the guard tests that matter most. ++6. [REST API](Developers--REST-API) — endpoint groups, SSE, Plotly JSON. ++ ++## Before your first commit ++ ++- `just ci` must pass (it is what CI runs). `CONTRIBUTING.md` has commit and ++ branch conventions (the commit gate is real and grades every non-merge ++ commit). ++- Three invariants this fork will not trade away: ++ 1. **No placeholder ever returns as a result.** Unsuccessful states carry ++ names; numbers carry provenance. A guard test enforces this. ++ 2. **Method conditions are published before implementation.** If you add a ++ statistical method and there is no conditions document, you are doing it ++ in the wrong order (`docs/statistics/method-conditions/`). ++ 3. **Never `any`, never `@ts-ignore`.** `unknown` and narrow; the two ++ ambient stubs are the only exception and they are marked ++ `FIXME(types)`. ++ ++## Here now vs coming (developer view) ++ ++**IN PLACE:** the Julia engine (typed stages, config cascade, DuckDB store, ++REST+SSE), the analysis layer (diversity, estimation, exact summaries, ++numeric policy, scaling, AnalysisConfig, epistemic), the strict-typed React ++frontend, tests/benchmarks/CI, the compliance estate. ++ ++**COMING:** exact tests and the milestone-3 analysis suite (#3, #17–21) — ++each gated on pre-published conditions; CladeCumulus frontend (#6); Evidence ++Mode UI (#7); Zenodo integration (#8, a recoverable publication slice is in ++draft PR #74); Stipple UI parity (`ui/`, `docs/migration/STATUS.md`). ++ ++**BLOCKED:** the symbolic engine (#2) — by policy, until the numeric layer ++passes real-data validation. If you want formula manipulation, you want #1's ++review first. +diff --git a/docs/wikis/Home.md b/docs/wikis/Home.md +new file mode 100644 +index 0000000..796f288 +--- /dev/null ++++ b/docs/wikis/Home.md +@@ -0,0 +1,84 @@ ++ ++ ++# MetaManifold ++ ++Amplicon metabarcoding from raw paired-end Illumina reads to filtered, ++taxonomy-annotated ASV/OTU tables — one Julia engine, one browser workbench, ++and statistics that say what they cannot claim. ++ ++This wiki is the long-form documentation for ++[MetaManifold-WebUI](https://github.com/hyperpolymath/MetaManifold-WebUI) ++(origin design by Joshua Jewell; maintained and extended on the ++hyperpolymath fork). It is written in the ++[berrywiki](https://github.com/metadatastician/berrywiki) page format: plain ++Markdown with hidden tree metadata, so it survives with nothing but `git` and ++a browser. ++ ++## Three doors — pick the one that is yours ++ ++| If you are… | Start here | What you get | ++|---|---|---| ++| **A user** — running analyses (academic or lab professional) | [Users](Users) | Install, your first study, the configuration reference, the analysis surface — with separate tracks for [academics](Users--For-Academics) (reproducibility, citing, exactness) and [lab professionals](Users--For-Lab-Professionals) (routine throughput, QC, curation). | ++| **A platform maintainer** — running the platform or the repository | [Maintainers](Maintainers) | [Operator track](Maintainers--Operator-Track) (deployment, databases, toolchains, upgrades) and [steward track](Maintainers--Steward-Track) (CI gates, reviews, upstream relations, estate compliance). | ++| **A developer** — changing the code | [Developers](Developers) | Architecture tour, the type system, statistics internals, extension points, tests and the REST surface. | ++ ++Beyond the three doors, two cross-cutting sections: ++ ++- **[Deep Dives](Deep-Dives)** — the mathematics and engineering receipts that ++ the README and [EXPLAINME](https://github.com/hyperpolymath/MetaManifold-WebUI/blob/main/EXPLAINME.adoc) ++ deliberately keep short: the three-layer design progression, type theory ++ meets statistics, exact arithmetic, maximum likelihood and refusals, ++ compositional statistics and offsets, epistemic status, and the advanced ++ functionality that is coming. ++- **[Status and Roadmap](Status-and-Roadmap)** — the single board that marks, ++ for every area, what is **here now** and what is **coming**. ++ ++## Status legend — used on every page ++ ++Every section that could be mistaken for a promise carries one of four ++markers. They are used strictly: ++ ++- **IN PLACE** — implemented and verifiable in the tree today. ++- **PARTIAL** — usable now, with the limits named on the spot. ++- **COMING** — specified and tracked (issue numbers given), *not* ++ implemented. Never listed as a feature of the present. ++- **BLOCKED** — deliberately halted pending a named condition. ++ ++If a claim has no marker, it is a description of what the software does today. ++When in doubt, the [Status and Roadmap](Status-and-Roadmap) board is ++authoritative, and the repository tree is the final arbiter. ++ ++## Orientation in ninety seconds ++ ++MetaManifold is three designs stacked on each other (full story in ++[Deep Dives — Design Progression](Deep-Dives--Design-Progression)): ++ ++1. **The base design (R/Python):** raw DADA2 (R) for ASVs and the ++ swarm/vsearch shell tradition for OTUs — the science, used as published. ++2. **The MetaManifold augmentation (JoshuaJewell):** a Julia orchestrator that ++ runs both lanes per run and a web workbench (config cascade, QC, results, ++ annotation, composition) over a per-run DuckDB store. ++3. **The hyperpolymath steps (this fork):** honest statistics (real ++ maximum-likelihood fits or explicit refusals; exact counts and rationals; ++ exact TSS/CSS/RSS offsets), a typed frontend estate, and CI/tests/ ++ benchmarks that keep every claim checkable. ++ ++## Conventions ++ ++- **Academics vs lab professionals.** Where the two audiences need different ++ guidance, pages carry **Academics:** and **Lab professionals:** callouts; ++ where the guidance is shared, nothing is split. ++- **Issue references** use `#n` for ++ [this repository's tracker](https://github.com/hyperpolymath/MetaManifold-WebUI/issues). ++ Upstream references are named in full (e.g. JoshuaJewell/MetaManifold-WebUI). ++- **Source of truth.** These pages are sourced from `docs/wikis/` in the ++ repository and published here; see `docs/wikis/README.md` for the sync ++ convention. +diff --git a/docs/wikis/Maintainers--Compliance-and-Estate.md b/docs/wikis/Maintainers--Compliance-and-Estate.md +new file mode 100644 +index 0000000..985cd17 +--- /dev/null ++++ b/docs/wikis/Maintainers--Compliance-and-Estate.md +@@ -0,0 +1,82 @@ ++ ++ ++# Compliance and estate ++ ++**Status: IN PLACE — compliant-with-documented-deviations** (the deviations ++are the point of the record: they are written down, dated, and reasoned). ++This page orients maintainers in the hyperpolymath estate's machinery. ++ ++## Where this repository sits in the estate ++ ++| Estate artefact | Role here | Record | ++|---|---|---| ++| [standards](https://github.com/hyperpolymath/standards) | Organisation-wide specs: README/EXPLAINME authoring, A2ML metadata family, RSR canon, licence policy | `docs/compliance/standards-alignment.md` | ++| [rsr-template-repo](https://github.com/hyperpolymath/rsr-template-repo) | The Rhodium Standard Repository template this tree is aligned against | `docs/compliance/rsr-alignment.md` | ++| LICENCE-POLICY (in standards) | Five-rule register; the AGPL/MPL split applied here | `NOTICE`, `docs/compliance/standards-alignment.md` | ++| Language policy | TypeScript is estate-banned but **fork-exempt** here (the strict compiler is declared the lint dialect); Julia/R are the science stack | `docs/compliance/standards-alignment.md` | ++ ++## The licence split (do not "tidy" this) ++ ++- Upstream-authored files: `AGPL-3.0-only` (inherited work licence — the ++ origin's choice, untouchable under fork rules). ++- Fork-authored files: `MPL-2.0` (code) / `CC-BY-SA-4.0` (prose). ++- Authorship classification is mechanical (first-commit author), stated in ++ `NOTICE`; `scripts/check-spdx.sh` gates headers across the covered set ++ (machine metadata and generated files excluded, list in the compliance ++ record). ++- `LICENSES/` holds the canonical texts (AGPL-3.0-only, MPL-2.0, ++ CC-BY-SA-4.0); `LICENSE` is the AGPL root. ++ ++## Documentation standards in force ++ ++- **README + EXPLAINME authoring standard** (`standards:docs/README-EXPLAINME-STANDARD.adoc`) ++ — the root `README.adoc` + `EXPLAINME.adoc` pair follows it: README sells, ++ EXPLAINME proves with claim→implementation receipts, the mathematics lives ++ in this wiki (deliberately, so neither file becomes unreadable). ++- **BerryWiki page format** for this wiki ([metadatastician/berrywiki](https://github.com/metadatastician/berrywiki)) ++ — plain Markdown with hidden tree metadata; sourced from `docs/wikis/`. ++- GitHub-required files stay `.md` (SECURITY, CONTRIBUTING, CODE_OF_CONDUCT, ++ CHANGELOG) — the standard's platform exception. ++ ++## Repository settings (the manual surface) ++ ++- **Autolink references** — the complete four-tier elaboration (lineage and ++ estate; upstream bioinformatics tools; runtime/toolchain; registries) is ++ specified paste-ready in `docs/integration/autolink-references.md`. ++ Applying it needs Settings → Autolink references (Administration ++ permission). When estate automation gains that permission, the same file ++ is the machine source of truth. **Action pending: one human pass.** ++- Rulesets (`Immutable-Tags` style) and CODEOWNERS (`@hyperpolymath`) are in ++ force; Dependabot scoping is deliberate (see ++ [Steward Track](Maintainers--Steward-Track)). ++ ++## Estate patterns this repository dogfoods ++ ++Listed with their receipts in `EXPLAINME.adoc` ("Dogfooded across the ++account"): the README/EXPLAINME pair standard, BerryWiki page format, the ++Justfile command surface (Makefiles banned estate-wide), the pinned ++dual-lane toolchain (mise + Guix), and publish-conditions-before-implementation ++for statistics methods. The last is the pattern the estate's statistics ++track ([statistikles](https://github.com/hyperpolymath/statistikles)) adopts ++in turn — the method catalogue and method-conditions documents here are its ++working exemplar. ++ ++## Reading the alignment records ++ ++Both alignment checklists are **dated snapshots** ("compliant-with- ++documented-deviations, 2026-09-17"). When the estate standards move (they ++are a moving target by design — e.g. the language policy's ReScript ban ++date), re-run the comparison and update the checklist with a new date rather ++than editing history in place. The fixme index ++(`docs/compliance/fixme-index.md`) is the living exception register ++(`skipLibCheck`, `alphaFig`, `FIXME(types)` stubs) with exit criteria. +diff --git a/docs/wikis/Maintainers--Operator-Track.md b/docs/wikis/Maintainers--Operator-Track.md +new file mode 100644 +index 0000000..a1fab44 +--- /dev/null ++++ b/docs/wikis/Maintainers--Operator-Track.md +@@ -0,0 +1,113 @@ ++ ++ ++# Operator track ++ ++**Status: PARTIAL — single-user local operation is IN PLACE and supported; ++server/multiuser operation is explicitly NOT (yet).** This page is for ++whoever keeps an instance healthy. ++ ++## Deployment posture (read this first) ++ ++The present server is a **local single-user** application. It binds `:8080` ++and serves the workbench to a browser on the same machine. Do **not** expose ++it to a network as if it were multiuser-ready — no authn/authz layer exists, ++and job execution is not isolated between users. A standing line in ++`docs/migration/STATUS.md` says exactly this; repeat it to stakeholders. ++ ++What "local" still gives you in a lab: multiple analysts can use one machine ++(same OS account), and the service answer "the analysis station" is the ++supported shape. ++ ++**COMING:** authenticated remote/multiuser deployment (agreed requirement, ++not delivered), standalone archives (below), coordinated signed updater ++(below). ++ ++## The install you operate ++ ++Source install is the supported lane (see ++[Install and First Run](Users--Install-and-First-Run)). The operator's pieces: ++ ++| Artefact | You operate | Notes | ++|---|---|---| ++| `mise.toml` | exact toolchain pins | source of truth; `.bun-version` is generated from it (`just sync-pins`) | ++| `guix.scm` + `channels.scm` | the peer dev lane | time-machine-pinned; functional equivalents for tools, not binary identity | ++| `renv.lock` | exact R package set | restore via `renv::restore()`; `.Rprofile` activates the project library | ++| `config/defaults/tool_versions.yml` | sha256-pinned tool downloads | `install.sh` fetches byte-exact; preflight asserts | ++| `config/tools.yml` | resolved tool paths | including SSH-hosted binaries for vsearch etc. | ++ ++## Reference databases (the real operational load) ++ ++`config/databases.yml` is the one place to manage URIs, local overrides and ++remote paths. Operator rules that matter in production: ++ ++1. **Mirror the reference FASTAs locally** (`local:` override) before a busy ++ period — first use downloads, and you do not want that at 09:00 on batch ++ day. The Databases page can trigger downloads explicitly. ++2. **Pin the release** you validated for the assay, for **both** formats ++ (dada2 + vsearch) of a database. The consensus comparator is string ++ equality across the two classifiers — mixed releases silently degrade ++ agreement scores. ++3. **Renaming/removing a database or its `levels` reports blast radius** — ++ which studies it affects, resolved through the real cascade including ++ inheritors. Act on the report. ++4. For the SSH taxonomy offload: `remote_path` in `databases.yml` keeps the ++ database resident on the remote host (no per-run transfer). **Authorisation ++ to the remote host is solely yours to ensure** — the disclaimer in ++ `config/defaults/pipeline.yml` is deliberate and non-negotiable. ++ ++## Capacity and performance ++ ++- Memory pressure concentrates in `assignTaxonomy()` (DADA2, large ++ databases). Levers: `dada2.taxonomy.multithread` (higher = more memory), ++ the SSH offload, or smaller/custom reference sets. ++- The OTU lane (swarm) is CPU-scalable (`swarm.threads`); cd-hit-est likewise ++ (`cdhit.threads`). ++- Benchmarks (`bench/`) carry recorded baselines (table loading, duckdb ++ aggregation, tree rendering, PERMANOVA/NMDS, epistemic parsing) — use them ++ to detect that *your* host is the regression, not the code. ++- Storage: budget per run ≈ trimmed FASTQs + QC reports + tables; `projects/` ++ grows monotonically (checkpoints included). Back up `data/` + `projects/`; ++ `config/` is small and precious (presets, primers, databases, composition ++ library). ++ ++## Day-2 operations ++ ++| Task | How | ++|---|---| ++| Update tools | `bash install.sh --update` (re-resolves against the pinned records) | ++| Update R packages | **don't**, casually — `renv.lock` is the reproducibility contract; changes go through the steward track with a lockfile diff | ++| Backup | `data/`, `projects/`, `config/` — everything else is reconstructable | ++| Health | `GET /api/v1/capabilities` (e.g. R availability); jobs panel + SSE stream during runs | ++| Port/root changes | `JULIA_METAMANIFOLD_PORT`, `JULIA_METAMANIFOLD_ROOT`, `JULIA_THREADS` | ++| Suspected stale outputs | trust the staleness flags (they name changed keys); `run_config.yml` is ground truth for what ran | ++ ++## When something refuses at 09:00 ++ ++Refusals name themselves (unsuccessful states, "Not Implemented", "unknown" ++significance). [Troubleshooting](Users--Troubleshooting) decodes them. The ++one operational trap: **"Not Implemented" is not an outage** — the method or ++package is genuinely absent from the locked environment and the software ++refuses rather than guesses. Escalating that means a feature request, not a ++restart. ++ ++## Coming on this track ++ ++- **COMING:** standalone offline archives (Linux x86-64 + ARM64) — unprivileged ++ launcher, writable state outside the immutable release, no toolchain on the ++ host, lazy science downloads proven, then explicit WSL2 tests. Policy ++ exists (`packaging/`, `docs/migration/STATUS.md` workstream B); the builder ++ does not exist yet. ++- **COMING:** the coordinated updater — one tested pinned combination, ++ signed metadata, transactional switch with rollback, offline startup. ++- **COMING:** multiuser/authenticated serving — the precondition for ++ network deployment. +diff --git a/docs/wikis/Maintainers--Releases-and-Distribution.md b/docs/wikis/Maintainers--Releases-and-Distribution.md +new file mode 100644 +index 0000000..de88369 +--- /dev/null ++++ b/docs/wikis/Maintainers--Releases-and-Distribution.md +@@ -0,0 +1,89 @@ ++ ++ ++# Releases and distribution ++ ++**Status: PARTIAL.** What ships today is honest and small: tags, source, and ++release-served archives. The standalone product is **COMING**. This page ++draws the line between the two precisely — `docs/migration/STATUS.md` is the ++authoritative record. ++ ++## What a release is today ++ ++- **Git tags + CHANGELOG.md** on `hyperpolymath/MetaManifold-WebUI`, source ++ archives generated by the forge (`v0.1.0` notes in ++ `docs/release-notes/v0.1.0.md`). ++- **Release-served pinned tool archives** — e.g. the FastQC archive is served ++ from this repository's releases (issue #54) so `install.sh` fetches a ++ byte-exact, sha256-verified binary rather than chasing upstream URLs. ++- Users install from source ([Install and First Run](Users--Install-and-First-Run)). ++ ++A source ZIP is **not** the standalone product — the distinction is called ++out in the migration record precisely because it will matter to regulators ++and core facilities. ++ ++## The standalone product (COMING — agreed, not built) ++ ++**Target shape** (agreed requirements; none validated end-to-end yet): ++ ++- Standalone offline-first release archives, native Linux x86-64 and ARM64. ++ Guix builds the environment; end users never install Guix. ++- Bundled: mise, just, Bun, Julia + packages, R + Bioconductor (RCall), the ++ Python/native/Java scientific tools (cutadapt, swarm, vsearch, cd-hit-est, ++ FastQC/MultiQC), and the prebuilt web assets. Users supply sequencing data ++ and reference databases only. ++- Unprivileged launcher; writable state outside the immutable release; ++ existing browser remains the UI. Nothing is built or installed at first ++ launch. ++- Proven relocatable offline execution — including paths with spaces, lazy ++ scientific functionality, no `/gnu/store` dependency — with measured ++ artefact sizes and published kernel/CPU baselines. ++- Published per-architecture: archives, hashes, **signed metadata**, and a ++ component/licence inventory on GitHub Releases. ++ ++**Workstream B's remaining steps** (numbered in `docs/migration/STATUS.md`): ++runtime-closure audit (every lazy-download path and its licence obligation), ++pinned Guix recipes + mise/just build tasks, architecture-specific assembly, ++the unprivileged launcher, relocatability proof, native verification then ++explicit WSL2 tests (Linux-filesystem install, Windows-browser localhost ++access), then publication. WSL1 and native Windows are **out of scope**; ++WSL2 is secondary and currently **unvalidated**. ++ ++## The coordinated updater (COMING) ++ ++Agreed shape: one tested, pinned combination (no independent in-place ++component upgrades), stable-version discovery through release tooling, ++signed trust metadata, transactional switch with health checks and rollback, ++and startup that works offline. Signing/trust-key provisioning and the ++release metadata format are open workstreams — publication secrets must be ++configured securely before anything is signed. ++ ++## Distribution obligations already honoured ++ ++- No third-party binaries are vendored in the repository; tools are fetched ++ from upstream under their own licences (table in `NOTICE`: MIT/GPL v2+/LGPL ++ v3 as applicable). ++- The archive-served FastQC lane still respects FastQC's GPL v3 — the ++ release asset and its licence obligations travel together (see ++ `docs/compliance/vendored-archives.md`). ++- The standalone release must ship the same component/licence inventory ++ discipline — that inventory is a release artefact, not an afterthought. ++ ++## Maintainer checklist for a release (today) ++ ++1. `just ci` green on the release commit (the proof lane). ++2. CHANGELOG.md entry; `docs/release-notes/` note for significant releases. ++3. Tag (immutable by ruleset), push, verify the source archive. ++4. If a pinned tool archive changed: re-verify sha256 records in ++ `config/defaults/tool_versions.yml` match what the release serves. ++5. Wiki [Status and Roadmap](Status-and-Roadmap) still matches reality ++ (update IN PLACE/COMING markers if the release moves any). +diff --git a/docs/wikis/Maintainers--Steward-Track.md b/docs/wikis/Maintainers--Steward-Track.md +new file mode 100644 +index 0000000..cc488ba +--- /dev/null ++++ b/docs/wikis/Maintainers--Steward-Track.md +@@ -0,0 +1,111 @@ ++ ++ ++# Steward track ++ ++**Status: IN PLACE** as a working maintenance regime (since Milestone 2, ++2026-09). This page is for whoever merges, releases, and answers for the ++repository. ++ ++## The gates (all run locally and in CI) ++ ++| Gate | Command | Authority | ++|---|---|---| ++| Strict typecheck | `bun run typecheck` (in `frontend/`) | 0 errors, gated | ++| Unit + integration tests | `bun test` | gated (DOM-less lane) | ++| Benchmarks | `bun run bench/` | informational, no gate | ++| Combined pre-push | `bun run check` | the everything lane | ++| Licence headers | `scripts/check-spdx.sh` | gated | ++| Formatting | `scripts/check-format.sh` | gated | ++| Lint (tsc semantics + shell) | `scripts/check-lint.sh` | gated | ++| Julia suite | `just` test recipes | gated | ++| Everything CI runs | `just ci` | the proof lane | ++ ++CI (`.github/workflows/ci.yml`) runs repo-hygiene first, then the pinned ++Julia and frontend jobs — the same commands as local, by design. ++ ++## Review and merge conventions ++ ++- Commit convention is enforced (the commit gate grades every non-merge ++ commit a PR proposes — merge commits themselves are exempt since #37/#45's ++ fix, so GitHub's "Update branch" button no longer reds a conforming PR). ++- Branch/PR shape: `CONTRIBUTING.md`; the PR template adapts the RSR one to ++ the bun/frontend gates (ABI/FFI items dropped as not applicable). ++- Dependabot (github-actions + bun/frontend ecosystems) auto-merges **on ++ green** — a broken bump fails CI and stays open, never lands red. Do not ++ add Dependabot as a ruleset bypass actor (that would skip required ++ security checks — a documented anti-pattern in the estate records). ++- Required status checks embed toolchain names; when bumping the Julia pin ++ remember the known trap (issue #38): update the required-check names in the ++ same change or every PR deadlocks. ++ ++## Upstream relations (the fork discipline) ++ ++This repository is a **fork of JoshuaJewell/MetaManifold-WebUI** and the ++relationship is contractual in practice: ++ ++1. **Base is always `hyperpolymath`** for day-to-day work. Never push ++ directly to `joshuajewell`. ++2. **The fork tracks application changes; it does not fork the science.** ++ Pipeline semantics (DADA2/swarm/vsearch behaviour) are upstream's domain; ++ fixes there flow *to* upstream as patch series/PRs. ++3. Upstream PRs are prepared as a reviewed series (the owner-review record in ++ `docs/owner-review-2026-09-25.md` documents the PR flow and the four ++ options offered to the origin owner). Upstream-facing commitments: ++ - Issues #1 (statistics umbrella) and #2 (symbolic engine, BLOCKED) must ++ never be closed prematurely. ++ - Closed milestone issues mean *the milestone closed* — deferred designs ++ are preserved in `docs/milestones/02-deferred-issues.md` and ++ `docs/issues/milestone3/`, not discarded. ++ - Nothing from the origin repository was ever rewritten; divergence is ++ additive. ++4. Licence asymmetry is deliberate (`NOTICE`): upstream-authored = ++ `AGPL-3.0-only`, fork-authored = `MPL-2.0` / `CC-BY-SA-4.0`. Do not ++ "harmonise" headers. ++ ++## Milestone and issue discipline ++ ++- Milestones document what the repository *does* (audit close-outs in ++ `docs/audit/`), not what it wishes. ++- The project board is maintained via GraphQL (the classic-projects deprecation ++ is already absorbed — see `docs/milestones/01-project-board-graphql.md`). ++- Scientific-value and difficulty labels on analysis issues are load-bearing ++ for sequencing the deferred queue ([Status and Roadmap](Status-and-Roadmap) ++ "Near-term order of work"). ++ ++## Release mechanics (today) ++ ++Tags + CHANGELOG + source. The full release/distribution story is ++[Releases and Distribution](Maintainers--Releases-and-Distribution). `v0.1.0` ++notes: `docs/release-notes/v0.1.0.md`. ++ ++## Settings and repository surface ++ ++- **Autolink references** are fully specified (four tiers: lineage/estate, ++ upstream tools, toolchain, registries) in ++ `docs/integration/autolink-references.md` — paste-ready for ++ Settings → Autolink references. Autolink application needs Administration ++ permission; the specification file is the source of truth for whoever has ++ it. ++- Issue templates carry the fork/upstream scope split; CODEOWNERS routes to ++ `@hyperpolymath`. ++- The wiki (this documentation) is sourced from `docs/wikis/` and published ++ to `MetaManifold-WebUI.wiki.git` — review wiki content in PRs like any ++ other change (`docs/wikis/README.md`). ++ ++## Coming on this track ++ ++- **COMING:** coverage gate — lands against a recorded baseline ++ (`docs/testing/coverage.md`), not an arbitrary number; OpenSSF Best ++ Practices + Scorecard enrolment (badges only on the day); the DOM test-lane ++ decision (playwright vs harness) before growing the e2e set; benchmark ++ stability window → non-gating regression alert. +diff --git a/docs/wikis/Maintainers.md b/docs/wikis/Maintainers.md +new file mode 100644 +index 0000000..8f9c470 +--- /dev/null ++++ b/docs/wikis/Maintainers.md +@@ -0,0 +1,54 @@ ++ ++ ++# Maintainers ++ ++This section has **two tracks**, because "maintaining the platform" means two ++different jobs with two different risk registers: ++ ++- **[Operator track](Maintainers--Operator-Track)** — you run MetaManifold ++ *for people*: installs, reference databases, toolchain updates, backups, ++ capacity, the day something refuses at 09:00 before a batch is due. Lab ++ IT, platform engineers, self-hosting PIs. ++- **[Steward track](Maintainers--Steward-Track)** — you maintain the ++ *repository*: CI gates, reviews and merges, releases, the relationship with ++ the upstream origin design, and the estate's compliance machinery. Repo ++ maintainers, OSS stewards, the hyperpolymath estate. ++ ++A small deployment (one lab machine, one maintainer) may wear both hats — ++read both tracks; the risk registers are simply separate. ++ ++Shared across both: ++ ++- **[Releases and Distribution](Maintainers--Releases-and-Distribution)** — ++ what a "release" is today (honest answer: tags + source) and the **COMING** ++ standalone archives, WSL2 validation and coordinated updater. ++- **[Compliance and Estate](Maintainers--Compliance-and-Estate)** — RSR and ++ standards alignment, the licence split (AGPL upstream / MPL+CC-BY-SA ++ fork), SPDX gates, and the autolink/settings elaboration. ++ ++## What is here now vs what is coming (maintainer view) ++ ++**IN PLACE:** pinned dual-lane toolchain (mise + Guix, sha256 pipeline ++tools); CI (hygiene, pinned tools, commit gate, Dependabot auto-merge-on-green); ++the `Justfile` command surface; the compliance record in `docs/compliance/`; ++single-user local serving. ++ ++**COMING:** standalone offline release builders (policy exists in ++`packaging/`, no builder), WSL2 validation, the coordinated signed updater, ++multiuser/authenticated deployment, coverage gate (deliberately deferred to a ++recorded baseline), OpenSSF enrolment (badges only after), the DOM test lane ++decision. ++ ++**BLOCKED / held:** nothing on the maintainer side is blocked; the analysis ++layer's independent review (#1) is the estate-level gate that keeps ++statistical claims at PARTIAL. Board: ++[Status and Roadmap](Status-and-Roadmap). +diff --git a/docs/wikis/README.md b/docs/wikis/README.md +new file mode 100644 +index 0000000..5252536 +--- /dev/null ++++ b/docs/wikis/README.md +@@ -0,0 +1,60 @@ ++ ++ ++# Project wikis — source of truth ++ ++This directory is the **source** of the GitHub wiki for MetaManifold-WebUI ++(`hyperpolymath/MetaManifold-WebUI.wiki.git`), per the RSR convention that ++wiki content is versioned in the repository and synchronised to the forge. ++ ++## Format — BerryWiki page format ++ ++Every page is ordinary GitHub-flavoured Markdown that **may** begin with a ++hidden metadata comment in the ++[berrywiki](https://github.com/metadatastician/berrywiki) page format: ++ ++```markdown ++ ++``` ++ ++GitHub strips the comment when rendering, so the wiki stays fully usable with ++nothing but `git` and a browser. Identity lives in `id` (stable across ++renames); hierarchy lives in `parent`; the `--` in filenames ++(`Users--Install-and-First-Run.md`) is a human-readable title path, not ++structure. `_Sidebar.md` is generated from the tree and must not be hand-edited ++beyond regenerating it. ++ ++The structure and status conventions used here (IN PLACE / PARTIAL / COMING / ++BLOCKED) are described on the wiki's `Home` page. ++ ++## Layout of this source ++ ++- One `.md` file per wiki page, named exactly as the wiki page name ++ (GitHub wiki page = filename without `.md`). ++- `_Sidebar.md` — the generated sidebar (GitHub control file). ++- No other formats; the wiki is plain Markdown by design. ++ ++## Synchronisation ++ ++Changes made here must be pushed to the forge wiki: ++ ++```bash ++# from a checkout of https://github.com/hyperpolymath/MetaManifold-WebUI.wiki.git ++cp docs/wikis/*.md /path/to/MetaManifold-WebUI.wiki/ ++cd /path/to/MetaManifold-WebUI.wiki ++git add -A && git commit -m "docs(wiki): sync from docs/wikis" && git push ++``` ++ ++The two sides must stay byte-identical; `docs/wikis/` is the reviewable copy ++(pull requests review wiki content like any other change) and the `.wiki.git` ++push is the publish step. Last sync: 2026-09-26. +diff --git a/docs/wikis/Status-and-Roadmap.md b/docs/wikis/Status-and-Roadmap.md +new file mode 100644 +index 0000000..f9356d1 +--- /dev/null ++++ b/docs/wikis/Status-and-Roadmap.md +@@ -0,0 +1,107 @@ ++ ++ ++# Status and roadmap ++ ++**The board.** Everything MetaManifold claims to do is marked here as IN ++PLACE, PARTIAL, COMING or BLOCKED (legend on [Home](Home)). Dates are ++Europe/London. The repository tree is the final arbiter; where this page and ++the code disagree, the code has moved and this page owes an update. ++ ++## Pipeline — raw reads to tables ++ ++| Capability | Status | Notes | ++|---|---|---| ++| cutadapt primer trimming (named primer pairs, IUPAC-validated) | **IN PLACE** | `src/pipeline/tools.jl`, `config/primers.yml`; UI editor under SYSTEM | ++| DADA2 ASV lane: filter/trim, learn errors, denoise, merge, chimera cull, taxonomy | **IN PLACE** | `src/pipeline/dada2/`; R bridge pinned by `renv.lock` | ++| SWARM OTU lane: merge, dereplicate, cluster, chimera check, vsearch taxonomy | **IN PLACE** | `src/pipeline/swarm.jl`, `src/pipeline/vsearch` stage | ++| Optional cd-hit-est pre-clustering (multiplex inflation control) | **IN PLACE** | `cdhit:` config block | ++| merge_taxa + named taxonomic filter sets | **IN PLACE** | `src/pipeline/merge_taxa.jl`, `config/composition.yml` library | ++| Per-run DuckDB results store | **IN PLACE** | `src/core/duckdb_store.jl` | ++| Stage freshness skipping (mtime for files, content hash for config) | **IN PLACE** | stale stages flagged with the exact changed keys | ++| FastQC + MultiQC prefilter QC | **PARTIAL** | implemented and CI-exercised since #44; reports embed in the run view | ++| Remote (SSH) offload of the taxonomy step | **PARTIAL** | works; authorisation is solely the operator's responsibility | ++| DADA2 single-end / forward / reverse modes | **IN PLACE** | `file_patterns.mode` | ++ ++## Analysis and statistics ++ ++| Capability | Status | Notes | ++|---|---|---| ++| Alpha diversity (richness, Shannon, Simpson) + group comparisons | **IN PLACE** | significance reports its status; never silently degrades (#31's rule) | ++| Taxonomic composition bars; organism composition categories | **IN PLACE** | category sets editable in the UI | ++| Taxon overlap (Euler/UpSet), pipeline-stage read summaries | **IN PLACE** | | ++| NMDS + PERMANOVA (vegan, locked R) | **IN PLACE** | permutation exchangeability is the standing caveat | ++| Normalisation: none, rarefaction | **IN PLACE** | | ++| TSS/CSS/RSS size-factor **offsets** (exact, not aliases) | **IN PLACE** | issues #16, #61, #62 — conditions in `docs/statistics/method-conditions/scaling-and-offsets.md` | ++| Exact descriptive summaries (counts, rational proportions) | **IN PLACE** | catalogue item 1; `src/analysis/exact_summaries.jl` | ++| Numeric policy: exact / approximate / rounded, explicit modes | **IN PLACE** | `src/analysis/numeric_policy.jl` | ++| Parametric fits — NB-GLM, CLR/ILR-LM, logistic (ML with refusal states) | **PARTIAL** | implemented with R-reference tests; **independent statistical review (#1) outstanding** | ++| Zero-depth sample healing before transforms | **IN PLACE** | #58 | ++| Nonparametric tests (permutation/bootstrap) | **COMING** | catalogue item 3, approved scope only | ++| Exact statistical tests (Fisher, exact NB, permutation PERMANOVA) | **COMING** | issue #3, catalogue item 4 | ++| Multinomial / Dirichlet-Multinomial regression | **COMING** | issue #17 | ++| Occupancy models, ZINB, hurdle | **COMING** | issue #18 | ++| Constrained ordinations (RDA/CCA/CAP/dbRDA) | **COMING** | issue #19 | ++| PhILR / SBP ILR bases, balance dendrograms | **COMING** | issue #20 | ++| Advanced zero handling (glmGamPoi dispersion, Bayesian multiplicative) | **COMING** | issue #21 | ++| ANCOM-BC, ALDEx2, Songbird-style methods | **COMING** | issue #5 | ++| Symbolic formula engine | **BLOCKED** | issue #2 — waits on real-data validation of the numeric layer | ++ ++## Workbench and application ++ ++| Capability | Status | Notes | ++|---|---|---| ++| Config cascade (instance → study → group → run) + `run_config.yml` provenance | **IN PLACE** | UI editors at every level | ++| Studies / groups / runs management, jobs panel, SSE live progress | **IN PLACE** | | ++| Results explorer (filters, presets, column presets, OTU drill-down, xlsx export) | **IN PLACE** | | ++| Functional annotation, consensus rank, contamination curation, FuncDB ledger | **IN PLACE** | composite confidence is explicitly curiosity-grade | ++| Primers / databases / composition editors with validation and impact warnings | **IN PLACE** | | ++| CladeCumulus (cumulative cladistic explorer) | **COMING** | issue #6 — data structures and validation scaffolded (`src/analysis/clade_cumulus.jl`) | ++| Full Evidence Mode (epistemic editor, fibre visualiser) | **COMING** | issue #7 — epistemic core in place (`src/core/epistemic.jl`) | ++| Zenodo DOI minting | **COMING** | issue #8 | ++| Julia-authored Stipple/Vue UI | **PARTIAL** | first read-only slice landed; React remains default (`docs/migration/STATUS.md`) | ++| Standalone offline release archives (Linux x64/ARM64, WSL2) | **COMING** | agreed requirement, no builder yet (`docs/migration/STATUS.md`) | ++| Coordinated signed updater | **COMING** | same source | ++| Multi-user / authenticated remote deployment | **COMING** | local single-user only until then — do not expose the server | ++ ++## Engineering estate ++ ++| Capability | Status | Notes | ++|---|---|---| ++| Pinned toolchain (mise exact pins + Guix peer lane; sha256 pipeline tools) | **IN PLACE** | R is a documented system exception via `renv.lock` | ++| Strict TypeScript estate (0 errors; `unknown`-safe stubs) | **PARTIAL** | `skipLibCheck` exception + `alphaFig` narrowing tracked in `docs/compliance/fixme-index.md` | ++| Julia + frontend test suites, type-level assertions | **IN PLACE** | DOM-lane for Plotly-chain modules **COMING** (decision queued) | ++| Benchmarks with baselines (informational lane) | **IN PLACE** | promotion to non-gating regression alert after a stability window | ++| CI: hygiene, pinned tools, commit gate, Dependabot auto-merge-on-green | **IN PLACE** | | ++| Coverage gate | **COMING** | deliberate: gates on a recorded baseline, not an arbitrary number | ++| OpenSSF Best Practices / Scorecard badges | **COMING** | enrolment first, badges on the day (never before) | ++| Independent statistical review of the analysis layer | **COMING** | issue #1's acceptance gate; the reason PARTIAL marks exist above | ++ ++## Near-term order of work (as tracked) ++ ++1. Independent statistical review gate (#1) — unlocks the rest of the ++ catalogue honestly. ++2. Exact tests (#3) and the milestone-3 deferred suite (#17–21) in their ++ approved dependency order (`docs/issues/milestone3/`). ++3. Stipple UI parity and the standalone-release workstream ++ (`docs/migration/STATUS.md` workstreams A and B). ++4. CladeCumulus + Evidence Mode (#6–7), Zenodo (#8). ++5. Symbolic engine (#2) only after the numeric layer is validated — the ++ BLOCKED sign is policy, not neglect. ++ ++## Reading the history ++ ++- Closed milestone issues mean *the milestone closed*, not "every deferred ++ idea shipped" — deferred designs are preserved in ++ `docs/milestones/02-deferred-issues.md` and `docs/issues/milestone3/`. ++- Issues #1 and #2 are the two that must never be closed prematurely (the ++ owner-review record makes this explicit). +diff --git a/docs/wikis/Users--Analysis-and-Statistics-Today.md b/docs/wikis/Users--Analysis-and-Statistics-Today.md +new file mode 100644 +index 0000000..795a653 +--- /dev/null ++++ b/docs/wikis/Users--Analysis-and-Statistics-Today.md +@@ -0,0 +1,112 @@ ++ ++ ++# Analysis and statistics today ++ ++**Status: PARTIAL by policy** — everything on this page is implemented, but ++the inference layer awaits issue #1's independent statistical review. Read ++results as "computed as documented", not "reviewed". The deep reasoning lives ++in [Deep Dives](Deep-Dives); this page is the user-facing contract. ++ ++## The honesty rule ++ ++> If a method cannot run validly on your data, MetaManifold returns an ++> **unsuccessful state** — a named refusal — and no number at all. It will ++> never substitute a plausible-looking statistic. ++ ++This is enforced in code and by test (a guard test fails if placeholder ++statistics are ever reintroduced). When a significance test is requested ++while R is busy, the answer is "unknown", not "not significant" (#31's rule). ++ ++## What is computed today ++ ++| Analysis | What it gives you | Where it comes from | ++|---|---|---| ++| Richness, Shannon, Simpson | per-sample descriptive indices | `src/analysis/diversity.jl` | ++| Group comparisons of alpha diversity | test statistic + p-value **with status** (Kruskal–Wallis default; paired option) | `src/analysis/analysis.jl` | ++| Exact descriptive summaries | counts as exact integers; proportions as exact rationals — no inference claimed | `src/analysis/exact_summaries.jl` | ++| Parametric fits — `nb_glm`, `clr_lm`, `ilr_lm`, `logistic` | per-feature maximum-likelihood estimates with diagnostics **or** refusal | `src/analysis/estimation.jl` | ++| Normalisation: none, rarefaction | chart-facing count scaling | `diversity.jl` | ++| TSS / CSS / RSS | size-factor **offsets** for the fits (depth modelled, counts unchanged) | `src/analysis/scaling.jl` | ++| Taxonomic composition bars, organism composition categories | per-sample or pooled composition | analysis + composition modules | ++| Taxon overlap | proportional Euler / UpSet | `analysis.jl` | ++| NMDS (Bray–Curtis), PERMANOVA | ordination and group test via the locked R runtime (vegan) | `analysis.jl` | ++| Pipeline-stage read summaries | retention per stage | pipeline stats | ++ ++## Reading a parametric fit ++ ++Four methods, four response types (never mixed): `nb_glm` takes raw counts ++(negative binomial, log link, R `MASS::glm.nb`); `clr_lm`/`ilr_lm` take ++CLR/ILR-transformed values (Gaussian LM); `logistic` takes 0/1 presence. The ++fit reports, per feature: ++ ++- convergence and identifiability status — an **unsuccessful state** replaces ++ the estimates when either fails; ++- boundary estimates (e.g. fitted dispersion at a limit) flagged as such; ++- effect sizes and intervals by the method's published conditions; ++- Benjamini–Hochberg adjusted p-values wherever several tests are reported — ++ **mandatory**, with a DANGER banner if anyone tries to turn it off. ++ ++**Academics:** the conditions each method is held to are published *before* ++implementation and live in `docs/statistics/method-conditions/` — those ++documents are what you cite (or reproduce) when writing up which model, which ++response, which refusals. The catalogue that gates them: ++`docs/statistics/method-catalogue-v1.md`. ++ ++**Lab professionals:** the practical version is — a table with `refused` in a ++status column is telling you the assay or design does not support that ++analysis. Route it to a human decision; do not fish for a tool that will ++give you a number. ++ ++## Offsets, not substitutions ++ ++TSS/CSS/RSS here are **offsets**: one positive size factor per sample (its ++log after a stated centring) handed to the count model so sequencing depth is ++*modelled*, not analysed away. The counts themselves are not replaced. An ++earlier build silently aliased these to relative abundances — that behaviour ++was removed and a regression test now forbids it. Offsets are **not** a ++compositional solution; for that, see "coming" below. ++ ++## Exact means exact — and only where it can be ++ ++Counts and proportions in the descriptive layer carry exact precision ++(integers beyond 2⁵³−1 included; proportions as rationals). Nothing floats ++into that claim silently: a float-derived input is labelled approximate ++wherever shown. Higher precision (any `BigFloat`) is *not* exactness. What ++exactness does **not** extend to: inference. P-values and fits are ++approximate by nature and say so. Deep end: ++[Deep Dives — Exact Arithmetic](Deep-Dives--Exact-Arithmetic). ++ ++## What is coming — and how it is gated ++ ++Every method below is **COMING** (specified in `docs/issues/milestone3/`, ++approved in scope, not implemented). Each must publish its supported / ++unsupported conditions *before* implementation — response types, design ++assumptions, zero handling, overdispersion and depth policy, uncertainty ++method, multiple-testing policy (BH remains mandatory), diagnostics and ++computational limits — and only then be written: ++ ++1. Nonparametric tests (permutation/bootstrap; catalogue item 3). ++2. Exact statistical tests — Fisher's exact, exact NB, permutation PERMANOVA ++ (#3, catalogue item 4) — for small-n and rare-taxon work. ++3. Multinomial / Dirichlet-Multinomial regression (#17). ++4. Occupancy models, zero-inflated NB, hurdle (#18). ++5. Constrained ordinations — RDA/CCA/CAP/dbRDA with permutation tests (#19). ++6. PhILR / SBP ILR bases with balance dendrograms (#20). ++7. Advanced zero handling — glmGamPoi dispersion, Bayesian multiplicative ++ replacement (#21). ++8. ANCOM-BC, ALDEx2, Songbird-style compositional methods (#5). ++ ++**BLOCKED:** a symbolic formula engine (#2) — deliberately, until the numeric ++layer has passed real-data validation. ++ ++The board: [Status and Roadmap](Status-and-Roadmap). +diff --git a/docs/wikis/Users--Configuration-Reference.md b/docs/wikis/Users--Configuration-Reference.md +new file mode 100644 +index 0000000..3b4abe8 +--- /dev/null ++++ b/docs/wikis/Users--Configuration-Reference.md +@@ -0,0 +1,254 @@ ++ ++ ++# Configuration reference ++ ++**Status: IN PLACE** — this is the complete key-by-key reference for ++pipeline configuration. (This page replaces the configuration chapters the ++origin README used to carry, so the README can stay readable.) ++ ++## The cascade ++ ++Settings cascade: each level overrides the one above it; any key you omit is ++inherited from the nearest ancestor. The fully merged result is written to ++`projects/{name}/{run}/run_config.yml` at runtime — **that file is the single ++place to see exactly what a run used.** ++ ++| File | Purpose | ++|---|---| ++| `config/defaults/` | canonical defaults for every setting; do not edit | ++| `config/composition.yml` | composition library: named taxonomic filters + category sets | ++| `config/presets/` | saved table-view filter presets (written from the Tables view) | ++| `config/databases.yml` | database URIs and optional local paths (Databases page under SYSTEM) | ++| `config/primers.yml` | primer sequences and pair definitions (Primers page under SYSTEM) | ++| `config/tools.yml` | tool binary paths (cutadapt, FastQC, MultiQC, vsearch, cd-hit-est) | ++| `config/pipeline.yml` | machine-level overrides (lowest user-editable precedence) | ++| `data/{name}/pipeline.yml` | study-level overrides | ++| `data/{name}/{group}/pipeline.yml` | group-level overrides (intermediate directories) | ++| `data/{name}/{run}/pipeline.yml` | run-level overrides (highest precedence) | ++| `projects/{name}/{run}/run_config.yml` | generated merged config (provenance); do not edit | ++ ++Each `pipeline.yml` stub carries a comment block explaining its level's role. ++Write only the keys you want to change. ++ ++## Databases (`config/databases.yml`) ++ ++The one place database URIs are managed (also editable from the Databases ++page): ++ ++```yaml ++databases: ++ dir: "./databases" ++ pr2: ++ dada2: ++ uri: "https://..." # DADA2-format FASTA (downloaded on first use) ++ local: ~ # set to a local path to skip download ++ vsearch: ++ uri: "https://..." # vsearch-format FASTA ++ local: ~ ++``` ++ ++Per database: the `dada2` and `vsearch` source URIs, `local:` override, ++`remote_path` (dada2 only, file already on the remote taxonomy host), the ++ordered taxonomy `levels`, the `vsearch_format` parser selector (`pr2` = ++pipe-separated; anything else parses generically), and `corrections`. ++Adding/removing whole databases is supported. Removing or renaming one (or ++changing `levels`) is allowed but the save **reports which studies it ++affects** — including studies that inherit the database without naming it. ++ ++> Both formats of one database must come from the same reference release: the ++> dual-classifier consensus compares labels by string equality, so mixed ++> releases score real agreements as disagreements. The editor warns on ++> version-token mismatches (a filename heuristic — cannot warn when URIs ++> carry no version). ++ ++## Primers (`config/primers.yml`) ++ ++```yaml ++Forward: ++ PrimerF: "CCAGCASCYGCGGTAATTCC" ++ ++Reverse: ++ Primer1R: "ACTTTCGTTCTTGATYRA" ++ Primer2R: "DCTKTCGTYCTTGATYRA" ++ ++Pairs: ++ - PrimerPair1: [PrimerF, Primer1R] ++ - PrimerPair2: [PrimerF, Primer2R] ++``` ++ ++Shared primers are deduplicated in the cutadapt invocation (cutadapt ++complains about duplicates — if you need duplicates, define the same sequence ++under a second name). The Primers page validates each sequence against the ++IUPAC base set as you type and validates the whole document before writing: ++a pair naming a missing primer is rejected and the file is left untouched. ++Removing/renaming a pair that studies still reference is permitted but the ++save reports every dangling reference. Renaming a primer carries its pairs. ++ ++## cutadapt ++ ++```yaml ++cutadapt: ++ primer_pairs: [PrimerPair1, PrimerPair2] # names from primers.yml ++ min_length: 200 # -m: discard shorter reads after trimming ++ discard_untrimmed: true # --discard-untrimmed ++ cores: 0 # -j: 0 = auto ++ quality_cutoff: ~ # -q 3' trim; null disables ++ error_rate: ~ # -e adapter mismatch rate; null = cutadapt default ++ overlap: ~ # -O min adapter overlap; null = default ++ optional_args: "" # passed verbatim ++``` ++ ++## DADA2 ++ ++```yaml ++dada2: ++ file_patterns: ++ mode: "paired" # paired | forward | reverse ++ ++ filter_trim: # DADA2 filterAndTrim() ++ trunc_q: 2 ++ trunc_len: [220, 220] # [forward, reverse] ++ max_ee: [3, 3] ++ min_len: 175 ++ max_n: 0 ++ match_ids: true ++ rm_phix: true ++ ++ dada: # learnErrors() and dada() ++ seed: 123 ++ nbases: 200000000 ++ max_consist: 15 ++ pool_method: "pseudo" # none | pseudo | true ++ ++ merge: # mergePairs(), paired mode only ++ min_overlap: 20 ++ max_mismatch: 0 ++ trim_overhang: true ++ ++ asv: # length filtering + chimera removal ++ band_size_min: 200 # null skips length filtering ++ band_size_max: 430 ++ denovo_method: "consensus" # consensus | pooled | per-sample ++ ++ taxonomy: # assignTaxonomy() ++ database: pr2 # key into databases.yml ++ multithread: 4 ++ min_boot: 0 # 0-100 bootstrap floor ++ remote: # optional SSH offload of assignTaxonomy() ++ host: ~ # user@hostname — authorisation is YOURS to ensure ++ identity_file: ~ ++ rscript: "Rscript" ++ staging_dir: "/absolute/path/on/server" ++ ++ output: ++ seq_table_prefix: "seqtab_nochim" ++ fasta_prefix: "asvs" ++ taxa_prefix: "taxonomy" ++``` ++ ++Rank names come from `databases.yml` (`levels:`), not from here. ++ ++## vsearch, cd-hit-est, swarm ++ ++```yaml ++vsearch: ++ identity: 0.75 # --id ++ query_cov: 0.8 # --query_cov ++ maxaccepts: ~ ++ maxrejects: ~ ++ strand: ~ # "plus" | "both" ++ optional_args: "" ++ ++cdhit: # optional pre-clustering (multiplex inflation control) ++ identity: 1 # -c ++ threads: 0 # -T: 0 = all ++ optional_args: "" ++ ++swarm: # the OTU lane ++ differences: 1 # -d ++ threads: 0 ++ chimera_check: true # vsearch --uchime_denovo first ++ min_abundance: 2 # --minsize ++ fastq_minovlen: 20 ++ identity: 0.97 # mapping reads back to OTU seeds ++ optional_args: "" ++``` ++ ++The `vsearch:` block configures taxonomy alignment for **both** lanes. ++ ++## merge_taxa and the filter library ++ ++```yaml ++merge_taxa: ++ filters: ++ - "protist_filter.yml" # -> merged/protist_filter.csv ++``` ++ ++`merged.csv` (unfiltered) is always written; each `filters` entry produces ++one more CSV. Entries name filters from the `filters:` library of ++`config/composition.yml`. A filter file: ++ ++```yaml ++databases: [pr2] # omit to apply regardless of active database ++mappings: # optional column remapping before filtering ++ - source_column: Division ++ target_column: Supergroup ++ values: { Rhizaria: Rhizaria, Alveolata: Alveolata } ++filters: ++ - column: Domain ++ pattern: "Bacteria|Archaea" ++ regex: true # false (default) = substring ++ action: exclude # exclude (default) | keep ++remove_empty: [Domain] # drop rows blank/"NA" in these columns ++``` ++ ++The bundled library (PR2-shaped): `bacteria_archaea`, `environmental_protozoa`, ++`fungi`, `helminths`, `parasitic_protozoa`, `plants_invertebrates`, `protist`, ++`vertebrates`. Category sets (protozoa, helminths, fungi, host, …) reference ++these filters by name and back the Composition view. ++ ++## analysis and annotation defaults ++ ++```yaml ++analysis: ++ exclude_categories: ++ - {set: contamination, category: Contaminant, apply_to: [diversity, taxa, venn]} ++ normalisation: none # none | rarefaction | rss ++ normalisation_depth: 0 # 0 = auto (min positive library size) ++ alpha: ++ show_points: true ++ annotate_significance: false ++ pairwise_brackets: false ++ paired_samples: false ++ significance_test: "kruskal_wallis" ++ nmds: ++ max_stress: 0.2 # warn above this stress ++ ++annotation: ++ max_rank: "species" ++``` ++ ++Per-chart choices (rank, relative/absolute) are interactive UI state, not ++config keys. **Note the `normalisation` key above is the chart-facing ++normalisation** (none/rarefaction/relative sum scaling for display); the ++statistical layer's TSS/CSS/RSS **offsets** are a separate, analysis-config ++concern — see [Analysis and Statistics Today](Users--Analysis-and-Statistics-Today) ++and [Deep Dives — Compositional Statistics](Deep-Dives--Compositional-Statistics). ++ ++## Coming in configuration ++ ++- **COMING:** analysis-config keys for the deferred methods (exact tests, ++ multinomial/DM, occupancy, constrained ordination, PhILR/SBP, advanced zero ++ handling) — the schemas will grow with issues #3 and #17–21. New keys will ++ be refused until their method conditions are published, by the same rule ++ that governed the current set. +diff --git a/docs/wikis/Users--Exploring-Results.md b/docs/wikis/Users--Exploring-Results.md +new file mode 100644 +index 0000000..3dc0cd1 +--- /dev/null ++++ b/docs/wikis/Users--Exploring-Results.md +@@ -0,0 +1,97 @@ ++ ++ ++# Exploring results ++ ++**Status: IN PLACE** — these views are the daily working surface of the ++application. (Screenshots are **COMING** to this page; the origin README ++referenced UI captures that were never committed to the repository. The ++descriptions below are complete without them.) ++ ++A run's page is the working surface: run-level configuration, per-stage ++launch controls (full pipeline or any single stage, DADA2 substages ++included), and per-stage status. Long jobs report progress over a live ++server-sent event stream; the jobs panel watches or cancels them. ++ ++## QC ++ ++Raw-read QC (FastQC aggregated by MultiQC) and the DADA2 quality, ++denoising, merging and taxonomy diagnostics are embedded in the UI, each ++with its relevant configuration alongside and a re-run control. This is the ++first place to look when retention numbers surprise you. ++ ++## Results explorer ++ ++The run's merged taxonomy-and-count table, and derived tables, browsed ++interactively: ++ ++- per-column filters — text search, numeric range, include/exclude lists — ++ and a global text filter; ++- sorting, configurable pagination, column visibility with taxonomy-source ++ presets (VSEARCH-only, DADA2-only, all) and a switch for the per-sample ++ count columns; ++- sequences carry BLAST links; OTU rows expand to their constituent ++ sequences; ++- filters save as named presets (`config/presets/`) and reapply; tables ++ export to Excel (`.xlsx`) or save back into the run. ++ ++**Lab professionals:** column presets + saved filters give you standard ++views per assay — build the view once, share the preset. The OTU ++drill-down is how you audit what a cluster contains before reporting it. ++ ++## Annotation ++ ++The annotation view applies a functional database (`funcdb`) to the merged ++table, independently per classifier (VSEARCH and DADA2). For each sequence: ++functional metadata (function, associated organism and material, ++environment, pathogen status, notes) down to a configurable maximum rank; a ++**consensus rank** (finest rank where the classifiers agree); and a composite ++confidence score (DADA2 bootstrap at that rank × vsearch percent identity) — ++which is deliberately labelled curiosity-grade, not a calibrated probability. ++ ++Curation, in place: ++ ++- **Contamination tagging** — yes / no / unassigned per taxon at a rank, with ++ a live summary of affected reads. ++- **Manual BLAST assignment** — override one sequence inline. ++- **FuncDB ledger** — new functional entries, prefilled from the selected row, ++ append to a ledger available to later annotation runs. Your edits survive ++ re-annotation. ++ ++**Lab professionals:** contamination tagging is your decontamination audit ++trail — flags apply to every row sharing the rank and taxon, and the ++affected-read summary is the number to put in the QC record. ++ ++## Composition ++ ++Per-sample or pooled stacked bars of ASV/OTU counts by biological category ++(protozoa, helminths, fungi, host, plants, invertebrates — the `default` ++set in `config/composition.yml`; sets and their named filters are editable ++from the Compositions page under SYSTEM). A category summary precedes the ++chart; a quality filter caps unresolved taxonomic placeholders. ++ ++## Cross-run comparison ++ ++Across runs within a study: alpha-diversity comparison boxplots with ++significance testing, taxon overlap as proportional Euler or UpSet plots, ++NMDS ordination and PERMANOVA, and pipeline-stage read summaries — see ++[Analysis and Statistics Today](Users--Analysis-and-Statistics-Today) for ++what the statistics behind those charts claim. ++ ++## Coming to this area ++ ++- **COMING:** CladeCumulus (#6) — a cumulative cladistic explorer with ++ epistemic colour coding (scaffolded in `src/analysis/clade_cumulus.jl`); ++ Full Evidence Mode (#7) — epistemic editor and fibre visualiser. ++- **COMING:** Euler/UpSet visualisations and chart-editor parity in the ++ Stipple UI migration (`docs/migration/STATUS.md` names these as the major ++ parity risks). +diff --git a/docs/wikis/Users--For-Academics.md b/docs/wikis/Users--For-Academics.md +new file mode 100644 +index 0000000..67c606c +--- /dev/null ++++ b/docs/wikis/Users--For-Academics.md +@@ -0,0 +1,107 @@ ++ ++ ++# For academics ++ ++**Status: IN PLACE** as a working method, with the review caveat below. This ++page is for readers who will publish, teach, or peer-review work produced ++with MetaManifold: what to cite, what to disclose, and what the software will ++not let you overclaim. ++ ++## The reproducibility record is already written for you ++ ++Every run materialises its complete configuration to ++`projects/{study}/{run}/run_config.yml` — the merged truth of the cascade ++(instance → study → group → run). For a methods section you need three ++artefacts, and two are automatic: ++ ++1. **`run_config.yml`** — every parameter the pipeline used (attach as ++ supplementary material). ++2. **The toolchain pins** — `mise.toml` (Julia/Bun/Node/just exact ++ versions), `renv.lock` (exact R package versions, restored byte-locked), ++ `config/defaults/tool_versions.yml` (sha256-pinned binaries for cutadapt, ++ FastQC, MultiQC, vsearch, cd-hit-est, swarm). ++3. **Your taxonomy reference releases** — `config/databases.yml` records the ++ URIs; record the release versions in your own lab book too (the software ++ warns on mismatched releases between the DADA2 and vsearch formats of one ++ database precisely because consensus comparison requires one release). ++ ++`docs/reproducibility.md` is the toolchain source of truth; honest gaps ++(e.g. the Guix lane carrying functional equivalents rather than binary ++identity) are documented there rather than smoothed over. ++ ++## Citing ++ ++- `CITATION.cff` at the repository root is the citable record (with ORCID ++ and references) — most reference managers ingest it directly. ++- **Origin design:** always credit Joshua Jewell's MetaManifold design; this ++ fork extends it. The lineage is stated in `NOTICE` and drawn visually in ++ [Deep Dives — Design Progression](Deep-Dives--Design-Progression). ++- **The science you are running** is the upstream tools' as well: DADA2 ++ (Callahan et al.), swarm, vsearch, cutadapt — cite their papers; the ++ software orchestrates, it does not replace their methods. ++- **COMING:** Zenodo DOI minting (#8) so a completed study can be released ++ with a citable, archived record from the workbench. Until then, archive the ++ `projects/` artefacts yourself. ++ ++## Method disclosure — how to word it ++ ++MetaManifold's statistical layer distinguishes three kinds of number, and ++your paper should too (the distinctions are load-bearing, see ++[Deep Dives — Exact Arithmetic](Deep-Dives--Exact-Arithmetic)): ++ ++- **Exact** — integer counts and rational proportions (descriptive layer). ++- **Approximate** — every fitted or floating quantity, including ++ high-precision ones. Higher precision is not exactness. ++- **Rounded** — a display rendering of an approximation. ++ ++For inference, report *which* model (`nb_glm`, `clr_lm`, `ilr_lm`, ++`logistic`), *which* normalisation story (none/rarefaction for display; ++TSS/CSS/RSS **offsets** for depth modelling — never "normalised to relative ++abundances" for these fits, because that is not what happens), and BH ++correction (mandatory whenever multiple tests are reported). ++ ++> **Disclose refusals.** When the software returns an unsuccessful state ++> (non-convergence, non-identifiable design, boundary pathology, violated ++> precondition), report the feature as *not tested* — not by dropping it ++> silently and not by switching to a method that will produce a number. The ++> refusal is the result. ++ ++## Small-n and rare-taxon work — the honest limit ++ ++**Academics (particularly clinical and low-biomass):** the asymptotic tests ++available today degrade with tiny n and sparse features. The exact layer ++that is designed for this — Fisher's exact, exact negative binomial, ++permutation PERMANOVA (#3) — is **COMING, not present**. Until it lands, the ++supported response to "2/3 cases vs 0/20 controls" is a descriptive summary ++(exact counts, exact proportions) plus the words "no valid inferential test ++was computed", which the method catalogue explicitly endorses as the correct ++default when valid inference information is missing. ++ ++## The review caveat, stated once ++ ++The inference layer (fits, offsets) is implemented with strong tests — known ++answers, an independent R reference comparison, negative controls for every ++refusal path — but has **not** passed issue #1's independent statistical ++review. For high-stakes claims, either wait for the review or have a ++statistician read `docs/statistics/method-conditions/` against your design ++before you trust a p-value. The software will compute exactly what the ++conditions say; whether the model suits your study design remains your ++judgement (and the conditions list the assumption most likely to be wrong). ++ ++## Teaching ++ ++The bundled `data/MiSeq_SOP/` dataset (mothur's MiSeq SOP) runs the full ++path on known data — good for practicals. The refusal behaviour is also ++pedagogically useful: students can watch a method decline rather than ++fabricate. CladeCumulus (#6, **COMING**) is aimed at teaching cladistic ++reading of composition. +diff --git a/docs/wikis/Users--For-Lab-Professionals.md b/docs/wikis/Users--For-Lab-Professionals.md +new file mode 100644 +index 0000000..fc09d7d +--- /dev/null ++++ b/docs/wikis/Users--For-Lab-Professionals.md +@@ -0,0 +1,105 @@ ++ ++ ++# For lab professionals ++ ++**Status: IN PLACE** as a routine production tool (single-user, local ++machine). This page is for people running samples as a service: validated ++routine paths, QC gates before results leave the lab, and the curation ++records auditors ask for. ++ ++## The routine loop ++ ++1. **Receive** FASTQs → drop them into `data/{Study}/{run}/` with Illumina ++ naming. Group runs by batch/site/cohort in intermediate directories. ++2. **Configure once per assay.** Set primer pairs and database at study ++ level (or in `config/pipeline.yml` for the whole machine); everything else ++ inherits. The [Configuration Reference](Users--Configuration-Reference) ++ has the keys; the SYSTEM pages (Primers, Databases, Compositions) edit ++ them with validation. ++3. **Launch and watch the jobs panel.** Per-stage status and live progress; ++ stages can be re-run individually after a fix. ++4. **Gate on QC** (below), curate, export. ++ ++## QC gates before results leave the lab ++ ++| Gate | Where | What "pass" looks like | ++|---|---|---| ++| Raw-read quality | QC view (FastQC + MultiQC) | per-base quality sane; adapter content low after trimming | ++| Trimming retention | `pipeline_stats.csv`, stage summary | cutadapt retention matches assay expectations (primer mismatch shows here first) | ++| Denoising/merging retention | stage summary | no cliff between denoise → merge for the amplicon length | ++| Chimera rate | DADA2 diagnostics | consistent with batch history, not spiking | ++| Taxonomy hit rate | taxonomy stage outputs | assignable fraction stable vs previous runs on the same assay | ++| Classification agreement | annotation view consensus rank | disagreement between DADA2 and vsearch labels investigated before reporting | ++| Contamination | annotation view, contamination tagging | flagged taxa summarised per rank; affected-read count in the QC record | ++ ++The per-stage read accounting is the retention curve of your assay; keep a ++running record per batch and the outliers name themselves. ++ ++## Curation records that survive ++ ++- **Contamination flags** (yes / no / unassigned) attach to rank+taxon and ++ apply to every matching row, with a live affected-read summary — export the ++ summary into the batch QC record. ++- **Manual BLAST overrides** are per-sequence and explicit. ++- **FuncDB ledger** entries are append-only and persist across re-annotation; ++ user edits survive regeneration. That persistence is the audit trail: ++ what was decided, when, by whom, survives the next analyst's re-run. ++ ++## Primers and databases — the lab-owned layer ++ ++Assays live in `primers.yml` (sequences validated against the IUPAC base set ++as you type; whole-document validation before write). Reference databases ++live in `databases.yml` (URIs + optional local paths; download on first use; ++SSH `remote_path` when the taxonomy host already has the file). ++ ++- Renaming/removing a primer pair or a database that studies still reference ++ is **permitted but reported** — the save names every affected study, ++ including those inheriting a database without naming it. Read the report; ++ it is the difference between a rename and an outage. ++- **Keep the two formats of one database on the same reference release.** ++ The dual-classifier consensus is string equality of labels; mixed releases ++ turn agreements into disagreements. The editor warns on version-token ++ mismatch (filename heuristic only). ++ ++## Throughput notes ++ ++- Multi-run studies (groups) compare across runs: alpha-diversity boxplots ++ with significance, taxon overlap (Euler/UpSet), NMDS, PERMANOVA. ++- Saved table-filter presets (`config/presets/`) standardise the view per ++ assay; `.xlsx` export feeds LIMS-side reporting. ++- For heavy runs, the memory-intensive `assignTaxonomy()` can be offloaded to ++ a server via SSH (`dada2.remote:`). **You** are responsible for ++ authorisation to that host — the disclaimer in ++ `config/defaults/pipeline.yml` is deliberate. ++ ++## Deployment posture — read before putting this on a server ++ ++**IN PLACE:** local single-user operation. The current server is **not** ++multiuser-ready and must not be exposed as if it were. On a shared lab ++machine, bind to localhost (default) and treat `data/` + `projects/` as the ++backup unit. Operator-grade deployment guidance: ++[Operator Track](Maintainers--Operator-Track). ++ ++**COMING:** authenticated remote/multiuser deployment, standalone offline ++archives for managed installs, and a coordinated signed updater ++(`docs/migration/STATUS.md`). None exist today; plan for the source install. ++ ++## Statistics your service reports vs research claims ++ ++Routine service work mostly needs the descriptive layer (exact counts and ++proportions), composition views and QC — all **IN PLACE** and honest under ++review. When a client asks for differential-abundance claims, check ++[Analysis and Statistics Today](Users--Analysis-and-Statistics-Today): the ++fits exist, the review caveat applies, and the exact small-n layer (#3) is ++still **COMING**. "We cannot make that claim validly yet" is a defensible ++service answer; a fabricated p-value is not. +diff --git a/docs/wikis/Users--Install-and-First-Run.md b/docs/wikis/Users--Install-and-First-Run.md +new file mode 100644 +index 0000000..d855397 +--- /dev/null ++++ b/docs/wikis/Users--Install-and-First-Run.md +@@ -0,0 +1,106 @@ ++ ++ ++# Install and first run ++ ++**Status: IN PLACE** (the install path; standalone offline installers are ++**COMING** — `docs/migration/STATUS.md` — until then this source-install path ++is the supported one). ++ ++MetaManifold runs locally on one machine, single-user, and serves its ++workbench to your browser at `http://localhost:8080`. ++ ++## What you need before starting ++ ++| Requirement | Pin | Why | ++|---|---|---| ++| Linux or macOS (WSL2 works **PARTIAL**ly, unvalidated) | — | the orchestrator shells out to native tools | ++| Julia | 1.12.5 (pinned; auto-installed) | the pipeline engine and API server | ++| R ≥ 4.0 | system install (documented exception) | DADA2, vegan — the science lane | ++| Bun | 1.3.10 (pinned) | builds the browser UI | ++| Node | 20.20.2 (pinned) | toolchain peer | ++| just | 1.43.1 (pinned) | task runner for the dev lanes (not needed to *use* the app) | ++| Internet (first run only) | — | downloads sha256-pinned pipeline tools and reference databases | ++ ++The exact pins live in `mise.toml` (the toolchain source of truth is ++`docs/reproducibility.md`). Pipeline tools — cutadapt, FastQC, MultiQC, ++vsearch, cd-hit-est, swarm — are fetched **byte-exact** (sha256-verified) by ++`install.sh`; no third-party binary is vendored in the repository. ++ ++## Install — one command lane ++ ++```bash ++git clone https://github.com/hyperpolymath/MetaManifold-WebUI.git ++cd MetaManifold-WebUI ++bash install.sh ++``` ++ ++`install.sh` checks for Julia and R, installs the Julia and R dependencies ++(the R side through `renv.lock` — exact package versions), and locates or ++downloads each pipeline tool. It writes `config/tools.yml` with resolved ++paths. ++ ++**Alternative (reproducible shell):** with `mise` installed, ++`just bootstrap && just setup-full` gives the CI-pinned toolchain plus ++tools; with GNU Guix, `guix time-machine -C channels.scm -- shell -D -f ++guix.scm` opens the peer-lane environment. ++ ++## Start the server ++ ++```bash ++bash start.sh # builds the frontend on first run ++# → open http://localhost:8080 ++``` ++ ++Environment knobs (all optional): ++ ++| Variable | Default | Meaning | ++|---|---|---| ++| `JULIA_METAMANIFOLD_PORT` | `8080` | server port | ++| `JULIA_METAMANIFOLD_ROOT` | working directory | where `data/` and `projects/` live | ++| `JULIA_THREADS` | `8` | Julia threads | ++ ++**Academics:** pin the toolchain exactly as above and your methods section ++can honestly say "versions pinned in `mise.toml` and `renv.lock`". The ++reproducibility statement for a paper is mostly written for you — see ++[For Academics](Users--For-Academics). ++ ++**Lab professionals:** `install.sh --update` re-resolves tools when you ++refresh the install. For lab servers, read the operator track — ++[Operator Track](Maintainers--Operator-Track) — before running this on a ++shared machine. ++ ++## Verify the install ++ ++The workbench's SYSTEM sidebar (Databases, Primers, Compositions pages) is ++populated, and `GET /api/v1/capabilities` reports what the server can see ++(e.g. whether R answered). From a checkout with the dev toolchain, ++`just ci` runs every gate CI runs — useful after a move to new hardware. ++ ++## Known first-run snags ++ ++- **R not found.** R is the one documented system install (it is absent from ++ the `mise` registry). Install R, re-run `install.sh`; `renv` restores the ++ pinned packages into a project-local library via the committed `.Rprofile`. ++- **Tool download failures** (firewalled hosts): re-run — the fetcher ++ retries and names TLS failures explicitly (#51's behaviour). For air-gapped ++ hosts see the operator track. ++- **Port 8080 busy:** set `JULIA_METAMANIFOLD_PORT`. ++- Everything else: [Troubleshooting](Users--Troubleshooting). ++ ++## What is coming at this layer ++ ++- **COMING:** standalone offline release archives (Linux x86-64 and ARM64) ++ with everything bundled — users then provide only sequencing data and ++ reference databases; and the coordinated signed updater. Not built yet; ++ tracked in `docs/migration/STATUS.md`. When they land, this page gets a ++ second, shorter install path above the source one. +diff --git a/docs/wikis/Users--Troubleshooting.md b/docs/wikis/Users--Troubleshooting.md +new file mode 100644 +index 0000000..e02dfd9 +--- /dev/null ++++ b/docs/wikis/Users--Troubleshooting.md +@@ -0,0 +1,69 @@ ++ ++ ++# Troubleshooting ++ ++**Status: IN PLACE.** Failure modes with known causes, and what each refusal ++actually means. Also see `SECURITY.md` for reporting, and `CONTRIBUTING.md` ++for bug-report shape. ++ ++## Installation and startup ++ ++| Symptom | Cause | Fix | ++|---|---|---| ++| `install.sh` stops at "R not found" | R is the one documented system dependency (absent from the mise registry) | install R ≥ 4.0, re-run `install.sh`; `renv` restores pinned packages via `.Rprofile` | ++| Tool download fails | flaky network or TLS interception | re-run — the fetcher retries and names TLS failures explicitly; check `config/defaults/tool_versions.yml` sha256s if it persists | ++| Port already in use | default is 8080 | set `JULIA_METAMANIFOLD_PORT` | ++| Frontend missing / blank | first-run build did not complete | `bash start.sh` rebuilds; dev toolchain users: `just setup-full` | ++| `just ci` fails on a fresh clone | a lane missing its tool (Julia lanes without Julia, e2e without browsers) | fail-loud by design — install the named tool or run the individual lanes (`just` lists all) | ++ ++## Pipeline runs ++ ++| Symptom | Cause | Fix | ++|---|---|---| ++| A stage is marked stale after an edit | config change at a finer cascade level | intended: the tooltip lists exactly which keys changed and at which level; re-run the flagged stages | ++| cutadapt discards everything | wrong primer pair selected, or `discard_untrimmed: true` with non-matching primers | check `cutadapt.primer_pairs` against `primers.yml`; IUPAC codes matter | ++| Merging collapses in paired mode | `trunc_len` too short for the amplicon to overlap | DADA2 needs ~20 bp overlap after truncation: `trunc_len F + R ≥ amplicon + min_overlap` | ++| Taxonomy step is slow / memory-heavy | `assignTaxonomy()` with large databases | raise `dada2.taxonomy.multithread` within memory, or use the SSH offload (`dada2.remote`) | ++| Remote taxonomy fails | SSH authorisation is the operator's responsibility | verify `host`, `identity_file`, `staging_dir` and `rscript` manually as the same user | ++| OTU lane produces tiny clusters | `swarm.min_abundance` / `differences` at odds with depth | defaults (`min_abundance: 2`, `differences: 1`) are sane; deep data can raise `differences` to 2 with justification | ++| vsearch taxonomy hits everything at `identity: 0.75` | that floor is deliberately permissive for exploratory work | raise `vsearch.identity` per assay validation; the consensus rank still demands classifier agreement | ++ ++## Analysis and refusals (these are features) ++ ++| What you see | What it means | What to do | ++|---|---|---| ++| An **unsuccessful state** instead of estimates | a precondition failed: non-convergence, non-identifiable design, boundary pathology, wrong response type | read the named state; adjust the design/method per `docs/statistics/method-conditions/`; do not fish | ++| "Not Implemented" for a method | the method or its R package is genuinely absent from `renv.lock` | it is refused rather than faked (by design); the method you want may be **COMING** — see [Status and Roadmap](Status-and-Roadmap) | ++| Significance "unknown" | R was busy / unavailable when the test was requested | re-run; an empty result is never silently "not significant" (#31's rule) | ++| DANGER banner on analysis config | someone attempted to disable BH correction or set an out-of-policy advanced parameter | BH is mandatory; the banner logs the attempt. Advanced flags (custom pseudocount, epsilon, zero policy) require the documented justification | ++| A requested normalisation was "not applied" | the method name or mode did not survive validation (silent substitution is forbidden) | fix the spelling/mode in config; methods compare case-insensitively since #62 | ++| Zero-depth samples disappeared | they are healed/removed **before** transforms by design (#58) | expected; `docs/statistics/behaviour-change-zero-depth-samples.md` describes the change | ++ ++## Results and views ++ ++| Symptom | Cause | Fix | ++|---|---|---| ++| Table edits lost after re-annotation | curation is stored separately, but re-annotation with different `max_rank` reshapes rows | re-apply saved presets; contamination flags survive (they key to rank+taxon) | ++| Consensus rank very coarse | the classifiers disagree below that rank | often a reference-release mismatch — put both database formats on the same release | ++| Composition mostly "unresolved" | quality filter is capping placeholders, or the filter library does not match your taxa | edit the filter library on the Compositions page; loosen the unresolved cap for exploration | ++| `run_config.yml` differs from what I set in a UI | finer-level override wins in the cascade | the cascade is instance → study → group → run; check the finer level, and `GET .../config/overrides` lists downstream overrides | ++ ++## Getting help ++ ++- Bug reports: use the repository's issue templates (they ask for scope: ++ fork vs upstream science). The upstream science lane belongs to ++ JoshuaJewell/MetaManifold-WebUI; application behaviour to this fork. ++- Security: `SECURITY.md`. ++- "Is this a bug or a refusal?" — if the output names an unsuccessful state, ++ it is the software working. If it produced a number you do not trust, ++ that *is* a bug worth filing. +diff --git a/docs/wikis/Users--Your-First-Study.md b/docs/wikis/Users--Your-First-Study.md +new file mode 100644 +index 0000000..070042a +--- /dev/null ++++ b/docs/wikis/Users--Your-First-Study.md +@@ -0,0 +1,101 @@ ++ ++ ++# Your first study ++ ++**Status: IN PLACE** — the whole path from FASTQs to merged tables is the ++oldest part of the software and the most exercised. ++ ++## 1. Lay out your data ++ ++Paired-end Illumina files, standard naming, under `data/`: ++ ++``` ++data/MyProject/run_A/ ++ SampleName_S1_L001_R1_001.fastq.gz ++ SampleName_S1_L001_R2_001.fastq.gz ++ SampleName_S2_L001_R1_001.fastq.gz ++ SampleName_S2_L001_R2_001.fastq.gz ++ ... ++``` ++ ++- Any directory containing `.fastq.gz` files is a **run** (a leaf). ++- Directories between the study and the run are **groups** — use them for ++ multi-run studies (cohorts, sequencing batches, sites). ++- A matching project directory is created under `projects/` for outputs. ++ ++Not sure yet? The repository ships `data/MiSeq_SOP/` (the mothur MiSeq SOP ++dataset, two runs) — run it first to see the whole path on known data. ++ ++## 2. Create the study and launch ++ ++Open `http://localhost:8080`: ++ ++1. Create the study `MyProject` (the UI shows the runs it found on disk). ++2. Optionally set study-level configuration (e.g. which primer pairs cutadapt ++ should use) — leave everything else at defaults for the first run. The ++ full key-by-key reference is ++ [Configuration Reference](Users--Configuration-Reference). ++3. Launch the pipeline. The jobs panel and the live event stream show ++ progress; any individual stage (including DADA2 substages) can be ++ launched alone. ++ ++## 3. What happens, in order ++ ++``` ++FASTQs → cutadapt (primers) → [QC: FastQC/MultiQC] ++ → DADA2 lane: filter/trim, error learning, denoise, merge, ++ length filter, chimera removal, taxonomy, (cd-hit-est), vsearch ++ → SWARM lane: merge pairs, dereplicate, cluster, chimera check, ++ vsearch taxonomy ++ → merge_taxa: join tables, apply named taxonomic filters ++ → DuckDB results store (per run) ++``` ++ ++Both lanes run side by side: you get ASVs (DADA2) **and** OTUs (SWARM) from ++the same run. Stages skip themselves when their outputs are already current; ++changing a setting flags exactly the stages that would regenerate, and names ++the keys you changed. ++ ++**Lab professionals:** the per-stage read accounting (`pipeline_stats.csv` ++and the stage summary view) is your QC gate — the retention curve across ++stages is where library problems show up first. See ++[For Lab Professionals](Users--For-Lab-Professionals). ++ ++**Academics:** the merged configuration used for the run is written to ++`projects/{study}/{run}/run_config.yml` — that file *is* your provenance ++record for methods sections. See [For Academics](Users--For-Academics). ++ ++## 4. Collect your outputs ++ ++Everything for a run lives under `projects/{study}/{run}/`: ++ ++| Path | Contents | ++|---|---| ++| `cutadapt/` | trimmed FASTQ pairs + logs | ++| `QC/` | per-file FastQC reports, `multiqc_report.html`, logs | ++| `dada2/Tables/` | `seqtab_nochim.csv` (ASV counts), `asvs.fasta`/`asvs.csv`, `taxonomy.csv`, `taxonomy_bootstraps.csv`, `taxonomy_combined.csv`, `tax_counts.csv`, `asv_counts.csv`, `pipeline_stats.csv` | ++| `dada2/Figures/`, `Checkpoints/`, `Logs/` | quality/error PDFs, RData stage checkpoints, per-stage R logs | ++| `cdhit/` | clustered ASVs + cluster membership (when enabled) | ++| `swarm/` | `otus.fasta`, `otus.count_table.csv`, logs | ++| `vsearch/` | `taxonomy.tsv` top-hit assignments | ++| `merged/` | `merged.csv` (all taxa), one filtered CSV per configured filter, `results.duckdb` | ++ ++CSVs travel into R, Python, or spreadsheets directly; the DuckDB file is what ++the workbench queries. ++ ++## What is coming on this path ++ ++- **COMING:** nothing about the path itself — it is stable. What will change ++ is delivery: standalone archives so "install" becomes "unzip and launch" ++ (`docs/migration/STATUS.md`), and Zenodo DOI minting (#8) so a completed ++ study can be published with a citable record straight from the workbench. +diff --git a/docs/wikis/Users.md b/docs/wikis/Users.md +new file mode 100644 +index 0000000..16e8852 +--- /dev/null ++++ b/docs/wikis/Users.md +@@ -0,0 +1,71 @@ ++ ++ ++# Users ++ ++This section is for people who **run analyses** with MetaManifold: from raw ++paired-end FASTQs to filtered, taxonomy-annotated tables and interactive ++figures. You do not need to be a programmer. You do need sequencing data. ++ ++Two audiences share this section, and they want different things from the ++same software: ++ ++- **Academics** — you are answering a research question and will publish. ++ You care about reproducibility, method disclosure, exactness of what is ++ claimed, and citing the software. Your home is ++ [For Academics](Users--For-Academics), but read the install and ++ first-study pages first like everyone else. ++- **Lab professionals** — you are running samples as a service: routine ++ panels, QC gates before results leave the lab, primer and database ++ management, contamination curation. Your home is ++ [For Lab Professionals](Users--For-Lab-Professionals). ++ ++Where guidance differs, pages carry **Academics:** and **Lab professionals:** ++callouts. Where it does not, they just tell you how the thing works. ++ ++## Learning path ++ ++1. [Install and First Run](Users--Install-and-First-Run) — prerequisites, ++ `install.sh`, starting the server. (~20 minutes including downloads.) ++2. [Your First Study](Users--Your-First-Study) — put FASTQs where they ++ belong, launch a run, find your tables. (~15 minutes plus compute time.) ++3. [Exploring Results](Users--Exploring-Results) — QC, the results explorer, ++ annotation and composition views. ++4. [Analysis and Statistics Today](Users--Analysis-and-Statistics-Today) — ++ what each analysis computes, and what it refuses to compute. ++5. Your track: [For Academics](Users--For-Academics) or ++ [For Lab Professionals](Users--For-Lab-Professionals). ++6. [Configuration Reference](Users--Configuration-Reference) — the full YAML ++ reference, when you need the exact key. ++7. [Troubleshooting](Users--Troubleshooting) — when something refuses. ++ ++## What is here now vs what is coming ++ ++**IN PLACE today:** the full pipeline (both ASV and OTU lanes), the browser ++workbench (config editors, runs and jobs, QC, results explorer, annotation ++and curation, composition), and the analysis surface: alpha diversity with ++honest significance statuses, composition bars, organism categories, taxon ++overlap, NMDS and PERMANOVA, exact descriptive summaries, ML fits ++(NB-GLM / CLR-ILR linear models / logistic) with refusal states, and exact ++TSS/CSS/RSS offsets. ++ ++**COMING (tracked, not yet available to you):** exact statistical tests (#3), ++compositional differential-abundance methods (#5), occupancy models (#17), ++constrained ordinations (#18→#19), PhILR/SBP balances (#20), advanced zero ++handling (#21), CladeCumulus (#6), Full Evidence Mode (#7), Zenodo DOI ++minting (#8), standalone offline installers and the coordinated updater ++(`docs/migration/STATUS.md`). The full board: ++[Status and Roadmap](Status-and-Roadmap). ++ ++**A promise the software keeps:** when it cannot compute something validly, ++it says so and returns an unsuccessful state — never a plausible-looking ++number. Refusals are a feature. See ++[Analysis and Statistics Today](Users--Analysis-and-Statistics-Today). +diff --git a/docs/wikis/_Sidebar.md b/docs/wikis/_Sidebar.md +new file mode 100644 +index 0000000..c77a48b +--- /dev/null ++++ b/docs/wikis/_Sidebar.md +@@ -0,0 +1,33 @@ ++# Notebook ++ ++- [MetaManifold](Home) ++- [Users](Users) ++ - [Install and First Run](Users--Install-and-First-Run) ++ - [Your First Study](Users--Your-First-Study) ++ - [Configuration Reference](Users--Configuration-Reference) ++ - [Exploring Results](Users--Exploring-Results) ++ - [Analysis and Statistics Today](Users--Analysis-and-Statistics-Today) ++ - [For Academics](Users--For-Academics) ++ - [For Lab Professionals](Users--For-Lab-Professionals) ++ - [Troubleshooting](Users--Troubleshooting) ++- [Maintainers](Maintainers) ++ - [Operator Track](Maintainers--Operator-Track) ++ - [Steward Track](Maintainers--Steward-Track) ++ - [Releases and Distribution](Maintainers--Releases-and-Distribution) ++ - [Compliance and Estate](Maintainers--Compliance-and-Estate) ++- [Developers](Developers) ++ - [Architecture Tour](Developers--Architecture-Tour) ++ - [Type System](Developers--Type-System) ++ - [Statistics Internals](Developers--Statistics-Internals) ++ - [Extending the Pipeline](Developers--Extending-the-Pipeline) ++ - [Testing and Benchmarks](Developers--Testing-and-Benchmarks) ++ - [REST API](Developers--REST-API) ++- [Deep Dives](Deep-Dives) ++ - [Design Progression](Deep-Dives--Design-Progression) ++ - [Type Theory Meets Statistics](Deep-Dives--Type-Theory-Meets-Statistics) ++ - [Exact Arithmetic](Deep-Dives--Exact-Arithmetic) ++ - [Maximum Likelihood](Deep-Dives--Maximum-Likelihood) ++ - [Compositional Statistics](Deep-Dives--Compositional-Statistics) ++ - [Epistemic Status](Deep-Dives--Epistemic-Status) ++ - [Advanced Functionality](Deep-Dives--Advanced-Functionality) ++- [Status and Roadmap](Status-and-Roadmap) diff --git a/docs/triage/pr7-split/RUNBOOK.md b/docs/triage/pr7-split/RUNBOOK.md new file mode 100644 index 0000000..6a2ffdf --- /dev/null +++ b/docs/triage/pr7-split/RUNBOOK.md @@ -0,0 +1,144 @@ + + +# Runbook — landing parent PR #7 in reviewable pieces + +Parent PR [JoshuaJewell/MetaManifold-WebUI#7][p7] is **CLEAN and mergeable +as-is**. The 14 stacks in this kit exist so a human can *understand* the +319-file change — you do not have to split the PR to land it. Pick one: + +- **Option A (recommended): merge #7 as-is**, reviewing stack-by-stack + from this kit. Least process, same comprehension. +- **Option B: 14 stacked PRs.** Close #7, open one PR per stack, merge in + order. Most review surface per PR, most process. +- **Option C: close #7 and re-do the work.** Not recommended — the branch + is the only copy of this work and it merges cleanly. + +All commands run on **your** machine with **your** credentials (the Arena +bot cannot push branches or act on the parent repo). + +## 0. One-time setup + +```sh +git clone https://github.com/JoshuaJewell/MetaManifold-WebUI.git mm-parent +cd mm-parent +# The split kit lives on the fork until it merges; fetch it: +git fetch https://github.com/hyperpolymath/MetaManifold-WebUI.git \ + arena/01a0dd12-metamanifold-webui +git checkout FETCH_HEAD -- docs/triage/pr7-split +# The two endpoints of PR #7: +git fetch origin main:refs/remotes/upstream/main +git fetch https://github.com/hyperpolymath/MetaManifold-WebUI.git \ + feat/stipple-typed-studies-ui:refs/remotes/fork/stipple +# Regenerate + verify the 14 patches (checks coverage AND tree equality): +UPSTREAM_REF=refs/remotes/upstream/main BRANCH_REF=refs/remotes/fork/stipple \ + docs/triage/pr7-split/make-stacks.sh +``` + +Expected: `OK: all stacks, disjoint, cover all 319 files.` then +`OK: stacks applied in order reproduce the branch tree exactly.` + +## Option A — review by stack, merge #7 whole (recommended) + +1. Read `STACKS.md`, then review each stack's patch in order: + `patches/01-repo-roots.patch` … `patches/14-bench-scripts.patch`. + Suggested focus per stack is in `STACKS.md`; anything marked + "1-line" is almost always just the SPDX header. +2. Run the checks the branch itself carries (upstream CI runs `ci.yml` + from the PR): + ```sh + git checkout -b review/pr7 refs/remotes/fork/stipple + # …run the project's documented checks… + ``` +3. Merge #7 on GitHub (or `gh pr merge 7 --repo + JoshuaJewell/MetaManifold-WebUI --squash|--merge` from your account). +4. Delete the `feat/stipple-typed-studies-ui` branch on the fork once + merged, so it stops shadowing future diffs. + +## Option B — 14 stacked PRs + +```sh +cd mm-parent +base=upstream/main # after the fetches in §0 +prev="$base" +for n in 01 02 03 04 05 06 07 08 09 10 11 12 13 14; do + spec="docs/triage/pr7-split/stacks/$n-"*.paths + name="$(basename "$spec" .paths)" + branch="pr7-stack/$name" + git checkout -b "$branch" "$prev" + git apply --index "docs/triage/pr7-split/patches/$name.patch" + git -c commit.gpgsign=false commit -m "pr7-stack: $name" --no-verify + git push -u origin "$branch" + gh pr create --repo JoshuaJewell/MetaManifold-WebUI \ + --base "${prev#upstream/}" --head "$branch" \ + --title "$(grep "| $n |" docs/triage/pr7-split/STACKS.md \ + | sed 's/.*| //')" \ + --body "Stack $n of 14 splitting #7 (see STACKS.md in +hyperpolymath/MetaManifold-WebUI \`docs/triage/pr7-split/\`). +Base: \`${prev}\`. Merge in numeric order; closes part of #7." + prev="$branch" +done +``` + +Notes: + +- Pushes go to `origin` = the **parent** repo, so run this with an + account that can push there (yours). +- The `--title` extraction reads the suggested title from the `STACKS.md` + table; adjust wording per PR before sending. +- Merge 01→14 in order (each PR's base is the previous stack's branch; + retarget to `main` as each lands, or merge the chain at the end). +- Close #7 with a comment pointing at the stack once stack 01 opens — + or keep #7 open as the tracking issue until 14 lands. +- If a stack needs fixes during review, commit on that stack's branch + and `git rebase --update-refs` the children (they are disjoint file + sets, so rebases are conflict-free by construction). + +## Appendix 1 — closing the three dead parent PRs + +Parent PRs #8, #10 and #11 have **deleted head branches**, render as +60k-line whole-tree noise, and their real deltas (2 commits, 0 commits, +1 commit respectively) are already on fork `main` (verified +patch-identical: `6c29b53b`, `c0be9516`, `0ef5eeec`). They can never +merge. The bot has no write access to the parent repo, so close them +from your account: + +```sh +for n in 8 10 11; do + gh pr close "$n" --repo JoshuaJewell/MetaManifold-WebUI \ + --comment "Closing: head branch deleted and the real delta is already +on hyperpolymath/MetaManifold-WebUI main (verified patch-identical; +see fork triage report docs/triage/2026-09-26-notification-backlog.md). +This PR rendered as whole-tree noise after the fork's git-filter-repo +rewrite and can never merge. Reopen if you disagree." +done +``` + +## Appendix 2 — unblocking CI (owner action, both repos) + +Since 2026-09-25 22:19 UTC **every file-based workflow run** +(`CI`, `DOI publication contracts`, `Stipple UI contracts`, incl. the +long-untouched `ui.yml`) ends in `startup_failure` with zero jobs, for +bot actors and Dependabot alike — while dynamic CodeQL/Dependabot runs +succeed. The workflow files are active and unchanged-at-the-break, so +this is a settings/platform-side block, not a code break. On each repo: + +1. Open any failed run (e.g. fork run `36210537997`) and read the exact + banner text (a previous session reported + "Actor is not allowed to trigger Actions workflows"). +2. Settings → Actions → General: confirm the Actions permissions level + allows these workflows, and check fork-PR / third-party-action + restrictions if the repos sit under an organisation policy. +3. Push an empty commit from your **human** account and confirm a `CI` + run starts — this discriminates "bot actors blocked" from + "Actions broken for everyone". +4. If human pushes fail identically, raise it with GitHub Support + (platform-side incident) — and re-run the failed Dependabot Updates + run on fork `main` once green (`gh run rerun 36210545146`). +5. Until CI runs again, PRs cannot go green: any stack PRs from this + kit are gated on this fix — and fork PR #75 (merged 2026-09-26 + without green CI, closing fork issue #20) still owes its + post-merge verification (first real Julia execution of its tests). + +[p7]: https://github.com/JoshuaJewell/MetaManifold-WebUI/pull/7 diff --git a/docs/triage/pr7-split/STACKS.md b/docs/triage/pr7-split/STACKS.md new file mode 100644 index 0000000..aea607e --- /dev/null +++ b/docs/triage/pr7-split/STACKS.md @@ -0,0 +1,153 @@ + + +# Parent PR #7 — the 319-file change as 14 reviewable stacks + +Parent PR [JoshuaJewell/MetaManifold-WebUI#7][p7] ("Stipple/Vue migration +with isolated typed studies UI") is one squashed commit plus two tiny +CodeQL-permission fixes: **319 files, +22,662 / −404** against upstream +`main`. It merges cleanly, but nobody can review 319 files at once. + +This kit partitions the exact same change into **14 disjoint stacks** whose +union is byte-identical to the branch (verified: applying 01→14 in order +onto upstream `main` reproduces the branch tree exactly — see +`make-stacks.sh`, which regenerates and re-verifies everything). + +- Base: upstream `main` @ `ecefb1c` ("Real real CI fix.", 2026-07-21) +- Head: `feat/stipple-typed-studies-ui` @ `eab8ea0` +- Stack definitions: `stacks/*.paths` (one git pathspec per line) +- Generated patches: `patches/*.patch` (git-ignored; regenerate locally) + +Reading guide: each stack below lists a suggested PR title, its size, and +the files that actually matter. A recurring pattern: **1-line files are +almost always the added `SPDX-License-Identifier` header** — skim those, +read the big files. + +| # | Stack | Files | +/− | Suggested PR title | +|---|-------|------:|-----|--------------------| +| 01 | repo-roots | 27 | +2,822/−43 | chore(repo): licences, manifests, installers and repo roots | +| 02 | ci-governance | 12 | +880/−18 | ci: overhaul workflows, templates, hooks and release policy | +| 03 | backend-core | 37 | +773/−7 | feat(core): epistemic core, lint tooling, SPDX headers | +| 04 | pipeline | 15 | +17/−2 | chore(pipeline): SPDX headers and merge_taxa touch-up | +| 05 | analysis-engine | 8 | +4,357/−0 | feat(analysis): AnalysisConfig v1 engine and execution | +| 06 | analysis-config | 19 | +1,675/−0 | feat(analysis): diversity/clade modules, contracts, routes | +| 07 | server | 18 | +145/−14 | feat(server): studies/runs/jobs routes and integration tests | +| 08 | stipple-ui | 11 | +460/−0 | feat(ui): first Stipple/Vue slice (MetaManifoldUI.jl) | +| 09 | frontend-shell | 41 | +884/−72 | feat(frontend): typed app shell, API client; drop legacy web/ | +| 10 | frontend-views | 11 | +92/−47 | feat(frontend): RunView rework and view headers | +| 11 | frontend-components | 33 | +1,289/−197 | feat(frontend): analysis components (config editor, tables) | +| 12 | frontend-tests | 23 | +1,947/−0 | test(frontend): unit/integration suites, benches, gates | +| 13 | docs | 40 | +5,653/−0 | docs: milestones, migration log, deferred-issue specs | +| 14 | bench-scripts | 24 | +1,668/−4 | chore(bench): Julia benchmarks and repo scripts | + +## 01 — repo-roots (27 files, +2,822/−43) + +Everything at the repository root plus the licence texts: the three +licence files (AGPL 661 + CC-BY-SA 428 + MPL 373 lines), `Justfile` (391), +editor/git/Guix/Mise/Direnv configs, install scripts, `Project.toml`, +`README`/`CHANGELOG`/`CONTRIBUTING`/`SECURITY`/`CODE_OF_CONDUCT`, and the +`codecov.yml` deletion. + +Review notes: `Manifest.toml` renders as a binary patch (estate +`.gitattributes` marks manifests binary) — verify it with +`julia --project=. -e 'using Pkg; Pkg.instantiate()'` rather than by +reading the diff. The rest is new-file prose/config; the signal is +`install.jl`/`install.sh`/`Justfile`. + +## 02 — ci-governance (12 files, +880/−18) + +`.github/` (both workflows, issue templates, PR template, CODEOWNERS, +dependabot), `.githooks/`, `packaging/`. The one file to read carefully is +`.github/workflows/ci.yml` (+488/−18); `ui.yml` (+31) and the templates are +small; the rest is short new files. + +## 03 — backend-core (37 files, +773/−7) + +Julia core (`src/core/`, `src/annotation/`, `src/MetaManifold.jl`), the R +dependency files (`R/`, `renv/`), `config/ci/`, and the 15 core test +files. Two files carry the change: `src/core/epistemic.jl` (+352, new) +and `config/ci/lint_source.jl` (+364, new), plus a 27-line touch to +`download_databases.jl`. Nearly every other file is a 1-line SPDX header. + +## 04 — pipeline (15 files, +17/−2) + +`src/pipeline/`, the MiSeq pipeline fixture, and 5 pipeline tests. Almost +entirely 1-line SPDX headers; the only logic is a 5-line touch to +`src/pipeline/merge_taxa.jl`. Fastest stack to review. + +## 05 — analysis-engine (8 files, +4,357/−0) + +The core backend change, all new code: `AnalysisConfig.jl` (+1,615), +`Execution.jl` (+1,497), and their three test files (+479/+397/+354). +`analysis.jl` and `analysis_config.jl` gain one line each. Read this stack +as a unit: engine + execution + tests. + +## 06 — analysis-config (19 files, +1,675/−0) + +Diversity/clade modules (`clade_cumulus.jl` +510, `diversity.jl`), the +config contracts (Nickel +313, JSON Schema +297, DEED +84, defaults), +the two analysis HTTP routes (`analysis_config.jl` +453), and 4 small +tests. This is "configuration as contracts": read the Nickel/Schema/DEED +trio together with the route that serves them. + +## 07 — server (18 files, +145/−14) + +The remaining `src/server/` routes, `test/unit/test_routes.jl`, +`test/unit/test_jobs.jl`, both integration tests, and `test/runtests.jl`. +Bulk of the lines is `test/integration/test_server.jl` (+105); every route +file changes by ≤9 lines. A small, safe stack. + +## 08 — stipple-ui (11 files, +460/−0) + +The headline migration piece, isolated as promised: `ui/` — +`MetaManifoldUI.jl` (+125), `Contracts.jl`, `BackendClient.jl`, `serve.jl`, +project files, styles, README, and three tests. `ui/Manifest.toml` is a +binary-rendered manifest (verify by instantiate, not by reading). + +## 09 — frontend-shell (41 files, +884/−72) + +Typed foundation of the new frontend: configs (`package.json`, +`vite.config.ts`, tsconfigs, playwright), app shell (`App.tsx`, +`main.tsx`, layout, styles), `api/`, `types/` (12 files, incl. the +`analysis_config.ts` contract at +268), `hooks/`, `utils/` — and the +deletion of the 7 legacy `web/dist/` build artefacts plus one stale +`.d.ts`. `frontend/bun.lock` is a binary-rendered lockfile (verify with +`bun install --frozen-lockfile`, not by reading). + +## 10 — frontend-views (11 files, +92/−47) + +`RunView.tsx` (+76/−41) is the only substantial change; the other ten +views gain a line or two each (headers). Quick stack. + +## 11 — frontend-components (33 files, +1,289/−197) + +The component catalogue. Three large new components +(`AnalysisConfigEditor` +342, `AdvancedAnalysisExpander` +251, +`CladeCumulus` +218), two reworks (`DataTable` +96/−79, `PipelineStages` ++86/−58) worth reading line-by-line, plus `DangerBanner` (+90) and +`EvidenceModeToggle` (+81). The long tail is 1–14 line touches. + +## 12 — frontend-tests (23 files, +1,947/−0) + +All new: 20 frontend test files (unit + integration + e2e + fixtures) +and the 3-file `frontend/bench/` harness. Review alongside stacks 09–11; +nothing here changes shipped code. + +## 13 — docs (40 files, +5,653/−0) + +Prose only, all additive: milestone records (`docs/milestones/`, incl. +the deferred-issues spec that backs fork issues #3–#8), the +`docs/issues/milestone3/` per-issue specs, the `docs/migration/` log, +compliance/testing/type-system notes, and the v0.1.0 release notes. +Largest line count, lightest review weight — read the migration log and +the deferred-issues spec first. + +## 14 — bench-scripts (24 files, +1,668/−4) + +Eight Julia benchmark suites (`bench/`, incl. `comprehensive_benchmark.jl`) +and seven repo scripts (format/lint/SPDX gates, history-hygiene tooling, +`migrate_composition.jl`). Self-contained; the scripts mirror the CI gates +from stack 02. + +[p7]: https://github.com/JoshuaJewell/MetaManifold-WebUI/pull/7 diff --git a/docs/triage/pr7-split/make-stacks.sh b/docs/triage/pr7-split/make-stacks.sh new file mode 100755 index 0000000..be79614 --- /dev/null +++ b/docs/triage/pr7-split/make-stacks.sh @@ -0,0 +1,107 @@ +#!/usr/bin/env bash +# SPDX-License-Identifier: MPL-2.0 +# SPDX-FileCopyrightText: 2026 Jonathan D.A. Jewell +# +# make-stacks.sh — split parent PR #7 (feat/stipple-typed-studies-ui) into +# reviewable stacked patches. +# +# Reads stacks/*.paths (one git pathspec per line; disjoint file sets whose +# union is the whole upstream/main..branch diff) and writes one +# patches/.patch per stack, then verifies: +# 1. the stacks are disjoint and jointly cover all 319 files, and +# 2. applying every stack in order onto upstream/main reproduces the branch +# tree byte-for-byte (empty final diff). +# +# Usage (from the repository root): +# docs/triage/pr7-split/make-stacks.sh +# +# Env overrides (the defaults name the refs this kit was built against): +# UPSTREAM_REF base of parent PR #7 (default: refs/remotes/upstream/main) +# BRANCH_REF head of feat/stipple-typed-studies-ui (default: eab8ea0...) +# +# To point the kit at live refs on your own machine: +# git fetch https://github.com/JoshuaJewell/MetaManifold-WebUI.git \ +# main:refs/remotes/upstream/main +# git fetch https://github.com/hyperpolymath/MetaManifold-WebUI.git \ +# feat/stipple-typed-studies-ui +# UPSTREAM_REF=refs/remotes/upstream/main \ +# BRANCH_REF=FETCH_HEAD docs/triage/pr7-split/make-stacks.sh +set -eu +cd "$(dirname "$0")" +HERE="$(pwd)" + +UPSTREAM_REF="${UPSTREAM_REF:-refs/remotes/upstream/main}" +BRANCH_REF="${BRANCH_REF:-eab8ea09834a90547d29ea3ba5730281313a3ba1}" + +for ref in "$UPSTREAM_REF" "$BRANCH_REF"; do + git rev-parse --verify --quiet "$ref^{commit}" >/dev/null \ + || { echo "make-stacks: ref not found: $ref" >&2; exit 1; } +done + +mkdir -p patches +rm -f patches/*.patch + +echo "kit dir: $HERE" +echo "base: $UPSTREAM_REF ($(git rev-parse --short "$UPSTREAM_REF"))" +echo "head: $BRANCH_REF ($(git rev-parse --short "$BRANCH_REF"))" +echo + +: > /tmp/stacks_union.txt +fail=0 +for spec in stacks/*.paths; do + name="$(basename "$spec" .paths)" + mapfile -t paths < "$spec" + # Anchor every pathspec at the repo root (":/") so this script works no + # matter which directory it runs from; drop blank lines defensively. + filtered=() + for p in "${paths[@]}"; do [ -n "$p" ] && filtered+=(":/$p"); done + git diff --binary "$UPSTREAM_REF" "$BRANCH_REF" -- "${filtered[@]}" \ + > "patches/$name.patch" + git diff --name-only "$UPSTREAM_REF" "$BRANCH_REF" -- "${filtered[@]}" \ + > /tmp/stacks_names.txt + cat /tmp/stacks_names.txt >> /tmp/stacks_union.txt + n=$(wc -l < /tmp/stacks_names.txt) + stat=$(git diff --numstat "$UPSTREAM_REF" "$BRANCH_REF" -- "${filtered[@]}" \ + | awk '$1 != "-" {a+=$1; d+=$2} END {printf "+%d/-%d", a, d}') + printf '%-24s %3s files %s\n' "$name" "$n" "$stat" +done + +echo +echo "--- coverage check ---" +total=$(git diff --name-only "$UPSTREAM_REF" "$BRANCH_REF" | sort > /tmp/stacks_full.txt; wc -l < /tmp/stacks_full.txt) +sort /tmp/stacks_union.txt | uniq -d > /tmp/stacks_dupes.txt +sort -u /tmp/stacks_union.txt > /tmp/stacks_union_sorted.txt +if [ -s /tmp/stacks_dupes.txt ]; then + echo "OVERLAP between stacks:"; cat /tmp/stacks_dupes.txt; fail=1 +fi +if ! cmp -s /tmp/stacks_full.txt /tmp/stacks_union_sorted.txt; then + echo "COVERAGE MISMATCH:"; comm -3 /tmp/stacks_full.txt /tmp/stacks_union_sorted.txt; fail=1 +fi +[ "$fail" -eq 0 ] && echo "OK: all stacks, disjoint, cover all $total files." + +echo +echo "--- sequential-apply check (temp worktree, deleted afterwards) ---" +tmp="$(mktemp -d)" +git worktree add --detach "$tmp" "$UPSTREAM_REF" >/dev/null 2>&1 +# shellcheck disable=SC2064 +trap "git worktree remove --force '$tmp'" EXIT +i=0 +for spec in stacks/*.paths; do + name="$(basename "$spec" .paths)" + i=$((i + 1)) + git -C "$tmp" apply --check "$HERE/patches/$name.patch" \ + || { echo "APPLY-CHECK FAILED: $name" >&2; fail=1; break; } + git -C "$tmp" apply --index "$HERE/patches/$name.patch" + git -C "$tmp" -c user.name=triage -c user.email=triage@local \ + -c commit.gpgsign=false commit -qm "stack $name" --no-verify +done +if [ "$fail" -eq 0 ]; then + if [ "$(git -C "$tmp" rev-parse 'HEAD^{tree}')" = "$(git rev-parse "$BRANCH_REF^{tree}")" ]; then + echo "OK: stacks applied in order reproduce the branch tree exactly." + else + echo "TREE MISMATCH after applying all stacks:" >&2 + git -C "$tmp" diff --stat HEAD "$BRANCH_REF" | tail -n 5 >&2; fail=1 + fi +fi + +exit "$fail" diff --git a/docs/triage/pr7-split/stacks/01-repo-roots.paths b/docs/triage/pr7-split/stacks/01-repo-roots.paths new file mode 100644 index 0000000..2ddb2aa --- /dev/null +++ b/docs/triage/pr7-split/stacks/01-repo-roots.paths @@ -0,0 +1,25 @@ +.bun-version +.editorconfig +.envrc +.gitattributes +.gitignore +.gitmessage +CHANGELOG.md +CODE_OF_CONDUCT.md +CONTRIBUTING.md +Justfile +Manifest.toml +NOTICE +Project.toml +README.md +ROADMAP.md +SECURITY.md +channels.scm +codecov.yml +guix.scm +install.jl +install.sh +mise.toml +precompile_exec.jl +start.sh +LICENSES/ diff --git a/docs/triage/pr7-split/stacks/02-ci-governance.paths b/docs/triage/pr7-split/stacks/02-ci-governance.paths new file mode 100644 index 0000000..b6ff779 --- /dev/null +++ b/docs/triage/pr7-split/stacks/02-ci-governance.paths @@ -0,0 +1,3 @@ +.github/ +.githooks/ +packaging/ diff --git a/docs/triage/pr7-split/stacks/03-backend-core.paths b/docs/triage/pr7-split/stacks/03-backend-core.paths new file mode 100644 index 0000000..5667b83 --- /dev/null +++ b/docs/triage/pr7-split/stacks/03-backend-core.paths @@ -0,0 +1,21 @@ +src/core/ +src/annotation/ +src/MetaManifold.jl +R/ +renv/ +config/ci/ +test/unit/test_categories.jl +test/unit/test_composition.jl +test/unit/test_composition_library.jl +test/unit/test_migrate_composition.jl +test/unit/test_config.jl +test/unit/test_config_hashing.jl +test/unit/test_databases.jl +test/unit/test_databases_library.jl +test/unit/test_funcdb.jl +test/unit/test_log.jl +test/unit/test_provenance.jl +test/unit/test_validation.jl +test/unit/test_project.jl +test/unit/test_primers_library.jl +test/unit/test_install_pins.jl diff --git a/docs/triage/pr7-split/stacks/04-pipeline.paths b/docs/triage/pr7-split/stacks/04-pipeline.paths new file mode 100644 index 0000000..7a7dbac --- /dev/null +++ b/docs/triage/pr7-split/stacks/04-pipeline.paths @@ -0,0 +1,7 @@ +src/pipeline/ +data/MiSeq_SOP/pipeline.yml +test/unit/test_dada2_commands.jl +test/unit/test_merge_taxa.jl +test/unit/test_merge_taxa_mappings.jl +test/unit/test_read_conservation.jl +test/unit/test_tools.jl diff --git a/docs/triage/pr7-split/stacks/05-analysis-engine.paths b/docs/triage/pr7-split/stacks/05-analysis-engine.paths new file mode 100644 index 0000000..4906145 --- /dev/null +++ b/docs/triage/pr7-split/stacks/05-analysis-engine.paths @@ -0,0 +1,8 @@ +src/analysis/AnalysisConfig.jl +src/analysis/Execution.jl +src/analysis/analysis.jl +src/analysis/analysis_config.jl +test/unit/test_analysis.jl +test/unit/test_analysis_config.jl +test/unit/test_analysis_config_milestone3.jl +test/unit/test_execution.jl diff --git a/docs/triage/pr7-split/stacks/06-analysis-config.paths b/docs/triage/pr7-split/stacks/06-analysis-config.paths new file mode 100644 index 0000000..1b6b778 --- /dev/null +++ b/docs/triage/pr7-split/stacks/06-analysis-config.paths @@ -0,0 +1,12 @@ +src/analysis/clade_cumulus.jl +src/analysis/diversity.jl +config/defaults/ +config/schemas/ +config/templates/ +config/global_configs.md +src/server/routes/analysis.jl +src/server/routes/analysis_config.jl +test/unit/test_analysis_duckdb.jl +test/unit/test_diversity.jl +test/unit/test_duckdb_store.jl +test/unit/test_r_runtime.jl diff --git a/docs/triage/pr7-split/stacks/07-server.paths b/docs/triage/pr7-split/stacks/07-server.paths new file mode 100644 index 0000000..f096c09 --- /dev/null +++ b/docs/triage/pr7-split/stacks/07-server.paths @@ -0,0 +1,17 @@ +src/server/jobs.jl +src/server/server.jl +src/server/routes/annotations.jl +src/server/routes/composition.jl +src/server/routes/config.jl +src/server/routes/databases.jl +src/server/routes/duckdb_helpers.jl +src/server/routes/events.jl +src/server/routes/jobs.jl +src/server/routes/pipeline.jl +src/server/routes/results.jl +src/server/routes/runs.jl +src/server/routes/studies.jl +test/unit/test_routes.jl +test/unit/test_jobs.jl +test/integration/ +test/runtests.jl diff --git a/docs/triage/pr7-split/stacks/08-stipple-ui.paths b/docs/triage/pr7-split/stacks/08-stipple-ui.paths new file mode 100644 index 0000000..ec1ca5c --- /dev/null +++ b/docs/triage/pr7-split/stacks/08-stipple-ui.paths @@ -0,0 +1 @@ +ui/ diff --git a/docs/triage/pr7-split/stacks/09-frontend-shell.paths b/docs/triage/pr7-split/stacks/09-frontend-shell.paths new file mode 100644 index 0000000..7ac10f0 --- /dev/null +++ b/docs/triage/pr7-split/stacks/09-frontend-shell.paths @@ -0,0 +1,16 @@ +frontend/bun.lock +frontend/package.json +frontend/playwright.config.ts +frontend/tsconfig.build.json +frontend/tsconfig.json +frontend/vite.config.ts +frontend/src/App.tsx +frontend/src/main.tsx +frontend/src/layout/ +frontend/src/styles/ +frontend/src/vite-env.d.ts +frontend/src/api/ +frontend/src/types/ +frontend/src/hooks/ +frontend/src/utils/ +web/dist/ diff --git a/docs/triage/pr7-split/stacks/10-frontend-views.paths b/docs/triage/pr7-split/stacks/10-frontend-views.paths new file mode 100644 index 0000000..80b555d --- /dev/null +++ b/docs/triage/pr7-split/stacks/10-frontend-views.paths @@ -0,0 +1 @@ +frontend/src/views/ diff --git a/docs/triage/pr7-split/stacks/11-frontend-components.paths b/docs/triage/pr7-split/stacks/11-frontend-components.paths new file mode 100644 index 0000000..a6d7174 --- /dev/null +++ b/docs/triage/pr7-split/stacks/11-frontend-components.paths @@ -0,0 +1 @@ +frontend/src/components/ diff --git a/docs/triage/pr7-split/stacks/12-frontend-tests.paths b/docs/triage/pr7-split/stacks/12-frontend-tests.paths new file mode 100644 index 0000000..1e99e7a --- /dev/null +++ b/docs/triage/pr7-split/stacks/12-frontend-tests.paths @@ -0,0 +1,2 @@ +frontend/tests/ +frontend/bench/ diff --git a/docs/triage/pr7-split/stacks/13-docs.paths b/docs/triage/pr7-split/stacks/13-docs.paths new file mode 100644 index 0000000..77f12ae --- /dev/null +++ b/docs/triage/pr7-split/stacks/13-docs.paths @@ -0,0 +1 @@ +docs/ diff --git a/docs/triage/pr7-split/stacks/14-bench-scripts.paths b/docs/triage/pr7-split/stacks/14-bench-scripts.paths new file mode 100644 index 0000000..0677f7d --- /dev/null +++ b/docs/triage/pr7-split/stacks/14-bench-scripts.paths @@ -0,0 +1,2 @@ +bench/ +scripts/