From dc5a1f8d7f8c927597e3e4f6091bd23930ec3811 Mon Sep 17 00:00:00 2001 From: Gavin Borges Date: Tue, 25 Aug 2026 23:39:18 -0400 Subject: [PATCH] docs: correct a coverage overstatement in the D34 footnote, and repair the D34 row Three defects, all shipped by PR #291 and all found by an adversarial re-read of that PR's own output against the row it cites. 1. README.md's competitor-table footnote claimed "Every PredIG/PRIME/NetMHCpan/ pVACtools cell above ... is bound to a primary source read in full" and named ONE exception. D34 describes TWO. The second is pVACtools's "Pan-allele training" cell, which D34 records as excluded from the verification pass entirely - one research pass had transcribed SESTRAV's own column text into that row - and states plainly is "unchanged and unverified by this pass". The footnote now says 34 of 36 are source-bound and names both exceptions, and the pVACtools cell is flagged inline in the table the way the NetMHCpan one already was. Note the direction: the overstatement claimed more verification coverage than the pass achieved, which is the flattering direction this repository's defects keep taking. 2. D34's own disposition arithmetic produced that overstatement and is the root cause, so it is fixed rather than patched around. The row read "seven cells were corrected" plus "Twenty-nine cells were checked and confirmed accurate", which sums to exactly the 36 cells in scope and so left no slot for the one unresolved cell and the one excluded cell the same row goes on to describe. Corrected to 7 + 27 + 1 + 1 = 36, with a dated note recording that the prior accounting double-booked two cells and that the README sentence was written from it. 3. The D34 table row was malformed: 8 cells against the table's 9 columns, so every field from "Scope boundary" onward rendered one column to the left - the 2026-08-25 date displayed under "Limitation / corrective language" and the file list under "Last verified". Counted escape-aware (the row contains escaped pipes, so a naive pipe count misreads it), and confirmed present on origin/main at fce43b5 before this change rather than introduced by it. The missing "Scope boundary" cell is supplied, stating only what the row already evidences: the 36 competitor cells across nine capability rows, with SESTRAV's own column and the Tier A metric row explicitly out of scope. Integrity harness re-run after these edits: 150 PASS / 0 WARN / 2 FAIL / 7 SKIP, unchanged from the pre-edit baseline; the two FAILs are the standing owner-owned C1/D20 and C2/D21 rulings and are untouched here. Signed-off-by: Gavin Borges --- README.md | 24 +++++++++++++++--------- docs/claims_register.md | 2 +- 2 files changed, 16 insertions(+), 10 deletions(-) diff --git a/README.md b/README.md index 3ac7ebf..71114da 100644 --- a/README.md +++ b/README.md @@ -43,19 +43,25 @@ SESTRAV carries the OpenSSF Best Practices **Passing** badge (project 13191), wh | OpenSSF Passing badge | ✓ | ✗ | ✗ | ✗ | ✗ | | Antigen processing as training features | `feature_mode=33` | ✓ | ✗ - binding + physicochemical + length only (PRIME2.0's 28-node input layer carries no proteasomal-cleavage or TAP feature) | ✗ | ✗ - NetChop/NetMHCstabPan run as optional post-hoc annotation on already-ranked output, not as input to a trained scorer | | Graph Neural Network scorer | ✓ (v2.3 GINEConv+ESM-2; research/ensemble component) | ✗ | ✗ | ✗ | ✗ | -| Pan-allele training | ✗ - ten fixed HLA-A/-B binding columns; allele identity is not a model feature, so the production model is allele-blind. A `feature_mode=166` allele-aware track exists but was not adopted (`ROADMAP.md`), and its pocket features are under review - see claims register D30 | ✓ - pan-HLA-I predictions (allele identity is not a direct model input; abstracted through upstream MHCflurry2/NOAH) | Partial - trained/validated on an expanding but still finite HLA-I allele list; its MixMHCpred binding feature is itself pan-allele | ✓ | ✓ | +| Pan-allele training | ✗ - ten fixed HLA-A/-B binding columns; allele identity is not a model feature, so the production model is allele-blind. A `feature_mode=166` allele-aware track exists but was not adopted (`ROADMAP.md`), and its pocket features are under review - see claims register D30 | ✓ - pan-HLA-I predictions (allele identity is not a direct model input; abstracted through upstream MHCflurry2/NOAH) | Partial - trained/validated on an expanding but still finite HLA-I allele list; its MixMHCpred binding feature is itself pan-allele | ✓ | ✓ (see note below - this cell is unverified) | | Multi-virus support | 9 viruses (v5 active), each a separately-validated within-virus panel - not cross-virus transfer, see LOO below: CMV, EBV, HBV, HCV, HPV, HIV-1, IAV, DENV, SARS-CoV-2 | Limited - only SARS-CoV-2 explicitly named/validated | Limited - trained mainly on tumor neoepitopes; validated prospectively on SARS-CoV-2 only | Pan-pathogen (see note below - this cell is flagged unresolved) | Tumor | | Wet-lab candidate protocol included | ✓ | ✗ | ✗ | ✗ | Partial | | AUC-PR on labeled benchmark (Tier A) | **0.828 (OOF, 30-feature, unweighted, 2026-05; not `mode_31` - see note below)** | capabilities only (see note) | capabilities only (see note) | N/A | N/A | -*Every PredIG/PRIME/NetMHCpan/pVACtools cell above (B4/L-1) is bound to a primary source read in -full - the tool's own paper and, where one exists, its GitHub README/docs - per -`docs/claims_register.md` **D34**, which lists every source and every cell's citation. **One cell -is deliberately left unresolved rather than published on an inferred claim:** NetMHCpan's -"Multi-virus support" (currently "Pan-pathogen") could not be traced to primary-source language -about pathogen/antigen-source breadth - the closest passage in its own paper only supports MHC -allele/species breadth, a different claim - so the label is retained pending a maintainer ruling -rather than being silently rewritten to a synthesized replacement (`docs/claims_register.md` D34).* +*Of the 36 PredIG/PRIME/NetMHCpan/pVACtools cells above (B4/L-1), **34 are bound to a primary +source read in full** - the tool's own paper and, where one exists, its GitHub README/docs - per +`docs/claims_register.md` **D34**, which lists every source consulted. **Two cells are not, and +both are flagged inline above rather than counted as verified.** (1) NetMHCpan's "Multi-virus +support" (currently "Pan-pathogen") is **unresolved**: it could not be traced to primary-source +language about pathogen/antigen-source breadth - the closest passage in its own paper only +supports MHC allele/species breadth, a different claim - so the label is retained pending a +maintainer ruling rather than being silently rewritten to a synthesized replacement. (2) +pVACtools's "Pan-allele training" cell is **unverified**: one research pass transcribed SESTRAV's +own column text into that row, so the row was excluded from the verification pass entirely rather +than corrected from a guess, and the cell stands as originally written with nothing behind it. +**An earlier version of this footnote said "every" cell was source-bound and named only the first +exception**, which overstated the pass's coverage by one cell in the flattering direction; it is +corrected here (`docs/claims_register.md` D34).* *Tier A 704-peptide labeled benchmark. SESTRAV RF is evaluated out-of-fold; external tools are fully scored on the same peptides. **That asymmetry does not favour the external tools as previously claimed here.** This Tier A arm's 720-peptide corpus has zero duplicate peptides (D16), so the exact-peptide leakage affecting the v5 figures below (D15) is a structural no-op here and does not apply. A different, unquantified risk does apply: 32.1% of the 704-peptide scored pool has a substring-level near-duplicate elsewhere in the pool, never filtered for this benchmark; whether it affected the score is not established (`docs/claims_register.md` D22). Every v5 cross-validation figure below is peptide-grouped as of 2026-08-10; this Tier A figure deliberately is not, and peptide-grouping would not address the substring-homology risk in any case (D16, D22). The certified head-to-head field is in External Benchmark Results below - BigMHC (0.822), MHCflurry binding-only (0.800), MixMHCpred 2.2 (0.795), and DeepImmuno (0.698), all bound to `results/table3_tier_a_metrics.csv`; the closest external tool is BigMHC (0.822). PredIG and PRIME are compared on capabilities only: their metric head-to-head is not reproducible from a certified results file and is not reported. pVACtools targets a different problem domain (patient-specific tumor neoantigens from somatic variant calls) rather than published viral epitopes from proteome sequence, so the "(neoantigens)"/"Tumor"/"N/A" cells in its column above are not a head-to-head with the viral-epitope rows. Separately, on the harder v5 generalization set (35,597 active rows, 9 viruses), canonical `mode_31` reports pooled **peptide-grouped** cross-validation AUC-PR **0.6058** (re-baselined 2026-08-10, closing D15; the prior ungrouped figure 0.8312 was leakage-inflated and is retracted) and same-pathogen (within-virus) discrimination per-virus (mean within-CV AUC-ROC **0.658**; prior ungrouped 0.751 retracted; the pooled AUC-ROC 0.9368 reported before that was separately decoy-inflated and is also retracted - see Paradigm 2 below). **0.828 is NOT a `full_31`/`mode_31` result** - it is a 30-feature, unweighted, 200-tree measurement from 2026-05, predating `feature_mode=31`'s introduction by 26 days (`docs/claims_register.md` D16); the extended `full_33` antigen-processing configuration is reported separately under Release Tracks and is not part of this certified field. Certified per-tool metrics and their scope boundaries: `results/table3_tier_a_metrics.csv` and `docs/claims_register.md`. (`results/external_benchmark_comparison.md` is a 2026-05-22 SESTRAV-vs-binding-only-only report carrying the same pre-D16 mislabel and an unresolved provenance gap - historical reference only, not a citable source for the 5-tool field above.)* diff --git a/docs/claims_register.md b/docs/claims_register.md index 1f4c823..c5f43e8 100644 --- a/docs/claims_register.md +++ b/docs/claims_register.md @@ -76,7 +76,7 @@ D28 added and the whole table recounted 2026-08-18; D29, D30 and D31 added 2026- --- -| D34 | PARTIAL | **The README "SESTRAV vs Field" table's 36 competitor cells (PredIG/PRIME/NetMHCpan/pVACtools x 9 capability rows) carried no citation of any kind - B4/L-1, "the largest reader-facing integrity gap on the board." Each tool's own paper and GitHub README/docs were read directly and every cell checked against it; seven cells were corrected.** **PredIG:** "Open source, pip-installable" Yes -> corrected - open source (GPL-2.0) but distributed only as R scripts via Docker/Singularity or a webserver, no pip package (github.com/BSC-CNS-EAPM/PredIG README; PMC12613480 Code/Data availability). "Pan-allele training" Partial -> corrected to Yes - the paper states directly "PredIG performs pan HLA-I allele predictions" and that allele identity is not fed to the model as a direct feature (PMC12613480 Methods). **PRIME:** "Open source, pip-installable" Yes -> corrected - academic/non-commercial license only (a separate license is required for for-profit use), distributed as a precompiled binary or built from source, no PyPI package (github.com/GfellerLab/PRIME README, License and Installation sections). "Antigen processing as training features" Partial -> corrected to No - PRIME2.0's 28-node input layer (one MixMHCpred binding node, twenty amino-acid-frequency nodes, seven length nodes) carries no proteasomal-cleavage or TAP-transport feature (PMC9811684, STAR Methods, verbatim architecture description). "Pan-allele training" Yes -> corrected to Partial - PRIME's own training/validation allele list is described in its changelog as "expanded" in v2.1, which is broader coverage of a still-finite list, not the generalize-to-unseen-alleles architecture the row's "Yes" cells (NetMHCpan, pVACtools) describe; its MixMHCpred binding feature is separately pan-allele by that tool's own README, which is a different claim than PRIME's own training being pan-allele. **NetMHCpan:** "End-to-end workflow" No -> corrected to Partial - a submitted FASTA proteome/protein is auto-digested into peptides and ranked by %Rank with Strong/Weak-Binder labels; this is binding/presentation prediction only, with no immunogenicity or candidate-selection step (PMC7319546; services.healthtech.dtu.dk/services/NetMHCpan-4.1/1-Submission.php). **pVACtools:** "End-to-end workflow" Yes (neoantigens) -> corrected to add a qualifier - the authors' own Abstract states pVACtools produces an end-to-end solution "when paired with a well-established genomics pipeline" for variant calling/annotation, which pVACtools does not itself perform (PMC7056579 Abstract; Methods confirms input is an already-VEP-annotated VCF). "Antigen processing as training features" Partial -> corrected to No - NetChop and NetMHCstabPan run as optional post-hoc annotation on already-ranked output under a section literally titled "Optional Downstream Analysis Tools," not as input to a trained scorer; the row asks about training features specifically, and the cited page's own language ("not as inputs to a trained scoring model") argues against "Partial" (pvactools.readthedocs.io/en/latest/pvacseq.html and pvacseq/run.html). **Twenty-nine cells were checked and confirmed accurate as already written; those are unchanged.** | Every corrected cell listed above was published with no citation and had not been checked against its subject's own primary source since the table was first written. Two of the seven corrections (PredIG open-source/pip and PRIME open-source/pip) were in the direction of overstating a competitor's installability - the opposite of this project's usual error direction, and worth naming because AUD-7/D32 note the direction is not always self-serving. | Per tool, every source actually fetched and read: **PredIG** - PMC12613480 (full text) and github.com/BSC-CNS-EAPM/PredIG (README, full). **PRIME** - PMC9811684 (Gfeller et al., Cell Systems 2023, STAR Methods read in full for architecture) and github.com/GfellerLab/PRIME README (full, including License and Installation sections); github.com/GfellerLab/MixMHCpred README consulted only for MixMHCpred's own pan-allele status, kept separate from PRIME's own architecture per the correction above. **NetMHCpan** - PMC7319546 (Reynisson et al. 2020, NAR, full text) and services.healthtech.dtu.dk/services/NetMHCpan-4.1/ (service page) plus the sw_request licensing page. **pVACtools** - PMC7056579 (pVACtools paper, full text) and pvactools.readthedocs.io (Installation, pvacseq, pvacseq/run.html, FAQ, output_files.html, optional_downstream_analysis_tools.html). | **One cell is deliberately NOT corrected and NOT left as originally drafted either: NetMHCpan's "Multi-virus support" ("Pan-pathogen") is UNRESOLVED, pending a maintainer ruling.** The paper's only pan-specificity sentence ("given the pan-specific nature of both methods, predictions can be run for any MHC molecule of known sequence...") supports MHC-molecule breadth, not pathogen/antigen-source breadth; neither primary source discusses peptide biological origin (viral/bacterial/tumor/self) or uses the term "pan-pathogen" anywhere. The label is retained in the README rather than silently rewritten, and flagged in a footnote there, because every synthesized replacement considered was itself an inference beyond what either source states (rule 3: when the source cannot be read, the claim does not get made - here the *right* claim, not just the current one, could not be sourced either). **A verification process defect is recorded here because it is exactly the class this file exists to catch: one research pass's raw output listed pVACtools's "Pan-allele training" current-state field as SESTRAV's OWN column text ("No - ten fixed HLA-A/-B binding columns"), a transcription error, not a finding about pVACtools.** That row was excluded from this pass entirely rather than corrected from a guess; pVACtools's "Pan-allele training" cell is unchanged and unverified by this pass. **Six further findings from the first research pass were independently re-verified and rejected or narrowed by a second, adversarial pass** before reaching this row - including two citations containing a fabricated or misattributed quotation attributed to NetMHCpan's own paper (a "length distribution of naturally presented peptides" clause, and an "any peptide of known sequence" clause, neither of which appears in PMC7319546) - which is why the Multi-virus cell above is unresolved rather than published on the first pass's citation. | 2026-08-25 | `README.md` ("SESTRAV vs Field" table and its two footnotes); `docs/claims_register.md` (this row) | +| D34 | PARTIAL | **The README "SESTRAV vs Field" table's 36 competitor cells (PredIG/PRIME/NetMHCpan/pVACtools x 9 capability rows) carried no citation of any kind - B4/L-1, "the largest reader-facing integrity gap on the board." Each tool's own paper and GitHub README/docs were read directly and every cell checked against it; seven cells were corrected.** **PredIG:** "Open source, pip-installable" Yes -> corrected - open source (GPL-2.0) but distributed only as R scripts via Docker/Singularity or a webserver, no pip package (github.com/BSC-CNS-EAPM/PredIG README; PMC12613480 Code/Data availability). "Pan-allele training" Partial -> corrected to Yes - the paper states directly "PredIG performs pan HLA-I allele predictions" and that allele identity is not fed to the model as a direct feature (PMC12613480 Methods). **PRIME:** "Open source, pip-installable" Yes -> corrected - academic/non-commercial license only (a separate license is required for for-profit use), distributed as a precompiled binary or built from source, no PyPI package (github.com/GfellerLab/PRIME README, License and Installation sections). "Antigen processing as training features" Partial -> corrected to No - PRIME2.0's 28-node input layer (one MixMHCpred binding node, twenty amino-acid-frequency nodes, seven length nodes) carries no proteasomal-cleavage or TAP-transport feature (PMC9811684, STAR Methods, verbatim architecture description). "Pan-allele training" Yes -> corrected to Partial - PRIME's own training/validation allele list is described in its changelog as "expanded" in v2.1, which is broader coverage of a still-finite list, not the generalize-to-unseen-alleles architecture the row's "Yes" cells (NetMHCpan, pVACtools) describe; its MixMHCpred binding feature is separately pan-allele by that tool's own README, which is a different claim than PRIME's own training being pan-allele. **NetMHCpan:** "End-to-end workflow" No -> corrected to Partial - a submitted FASTA proteome/protein is auto-digested into peptides and ranked by %Rank with Strong/Weak-Binder labels; this is binding/presentation prediction only, with no immunogenicity or candidate-selection step (PMC7319546; services.healthtech.dtu.dk/services/NetMHCpan-4.1/1-Submission.php). **pVACtools:** "End-to-end workflow" Yes (neoantigens) -> corrected to add a qualifier - the authors' own Abstract states pVACtools produces an end-to-end solution "when paired with a well-established genomics pipeline" for variant calling/annotation, which pVACtools does not itself perform (PMC7056579 Abstract; Methods confirms input is an already-VEP-annotated VCF). "Antigen processing as training features" Partial -> corrected to No - NetChop and NetMHCstabPan run as optional post-hoc annotation on already-ranked output under a section literally titled "Optional Downstream Analysis Tools," not as input to a trained scorer; the row asks about training features specifically, and the cited page's own language ("not as inputs to a trained scoring model") argues against "Partial" (pvactools.readthedocs.io/en/latest/pvacseq.html and pvacseq/run.html). **Twenty-seven cells were checked and confirmed accurate as already written; those are unchanged.** **Disposition arithmetic, corrected 2026-08-25: 7 corrected + 27 confirmed + 1 unresolved (NetMHCpan "Multi-virus support") + 1 excluded-and-unverified (pVACtools "Pan-allele training") = 36. This row previously read "Twenty-nine", which with the seven corrections exhausted all 36 cells and so double-booked the two cells this same row describes as NOT verified - the arithmetic asserted full coverage while the prose disclosed two gaps. The error propagated: `README.md`'s footnote was written from it and claimed "every" cell was source-bound, naming only one of the two exceptions. Both are fixed in the same pass; the pVACtools cell is now flagged inline in the README table alongside the NetMHCpan one.** | Every corrected cell listed above was published with no citation and had not been checked against its subject's own primary source since the table was first written. Two of the seven corrections (PredIG open-source/pip and PRIME open-source/pip) were in the direction of overstating a competitor's installability - the opposite of this project's usual error direction, and worth naming because AUD-7/D32 note the direction is not always self-serving. | Per tool, every source actually fetched and read: **PredIG** - PMC12613480 (full text) and github.com/BSC-CNS-EAPM/PredIG (README, full). **PRIME** - PMC9811684 (Gfeller et al., Cell Systems 2023, STAR Methods read in full for architecture) and github.com/GfellerLab/PRIME README (full, including License and Installation sections); github.com/GfellerLab/MixMHCpred README consulted only for MixMHCpred's own pan-allele status, kept separate from PRIME's own architecture per the correction above. **NetMHCpan** - PMC7319546 (Reynisson et al. 2020, NAR, full text) and services.healthtech.dtu.dk/services/NetMHCpan-4.1/ (service page) plus the sw_request licensing page. **pVACtools** - PMC7056579 (pVACtools paper, full text) and pvactools.readthedocs.io (Installation, pvacseq, pvacseq/run.html, FAQ, output_files.html, optional_downstream_analysis_tools.html). | The 36 competitor cells ONLY - PredIG/PRIME/NetMHCpan/pVACtools across the nine capability rows of `README.md`'s "SESTRAV vs Field" table. **SESTRAV's own column is out of scope** and was not re-verified by this pass. The table's tenth row ("AUC-PR on labeled benchmark (Tier A)") is also out of scope: it carries metric and N/A values rather than capability claims, and is governed by D16 and D22 instead. Nothing in this row rests on any source not named in the citation column. | **One cell is deliberately NOT corrected and NOT left as originally drafted either: NetMHCpan's "Multi-virus support" ("Pan-pathogen") is UNRESOLVED, pending a maintainer ruling.** The paper's only pan-specificity sentence ("given the pan-specific nature of both methods, predictions can be run for any MHC molecule of known sequence...") supports MHC-molecule breadth, not pathogen/antigen-source breadth; neither primary source discusses peptide biological origin (viral/bacterial/tumor/self) or uses the term "pan-pathogen" anywhere. The label is retained in the README rather than silently rewritten, and flagged in a footnote there, because every synthesized replacement considered was itself an inference beyond what either source states (rule 3: when the source cannot be read, the claim does not get made - here the *right* claim, not just the current one, could not be sourced either). **A verification process defect is recorded here because it is exactly the class this file exists to catch: one research pass's raw output listed pVACtools's "Pan-allele training" current-state field as SESTRAV's OWN column text ("No - ten fixed HLA-A/-B binding columns"), a transcription error, not a finding about pVACtools.** That row was excluded from this pass entirely rather than corrected from a guess; pVACtools's "Pan-allele training" cell is unchanged and unverified by this pass. **Six further findings from the first research pass were independently re-verified and rejected or narrowed by a second, adversarial pass** before reaching this row - including two citations containing a fabricated or misattributed quotation attributed to NetMHCpan's own paper (a "length distribution of naturally presented peptides" clause, and an "any peptide of known sequence" clause, neither of which appears in PMC7319546) - which is why the Multi-virus cell above is unresolved rather than published on the first pass's citation. | 2026-08-25 | `README.md` ("SESTRAV vs Field" table and its two footnotes); `docs/claims_register.md` (this row) | ## Section 2: Honest Disclosure Matrix Claims