Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 15 additions & 9 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,19 +43,25 @@ SESTRAV carries the OpenSSF Best Practices **Passing** badge (project 13191), wh
| OpenSSF Passing badge | ✓ | ✗ | ✗ | ✗ | ✗ |
| Antigen processing as training features | `feature_mode=33` | ✓ | ✗ - binding + physicochemical + length only (PRIME2.0's 28-node input layer carries no proteasomal-cleavage or TAP feature) | ✗ | ✗ - NetChop/NetMHCstabPan run as optional post-hoc annotation on already-ranked output, not as input to a trained scorer |
| Graph Neural Network scorer | ✓ (v2.3 GINEConv+ESM-2; research/ensemble component) | ✗ | ✗ | ✗ | ✗ |
| Pan-allele training | ✗ - ten fixed HLA-A/-B binding columns; allele identity is not a model feature, so the production model is allele-blind. A `feature_mode=166` allele-aware track exists but was not adopted (`ROADMAP.md`), and its pocket features are under review - see claims register D30 | ✓ - pan-HLA-I predictions (allele identity is not a direct model input; abstracted through upstream MHCflurry2/NOAH) | Partial - trained/validated on an expanding but still finite HLA-I allele list; its MixMHCpred binding feature is itself pan-allele | ✓ | ✓ |
| Pan-allele training | ✗ - ten fixed HLA-A/-B binding columns; allele identity is not a model feature, so the production model is allele-blind. A `feature_mode=166` allele-aware track exists but was not adopted (`ROADMAP.md`), and its pocket features are under review - see claims register D30 | ✓ - pan-HLA-I predictions (allele identity is not a direct model input; abstracted through upstream MHCflurry2/NOAH) | Partial - trained/validated on an expanding but still finite HLA-I allele list; its MixMHCpred binding feature is itself pan-allele | ✓ | ✓ (see note below - this cell is unverified) |
| Multi-virus support | 9 viruses (v5 active), each a separately-validated within-virus panel - not cross-virus transfer, see LOO below: CMV, EBV, HBV, HCV, HPV, HIV-1, IAV, DENV, SARS-CoV-2 | Limited - only SARS-CoV-2 explicitly named/validated | Limited - trained mainly on tumor neoepitopes; validated prospectively on SARS-CoV-2 only | Pan-pathogen (see note below - this cell is flagged unresolved) | Tumor |
| Wet-lab candidate protocol included | ✓ | ✗ | ✗ | ✗ | Partial |
| AUC-PR on labeled benchmark (Tier A) | **0.828 (OOF, 30-feature, unweighted, 2026-05; not `mode_31` - see note below)** | capabilities only (see note) | capabilities only (see note) | N/A | N/A |

*Every PredIG/PRIME/NetMHCpan/pVACtools cell above (B4/L-1) is bound to a primary source read in
full - the tool's own paper and, where one exists, its GitHub README/docs - per
`docs/claims_register.md` **D34**, which lists every source and every cell's citation. **One cell
is deliberately left unresolved rather than published on an inferred claim:** NetMHCpan's
"Multi-virus support" (currently "Pan-pathogen") could not be traced to primary-source language
about pathogen/antigen-source breadth - the closest passage in its own paper only supports MHC
allele/species breadth, a different claim - so the label is retained pending a maintainer ruling
rather than being silently rewritten to a synthesized replacement (`docs/claims_register.md` D34).*
*Of the 36 PredIG/PRIME/NetMHCpan/pVACtools cells above (B4/L-1), **34 are bound to a primary
source read in full** - the tool's own paper and, where one exists, its GitHub README/docs - per
`docs/claims_register.md` **D34**, which lists every source consulted. **Two cells are not, and
both are flagged inline above rather than counted as verified.** (1) NetMHCpan's "Multi-virus
support" (currently "Pan-pathogen") is **unresolved**: it could not be traced to primary-source
language about pathogen/antigen-source breadth - the closest passage in its own paper only
supports MHC allele/species breadth, a different claim - so the label is retained pending a
maintainer ruling rather than being silently rewritten to a synthesized replacement. (2)
pVACtools's "Pan-allele training" cell is **unverified**: one research pass transcribed SESTRAV's
own column text into that row, so the row was excluded from the verification pass entirely rather
than corrected from a guess, and the cell stands as originally written with nothing behind it.
**An earlier version of this footnote said "every" cell was source-bound and named only the first
exception**, which overstated the pass's coverage by one cell in the flattering direction; it is
corrected here (`docs/claims_register.md` D34).*

*Tier A 704-peptide labeled benchmark. SESTRAV RF is evaluated out-of-fold; external tools are fully scored on the same peptides. **That asymmetry does not favour the external tools as previously claimed here.** This Tier A arm's 720-peptide corpus has zero duplicate peptides (D16), so the exact-peptide leakage affecting the v5 figures below (D15) is a structural no-op here and does not apply. A different, unquantified risk does apply: 32.1% of the 704-peptide scored pool has a substring-level near-duplicate elsewhere in the pool, never filtered for this benchmark; whether it affected the score is not established (`docs/claims_register.md` D22). Every v5 cross-validation figure below is peptide-grouped as of 2026-08-10; this Tier A figure deliberately is not, and peptide-grouping would not address the substring-homology risk in any case (D16, D22). The certified head-to-head field is in External Benchmark Results below - BigMHC (0.822), MHCflurry binding-only (0.800), MixMHCpred 2.2 (0.795), and DeepImmuno (0.698), all bound to `results/table3_tier_a_metrics.csv`; the closest external tool is BigMHC (0.822). PredIG and PRIME are compared on capabilities only: their metric head-to-head is not reproducible from a certified results file and is not reported. pVACtools targets a different problem domain (patient-specific tumor neoantigens from somatic variant calls) rather than published viral epitopes from proteome sequence, so the "(neoantigens)"/"Tumor"/"N/A" cells in its column above are not a head-to-head with the viral-epitope rows. Separately, on the harder v5 generalization set (35,597 active rows, 9 viruses), canonical `mode_31` reports pooled **peptide-grouped** cross-validation AUC-PR **0.6058** (re-baselined 2026-08-10, closing D15; the prior ungrouped figure 0.8312 was leakage-inflated and is retracted) and same-pathogen (within-virus) discrimination per-virus (mean within-CV AUC-ROC **0.658**; prior ungrouped 0.751 retracted; the pooled AUC-ROC 0.9368 reported before that was separately decoy-inflated and is also retracted - see Paradigm 2 below). **0.828 is NOT a `full_31`/`mode_31` result** - it is a 30-feature, unweighted, 200-tree measurement from 2026-05, predating `feature_mode=31`'s introduction by 26 days (`docs/claims_register.md` D16); the extended `full_33` antigen-processing configuration is reported separately under Release Tracks and is not part of this certified field. Certified per-tool metrics and their scope boundaries: `results/table3_tier_a_metrics.csv` and `docs/claims_register.md`. (`results/external_benchmark_comparison.md` is a 2026-05-22 SESTRAV-vs-binding-only-only report carrying the same pre-D16 mislabel and an unresolved provenance gap - historical reference only, not a citable source for the 5-tool field above.)*

Expand Down
Loading
Loading