A reproducible pipeline for planarian neural-fate-specific transcription factor discovery. Integrates five single-cell and regulatory atlases (Fincher 2018, Plass 2018, Cui 2023, King 2024, Perez 2025) — including the previously unused Fincher brain sub-atlas and the Cui regeneration time-course — into 11 evidence streams, with Bayesian Dirichlet uncertainty quantification, a cross-species ortholog benchmark, and ANANSE gene regulatory network validation to prioritize high-confidence targets for RNAi and functional validation.
2026-09-11 ground-truth correction (screened ≠ validated): King 2024 mmc5 is titled "All Transcription Factors Inhibited" — it lists every TF for which RNAi was performed and markers assayed, not only TFs with phenotypes. The table's own caption encodes phenotype status by font colour ("Transcription factors in red and marker genes in green have phenotypes"), but the distributed xlsx copy is monochrome. The historical
testedlabel therefore means screened, not RNAi-validated. The FISH-confirmed phenotype set (20 TFs, curated from the paper's Fig 3J/4E, S4, S7, S8) is now tracked explicitly via thephenotype_confirmedcolumn on every rank/shortlist table and viabioforge.evidence.groundtruth.PHENOTYPE_CONFIRMED_V6. All label-based statistics (ROC-AUC, calibration, effect sizes, negative controls) now report both labels:screened(legacy continuity) andphenotype_confirmed(the publishable "validated recovery" numbers). Under the stricter label the headline discrimination claims hold and strengthen: honest ROC-AUC 0.875, strict 0.833 (vs 0.846/0.807 on the screened label). The mmc7 green-font "previously published FSTF" annotation is also now captured (king_atlas.tsv:previously_reported_fstf,sig_color).
2026-09 hardening audit: the three prioritization methods (fixed / centered / uniform) now share one philosophy — the same all-candidate universe (rank.csv), the same annotation mask (one row per gene), the same +0.07 bonus layer (neural GO +0.03, TF GO +0.02, human ortholog +0.02), and the same Track-B DNA-binding-domain gate. The statistical suite reports circular AND circularity-controlled evaluations, and every figure is audited for non-emptiness (
projects/NeuralTF/scripts/audit_figures.py).2026-09-13 senior-review remediation: (1)
build_king_atlas.pywas rebuilt — value sheets are now keyed by subcluster NAME (the old positional join was off by one row, attaching a different gene's p-value/log2FC to ~836/1,046 atlas rows), X1's lead-TF column is captured (X1 has no Fincher-cluster column), and the paper-derived symbol map resolves all 20 phenotype-confirmed genes into the atlas. (2) The Tbx2/3b ground-truth entry was remapped to dd_Smed_v6_17143_0_1 — the paper's own RNAi table (mmc5) lists dd17143 with the GLIPR1 (dd210) marker, i.e. the actual Tbx2/3b phenotype experiment; the previous dd6470 assignment came from a PlanMine alias whose own descriptions are aminopeptidase-related. (3) All three prioritization methods now share ONE King catalog view (load_mmc4_tf_catalog: All-sheet TF?-filtered, 421 TFs incl. eya/meis) — the bonus mask and Track-B gate are now genuinely identical across methods. (4) Overlap/consensus p-values are stratification-corrected for the dual-track 5+5 design (the legacy uniform-subset nulls overstated significance by ~25–38 orders of magnitude; both are reported). (5) A king_free evaluation arm (all King-2024 streams excluded) was added alongside honest/strict, and the ortholog benchmark now reports the honest-score AUC with a 95% CI. (6) Figure 26's permutation null was rebuilt as a within-stream shuffle (the previous row-shuffle null was vacuous — the null multiset equaled the foreground distribution by (6) Figure 26's permutation null was rebuilt as a within-stream shuffle (the previous row-shuffle null was vacuous — the null multiset equaled the foreground distribution by construction); figures 06/07/24/27/28/32 fixed for stale cohort sizes, label conflation, metric mislabels, and circular contrasts; the figure catalog now lists all 26 figures including fig 37. Full re-run downstream: honest ROC-AUC 0.866 (screened) / 0.944 (phenotype_confirmed), strict 0.821 / 0.893, king-free 0.820 / 0.893; the three methods now emit an identical top-10; the flagship permutation artifact is a valid n=1000 run (74/143 BH-FDR significant) replacing a degenerate n=10 file.
2026-09 evidence model v2: candidate testing status is now reported as Tested / Not tested / Known FSTF (replacing the misleading "Track A/B" / "novel" framing; † = status read from the source paper's own tables), and the 9-stream model is extended to 11 streams — two independent streams from previously unused raw data (
fincher_brain,cui_temporal) — with weights rebalanced to sum to 1.0 (expression,rnai, andcorrelationwere down-weighted so no single stream dominates). An independent cross-species ortholog benchmark (ortholog_benchmark.py) is added.
# 1. Set up Python 3.12+ environment
git clone https://github.com/dhzso/neuraltf.git
cd neuraltf
python -m venv .venv
.venv\Scripts\activate # Windows
# source .venv/bin/activate # Linux/Mac
pip install -e ".[bio,streamlit]"
# 2. Place raw downloads in datasets/raw/ (see datasets/MANIFEST.md)
# and build all processed artifacts with ONE master orchestrator:
python scripts/generate_all.py
# 3. Or run the pipeline and downstream analyses step-by-step:
python scripts/run.py # Core pipeline (Fincher, Plass, Cui, King, Perez)
python scripts/run_downstream.py # Dirichlet UQ, ANANSE scan, tables & 26 figures
# 4. Launch the interactive Streamlit UI
bioforge ui # http://localhost:8501| File | Content |
|---|---|
rank.csv |
All 11,696 candidates with scores across all 11 evidence streams + phenotype_confirmed ground-truth flag |
rank_neural.csv |
Neural-enriched candidates with proof status + ground_truth_status |
evidence_cards.md |
Per-candidate markdown evidence summary |
pipeline_results.json |
Machine-readable candidate metadata |
de_pvalues.parquet |
Per-gene best DE p per atlas (meta-analysis input) |
checkpoint_0*.parquet |
6 incremental audit checkpoints (atlas loads, post-QC, post-scoring, King, Perez, stream matrix) |
| File | Content |
|---|---|
master_tf_catalog.csv |
Unified King + Perez TF catalog (14,682 unique v6 TFs) — in projects/NeuralTF/data/ |
dirichlet_centered_full_rank.csv |
Full Dirichlet-centered (k=40) composite rank (all candidates, 1 row/gene) |
dirichlet_uniform_full_rank.csv |
Full Dirichlet-uniform (α=1) composite rank (all candidates, 1 row/gene) |
dirichlet_centered_top10.csv |
Dual-track top-10 shortlist under centered Dirichlet (5 tested + 5 not-tested) |
dirichlet_uniform_top10.csv |
Dual-track top-10 shortlist under uniform Dirichlet (5 tested + 5 not-tested) |
dirichlet_*_draw_scores.csv |
Per-candidate draw-score matrices (bootstrap CIs, convergence) |
ananse_network_full.csv |
ANANSE GRN scan across all candidates & 9 cell fates (RBH-mapped, neuron-share normalized) |
ananse_top_regulators.csv |
Top planarian neural regulators (neuron-fate out-degree first) |
tf_ranked_neural_top19.csv |
Top 19 TFs: neural-filtered candidates |
tf_ranked_all_top43.csv |
Top 43 TFs: all expression-filtered candidates |
tf_ranked_catalog_top74.csv |
Top 74 TFs: full King mmc4 catalog |
supplementary_table_S1 – S7 |
Method comparison, and fixed/centered/uniform rank tables across all candidates |
de_pvalues.parquet |
Per-gene best DE p per atlas (pipeline checkpoint; drives meta-analysis) |
The pipeline seeds candidate TFs across five single-cell and regulatory atlases, computes an 11-stream multi-evidence matrix, evaluates scoring stability under Dirichlet uncertainty sampling, and maps candidates into cell-fate regulatory circuits.
| Atlas | Year | Modality / Data Type | Role in Pipeline |
|---|---|---|---|
| Fincher | 2018 | scRNA-seq (50.5K cells, v4 IDs) | Whole-animal cell-type expression & neural specificity (GSE111764) |
| Plass | 2018 | scRNA-seq (37.5K cells, v6 IDs) | Independent whole-anatomy replication & X1 dynamics (GSE103633) |
| Cui | 2023 | scRNA-seq (55.0K cells, 8 time points) | High-resolution regeneration time-course expression (OMIX) |
| King | 2024 | Single-cell TF catalog + RNAi screen | G0/X1 cluster enrichment, RNAi phenotypes (mmc5), TF pair correlations (mmc6) (Cell Reports) |
| Perez | 2025 | Lineage atlas + ANANSE GRNs | Lineage TF classification (MOESM5), ANANSE regulatory influence (MOESM19), and ANANSE GRN (MOESM22) (Nature Communications) |
Scoring utilizes a transparent, weighted multi-evidence integration model. Weights renormalize over streams present for each candidate:
| # | Stream | Default Weight ( |
Biological Basis & Computation |
|---|---|---|---|
| 1 | Expression | 0.100 |
|
| 2 | Specificity | 0.100 |
|
| 3 | Reproducibility | 0.100 |
|
| 4 | RNAi | 0.050 | Binary indicator (1.0) if the TF was RNAi-screened with markers assayed in King mmc5 ("All Transcription Factors Inhibited") — screened, phenotype NOT implied |
| 5 | Correlation | 0.050 |
|
| 6 | Neural Enriched | 0.100 | Binary indicator (1.0) for G0 neural subcluster log₂FC ≥ 1.5 (King mmc7) |
| 7 | Neural Specificity | 0.100 |
|
| 8 | Perez Lineage | 0.100 | Perez 2025 lineage TF class: 1.0 for neural-class, 0.5 for other TF classes, 0.0 if absent |
| 9 | Perez Influence | 0.100 | Perez 2025 ANANSE regulatory influence in neuron fate (MOESM19), normalized 0–1 rank |
| 10 | Fincher Brain | 0.100 | Independent Fincher 2018 BrainClustering sub-atlas enrichment ( |
| 11 | Cui Temporal | 0.100 | Cui 2023 regeneration time-course: neuronal temporal induction |
- Tier Assignment:
- HIGH: RNAi-screened OR (supporting streams ≥ 3 AND score ≥ 0.45)
- MEDIUM: supporting streams ≥ 2 AND score ≥ 0.25
- LOW: All other candidates
- Testing Status († = status read from the source paper's own tables):
tested— RNAi performed in the King et al. screen (†). Screened, phenotype NOT implied (see the ground-truth correction above).phenotype_confirmed(column, alongside) — FISH-confirmed loss-of-cell-type phenotype in the paper's figures (20 TFs) †not_tested— no RNAi record in the screen; untested rather than "novel" (many carry conserved orthologs)known_fstf— documented fate-specifying TF (FSTF) from literature, no RNAi data (†)
Consensus across all three unified methods (fixed / centered Dirichlet k=40 / uniform Dirichlet α=1; shared universe, bonus mask, and gates) under the 11-stream model. After the 2026-09-13 fixes all three methods now emit identical top-10 shortlists (10/10 consensus) — by design the methods share most streams, so this consensus measures weight-robustness, not method independence (see overlap_significance.json caveat; the stratification-corrected null is reported alongside the legacy upper bound).
| Testing status | Consensus candidates | Evidence |
|---|---|---|
| Tested #1 | dd_Smed_v6_38342_0_1 (dd38342 / POU3F4-class) |
RNAi-screened (†); consistent top-1 across methods |
| Tested #2 | dd_Smed_v6_34144_0_1 (dd34144 / TCF7-LEF) |
RNAi-screened (†) |
| Tested #3 | dd_Smed_v6_12722_0_1 (dd12722 / BHLHE23) |
RNAi-screened (†) |
| Tested #4 | dd_Smed_v6_29211_0_1 (dd29211 / PHOX2A) |
RNAi-screened (†); FISH phenotype-confirmed (dd8060+ neuron loss) |
| Tested #5 | dd_Smed_v6_22163_0_1 (dd22163 / UNCX) |
RNAi-screened (†); FISH phenotype-confirmed (sert+ neuron loss) |
| Not tested #1 | dd_Smed_v6_33456_0_1 (dd33456) |
Homeodomain; consensus top not-tested |
| Not tested #2 | dd_Smed_v6_16466_0_1 (dd16466 / FOXG1) |
Forkhead; oral/neural forebrain ortholog |
| Not tested #3 | dd_Smed_v6_18505_0_1 (dd18505 / MSH1) |
Homeodomain (msh1) |
| Not tested #4 | dd_Smed_v6_2442_0_1 (dd2442 / HOXC6) |
Homeobox |
| Not tested #5 | dd_Smed_v6_13704_0_1 (dd13704 / PTF-4) |
bHLH |
Per-method shortlists: dirichlet_centered_top10.csv, dirichlet_uniform_top10.csv, top10_neural_tfs_prioritized.csv (all carry phenotype_confirmed).
Ground-truth recovery under the 11-stream model (both labels, circularity-controlled — label-encoding streams excluded): screened label (68 genes): honest ROC-AUC 0.866, strict 0.821, king-free 0.820, vs 0.978 circular; phenotype_confirmed label (19 in-universe genes): honest ROC-AUC 0.944, strict 0.893, king-free 0.893. The king_free arm (2026-09-13) excludes every stream drawing on King 2024 — the study that also defines the labels — quantifying how much discrimination the independent atlases (Fincher, Plass, Cui, Perez-influence, Fincher-brain) provide alone; residual caveat: expression/specificity still carry King-mmc7 floors fused at pipeline time, so even king_free is an upper bound. An independent cross-species ortholog benchmark (neural-fate vs non-neural human orthologs, honest label-free score) gives ROC-AUC 0.634 (95% CI 0.52–0.74, one-sided Mann-Whitney p=0.011; the pre-fix 0.712 was computed on the full circular score and measured label re-encoding). Provisional, coverage-limited (~101 genes).
- "Tested" = screened. 48 of 68 screened TFs have no published phenotype; only 20 genes are FISH phenotype-confirmed (
bioforge.evidence.groundtruth). All "validated" claims must use thephenotype_confirmedflag. - Recall is imperfect.
dd_Smed_v6_48508_0_1has a confirmed neuron-loss phenotype but never entered the candidate universe (no seeded evidence) — the funnel misses it. - Cross-method consensus is not independence. The three methods share 9 of 11 streams, the bonus mask, and the King floors; the overlap p-values are an upper bound on significance, and both the legacy uniform-subset null and the stratification-corrected (dual-track 5+5) null are reported in
overlap_significance.json. fincher_brainpresence, not value, carries most of its signal. Its unconditional per-stream AUC (~0.41 against the screened label) is a missingness artifact: among genes where the stream is present, the value-only AUC is ~0.66. The per-stream table inprecision_recall.jsonnow reports both variants.- King-derived expression still enters the honest score. The honest/strict AUCs exclude label-encoding streams but keep King mmc7 expression floors; they are an upper bound on true label-free discrimination. The king_free arm removes the pure King streams but cannot un-fuse the King floors — see
precision_recall.json. - Ortholog benchmark is provisional (~101 classified genes from PlanMine BLAST descriptions; a full Compara/DIOPT table is needed for a definitive benchmark). The honest-score AUC (0.634) carries a wide CI (0.52–0.74) at this coverage.
To test sensitivity against arbitrary weighting assumptions, we employ Monte Carlo Dirichlet weight perturbation (1,000 draws, seed=2024 across all candidates in rank.csv):
-
Centered Dirichlet (
$k = 40$ ):$$\mathbf{w}^{(m)} \sim \text{Dirichlet}(k \cdot \mathbf{w}_{\text{default}})$$ Perturbs weights locally around default parameters ($\sim 95%$ of draws within$\pm 0.10$ of baseline).python projects/NeuralTF/scripts/dirichlet_centered.py
-
Uniform Dirichlet (
$\alpha_i = 1$ ):$$\mathbf{w}^{(m)} \sim \text{Dirichlet}(\mathbf{1}_{11})$$ Samples uniformly across the entire 11-simplex to discover robust data-driven signals without prior preference.python projects/NeuralTF/scripts/dirichlet_uniform.py
The pipeline includes a comprehensive statistical validation suite to ensure publication-grade rigor:
| # | Test | Script | Key Metrics |
|---|---|---|---|
| 1 | Full Permutation Test | scripts/stats/permutation_test_full.py |
Empirical p-values (n=1,000 permutations) |
| 2 | Bootstrap Confidence Intervals | scripts/stats/bootstrap_confidence.py |
95% CI on integrated scores |
| 3 | Overlap Significance | scripts/stats/overlap_significance.py |
Hypergeometric, Fisher's exact, binomial tests |
| 4 | Precision-Recall Analysis | scripts/stats/precision_recall.py |
Precision@5, Precision@10, PR-AUC |
| 5 | Negative Controls | scripts/stats/negative_controls.py |
Random non-TF & non-neural TF distributions |
| 6 | Effect Sizes | scripts/stats/effect_sizes.py |
Cliff's delta, Cohen's d, Mann-Whitney U |
| 7 | Leave-One-Atlas-Out | scripts/stats/leave_one_atlas_out.py |
Top-10 stability per excluded atlas |
| 8 | Meta-Analytic P-values | scripts/stats/meta_analytic_pvalue.py |
Fisher's & Stouffer's combined p-values |
| 9 | Power Analysis | scripts/stats/power_analysis.py |
Convergence & power curves |
| 10 | Mann-Whitney U (Top-10) | scripts/stats/mann_whitney_top10.py |
Rank-biserial correlation |
| 11 | Calibration Analysis | scripts/stats/calibration.py |
Empirical positive rates per decile |
| 12 | Brier Score | scripts/stats/brier_score.py |
Probabilistic classification accuracy |
| 13 | Cross-Method Correction | scripts/stats/cross_method_correction.py |
Bonferroni/BH-FDR for 3-method consensus |
| 14 | Score Shuffling Permutation | scripts/stats/score_shuffling_permutation.py |
Stream-assignment null model |
| 15 | Cross-Species Ortholog Benchmark | scripts/stats/ortholog_benchmark.py |
Conserved H. sapiens neural TF ROC-AUC (0.712) & Mann-Whitney test |
Run all tests:
python scripts/run_statistical_tests.pyThe authoritative, up-to-date catalog is
projects/NeuralTF/figures/FIGURE_CATALOG.md (26 publication figures: 25 single-panel + dual-panel 23/27/37, 500-DPI PNG
under the Nature Communications palette). Figure numbering reflects the active
set: data-integration (01, 03, 04, 20, 31), candidate atlas (05),
robustness/ablation (06, 07, 08, 09, 13, 15), scoring decomposition (18),
benchmark recovery (23, 24, 26, 27, 28, 30, 37), convergence (29), lineage (32),
method agreement (33), GRN (34), cross-atlas meta-analysis (35), and
regeneration temporal dynamics (36). Figures 02, 10–12, 14, 16, 17, 19, 21, 22,
25 were retired during the single-panel refactor.
Bioinformatics/
├── pyproject.toml Package configuration & dependencies
├── README.md Primary documentation
├── bioforge.md BioForge framework reference & operations
│
├── src/bioforge/ BioForge Core Framework
│ ├── evidence/ Multi-stream evidence engine
│ │ ├── schema.py EvidenceRecord & 11-stream EvidenceSource enum
│ │ ├── scoring.py Weighted score integration & DEFAULT_WEIGHTS
│ │ ├── confidence.py Tier classification (HIGH/MEDIUM/LOW)
│ │ ├── groundtruth.py King-2024 ground truth: screened vs phenotype-confirmed labels
│ │ └── cards.py Markdown evidence card generation
│ ├── projects/neuraltf/
│ │ ├── pipeline.py NeuralTFPipeline (5-atlas loader & 6 checkpoints)
│ │ ├── planmine.py PlanMine InterMine REST client & annotation parser
│ │ └── prioritize.py Dual-track candidate scoring & filtering
│ ├── omics/ scRNA-seq QC, log-norm, PCA, Leiden clustering
│ ├── workflow/ Declarative YAML workflow engine
│ ├── cli/ Command-line interface
│ └── ui/ Streamlit interactive dashboard
│
├── datasets/
│ ├── MANIFEST.md Download URLs & SHA256 checksums
│ ├── raw_data_manifest.md Complete 21-file raw data audit
│ ├── raw/ Raw downloads (gitignored)
│ └── processed/ Processed H5ADs & Parquet files (gitignored)
│
├── projects/NeuralTF/
│ ├── data/
│ │ ├── bridge.csv v4 ↔ v6 ↔ gene_name Rosetta Stone mapping
│ │ ├── king_atlas.tsv King 2024 G0 progenitor enrichment table
│ │ ├── perez_tf_summary.csv Perez 2025 TF lineage classification
│ │ └── master_tf_catalog.csv Unified King + Perez master TF catalog (14,682 TFs)
│ ├── scripts/
│ │ ├── convert_cui.py Cui SMED → v6 H5AD (optional --subsample N)
│ │ ├── preprocess_perez.py Parse Perez MOESM5 TF classes
│ │ ├── annotate_phenotype_groundtruth.py Stamp phenotype_confirmed on all outputs (idempotent)
│ │ ├── dirichlet_centered.py Centered Dirichlet k=40 (all candidates)
│ │ ├── dirichlet_uniform.py Uniform Dirichlet α=1 (all candidates)
│ │ ├── ananse_full_scan.py ANANSE GRN scan across all candidates
│ │ ├── export_fstf_ranked.py Export ranked TF tables
│ │ ├── create_supplementary_tables.py Generate supplementary tables S1–S4
│ │ ├── generate_publication_figures.py Generate 26 publication figures
│ │ └── figures/ 26 modular figure generation scripts & style.py
│ ├── results/ Dirichlet, ANANSE, and supplementary tables (gitignored)
│ ├── figures/ 26 publication + 4 supplementary PNG figures (gitignored)
│ └── runs/pipeline_run/ rank.csv, rank_neural.csv, 6 checkpoint parquets
│
└── scripts/ Master Orchestration & Build Scripts
├── generate_all.py End-to-end multi-step master pipeline runner
├── run_downstream.py Post-pipeline runner (Dirichlet, ANANSE, figures, stats)
├── run_statistical_tests.py Run all 14 statistical tests
├── run.py Core pipeline execution entry point
├── build_bridge.py Build v4↔v6 gene ID bridge from Rosetta Stone
├── build_king_atlas.py Build king_atlas.tsv from King mmc7
├── build_master_catalog.py Merge King mmc4 + Perez MOESM5 TF catalog
├── convert_fincher.py Convert Fincher DGE to H5AD
├── consolidate_plass.py Consolidate Plass RAW.tar to H5AD
└── stats/ 14 statistical test scripts
├── permutation_test_full.py
├── bootstrap_confidence.py
├── overlap_significance.py
├── precision_recall.py
├── negative_controls.py
├── effect_sizes.py
├── leave_one_atlas_out.py
├── meta_analytic_pvalue.py
├── power_analysis.py
├── mann_whitney_top10.py
├── calibration.py
├── brier_score.py
├── cross_method_correction.py
└── score_shuffling_permutation.py
MIT — see LICENSE file.