Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Omics-AutoEncoder-Classifier: Latent Space Representation for High-Dimensional Omics

Current release: v1.0 — the representation-learning half of the pipeline. Cross-modal alignment and the rest of the upgrade list are planned for v2.0; see the Roadmap.

Struggling with poor generalization performance or weak cross-domain transferability when handling high-dimensional, low-sample-size multi-omics data? This repository is about the representation layer of that problem: mapping a large biological feature matrix into a dense latent space with an Autoencoder (AE), so that the representation — not the raw feature columns — becomes the thing you transfer.

Why transfer the representation instead of the features?

Raw features do not travel. A model fitted on one cohort expects the same columns, in the same order, on the same scale. Move to another site, another platform, another batch, and that contract starts to break. Move to another modality and it breaks completely: a public gene-expression atlas and your own metabolomics or lipidomics table share no columns at all, so there is no feature-wise mapping to fit and nothing to re-scale. This is the part that makes cross-modal work hard — you cannot align what has no common coordinate system.

A latent space gives you coordinates that are not tied to individual peaks or probes. Each sample becomes a dense vector summarising covariation across the whole matrix, and that vector is the object you can freeze, adapt, and — for genuinely different modalities — align. The classifier on top is deliberately small and disposable; the encoder is the transferable asset.

On top of that sits the classic N ≪ P setting (thousands of features versus dozens to a few hundred samples) that makes plain classifiers overfit in metabolomics, lipidomics and clinical proteomics. A deterministic autoencoder is used rather than a variational one: there is no posterior collapse to tune around, and the objective stays a straightforward reconstruction term plus a classification term, combined by a learnable multi-task weighting layer so the two do not have to be balanced by hand.

Core Workflow

  • Source Domain Pre-training: Train the Autoencoder and its multi-task loss on a rich discovery cohort or a public database to learn a representation of the feature space.
  • Transfer & Linear Probing: Freeze the encoding backbone so the representation stays fixed, and apply strict preprocessing standards (variance thresholds and standardization) consistently to every split.
  • Target Fine-Tuning: Adapt only the classification head to your target cohort. The encoder defines the coordinate system, so source and target have to share it — see Scope below for what that does and does not cover.

Scope, and what this repository does not promise

This is deliberately honest about its limits, because over-claiming transfer performance is how this kind of tool ends up useless to other people.

  • No performance guarantee. There is no claim here that latent features beat raw features on your cohort. They do not always, and on some real cohorts we have tested, transfer to a held-out target cohort landed close to chance. What this repository gives you is the plumbing and the diagnostics — pre-training, strict preprocessing, a frozen encoder, head adaptation, stratified cross-validation, and a raw-versus-latent comparison — so you can find out on your own data instead of trusting a number from someone else.

  • Cross-modal alignment is v2.0 work, not a feature of this release. The transfer implemented in v1.0 is same feature space, different cohort: same modality, same columns, different site, batch, platform or time period. Transferring between modalities (for example a public gene-expression database into your metabolomics table) additionally requires an encoder per modality plus an objective that pulls the two latent spaces into alignment. That step is not implemented in v1.0, and the code does not pretend otherwise — it is the headline item on the Roadmap. What v1.0 does give you is the representation-learning half of the pipeline and a structure an alignment objective can be attached to:

    source modality (e.g. public gene-expression atlas)
          └─▶ encoder A ─┐
                         ├─▶ aligned latent space ─▶ classification head ─▶ target task
    target modality (e.g. your metabolomics table)
          └─▶ encoder B ─┘
                         ▲
                         └─ the alignment objective between A and B lives here: v2.0, not in this release
    

    What is in this repo is the solid half: one encoder trained on the source feature space, frozen, with a small head adapted to the target cohort, plus the diagnostics (raw-versus-latent comparison, stratified cross-validation, latent-space figures) that tell you whether the transfer worked before you invest in the alignment objective.

  • The demo data are random numbers. See About the synthetic data below. Anything the demo prints about AUC is meaningless by construction.

Let's Collaborate & Discuss

Whether you are tackling disease subtyping, cross-cohort batch effects, or multi-omics integration, feedback and collaborative ideas are warmly welcomed. Cross-modal latent alignment in particular — how to make a representation learned on one omics layer transferable to another — is the question we would most like to discuss. Feel free to open an issue, submit a pull request, or start a discussion if you want to work on omics representation learning together.


Roadmap

v1.0 — this release

What you can actually run today:

  • autoencoder with residual blocks, plus a learnable multi-task weighting between the reconstruction and classification terms, so the two do not have to be balanced by hand
  • source-cohort pre-training, frozen encoder, classification-head adaptation on the target cohort
  • binary / multiclass / multilabel heads, each with the loss function that matches how the classes are defined
  • strict preprocessing (variance filter, then standardisation) fitted on the source training split and applied unchanged everywhere else
  • stratified cross-validation, a raw-versus-latent comparison, latent-space t-SNE figures, and a run summary plus full log

This is same feature space, different cohort transfer. It is the half of the problem that can be built and validated properly right now.

v2.0 — planned, stay tuned

Cross-modal latent alignment is the headline of v2.0. The goal is the thing v1.0 explicitly does not do: making a representation learned on one omics layer usable on another, where no feature-wise mapping exists.

  • Cross-modal alignment. One encoder per modality (transcriptome, metabolome, lipidome, …) plus an objective that pulls their latent spaces into a shared coordinate system, so a model learned on one layer can be adapted to another.
  • Explicit domain-shift penalties — adversarial or moment-matching terms — layered on top of today's frozen-backbone transfer.
  • Multi-omics fusion of several aligned latent blocks into a single predictor.
  • Stronger model selection, for example choosing checkpoints and early stopping on a discrimination metric rather than on the reconstruction-dominated total loss. A first version of this already ships as --select-metric auc.

Implementing all of that at once is not realistic for a research-side project, and shipping a half-working alignment layer would be worse than shipping a solid v1.0. So v1.0 stays deliberately on the representation-learning half, and everything that is not there is labelled as not there rather than hinted at. If you want to work on the v2.0 alignment layer, that is the collaboration we would most like to have — please open an issue or start a discussion.

Repository layout

File Purpose
ae_classifier.py The pipeline: pre-training, transfer, latent export, evaluation, cross-validation, figures
make_synthetic_data.py Generates random demo cohorts, so the pipeline runs without any private data
requirements.txt Python dependencies

Classification task types

The classification head is configured with --task, and everything downstream — loss function, label encoding, head width, metrics, ROC export, cross-validation and figures — switches together.

--task Meaning Head output Loss Labels
binary control vs disease Linear(64, 1) BCEWithLogitsLoss one 0/1 column
multiclass mutually exclusive classes, e.g. 0 = control, 1 = early lesion, 2 = advanced lesion Linear(64, num_classes) CrossEntropyLoss one column of class indices 0..C-1, read as torch.long
multilabel several non-exclusive findings, e.g. marker A abnormal and marker B abnormal at once Linear(64, num_classes) BCEWithLogitsLoss several 0/1 columns given with --label-cols

The head is left linear on purpose: CrossEntropyLoss applies log-softmax internally and BCEWithLogitsLoss applies sigmoid internally, so adding either activation inside the model would make training less stable, not more.

Requirements

  • Python >= 3.9 (developed on 3.12)
  • PyTorch (CPU is enough; the pipeline runs on --device cpu)
  • scikit-learn, pandas, numpy, matplotlib
pip install -r requirements.txt

Quick start with synthetic data

The demo tables are random numbers, not biological measurements. They only exist so that the whole pipeline can be executed end to end without shipping any real cohort.

# 1. generate random demo cohorts (binary, multiclass and multilabel variants)
python make_synthetic_data.py --outdir synthetic --n-features 2000

# 2. run the pipeline on the binary demo
python ae_classifier.py \
    --data-dir synthetic/binary --outdir results/binary \
    --epochs 30 --finetune-epochs 60 --cv-epochs 40

# 3. the other two task types
python ae_classifier.py --task multiclass --num-classes 3 \
    --data-dir synthetic/multiclass --outdir results/multiclass \
    --epochs 30 --finetune-epochs 60 --cv-epochs 40

python ae_classifier.py --task multilabel \
    --label-cols abnormal_A abnormal_B abnormal_C --feature-start 5 \
    --data-dir synthetic/multilabel --outdir results/multilabel \
    --epochs 30 --finetune-epochs 60 --cv-epochs 40

The commands above are deliberately short so the demo finishes in a couple of minutes on a laptop CPU. The defaults used in a real run are larger (--epochs 50 --finetune-epochs 200 --cv-epochs 100 with 10 folds); raise them once you point the pipeline at your own data.

Input format

One CSV per split, rows are samples and columns are metadata followed by features:

sample_index, ID, Group, 80.9663, 81.0391, 81.9575, ...
S0000,       S0000, 1,   0.1777,  0.5476,  0.3124,  ...
  • the first --feature-start columns (default 3) are metadata; everything after that is a feature
  • --id-col and --label-col name the identifier and the label column
  • for --task multilabel the label block is several 0/1 columns, so --feature-start has to move past them (the demo uses 5)
  • the source and target cohorts must contain the same feature columns in the same order, and must not share samples — the source cohort is what the encoder learns from, so overlap would leak the target cohort into the representation

Preprocessing (variance filter, then standardisation) is fitted on the source training split only and then applied unchanged to every other split.

What the pipeline produces

Everything lands in --outdir:

Output Content
source_ae_classifier.pt, target_head_finetuned.pt checkpoints for the pre-trained model and the adapted head
source_training_history.csv loss and validation score per epoch
latent_source_train.csv, latent_target_val.csv, … latent features plus ground truth and predicted probabilities for every split
metrics_target_val.json, per_class_target_val.csv, confusion_matrix_target_val.csv evaluation on the target validation split
predictions_target_val.csv per-sample probabilities, ready for an ROC curve elsewhere
cv_head_folds.csv, cv_head_summary.csv stratified K-fold results over the target training cohort
downstream_folds.csv, downstream_summary.csv SVM / logistic regression on raw versus latent features
tsne_latent_domains.png/.pdf, tsne_latent_classes.png/.pdf latent space by domain and by class
run_summary.json, pipeline.log configuration, cohort composition and the full log

Useful options

Option Effect
--select-metric loss|auc how the best source checkpoint is chosen. loss uses the multi-task total loss, which is dominated by reconstruction; auc picks the epoch with the best validation AUC
--strict-freeze / --adapt-bn-stats whether BatchNorm running statistics inside the frozen encoder keep updating during head fine-tuning. Frozen means frozen by default
--skip-cv, --skip-tsne skip the slower stages for a quick smoke test
--device auto|cpu|cuda auto picks CUDA when it is available
--seed seeds Python, NumPy and PyTorch for reproducible runs

About the synthetic data

make_synthetic_data.py writes random Gaussian features with a weak planted class signal and a block of correlated features, which mimics the collinearity of real metabolomics and lipidomics tables. The default shape follows the situation this pipeline is built for: few samples, many features — on the order of 50–500 rows against up to a few thousand columns. Real cohorts are usually small (tens to a few hundred participants) while the feature matrix easily reaches thousands of mass peaks or metabolites, which is exactly the N ≪ P regime that makes plain classifiers overfit.

Because the data are random, the absolute performance numbers from the demo are meaningless. Use them to check that the mechanics work, then point --data-dir at your own tables — and read the Scope section above before drawing any conclusion from the numbers you get.

License

MIT — see LICENSE. You are free to use, modify, redistribute and even ship this commercially. A citation or a link back is appreciated but not required.

About

Latent-space representation learning for high-dimensional, low-sample-size omics. Pre-train an autoencoder on a discovery cohort, freeze the encoder, adapt the head. Binary/multiclass/multilabel heads. Cross-cohort transfer today; cross-modal latent alignment is the roadmap.

Topics

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages