Goal
v1.0 covers same feature space, different cohort transfer. v2.0 is about the case where the two tables share no columns at all — for example reusing a representation learned on a public transcriptome atlas on a metabolomics or lipidomics table.
Raw features cannot be aligned across modalities: column i of the source has no counterpart in the target, so there is nothing to fit and nothing to re-scale. The only object with a chance of transferring is the latent space, and getting there needs an explicit alignment objective on top of the encoders.
Why v1.0 does not do this
The encoder input dimension is fixed, so a model trained on one modality physically cannot consume another. Today's transfer is: same columns, different cohort. This issue is the plan for removing that restriction.
Proposed design
- One encoder per modality, sharing a latent dimensionality and (optionally) a decoder.
- An alignment objective on the latent space, added to the existing reconstruction + classification terms.
- An evaluation protocol that can actually tell whether alignment worked, rather than assuming it did.
Task breakdown
1. Multi-modality backbone
2. Alignment objective
3. Evaluation protocol
Alignment claims are easy to make and hard to check, so the metric comes with the feature:
4. Documentation and example
Non-goals for v2.0
- no claim of better classification performance; alignment is about making representations comparable, not about inflating AUC
- no large-scale benchmark suite; the goal is one honest, reproducible example
- no change to the v1.0 interface
How to contribute
Pick any checkbox and open a pull request, or comment here with a design objection. The two questions we would most like input on are:
- Paired or unpaired alignment as the default — most real cross-modal pairs are unpaired, but paired designs are much easier to evaluate honestly.
- How to keep the alignment metric from being gamed: a domain classifier AUC of 0.5 is necessary but not sufficient for a useful shared space.
Goal
v1.0 covers same feature space, different cohort transfer. v2.0 is about the case where the two tables share no columns at all — for example reusing a representation learned on a public transcriptome atlas on a metabolomics or lipidomics table.
Raw features cannot be aligned across modalities: column i of the source has no counterpart in the target, so there is nothing to fit and nothing to re-scale. The only object with a chance of transferring is the latent space, and getting there needs an explicit alignment objective on top of the encoders.
Why v1.0 does not do this
The encoder input dimension is fixed, so a model trained on one modality physically cannot consume another. Today's transfer is: same columns, different cohort. This issue is the plan for removing that restriction.
Proposed design
Task breakdown
1. Multi-modality backbone
AEWithClassifierto be instantiated per modality from a config, instead of assuming one feature matrix--modality-a-dir/--modality-b-dir, with the v1.0 single-modality path kept as the defaultpython ae_classifier.py --data-dir ...must behave exactly as it does today2. Alignment objective
MultiTaskLossWrapperfrom two tasks to N, so an alignment term gets its own learnable weight--align-loss {none,mmd,coral,adversarial,contrastive}withnoneas the default3. Evaluation protocol
Alignment claims are easy to make and hard to check, so the metric comes with the feature:
run_summary.jsonplus a before/after t-SNE figure4. Documentation and example
make_synthetic_data.py --multi-modal: two synthetic modalities with a shared latent factor, so the example is runnableNon-goals for v2.0
How to contribute
Pick any checkbox and open a pull request, or comment here with a design objection. The two questions we would most like input on are: