Skip to content

v2.0 roadmap: cross-modal latent alignment #1

Description

@recklesswater

Goal

v1.0 covers same feature space, different cohort transfer. v2.0 is about the case where the two tables share no columns at all — for example reusing a representation learned on a public transcriptome atlas on a metabolomics or lipidomics table.

Raw features cannot be aligned across modalities: column i of the source has no counterpart in the target, so there is nothing to fit and nothing to re-scale. The only object with a chance of transferring is the latent space, and getting there needs an explicit alignment objective on top of the encoders.

Why v1.0 does not do this

The encoder input dimension is fixed, so a model trained on one modality physically cannot consume another. Today's transfer is: same columns, different cohort. This issue is the plan for removing that restriction.

Proposed design

  1. One encoder per modality, sharing a latent dimensionality and (optionally) a decoder.
  2. An alignment objective on the latent space, added to the existing reconstruction + classification terms.
  3. An evaluation protocol that can actually tell whether alignment worked, rather than assuming it did.

Task breakdown

1. Multi-modality backbone

  • allow AEWithClassifier to be instantiated per modality from a config, instead of assuming one feature matrix
  • decide and document the decoder strategy: fully separate decoders (default) versus a shared decoder with per-modality heads
  • CLI: --modality-a-dir / --modality-b-dir, with the v1.0 single-modality path kept as the default
  • keep the v1.0 checkpoint format readable, or write a small converter
  • no breaking change: python ae_classifier.py --data-dir ... must behave exactly as it does today

2. Alignment objective

  • extend MultiTaskLossWrapper from two tasks to N, so an alignment term gets its own learnable weight
  • paired mode (sample-level correspondence exists): contrastive or MMD loss on matched pairs
  • unpaired mode (no correspondence, the common case): adversarial domain discriminator, or moment matching such as CORAL / MMD on latent moments
  • CLI: --align-loss {none,mmd,coral,adversarial,contrastive} with none as the default
  • log the alignment term separately so it can be seen competing with the other losses

3. Evaluation protocol

Alignment claims are easy to make and hard to check, so the metric comes with the feature:

  • domain classifier AUC on the pooled latents — 0.5 means the two latent spaces are indistinguishable
  • latent distribution distances: MMD and energy distance, before and after alignment
  • downstream transfer on a held-out target cohort, compared against the unaligned baseline
  • sanity check: shuffle the modality labels and confirm the alignment metric degrades; if it does not, the metric is broken
  • write all of it into run_summary.json plus a before/after t-SNE figure

4. Documentation and example

  • make_synthetic_data.py --multi-modal: two synthetic modalities with a shared latent factor, so the example is runnable
  • a worked example in the README showing aligned versus unaligned numbers on that synthetic pair
  • document what alignment can and cannot fix — it cannot create signal that is not present in either modality

Non-goals for v2.0

  • no claim of better classification performance; alignment is about making representations comparable, not about inflating AUC
  • no large-scale benchmark suite; the goal is one honest, reproducible example
  • no change to the v1.0 interface

How to contribute

Pick any checkbox and open a pull request, or comment here with a design objection. The two questions we would most like input on are:

  1. Paired or unpaired alignment as the default — most real cross-modal pairs are unpaired, but paired designs are much easier to evaluate honestly.
  2. How to keep the alignment metric from being gamed: a domain classifier AUC of 0.5 is necessary but not sufficient for a useful shared space.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions