Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions Deep-Generative-Models/CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
# Changelog

This module is versioned as a single drop within the parent repository.

## [Unreleased]

### Added
- **🧬 Deep Generative Models module.** A self-contained knowledge module on deep generative models: a landing README with the family map and the generative trilemma, seven concept notes (foundations; autoregressive models; normalizing flows; variational autoencoders; generative adversarial networks; energy-based models; diffusion and score-based models), one math page (objectives and transformations), one diagram (the family taxonomy), a glossary, and a canonical references list. All content is original and cites the primary papers.
16 changes: 16 additions & 0 deletions Deep-Generative-Models/LICENSE
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
Deep Generative Models module — Licensing
=========================================

Dual-licensed, consistent with the parent repository.

SPDX-License-Identifier: Apache-2.0 AND CC-BY-4.0

1. Source code (snippets, configuration) — Apache-2.0
https://www.apache.org/licenses/LICENSE-2.0
2. Prose and diagrams (Markdown, figures) — CC-BY-4.0
https://creativecommons.org/licenses/by/4.0/

Referenced papers remain under their own terms and are cited, not reproduced.
This module contains no text copied from any source document; the explanations
are original. If the parent repository carries top-level LICENSE files, those
govern in case of conflict.
46 changes: 46 additions & 0 deletions Deep-Generative-Models/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
# Deep Generative Models

A self-contained knowledge module on **deep generative models** — the families of models that learn the probability distribution behind a dataset well enough to sample new data from it. It mirrors the structure of the rest of this repository: a landing map, modular concept notes, a math page, a diagram, a glossary, and canonical references.

> Original educational content. The notes are written from scratch and cite the primary papers; no source text is reproduced. See [`STYLE.md`](./STYLE.md).

## What this covers

Generative modeling learns a model of $p(x)$ over data — images, text, audio — so it can do three things: **sample** new instances, **score** the likelihood of a point (density estimation, useful for anomaly detection), and **learn representations** (the latent structure it discovers is useful downstream). The hard part is that the joint distribution over high-dimensional data is intractable to write down directly, and every model family is a different bargain for getting around that.

## Concept notes

| # | Note | Family / topic |
|---|---|---|
| 0 | [Foundations](./concepts/foundations.md) | objectives; explicit vs. implicit; likelihood-based vs. -free; latent variables; MLE↔KL; the generative trilemma |
| 1 | [Autoregressive models](./concepts/autoregressive-models.md) | chain-rule factorization; exact likelihood; PixelCNN, WaveNet, GPT |
| 2 | [Normalizing flows](./concepts/normalizing-flows.md) | invertible maps; change of variables; NICE, RealNVP, Glow |
| 3 | [Variational autoencoders](./concepts/variational-autoencoders.md) | latent variables; the ELBO; the reparameterization trick |
| 4 | [Generative adversarial networks](./concepts/generative-adversarial-networks.md) | implicit density; the minimax game; mode collapse |
| 5 | [Energy-based models](./concepts/energy-based-models.md) | unnormalized densities; the partition function; score matching |
| 6 | [Diffusion and score-based models](./concepts/diffusion-and-score-based-models.md) | forward noising, reverse denoising; the current image SOTA |

## The generative trilemma

A useful lens for the whole field: three properties you want, and the fact that no single family gets all three for free.

1. **High sample quality and expressivity** — samples that are sharp and varied.
2. **Tractable, exact likelihood** — you can compute $p(x)$ and do density estimation.
3. **Fast, stable training and sampling.**

Autoregressive models and flows give exact likelihood; GANs give sharp samples; diffusion gives quality at the cost of slow sampling; VAEs give fast inference and representations at some cost to sharpness. Each family is a different corner of this triangle — the [model-families diagram](./diagrams/model-families.md) lays them out.

## Repository map

```text
Deep-Generative-Models/
├── concepts/ # one note per family, plus shared foundations
├── math/ # the objectives and transformations behind the families
├── diagrams/ # a Mermaid taxonomy of the families
├── glossary/ # shared vocabulary
└── references/ # canonical papers
```

## License

Dual-licensed: code under Apache-2.0, prose and diagrams under CC-BY-4.0. See [`LICENSE`](./LICENSE).
7 changes: 7 additions & 0 deletions Deep-Generative-Models/STYLE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
# Style notes

- **Original writing.** Every note is written from scratch. No text is copied from any source document, textbook, or article.
- **Cite primary sources.** Technical claims point to the original papers (collected in [`references/references.md`](./references/references.md)), not to secondary write-ups.
- **Notation.** Math uses standard notation with LaTeX; $p(x)$ is the data distribution, $z$ a latent variable, $\theta$ model parameters.
- **Audience.** Written for an engineer comfortable with probability and neural networks who wants the shape of each family — the assumption it makes, the objective it optimizes, and what it trades away — rather than a full derivation.
- **Plain language.** Define a term on first use; prefer plain prose over filler.
15 changes: 15 additions & 0 deletions Deep-Generative-Models/concepts/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
# Concepts

The shared foundations, then one note per model family. Each family note follows the same shape: the assumption it makes about $p(x)$, the objective it optimizes, the milestone models, and the tradeoff it accepts.

| # | Note |
|---|---|
| 0 | [Foundations](./foundations.md) |
| 1 | [Autoregressive models](./autoregressive-models.md) |
| 2 | [Normalizing flows](./normalizing-flows.md) |
| 3 | [Variational autoencoders](./variational-autoencoders.md) |
| 4 | [Generative adversarial networks](./generative-adversarial-networks.md) |
| 5 | [Energy-based models](./energy-based-models.md) |
| 6 | [Diffusion and score-based models](./diffusion-and-score-based-models.md) |

Math: [`../math/objectives-and-transformations.md`](../math/objectives-and-transformations.md). Map: [`../diagrams/model-families.md`](../diagrams/model-families.md).
38 changes: 38 additions & 0 deletions Deep-Generative-Models/concepts/autoregressive-models.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
# Autoregressive models

> Concept note. ~8 min. Builds on [foundations](./foundations.md).

Autoregressive models take the most direct route to an exact likelihood: apply the chain rule of probability and model the joint distribution as a product of conditionals, one variable at a time.

## The idea: factorize by the chain rule

For a data vector $x = (x_1, \dots, x_n)$, the chain rule gives, exactly and without approximation,

$$
p(x) = \prod_{i=1}^{n} p(x_i \mid x_{<i}),
$$

where $x_{<i}$ is everything before $x_i$ in a chosen ordering. Each conditional is a neural network that reads the previous variables and outputs the distribution of the next — a categorical for discrete data (pixels, tokens), a mixture for continuous. Because the factorization is exact, the model computes exact likelihoods and trains by plain maximum likelihood, which not every family can do.

The one imposition is an **ordering**: raster scan (top-to-bottom, left-to-right) for image pixels, time order for audio and text. The model's conditioning is strict — each variable depends only on those before it.

## The line of milestones

The principle is old; the architectures made it work at scale. Fully Visible Sigmoid Belief Networks expressed the conditionals with simple sigmoids; the Neural Autoregressive Density Estimator (NADE) added weight sharing for efficiency. The breakthrough for images was **PixelRNN** and **PixelCNN**, which modeled pixel dependencies — the recurrent version sequentially, the convolutional version with **masked convolutions** that preserve causality while training in parallel. **WaveNet** brought the idea to raw audio with dilated causal convolutions that reach far back in time. Then **Transformers** with masked self-attention — the architecture behind **GPT** — made autoregressive language modeling the dominant paradigm in NLP.

## The tradeoff

Autoregressive models sit firmly in the "exact likelihood" corner of the [trilemma](./foundations.md), and they produce high-quality samples. The price is **sequential sampling**: generating $n$ variables takes $n$ forward passes, because each depends on the last. Training parallelizes (all conditionals are known from the data), but sampling does not, which makes generation slow for long sequences or high-resolution images.

## What to remember

- Autoregressive models factorize $p(x)$ into a product of conditionals via the chain rule — exact, MLE-trainable.
- Masked convolutions (PixelCNN), dilated convolutions (WaveNet), and masked self-attention (GPT) are how the conditionals scale.
- Exact likelihood and strong samples, but sampling is inherently sequential and slow.

## References

- Larochelle, H. & Murray, I. (2011). *The Neural Autoregressive Distribution Estimator (NADE).*
- van den Oord, A., et al. (2016). *Pixel Recurrent Neural Networks.* arXiv:1601.06759.
- van den Oord, A., et al. (2016). *WaveNet.* arXiv:1609.03499.
- Vaswani, A., et al. (2017). *Attention Is All You Need.* arXiv:1706.03762. See [`../references/references.md`](../references/references.md).
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
# Diffusion and score-based models

> Concept note. ~9 min. Builds on [foundations](./foundations.md) and [energy-based models](./energy-based-models.md).

Diffusion models are the current state of the art for image synthesis, and the engines behind modern text-to-image systems. The idea is disarmingly simple: learn to reverse a process that gradually destroys data with noise.

## The idea: destroy, then learn to rebuild

The **forward process** takes a data point and adds a small amount of Gaussian noise, repeatedly, over many steps, until nothing is left but pure noise. This direction is fixed and requires no learning — it is just progressive corruption. The model learns the **reverse process**: at each step, given a noisier version, predict how to denoise it slightly toward the data. Chain enough learned denoising steps together, start from pure noise, and you walk back to a clean, realistic sample.

Training is stable because it reduces to a simple regression: at a random noise level, predict the noise that was added (equivalently, the **score** $\nabla_x \log p(x)$ — the same quantity from [energy-based models](./energy-based-models.md), which is why these are also called *score-based* models). There is no adversarial game and no intractable partition function — just denoising, which is what makes diffusion training far more stable than a GAN's.

## Why it wins on quality, and what it costs

Diffusion models produce samples that rival or exceed GANs in fidelity *and* diversity, without mode collapse — they cover the data distribution rather than fixating on a few modes. That combination is why they took over image generation. The cost lands squarely in the [trilemma](./foundations.md): **sampling is slow**, because generating one sample means running many sequential denoising steps (originally hundreds or thousands). A large body of recent work is about cutting the number of steps — faster samplers, distillation — to close that gap.

## How it relates to the other families

Diffusion ties the module together. It uses the **score** that energy-based models introduced, it can be cast as a (continuous) transformation of noise into data like a [flow](./normalizing-flows.md), and it shares the sequential-sampling cost of [autoregressive models](./autoregressive-models.md) — but trades exact likelihood for sample quality and training stability. It is the clearest example that the families are not isolated; each new one recombines ideas from the others to move to a different corner of the trilemma.

## What to remember

- Diffusion models learn to reverse a fixed noising process: predict the denoising (the score) at each step, then sample by denoising from pure noise.
- Training is a stable regression — no adversarial game, no partition function — which is why quality and diversity are high.
- The cost is slow, multi-step sampling; reducing the step count is the active frontier.

## References

- Sohl-Dickstein, J., et al. (2015). *Deep Unsupervised Learning Using Nonequilibrium Thermodynamics.* arXiv:1503.03585.
- Ho, J., Jain, A., Abbeel, P. (2020). *Denoising Diffusion Probabilistic Models.* arXiv:2006.11239.
- Song, Y. & Ermon, S. (2019). *Generative Modeling by Estimating Gradients of the Data Distribution.* arXiv:1907.05600.
- Song, Y., et al. (2020). *Score-Based Generative Modeling through Stochastic Differential Equations.* arXiv:2011.13456. See [`../references/references.md`](../references/references.md).
40 changes: 40 additions & 0 deletions Deep-Generative-Models/concepts/energy-based-models.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,40 @@
# Energy-based models

> Concept note. ~8 min. Builds on [foundations](./foundations.md).

Energy-based models (EBMs) are the most flexible explicit family and the most awkward to train. They define a distribution through an **energy function** that scores configurations, without committing to a tractable normalized form.

## The idea: an unnormalized density

An EBM assigns each point an energy $E_\theta(x)$ — low energy for likely data, high for unlikely — and defines

$$
p_\theta(x) = \frac{e^{-E_\theta(x)}}{Z_\theta}, \qquad Z_\theta = \int e^{-E_\theta(x)}\,dx.
$$

The energy can be *any* neural network, with no architectural constraints — no invertibility (as flows need), no ordering (as autoregressive models need). That freedom is the appeal. The problem is the denominator $Z_\theta$, the **partition function**: an integral over all of data space that is intractable to compute, which means you cannot evaluate $p_\theta(x)$ directly or take its gradient the easy way.

## Training around the partition function

The field's techniques are ways to learn $E_\theta$ without ever computing $Z_\theta$:

- **Contrastive divergence** pushes energy down on real data and up on samples drawn from the model (via short-run Markov-chain sampling), so the relative energies come out right even though the normalizer is unknown.
- **Score matching** sidesteps $Z_\theta$ entirely by matching the *gradient* of the log density, $\nabla_x \log p_\theta(x)$ — the **score** — which does not depend on the normalizing constant because the constant vanishes under the gradient. This idea is the bridge to [diffusion and score-based models](./diffusion-and-score-based-models.md).

Sampling is also hard: with no direct sampler, EBMs rely on iterative Markov-chain methods (such as Langevin dynamics, which follows the score plus noise), which can be slow to mix.

## The tradeoff

EBMs are maximally flexible in what they can represent, and the score-matching idea they motivate turned out to be foundational. But the intractable partition function makes both training and sampling difficult and often unstable, which kept them less practical than other families on their own — until score matching re-emerged at the heart of diffusion.

## What to remember

- An EBM defines $p(x) \propto e^{-E_\theta(x)}$ with an unconstrained energy network — very flexible.
- The partition function $Z_\theta$ is intractable; contrastive divergence and score matching train without it.
- Sampling needs iterative MCMC (e.g., Langevin); the score-matching idea leads directly to diffusion models.

## References

- Hinton, G. (2002). *Training Products of Experts by Minimizing Contrastive Divergence.*
- Hyvärinen, A. (2005). *Estimation of Non-Normalized Statistical Models by Score Matching.*
- Song, Y. & Kingma, D. P. (2021). *How to Train Your Energy-Based Models.* arXiv:2101.03288. See [`../references/references.md`](../references/references.md).
Loading
Loading