From 21e8251f95184dd08c82ebe6659f1ffc4a94c58d Mon Sep 17 00:00:00 2001 From: MHHamdan Date: Mon, 29 Jun 2026 23:15:05 -0400 Subject: [PATCH] docs(dgm): add Deep Generative Models knowledge module --- Deep-Generative-Models/CHANGELOG.md | 8 +++ Deep-Generative-Models/LICENSE | 16 +++++ Deep-Generative-Models/README.md | 46 +++++++++++++ Deep-Generative-Models/STYLE.md | 7 ++ Deep-Generative-Models/concepts/README.md | 15 ++++ .../concepts/autoregressive-models.md | 38 +++++++++++ .../diffusion-and-score-based-models.md | 32 +++++++++ .../concepts/energy-based-models.md | 40 +++++++++++ .../concepts/foundations.md | 49 +++++++++++++ .../generative-adversarial-networks.md | 29 ++++++++ .../concepts/normalizing-flows.md | 36 ++++++++++ .../concepts/variational-autoencoders.md | 38 +++++++++++ .../diagrams/model-families.md | 19 ++++++ Deep-Generative-Models/glossary/terms.md | 39 +++++++++++ .../math/objectives-and-transformations.md | 68 +++++++++++++++++++ .../references/references.md | 39 +++++++++++ 16 files changed, 519 insertions(+) create mode 100644 Deep-Generative-Models/CHANGELOG.md create mode 100644 Deep-Generative-Models/LICENSE create mode 100644 Deep-Generative-Models/README.md create mode 100644 Deep-Generative-Models/STYLE.md create mode 100644 Deep-Generative-Models/concepts/README.md create mode 100644 Deep-Generative-Models/concepts/autoregressive-models.md create mode 100644 Deep-Generative-Models/concepts/diffusion-and-score-based-models.md create mode 100644 Deep-Generative-Models/concepts/energy-based-models.md create mode 100644 Deep-Generative-Models/concepts/foundations.md create mode 100644 Deep-Generative-Models/concepts/generative-adversarial-networks.md create mode 100644 Deep-Generative-Models/concepts/normalizing-flows.md create mode 100644 Deep-Generative-Models/concepts/variational-autoencoders.md create mode 100644 Deep-Generative-Models/diagrams/model-families.md create mode 100644 Deep-Generative-Models/glossary/terms.md create mode 100644 Deep-Generative-Models/math/objectives-and-transformations.md create mode 100644 Deep-Generative-Models/references/references.md diff --git a/Deep-Generative-Models/CHANGELOG.md b/Deep-Generative-Models/CHANGELOG.md new file mode 100644 index 0000000..6521f7b --- /dev/null +++ b/Deep-Generative-Models/CHANGELOG.md @@ -0,0 +1,8 @@ +# Changelog + +This module is versioned as a single drop within the parent repository. + +## [Unreleased] + +### Added +- **🧬 Deep Generative Models module.** A self-contained knowledge module on deep generative models: a landing README with the family map and the generative trilemma, seven concept notes (foundations; autoregressive models; normalizing flows; variational autoencoders; generative adversarial networks; energy-based models; diffusion and score-based models), one math page (objectives and transformations), one diagram (the family taxonomy), a glossary, and a canonical references list. All content is original and cites the primary papers. diff --git a/Deep-Generative-Models/LICENSE b/Deep-Generative-Models/LICENSE new file mode 100644 index 0000000..ee4b9c3 --- /dev/null +++ b/Deep-Generative-Models/LICENSE @@ -0,0 +1,16 @@ +Deep Generative Models module β€” Licensing +========================================= + +Dual-licensed, consistent with the parent repository. + +SPDX-License-Identifier: Apache-2.0 AND CC-BY-4.0 + +1. Source code (snippets, configuration) β€” Apache-2.0 + https://www.apache.org/licenses/LICENSE-2.0 +2. Prose and diagrams (Markdown, figures) β€” CC-BY-4.0 + https://creativecommons.org/licenses/by/4.0/ + +Referenced papers remain under their own terms and are cited, not reproduced. +This module contains no text copied from any source document; the explanations +are original. If the parent repository carries top-level LICENSE files, those +govern in case of conflict. diff --git a/Deep-Generative-Models/README.md b/Deep-Generative-Models/README.md new file mode 100644 index 0000000..2ddc0ef --- /dev/null +++ b/Deep-Generative-Models/README.md @@ -0,0 +1,46 @@ +# Deep Generative Models + +A self-contained knowledge module on **deep generative models** β€” the families of models that learn the probability distribution behind a dataset well enough to sample new data from it. It mirrors the structure of the rest of this repository: a landing map, modular concept notes, a math page, a diagram, a glossary, and canonical references. + +> Original educational content. The notes are written from scratch and cite the primary papers; no source text is reproduced. See [`STYLE.md`](./STYLE.md). + +## What this covers + +Generative modeling learns a model of $p(x)$ over data β€” images, text, audio β€” so it can do three things: **sample** new instances, **score** the likelihood of a point (density estimation, useful for anomaly detection), and **learn representations** (the latent structure it discovers is useful downstream). The hard part is that the joint distribution over high-dimensional data is intractable to write down directly, and every model family is a different bargain for getting around that. + +## Concept notes + +| # | Note | Family / topic | +|---|---|---| +| 0 | [Foundations](./concepts/foundations.md) | objectives; explicit vs. implicit; likelihood-based vs. -free; latent variables; MLE↔KL; the generative trilemma | +| 1 | [Autoregressive models](./concepts/autoregressive-models.md) | chain-rule factorization; exact likelihood; PixelCNN, WaveNet, GPT | +| 2 | [Normalizing flows](./concepts/normalizing-flows.md) | invertible maps; change of variables; NICE, RealNVP, Glow | +| 3 | [Variational autoencoders](./concepts/variational-autoencoders.md) | latent variables; the ELBO; the reparameterization trick | +| 4 | [Generative adversarial networks](./concepts/generative-adversarial-networks.md) | implicit density; the minimax game; mode collapse | +| 5 | [Energy-based models](./concepts/energy-based-models.md) | unnormalized densities; the partition function; score matching | +| 6 | [Diffusion and score-based models](./concepts/diffusion-and-score-based-models.md) | forward noising, reverse denoising; the current image SOTA | + +## The generative trilemma + +A useful lens for the whole field: three properties you want, and the fact that no single family gets all three for free. + +1. **High sample quality and expressivity** β€” samples that are sharp and varied. +2. **Tractable, exact likelihood** β€” you can compute $p(x)$ and do density estimation. +3. **Fast, stable training and sampling.** + +Autoregressive models and flows give exact likelihood; GANs give sharp samples; diffusion gives quality at the cost of slow sampling; VAEs give fast inference and representations at some cost to sharpness. Each family is a different corner of this triangle β€” the [model-families diagram](./diagrams/model-families.md) lays them out. + +## Repository map + +```text +Deep-Generative-Models/ +β”œβ”€β”€ concepts/ # one note per family, plus shared foundations +β”œβ”€β”€ math/ # the objectives and transformations behind the families +β”œβ”€β”€ diagrams/ # a Mermaid taxonomy of the families +β”œβ”€β”€ glossary/ # shared vocabulary +└── references/ # canonical papers +``` + +## License + +Dual-licensed: code under Apache-2.0, prose and diagrams under CC-BY-4.0. See [`LICENSE`](./LICENSE). diff --git a/Deep-Generative-Models/STYLE.md b/Deep-Generative-Models/STYLE.md new file mode 100644 index 0000000..084daeb --- /dev/null +++ b/Deep-Generative-Models/STYLE.md @@ -0,0 +1,7 @@ +# Style notes + +- **Original writing.** Every note is written from scratch. No text is copied from any source document, textbook, or article. +- **Cite primary sources.** Technical claims point to the original papers (collected in [`references/references.md`](./references/references.md)), not to secondary write-ups. +- **Notation.** Math uses standard notation with LaTeX; $p(x)$ is the data distribution, $z$ a latent variable, $\theta$ model parameters. +- **Audience.** Written for an engineer comfortable with probability and neural networks who wants the shape of each family β€” the assumption it makes, the objective it optimizes, and what it trades away β€” rather than a full derivation. +- **Plain language.** Define a term on first use; prefer plain prose over filler. diff --git a/Deep-Generative-Models/concepts/README.md b/Deep-Generative-Models/concepts/README.md new file mode 100644 index 0000000..3ebe470 --- /dev/null +++ b/Deep-Generative-Models/concepts/README.md @@ -0,0 +1,15 @@ +# Concepts + +The shared foundations, then one note per model family. Each family note follows the same shape: the assumption it makes about $p(x)$, the objective it optimizes, the milestone models, and the tradeoff it accepts. + +| # | Note | +|---|---| +| 0 | [Foundations](./foundations.md) | +| 1 | [Autoregressive models](./autoregressive-models.md) | +| 2 | [Normalizing flows](./normalizing-flows.md) | +| 3 | [Variational autoencoders](./variational-autoencoders.md) | +| 4 | [Generative adversarial networks](./generative-adversarial-networks.md) | +| 5 | [Energy-based models](./energy-based-models.md) | +| 6 | [Diffusion and score-based models](./diffusion-and-score-based-models.md) | + +Math: [`../math/objectives-and-transformations.md`](../math/objectives-and-transformations.md). Map: [`../diagrams/model-families.md`](../diagrams/model-families.md). diff --git a/Deep-Generative-Models/concepts/autoregressive-models.md b/Deep-Generative-Models/concepts/autoregressive-models.md new file mode 100644 index 0000000..776b538 --- /dev/null +++ b/Deep-Generative-Models/concepts/autoregressive-models.md @@ -0,0 +1,38 @@ +# Autoregressive models + +> Concept note. ~8 min. Builds on [foundations](./foundations.md). + +Autoregressive models take the most direct route to an exact likelihood: apply the chain rule of probability and model the joint distribution as a product of conditionals, one variable at a time. + +## The idea: factorize by the chain rule + +For a data vector $x = (x_1, \dots, x_n)$, the chain rule gives, exactly and without approximation, + +$$ +p(x) = \prod_{i=1}^{n} p(x_i \mid x_{ Concept note. ~9 min. Builds on [foundations](./foundations.md) and [energy-based models](./energy-based-models.md). + +Diffusion models are the current state of the art for image synthesis, and the engines behind modern text-to-image systems. The idea is disarmingly simple: learn to reverse a process that gradually destroys data with noise. + +## The idea: destroy, then learn to rebuild + +The **forward process** takes a data point and adds a small amount of Gaussian noise, repeatedly, over many steps, until nothing is left but pure noise. This direction is fixed and requires no learning β€” it is just progressive corruption. The model learns the **reverse process**: at each step, given a noisier version, predict how to denoise it slightly toward the data. Chain enough learned denoising steps together, start from pure noise, and you walk back to a clean, realistic sample. + +Training is stable because it reduces to a simple regression: at a random noise level, predict the noise that was added (equivalently, the **score** $\nabla_x \log p(x)$ β€” the same quantity from [energy-based models](./energy-based-models.md), which is why these are also called *score-based* models). There is no adversarial game and no intractable partition function β€” just denoising, which is what makes diffusion training far more stable than a GAN's. + +## Why it wins on quality, and what it costs + +Diffusion models produce samples that rival or exceed GANs in fidelity *and* diversity, without mode collapse β€” they cover the data distribution rather than fixating on a few modes. That combination is why they took over image generation. The cost lands squarely in the [trilemma](./foundations.md): **sampling is slow**, because generating one sample means running many sequential denoising steps (originally hundreds or thousands). A large body of recent work is about cutting the number of steps β€” faster samplers, distillation β€” to close that gap. + +## How it relates to the other families + +Diffusion ties the module together. It uses the **score** that energy-based models introduced, it can be cast as a (continuous) transformation of noise into data like a [flow](./normalizing-flows.md), and it shares the sequential-sampling cost of [autoregressive models](./autoregressive-models.md) β€” but trades exact likelihood for sample quality and training stability. It is the clearest example that the families are not isolated; each new one recombines ideas from the others to move to a different corner of the trilemma. + +## What to remember + +- Diffusion models learn to reverse a fixed noising process: predict the denoising (the score) at each step, then sample by denoising from pure noise. +- Training is a stable regression β€” no adversarial game, no partition function β€” which is why quality and diversity are high. +- The cost is slow, multi-step sampling; reducing the step count is the active frontier. + +## References + +- Sohl-Dickstein, J., et al. (2015). *Deep Unsupervised Learning Using Nonequilibrium Thermodynamics.* arXiv:1503.03585. +- Ho, J., Jain, A., Abbeel, P. (2020). *Denoising Diffusion Probabilistic Models.* arXiv:2006.11239. +- Song, Y. & Ermon, S. (2019). *Generative Modeling by Estimating Gradients of the Data Distribution.* arXiv:1907.05600. +- Song, Y., et al. (2020). *Score-Based Generative Modeling through Stochastic Differential Equations.* arXiv:2011.13456. See [`../references/references.md`](../references/references.md). diff --git a/Deep-Generative-Models/concepts/energy-based-models.md b/Deep-Generative-Models/concepts/energy-based-models.md new file mode 100644 index 0000000..4e03a6d --- /dev/null +++ b/Deep-Generative-Models/concepts/energy-based-models.md @@ -0,0 +1,40 @@ +# Energy-based models + +> Concept note. ~8 min. Builds on [foundations](./foundations.md). + +Energy-based models (EBMs) are the most flexible explicit family and the most awkward to train. They define a distribution through an **energy function** that scores configurations, without committing to a tractable normalized form. + +## The idea: an unnormalized density + +An EBM assigns each point an energy $E_\theta(x)$ β€” low energy for likely data, high for unlikely β€” and defines + +$$ +p_\theta(x) = \frac{e^{-E_\theta(x)}}{Z_\theta}, \qquad Z_\theta = \int e^{-E_\theta(x)}\,dx. +$$ + +The energy can be *any* neural network, with no architectural constraints β€” no invertibility (as flows need), no ordering (as autoregressive models need). That freedom is the appeal. The problem is the denominator $Z_\theta$, the **partition function**: an integral over all of data space that is intractable to compute, which means you cannot evaluate $p_\theta(x)$ directly or take its gradient the easy way. + +## Training around the partition function + +The field's techniques are ways to learn $E_\theta$ without ever computing $Z_\theta$: + +- **Contrastive divergence** pushes energy down on real data and up on samples drawn from the model (via short-run Markov-chain sampling), so the relative energies come out right even though the normalizer is unknown. +- **Score matching** sidesteps $Z_\theta$ entirely by matching the *gradient* of the log density, $\nabla_x \log p_\theta(x)$ β€” the **score** β€” which does not depend on the normalizing constant because the constant vanishes under the gradient. This idea is the bridge to [diffusion and score-based models](./diffusion-and-score-based-models.md). + +Sampling is also hard: with no direct sampler, EBMs rely on iterative Markov-chain methods (such as Langevin dynamics, which follows the score plus noise), which can be slow to mix. + +## The tradeoff + +EBMs are maximally flexible in what they can represent, and the score-matching idea they motivate turned out to be foundational. But the intractable partition function makes both training and sampling difficult and often unstable, which kept them less practical than other families on their own β€” until score matching re-emerged at the heart of diffusion. + +## What to remember + +- An EBM defines $p(x) \propto e^{-E_\theta(x)}$ with an unconstrained energy network β€” very flexible. +- The partition function $Z_\theta$ is intractable; contrastive divergence and score matching train without it. +- Sampling needs iterative MCMC (e.g., Langevin); the score-matching idea leads directly to diffusion models. + +## References + +- Hinton, G. (2002). *Training Products of Experts by Minimizing Contrastive Divergence.* +- HyvΓ€rinen, A. (2005). *Estimation of Non-Normalized Statistical Models by Score Matching.* +- Song, Y. & Kingma, D. P. (2021). *How to Train Your Energy-Based Models.* arXiv:2101.03288. See [`../references/references.md`](../references/references.md). diff --git a/Deep-Generative-Models/concepts/foundations.md b/Deep-Generative-Models/concepts/foundations.md new file mode 100644 index 0000000..fbcbccd --- /dev/null +++ b/Deep-Generative-Models/concepts/foundations.md @@ -0,0 +1,49 @@ +# Foundations of generative modeling + +> Concept note. ~10 min. Math: [`../math/objectives-and-transformations.md`](../math/objectives-and-transformations.md). + +A generative model learns the probability distribution $p(x)$ behind a dataset β€” of images, text, audio β€” well enough to act on it. The motivation is the one captured in the line often attributed to Feynman, "what I cannot create, I do not understand": a model that can synthesize convincing data has, in some sense, internalized its structure. + +## Three objectives + +Generative models are asked to do three different things, and a given family may be good at some and not others: + +- **Sampling (generation).** Produce new instances that resemble the training data but are not copies β€” a new image, a coherent paragraph, a fresh audio clip. +- **Density estimation.** Assign a probability $p(x)$ to a given point. Low-probability points are unusual, which makes this directly useful for anomaly detection. +- **Representation learning.** In learning to generate, many models discover a compact latent code that captures the salient factors of the data, useful for downstream classification or clustering without labels. + +## Two axes that classify the families + +Almost every model family can be placed by answering two questions. + +**Explicit or implicit density?** An **explicit** model defines a mathematical form for $p(x; \theta)$ you can evaluate β€” autoregressive models and normalizing flows do this, with tractable likelihoods. An **implicit** model never writes $p(x)$ down; it defines a sampling process (a network that maps noise to data) and you can draw from it but not score it. The generative adversarial network is the canonical implicit model. + +**Likelihood-based or likelihood-free?** **Likelihood-based** models train by maximizing the likelihood of the data, or a tractable bound on it β€” autoregressive models, flows, and VAEs. **Likelihood-free** models avoid computing likelihood and train by other means: an adversarial loss (GANs), or score matching and contrastive divergence (energy-based models). + +## Latent variables + +Many families assume the data is generated from an unobserved **latent variable** $z$: draw $z$ from a simple prior $p(z)$ (often a standard Gaussian), then generate $x$ from $p(x \mid z)$. The latent is a compact set of hidden factors β€” pose, style, content β€” and recovering it (inferring the posterior $p(z \mid x)$) is what makes representation learning possible. VAEs are built entirely around this idea. + +## The training principle: MLE is KL minimization + +The default way to fit an explicit model is **maximum likelihood estimation** β€” choose $\theta$ to maximize the likelihood of the training data under $p(x; \theta)$. This is equivalent to minimizing the KL divergence from the data distribution to the model β€” so MLE is, precisely, pulling the model toward the data. The [math page](../math/objectives-and-transformations.md) makes the equivalence explicit. + +## Three challenges, and the trilemma + +The reason there are many families and not one is that high-dimensional $p(x)$ is hard along three axes at once: + +- **Representation** β€” the joint distribution over many variables is astronomically large (every $28\times28$ binary image is one of $2^{784}$ configurations), so you cannot tabulate it. Families impose structure: autoregressive factorization, or a low-dimensional latent manifold. +- **Learning** β€” you only have a finite sample from the true $p_{\text{data}}$; you must pick a divergence and optimize the model to minimize it. +- **Inference** β€” for latent-variable models, recovering $p(z \mid x)$ is its own hard, often intractable, problem. + +These pressures produce the **generative trilemma**: high sample quality, exact tractable likelihood, and fast stable training/sampling β€” pick the corner you need, because no family has all three without compromise. The rest of this module is the families, read as different answers to that trilemma. + +## What to remember + +- Generative models learn $p(x)$ to sample, score, and represent data. +- Classify a family by explicit-vs-implicit density and likelihood-based-vs-free training. +- MLE equals minimizing KL to the data; the generative trilemma (quality, exact likelihood, fast/stable) explains why no single family wins everywhere. + +## References + +- Goodfellow, I., Bengio, Y., Courville, A. (2016). *Deep Learning*, ch. 20 (generative models). See [`../references/references.md`](../references/references.md). diff --git a/Deep-Generative-Models/concepts/generative-adversarial-networks.md b/Deep-Generative-Models/concepts/generative-adversarial-networks.md new file mode 100644 index 0000000..e9e0b85 --- /dev/null +++ b/Deep-Generative-Models/concepts/generative-adversarial-networks.md @@ -0,0 +1,29 @@ +# Generative adversarial networks + +> Concept note. ~8 min. Builds on [foundations](./foundations.md). + +Generative adversarial networks (GANs) take the implicit, likelihood-free route. They never write $p(x)$ down. Instead, two networks compete, and the competition itself drives one of them to produce realistic data. + +## The idea: a two-player game + +A **generator** $G$ maps noise $z$ to data; a **discriminator** $D$ tries to tell real data from the generator's fakes. They train against each other in a minimax game: $D$ learns to classify real vs. fake, while $G$ learns to fool $D$. As $D$ gets sharper, $G$ is pushed to make its samples more realistic, and at the ideal equilibrium the generator's distribution matches the data and the discriminator can do no better than chance. $G$ never sees the data directly β€” it is trained only through $D$'s gradient signal, with no likelihood anywhere in the loop, which is what makes GANs **likelihood-free**. + +## Why they were a leap, and why they are hard + +GANs produce strikingly **sharp** samples β€” sharper than VAEs β€” which made them the image-synthesis workhorse for years (DCGAN for stable convolutional training, StyleGAN for high-fidelity faces). But the adversarial setup is **unstable**: you are looking for an equilibrium of two moving networks, not minimizing a single loss, so training can oscillate or diverge if the two sides fall out of balance. The signature failure is **mode collapse** β€” the generator finds a few outputs that reliably fool the discriminator and produces only those, abandoning the diversity of the data. Much of the GAN literature (alternative losses like Wasserstein GAN, regularizers, careful architectures) is about taming this instability. + +## The tradeoff + +In the [trilemma](./foundations.md), GANs buy sample quality at the cost of the other two corners: no tractable likelihood (you cannot score a point), and fragile training. They sample fast β€” one generator pass β€” but you give up density estimation entirely. + +## What to remember + +- A GAN pits a generator against a discriminator; the generator learns only through the discriminator's signal β€” no likelihood. +- They produce sharp samples (DCGAN, StyleGAN) but train unstably and can suffer mode collapse. +- The tradeoff: high sample quality and fast sampling, but no density estimate and fragile optimization. + +## References + +- Goodfellow, I., et al. (2014). *Generative Adversarial Nets.* arXiv:1406.2661. +- Radford, A., Metz, L., Chintala, S. (2015). *Unsupervised Representation Learning with Deep Convolutional GANs (DCGAN).* arXiv:1511.06434. +- Arjovsky, M., Chintala, S., Bottou, L. (2017). *Wasserstein GAN.* arXiv:1701.07875. See [`../references/references.md`](../references/references.md). diff --git a/Deep-Generative-Models/concepts/normalizing-flows.md b/Deep-Generative-Models/concepts/normalizing-flows.md new file mode 100644 index 0000000..6da901c --- /dev/null +++ b/Deep-Generative-Models/concepts/normalizing-flows.md @@ -0,0 +1,36 @@ +# Normalizing flows + +> Concept note. ~8 min. Builds on [foundations](./foundations.md). Math: [`../math/objectives-and-transformations.md`](../math/objectives-and-transformations.md). + +Normalizing flows get an exact likelihood a different way: start from a simple distribution and push it through a sequence of **invertible** transformations until it matches the data. + +## The idea: change of variables + +Let $z$ come from a simple prior $p_Z(z)$ (a standard Gaussian), and let $x = f_\theta(z)$ for an invertible, differentiable $f_\theta$. The change-of-variables formula gives the data density exactly: + +$$ +p_X(x) = p_Z\!\big(f_\theta^{-1}(x)\big)\;\Big| \det J_{f_\theta^{-1}}(x) \Big|, +$$ + +where $J$ is the Jacobian of the inverse map. The determinant accounts for how the transformation stretches or compresses volume, keeping the density normalized. Because the expression is exact, flows train by exact maximum likelihood β€” like autoregressive models, but with a single invertible map (or a stack of them) instead of a sequential factorization. + +## The engineering problem: a cheap Jacobian + +The catch is the determinant. For a general $n\times n$ Jacobian it costs $O(n^3)$, which is hopeless at image scale. The whole craft of flows is designing transformations that are both expressive and have a Jacobian determinant you can compute in $O(n)$ β€” typically by making the Jacobian **triangular**. **Coupling layers**, introduced in NICE, do this elegantly: split the variables in two, leave one half unchanged, and transform the other half using only the first β€” the Jacobian is triangular and the determinant is trivial. **RealNVP** generalized these to affine coupling for more flexibility; **Glow** added invertible $1\times1$ convolutions to mix channels; **MAF** and **IAF** built autoregressive structure into the flow. + +## The tradeoff + +Flows give exact likelihood and a meaningful latent space, with parallel sampling (unlike autoregressive models). Their structural constraint is that the latent $z$ must have the **same dimension** as the data $x$ β€” every transformation is a bijection β€” so they cannot compress to a lower-dimensional code the way a VAE can, and the invertibility requirement limits the architectures available. + +## What to remember + +- Flows transform a simple prior through invertible maps; the change-of-variables formula gives exact likelihood. +- The art is a transformation with a cheap (triangular) Jacobian determinant β€” coupling layers (NICE, RealNVP), $1\times1$ convolutions (Glow). +- Exact likelihood and parallel sampling, but the latent must match the data dimension β€” no built-in compression. + +## References + +- Dinh, L., Krueger, D., Bengio, Y. (2014). *NICE: Non-linear Independent Components Estimation.* arXiv:1410.8516. +- Dinh, L., Sohl-Dickstein, J., Bengio, S. (2016). *Density Estimation Using Real NVP.* arXiv:1605.08803. +- Kingma, D. P. & Dhariwal, P. (2018). *Glow.* arXiv:1807.03039. +- Papamakarios, G., et al. (2019). *Normalizing Flows for Probabilistic Modeling and Inference.* arXiv:1912.02762. See [`../references/references.md`](../references/references.md). diff --git a/Deep-Generative-Models/concepts/variational-autoencoders.md b/Deep-Generative-Models/concepts/variational-autoencoders.md new file mode 100644 index 0000000..d8fcf32 --- /dev/null +++ b/Deep-Generative-Models/concepts/variational-autoencoders.md @@ -0,0 +1,38 @@ +# Variational autoencoders + +> Concept note. ~9 min. Builds on [foundations](./foundations.md). Math: [`../math/objectives-and-transformations.md`](../math/objectives-and-transformations.md). + +A variational autoencoder (VAE) is the canonical latent-variable generative model. It assumes each data point $x$ is generated from a low-dimensional latent $z$, learns to map between the two with a pair of networks, and trains on a tractable bound on the likelihood. + +## The two networks and the intractable posterior + +A VAE has a **decoder** $p_\theta(x \mid z)$ that generates data from a latent (with a simple prior $p(z)$, usually a standard Gaussian), and an **encoder** $q_\phi(z \mid x)$ that infers a latent from data. The encoder exists because the true posterior $p_\theta(z \mid x)$ β€” the right latent for a given $x$ β€” is intractable to compute. The encoder is an *approximate* posterior, a network trained to stand in for it. + +## The objective: the ELBO + +Because the exact log-likelihood $\log p_\theta(x)$ is intractable, VAEs maximize a lower bound on it, the **evidence lower bound (ELBO)**: + +$$ +\log p_\theta(x) \;\ge\; \underbrace{\mathbb{E}_{q_\phi(z\mid x)}\big[\log p_\theta(x \mid z)\big]}_{\text{reconstruction}} \;-\; \underbrace{\mathrm{KL}\!\big(q_\phi(z\mid x)\,\|\,p(z)\big)}_{\text{regularizer}}. +$$ + +The two terms pull in useful tension. The **reconstruction** term wants latents that let the decoder rebuild the input. The **KL** term keeps the encoder's distribution close to the prior, which keeps the latent space smooth and regular β€” so you can sample a $z$ from the prior and decode it into something coherent. Maximizing the ELBO trains encoder and decoder together. + +## The reparameterization trick + +There is a snag: you need gradients to flow back through a *sampling* step ($z \sim q_\phi(z \mid x)$), and sampling is not differentiable. The **reparameterization trick** fixes it by moving the randomness outside the network: instead of sampling $z$ directly, sample noise $\epsilon \sim \mathcal{N}(0, I)$ and compute $z = \mu_\phi(x) + \sigma_\phi(x)\,\epsilon$. Now $z$ is a deterministic, differentiable function of the network outputs and an external noise source, and gradients pass through. This trick is what makes VAEs trainable by ordinary backpropagation. + +## The tradeoff + +VAEs give fast generation (one decoder pass), an explicit latent space useful for [representation learning](./foundations.md), and principled probabilistic training. Their well-known weakness is **sample sharpness**: samples are often blurrier than a GAN's or a diffusion model's, partly because the bound and the Gaussian assumptions smooth things out. In the [trilemma](./foundations.md), VAEs favor fast, stable training and useful inference over maximal sample fidelity. + +## What to remember + +- A VAE pairs a decoder $p_\theta(x\mid z)$ with an approximate-posterior encoder $q_\phi(z\mid x)$, because the true posterior is intractable. +- It maximizes the ELBO: a reconstruction term plus a KL regularizer that keeps the latent space smooth. +- The reparameterization trick makes sampling differentiable; the tradeoff is blurrier samples for fast inference and representations. + +## References + +- Kingma, D. P. & Welling, M. (2013). *Auto-Encoding Variational Bayes.* arXiv:1312.6114. +- Rezende, D. J., Mohamed, S., Wierstra, D. (2014). *Stochastic Backpropagation and Approximate Inference in Deep Generative Models.* arXiv:1401.4082. See [`../references/references.md`](../references/references.md). diff --git a/Deep-Generative-Models/diagrams/model-families.md b/Deep-Generative-Models/diagrams/model-families.md new file mode 100644 index 0000000..393511c --- /dev/null +++ b/Deep-Generative-Models/diagrams/model-families.md @@ -0,0 +1,19 @@ +# Diagram: the deep generative model families + +A taxonomy of the families by how they treat the density $p(x)$, with the tradeoff each accepts. Reused by [foundations](../concepts/foundations.md) and the family notes. + +```mermaid +flowchart TB + G["generative model
learn p(x)"] --> EX["explicit density"] + G --> IM["implicit density"] + EX --> TR["tractable likelihood"] + EX --> AP["approximate / intractable"] + TR --> AR["autoregressive
(PixelCNN, WaveNet, GPT)"] + TR --> FL["normalizing flows
(NICE, RealNVP, Glow)"] + AP --> VA["variational autoencoders
(ELBO bound)"] + AP --> EB["energy-based models
(score matching)"] + AP --> DF["diffusion / score-based
(DDPM, score SDE)"] + IM --> GA["GANs
(adversarial, likelihood-free)"] +``` + +Reading the tree: **explicit** families write a form for $p(x)$; among them, autoregressive models and flows give *tractable, exact* likelihood, while VAEs (a bound), energy-based models (an unnormalized form), and diffusion (a learned score) settle for *approximate or implicit* likelihood. **Implicit** models β€” GANs β€” never express $p(x)$ at all. Each leaf is a different corner of the [generative trilemma](../concepts/foundations.md): autoregressive and flows favor exact likelihood, GANs favor sharp samples, diffusion favors quality at the cost of slow sampling, and VAEs favor fast inference and representations. diff --git a/Deep-Generative-Models/glossary/terms.md b/Deep-Generative-Models/glossary/terms.md new file mode 100644 index 0000000..85cc347 --- /dev/null +++ b/Deep-Generative-Models/glossary/terms.md @@ -0,0 +1,39 @@ +# Glossary + +Vocabulary for the deep generative models module. Alphabetical; each entry links to where it is used. + +**Autoregressive model.** A generative model that factorizes $p(x)$ into a product of conditionals via the chain rule, one variable at a time, giving exact likelihood at the cost of sequential sampling. β†’ `concepts/autoregressive-models.md`. + +**Change of variables.** The formula relating the densities of $z$ and $x = f(z)$ for an invertible $f$, via the Jacobian determinant; the basis of normalizing flows. β†’ `concepts/normalizing-flows.md`. + +**Contrastive divergence.** A training method for energy-based models that lowers energy on data and raises it on model samples, avoiding the partition function. β†’ `concepts/energy-based-models.md`. + +**Diffusion model.** A generative model that learns to reverse a fixed noising process, sampling by denoising from pure noise; current state of the art for images. β†’ `concepts/diffusion-and-score-based-models.md`. + +**ELBO (evidence lower bound).** A tractable lower bound on the log-likelihood, maximized by VAEs; a reconstruction term minus a KL regularizer. β†’ `concepts/variational-autoencoders.md`. + +**Energy-based model (EBM).** A model defining $p(x) \propto e^{-E_\theta(x)}$ with an unconstrained energy network; flexible but burdened by an intractable partition function. β†’ `concepts/energy-based-models.md`. + +**Explicit vs. implicit density.** Explicit models write an evaluable form for $p(x)$; implicit models only define a sampling process (e.g. a GAN). β†’ `concepts/foundations.md`. + +**GAN (generative adversarial network).** An implicit, likelihood-free model in which a generator and discriminator compete; sharp samples, unstable training. β†’ `concepts/generative-adversarial-networks.md`. + +**Generative trilemma.** The observation that high sample quality, exact tractable likelihood, and fast stable training/sampling cannot all be had at once; each family is a different compromise. β†’ `concepts/foundations.md`. + +**Latent variable.** An unobserved variable $z$ assumed to generate the data $x$ via $p(x\mid z)$; recovering $p(z\mid x)$ enables representation learning. β†’ `concepts/foundations.md`. + +**Likelihood-based vs. likelihood-free.** Likelihood-based models train on the data likelihood or a bound; likelihood-free models use other signals (adversarial loss, score matching). β†’ `concepts/foundations.md`. + +**Maximum likelihood estimation (MLE).** Fitting $\theta$ to maximize the data likelihood; equivalent to minimizing the KL divergence from the data distribution to the model. β†’ `concepts/foundations.md`. + +**Mode collapse.** A GAN failure where the generator produces only a few outputs that fool the discriminator, losing the diversity of the data. β†’ `concepts/generative-adversarial-networks.md`. + +**Normalizing flow.** A generative model that transforms a simple prior through invertible maps, giving exact likelihood via the change-of-variables formula. β†’ `concepts/normalizing-flows.md`. + +**Partition function.** The normalizing constant $Z_\theta = \int e^{-E_\theta(x)}dx$ of an energy-based model; intractable, which is why EBMs are hard to train. β†’ `concepts/energy-based-models.md`. + +**Reparameterization trick.** Writing $z = \mu + \sigma\odot\epsilon$ with $\epsilon\sim\mathcal N(0,I)$ so that sampling becomes differentiable and gradients flow through a VAE. β†’ `concepts/variational-autoencoders.md`. + +**Score.** The gradient of the log density, $\nabla_x\log p(x)$; independent of the partition function, and the quantity diffusion and score-based models learn. β†’ `concepts/diffusion-and-score-based-models.md`. + +**Variational autoencoder (VAE).** A latent-variable model pairing an encoder (approximate posterior) with a decoder, trained by maximizing the ELBO. β†’ `concepts/variational-autoencoders.md`. diff --git a/Deep-Generative-Models/math/objectives-and-transformations.md b/Deep-Generative-Models/math/objectives-and-transformations.md new file mode 100644 index 0000000..f3e79b5 --- /dev/null +++ b/Deep-Generative-Models/math/objectives-and-transformations.md @@ -0,0 +1,68 @@ +# Objectives and transformations + +> Mathematical foundation. ~9 min. Supports the [concept notes](../concepts/). + +Four pieces of math recur across the generative families: the MLE–KL equivalence (how almost all of them train), the change-of-variables formula (flows), the ELBO (VAEs), and the score (energy-based and diffusion models). This page states each compactly. + +## Maximum likelihood is KL minimization + +Given data $x^{(1)}, \dots, x^{(m)}$ drawn from $p_{\text{data}}$, maximum likelihood chooses + +$$ +\theta^\* = \arg\max_\theta \frac{1}{m}\sum_{i=1}^m \log p_\theta\!\big(x^{(i)}\big). +$$ + +As $m \to \infty$ the average converges to $\mathbb{E}_{p_{\text{data}}}[\log p_\theta(x)]$, and + +$$ +\mathrm{KL}\!\big(p_{\text{data}} \,\|\, p_\theta\big) = \mathbb{E}_{p_{\text{data}}}[\log p_{\text{data}}(x)] - \mathbb{E}_{p_{\text{data}}}[\log p_\theta(x)]. +$$ + +The first term does not depend on $\theta$, so maximizing the log-likelihood is exactly minimizing $\mathrm{KL}(p_{\text{data}} \,\|\, p_\theta)$ β€” fitting by MLE pulls the model toward the data distribution in KL. + +## Change of variables (flows) + +If $x = f_\theta(z)$ is invertible and differentiable and $z \sim p_Z$, the density of $x$ is + +$$ +p_X(x) = p_Z\!\big(f_\theta^{-1}(x)\big)\,\Big|\det J_{f_\theta^{-1}}(x)\Big|, +\qquad J_{f_\theta^{-1}}(x) = \frac{\partial f_\theta^{-1}}{\partial x}. +$$ + +The log-likelihood a flow maximizes is then $\log p_Z(f_\theta^{-1}(x)) + \log|\det J|$. The engineering goal is a map whose $\log|\det J|$ costs $O(n)$, achieved by a triangular Jacobian (coupling layers). + +## The ELBO (VAEs) + +For a latent-variable model with prior $p(z)$, decoder $p_\theta(x\mid z)$, and approximate posterior $q_\phi(z\mid x)$, the marginal log-likelihood decomposes as + +$$ +\log p_\theta(x) = \underbrace{\mathbb{E}_{q_\phi}\!\Big[\log \tfrac{p_\theta(x,z)}{q_\phi(z\mid x)}\Big]}_{\text{ELBO}} + \mathrm{KL}\!\big(q_\phi(z\mid x)\,\|\,p_\theta(z\mid x)\big). +$$ + +The KL term is $\ge 0$ and involves the intractable true posterior, so the ELBO is a lower bound on $\log p_\theta(x)$. Rearranging the ELBO gives the trainable form, + +$$ +\text{ELBO} = \mathbb{E}_{q_\phi(z\mid x)}\big[\log p_\theta(x\mid z)\big] - \mathrm{KL}\!\big(q_\phi(z\mid x)\,\|\,p(z)\big), +$$ + +a reconstruction term minus a regularizer. The reparameterization $z = \mu_\phi(x) + \sigma_\phi(x)\odot\epsilon$, $\epsilon\sim\mathcal N(0,I)$, makes the expectation differentiable in $\phi$. + +## The score (energy-based and diffusion) + +The **score** of a density is the gradient of its log, $s(x) = \nabla_x \log p(x)$. For an energy-based model $p_\theta(x) = e^{-E_\theta(x)}/Z_\theta$, + +$$ +\nabla_x \log p_\theta(x) = -\nabla_x E_\theta(x), +$$ + +and the intractable normalizer $Z_\theta$ drops out because it does not depend on $x$ β€” which is why matching the score avoids the partition function. Diffusion models learn this score at every noise level and sample by following it (Langevin-style: $x \leftarrow x + \tfrac{\epsilon}{2}\nabla_x\log p(x) + \sqrt{\epsilon}\,\eta$), walking noise back to data. + +## What to remember + +- MLE $\equiv$ minimizing $\mathrm{KL}(p_{\text{data}}\,\|\,p_\theta)$ β€” the shared training principle. +- Change of variables gives flows an exact likelihood; a triangular Jacobian makes it cheap. +- The ELBO is a tractable lower bound for latent-variable models; the score removes the partition function and underlies diffusion. + +## See also + +- [`../concepts/`](../concepts/) β€” the families that use these objectives. diff --git a/Deep-Generative-Models/references/references.md b/Deep-Generative-Models/references/references.md new file mode 100644 index 0000000..5e214ca --- /dev/null +++ b/Deep-Generative-Models/references/references.md @@ -0,0 +1,39 @@ +# References + +Canonical primary sources for the deep generative models module. Explanations in this module are original; these are the works they draw on, cited rather than reproduced. + +## Foundations +- Goodfellow, I., Bengio, Y., & Courville, A. (2016). *Deep Learning* (ch. 20). + +## Autoregressive models +- Larochelle, H. & Murray, I. (2011). *The Neural Autoregressive Distribution Estimator.* AISTATS. +- van den Oord, A., Kalchbrenner, N., & Kavukcuoglu, K. (2016). *Pixel Recurrent Neural Networks.* arXiv:1601.06759. +- van den Oord, A., et al. (2016). *WaveNet: A Generative Model for Raw Audio.* arXiv:1609.03499. +- Vaswani, A., et al. (2017). *Attention Is All You Need.* arXiv:1706.03762. + +## Normalizing flows +- Dinh, L., Krueger, D., & Bengio, Y. (2014). *NICE: Non-linear Independent Components Estimation.* arXiv:1410.8516. +- Dinh, L., Sohl-Dickstein, J., & Bengio, S. (2016). *Density Estimation Using Real NVP.* arXiv:1605.08803. +- Kingma, D. P. & Dhariwal, P. (2018). *Glow: Generative Flow with Invertible 1Γ—1 Convolutions.* arXiv:1807.03039. +- Papamakarios, G., et al. (2019). *Normalizing Flows for Probabilistic Modeling and Inference.* arXiv:1912.02762. + +## Variational autoencoders +- Kingma, D. P. & Welling, M. (2013). *Auto-Encoding Variational Bayes.* arXiv:1312.6114. +- Rezende, D. J., Mohamed, S., & Wierstra, D. (2014). *Stochastic Backpropagation and Approximate Inference in Deep Generative Models.* arXiv:1401.4082. + +## Generative adversarial networks +- Goodfellow, I., et al. (2014). *Generative Adversarial Nets.* arXiv:1406.2661. +- Radford, A., Metz, L., & Chintala, S. (2015). *Unsupervised Representation Learning with Deep Convolutional GANs.* arXiv:1511.06434. +- Arjovsky, M., Chintala, S., & Bottou, L. (2017). *Wasserstein GAN.* arXiv:1701.07875. +- Karras, T., Laine, S., & Aila, T. (2018). *A Style-Based Generator Architecture for GANs (StyleGAN).* arXiv:1812.04948. + +## Energy-based models +- Hinton, G. (2002). *Training Products of Experts by Minimizing Contrastive Divergence.* Neural Computation. +- HyvΓ€rinen, A. (2005). *Estimation of Non-Normalized Statistical Models by Score Matching.* JMLR. +- Song, Y. & Kingma, D. P. (2021). *How to Train Your Energy-Based Models.* arXiv:2101.03288. + +## Diffusion and score-based models +- Sohl-Dickstein, J., et al. (2015). *Deep Unsupervised Learning Using Nonequilibrium Thermodynamics.* arXiv:1503.03585. +- Song, Y. & Ermon, S. (2019). *Generative Modeling by Estimating Gradients of the Data Distribution.* arXiv:1907.05600. +- Ho, J., Jain, A., & Abbeel, P. (2020). *Denoising Diffusion Probabilistic Models.* arXiv:2006.11239. +- Song, Y., et al. (2020). *Score-Based Generative Modeling through Stochastic Differential Equations.* arXiv:2011.13456.