Intro to Diffusion Models

Introduction Forward Diffusion Reverse Process and Loss Sampling Interactive Demo

Introduction

A diffusion model rests on an idea that is unusual at first glance. To generate, one starts from pure noise and gradually denoises it into a sample, reversing a process that incrementally adds Gaussian noise to data. Nothing in the construction is specific to images, and it applies to any data represented as vectors in \(\mathbb{R}^d\). This page develops the mathematics behind that picture.

Two curriculum threads converge here. The first comes from our treatment of the variational autoencoder, where we built a generative model around a learned encoder \(q_{\boldsymbol{\phi}}(\boldsymbol{z} \mid \boldsymbol{x})\), a decoder \(p_{\boldsymbol{\theta}}(\boldsymbol{x} \mid \boldsymbol{z})\), and the evidence lower bound as the training objective. The diffusion model inherits the ELBO formalism but rearranges the architecture. The analogue of the encoder is a fixed, non-learned Markov chain that progressively corrupts data into noise, while the analogue of the decoder is a learned Markov chain that reverses the corruption. Encoder learning, the central computational difficulty of the VAE, is absent by design.

The second thread is more subtle and was deliberately seeded earlier in the curriculum. Our page on PCA and autoencoders introduced the denoising autoencoder and stated, as a result taken on faith, that such a network implicitly learns the score function \(\nabla_{\mathbf{x}} \log p(\mathbf{x})\), the gradient of the log-density with respect to the data itself. The score-matching theory justifying this claim was deferred at the time. The diffusion model is where that thread is recovered, once we reach the reverse process.

The development proceeds in three movements.

  1. We first define the forward diffusion process (the fixed Markov chain that maps data to noise) and derive its closed-form marginals.
  2. We then construct the reverse process and its loss. We decompose the negative ELBO into a sum of Kullback-Leibler terms, arrive at the simplified noise-prediction objective of Ho, Jain, and Abbeel's denoising diffusion probabilistic models (DDPM, 2020), and connect this objective to the score function.
  3. We close with sampling, covering the ancestral sampler of DDPM, the deterministic accelerated DDIM sampler, and the classifier-free guidance mechanism for conditional generation.

A continuous-time formulation of diffusion exists and is mathematically rich. Taking the number of steps to infinity turns the discrete forward chain into a stochastic differential equation, the variance-preserving SDE. Anderson's 1982 reversal theorem yields a corresponding reverse SDE, and the deterministic DDIM sampler corresponds in this limit to a so-called probability-flow ODE. A rigorous treatment requires Brownian motion and Itô calculus, which the stochastic-analysis pages of the probability section now develop in full. We mention the continuous-time picture in passing where it illuminates the discrete-time mathematics, and the formal development lives in the study of the Fokker-Planck equation.

The Forward Diffusion Process

The forward process is the simpler of the two halves of a diffusion model. It is fixed, requires no training, and is specified by a single design choice, the rate at which Gaussian noise is added to the data at each step. Let \(\mathbf{x}_0 \in \mathbb{R}^d\) denote a data sample drawn from an unknown data distribution \(q_0(\mathbf{x}_0)\), and let \(T\) denote the total number of diffusion steps. We choose a sequence of small positive numbers \(\{\beta_t\}_{t=1}^T \subset (0, 1)\), called the noise schedule, and define a Markov chain \(\mathbf{x}_0 \to \mathbf{x}_1 \to \cdots \to \mathbf{x}_T\) by the following transition.

Definition: Forward Diffusion Kernel

Given a noise schedule \(\{\beta_t\}_{t=1}^T \subset (0, 1)\), the forward diffusion kernel is the Gaussian transition \[ q(\mathbf{x}_t \mid \mathbf{x}_{t-1}) = \mathcal{N}\!\left(\mathbf{x}_t; \sqrt{1 - \beta_t}\,\mathbf{x}_{t-1}, \beta_t \mathbf{I}\right), \quad t = 1, \ldots, T. \]

Three properties of this kernel matter.

  1. The chain is first-order Markov. Each \(\mathbf{x}_t\) depends only on its immediate predecessor \(\mathbf{x}_{t-1}\), so the joint conditional distribution factorizes as \[ q(\mathbf{x}_{1:T} \mid \mathbf{x}_0) = \prod_{t=1}^{T} q(\mathbf{x}_t \mid \mathbf{x}_{t-1}). \]
  2. The kernel contains no learned parameters. Here \(\beta_t\) is fixed by the schedule, and the mean and covariance are closed-form functions of \(\mathbf{x}_{t-1}\).
  3. The scaling \(\sqrt{1 - \beta_t}\) on the mean is not arbitrary. It is chosen so that, if \(\mathbf{x}_{t-1}\) has identity covariance, then \(\mathbf{x}_t\) does too: \(\operatorname{Cov}(\mathbf{x}_t) = (1-\beta_t)\mathbf{I} + \beta_t \mathbf{I} = \mathbf{I}\). For general data, independence of the injected noise gives \(\operatorname{Cov}(\mathbf{x}_t) = \bar{\alpha}_t \operatorname{Cov}(\mathbf{x}_0) + (1 - \bar{\alpha}_t)\mathbf{I}\), with \(\bar{\alpha}_t\) as defined below, so the marginal covariance moves from that of the data toward the identity along the chain.

Forward kernels with this scaling are called variance-preserving, and the form above is the variance-preserving construction of Ho, Jain, and Abbeel.

Closed-Form Marginal

Because the forward chain is a linear Gaussian Markov chain, the marginal distribution \(q(\mathbf{x}_t \mid \mathbf{x}_0)\) at any time \(t\) can be computed in closed form. We introduce the abbreviations \[ \alpha_t := 1 - \beta_t, \quad \bar{\alpha}_t := \prod_{s=1}^{t} \alpha_s, \] so that \(\bar{\alpha}_t\) is the cumulative product of the per-step retention factors.

Theorem: Closed-Form Diffusion Kernel

For every \(t = 1, \ldots, T\) and every \(\mathbf{x}_0\), \[ q(\mathbf{x}_t \mid \mathbf{x}_0) = \mathcal{N}\!\left(\mathbf{x}_t; \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0, (1 - \bar{\alpha}_t)\mathbf{I}\right). \]

Proof:

We argue by induction on \(t\). The base case \(t = 1\) is the kernel itself with \(\bar{\alpha}_1 = \alpha_1 = 1 - \beta_1\). For the inductive step, suppose \(q(\mathbf{x}_{t-1} \mid \mathbf{x}_0) = \mathcal{N}(\mathbf{x}_{t-1};\, \sqrt{\bar{\alpha}_{t-1}}\,\mathbf{x}_0,\, (1-\bar{\alpha}_{t-1})\mathbf{I})\). By the reparameterization trick, we can write \[ \mathbf{x}_{t-1} = \sqrt{\bar{\alpha}_{t-1}}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_{t-1}}\,\boldsymbol{\varepsilon}_{t-1}, \quad \boldsymbol{\varepsilon}_{t-1} \sim \mathcal{N}(\mathbf{0}, \mathbf{I}), \] and similarly \(\mathbf{x}_t = \sqrt{\alpha_t}\,\mathbf{x}_{t-1} + \sqrt{\beta_t}\,\boldsymbol{\varepsilon}_t\) with \(\boldsymbol{\varepsilon}_t \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\) independent of \(\boldsymbol{\varepsilon}_{t-1}\). Substituting, \[ \mathbf{x}_t = \sqrt{\alpha_t \bar{\alpha}_{t-1}}\,\mathbf{x}_0 + \sqrt{\alpha_t (1 - \bar{\alpha}_{t-1})}\,\boldsymbol{\varepsilon}_{t-1} + \sqrt{\beta_t}\,\boldsymbol{\varepsilon}_t. \]

The last two terms are independent zero-mean Gaussians, so their sum is Gaussian. Using \(\alpha_t \bar{\alpha}_{t-1} = \bar{\alpha}_t\), which holds by definition, its variance is \[ \begin{align*} \alpha_t(1 - \bar{\alpha}_{t-1}) + \beta_t &= \alpha_t - \alpha_t \bar{\alpha}_{t-1} + (1 - \alpha_t) \\\\ &= 1 - \bar{\alpha}_t, \end{align*} \] and the coefficient of \(\mathbf{x}_0\) is \(\sqrt{\alpha_t \bar{\alpha}_{t-1}} = \sqrt{\bar{\alpha}_t}\). Hence \[ \mathbf{x}_t = \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\varepsilon}, \quad \boldsymbol{\varepsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I}), \] which is the reparameterization of the claimed Gaussian.

The identity at the end of the proof deserves its own line. Any noisy sample at time \(t\) is produced in one shot from the clean data as \[ \mathbf{x}_t = \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\varepsilon}, \quad \boldsymbol{\varepsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I}). \] This one-shot reparameterization is the central computational identity of training. Rather than simulating the chain step by step, the trainer samples \(t\) uniformly and \(\boldsymbol{\varepsilon}\) from a standard Gaussian, and produces a noisy sample directly.

The Endpoint and the Choice of Schedule

The schedule \(\{\beta_t\}\) is chosen so that \(\bar{\alpha}_T\) is very close to zero. From the closed-form marginal, this forces \(q(\mathbf{x}_T \mid \mathbf{x}_0) \approx \mathcal{N}(\mathbf{0}, \mathbf{I})\), and the dependence on \(\mathbf{x}_0\) is washed out. Two schedules are most common in practice. The linear schedule interpolates \(\beta_t\) linearly between small endpoints (Ho, Jain, and Abbeel, 2020). The cosine schedule keeps \(\bar{\alpha}_t\) close to one for longer at small \(t\) and so destroys information less abruptly, which improved results on lower-resolution images (Nichol and Dhariwal, 2021). Typical values of \(T\) range from a few hundred to a few thousand. The schedule and \(T\) are hyperparameters of the model, not learned quantities.

The Reverse Process and Loss

The Reverse Posterior

To generate samples, we would like to run the chain backwards, starting from \(\mathbf{x}_T \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\) and successively sampling from \(q(\mathbf{x}_{t-1} \mid \mathbf{x}_t)\) until we recover \(\mathbf{x}_0\). The obstruction is that this reverse transition is not directly computable. Writing it as \[ q(\mathbf{x}_{t-1} \mid \mathbf{x}_t) = \frac{q(\mathbf{x}_t \mid \mathbf{x}_{t-1})\, q(\mathbf{x}_{t-1})}{q(\mathbf{x}_t)}, \] we see that both \(q(\mathbf{x}_{t-1})\) and \(q(\mathbf{x}_t)\) require integrating the forward chain against the unknown data distribution \(q_0(\mathbf{x}_0)\), and so are inaccessible.

A different conditional, however, is tractable. It is the reverse process conditioned on the clean data \(\mathbf{x}_0\). This auxiliary distribution will play the role of an oracle in the loss derivation of the next subsection. We cannot use it for sampling, since \(\mathbf{x}_0\) is precisely what we are trying to generate, but we can use it during training, where \(\mathbf{x}_0\) is drawn from the training set and is therefore available.

The derivation begins from Bayes' rule together with the Markov property of the forward chain. Because \(\mathbf{x}_t\) depends on \(\mathbf{x}_{t-1}\) only through the forward kernel and not through \(\mathbf{x}_0\), we have \(q(\mathbf{x}_t \mid \mathbf{x}_{t-1}, \mathbf{x}_0) = q(\mathbf{x}_t \mid \mathbf{x}_{t-1})\). Bayes' rule then gives \[ q(\mathbf{x}_{t-1} \mid \mathbf{x}_t, \mathbf{x}_0) = \frac{q(\mathbf{x}_t \mid \mathbf{x}_{t-1})\, q(\mathbf{x}_{t-1} \mid \mathbf{x}_0)}{q(\mathbf{x}_t \mid \mathbf{x}_0)}. \]

Every quantity on the right is a Gaussian whose mean and variance we already know. The denominator is independent of \(\mathbf{x}_{t-1}\) and serves only as a normalizing constant. The numerator is a product of two Gaussian densities in \(\mathbf{x}_{t-1}\), so for \(t \geq 2\) its logarithm is, up to terms free of \(\mathbf{x}_{t-1}\), \[ -\frac{1}{2} \left( \frac{\left\| \mathbf{x}_t - \sqrt{\alpha_t}\,\mathbf{x}_{t-1} \right\|^2}{\beta_t} + \frac{\left\| \mathbf{x}_{t-1} - \sqrt{\bar{\alpha}_{t-1}}\,\mathbf{x}_0 \right\|^2}{1 - \bar{\alpha}_{t-1}} \right). \] This is a quadratic in \(\mathbf{x}_{t-1}\) whose coefficient matrix is \[ \left( \frac{\alpha_t}{\beta_t} + \frac{1}{1 - \bar{\alpha}_{t-1}} \right) \mathbf{I} = \frac{1 - \bar{\alpha}_t}{\beta_t (1 - \bar{\alpha}_{t-1})}\,\mathbf{I}, \] where we used \(\alpha_t (1 - \bar{\alpha}_{t-1}) + \beta_t = 1 - \bar{\alpha}_t\). Completing the square therefore yields a multivariate normal distribution whose covariance is the inverse of this matrix and whose mean is that covariance applied to the linear coefficient \(\sqrt{\alpha_t}\,\mathbf{x}_t / \beta_t + \sqrt{\bar{\alpha}_{t-1}}\,\mathbf{x}_0 / (1 - \bar{\alpha}_{t-1})\).

Definition: Reverse Posterior

For \(t \geq 2\), the reverse posterior conditioned on \(\mathbf{x}_0\) is the Gaussian \[ q(\mathbf{x}_{t-1} \mid \mathbf{x}_t, \mathbf{x}_0) = \mathcal{N}\!\left(\mathbf{x}_{t-1}; \tilde{\boldsymbol{\mu}}_t(\mathbf{x}_t, \mathbf{x}_0), \tilde{\beta}_t \mathbf{I}\right), \] with mean and variance \[ \begin{align*} \tilde{\boldsymbol{\mu}}_t(\mathbf{x}_t, \mathbf{x}_0) &= \frac{\sqrt{\bar{\alpha}_{t-1}}\,\beta_t}{1 - \bar{\alpha}_t}\,\mathbf{x}_0 + \frac{\sqrt{\alpha_t}\,(1 - \bar{\alpha}_{t-1})}{1 - \bar{\alpha}_t}\,\mathbf{x}_t, \\\\ \tilde{\beta}_t &= \frac{1 - \bar{\alpha}_{t-1}}{1 - \bar{\alpha}_t}\,\beta_t. \end{align*} \]

First, the mean \(\tilde{\boldsymbol{\mu}}_t\) is a linear combination of \(\mathbf{x}_0\) and \(\mathbf{x}_t\) with positive coefficients. Both contributions are necessary because \(\mathbf{x}_{t-1}\) lies between them along the chain. Second, the variance \(\tilde{\beta}_t\) depends only on the schedule, not on \(\mathbf{x}_0\) or \(\mathbf{x}_t\), a consequence of the linear-Gaussian structure that we will exploit when matching this distribution against a learned reverse kernel. This distribution is the per-step target that the learned model \(p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t)\) is asked to approximate, as the next subsection makes precise.

The Evidence Lower Bound

With the reverse posterior in hand, we turn to the training objective. The evidence lower bound of the diffusion model takes a particularly clean form. It is a sum of Kullback-Leibler terms, one for each step of the chain.

Following the convention of the diffusion literature, we work with the negative ELBO, written as a variational upper bound on the negative log-likelihood: \[ \begin{align*} -\log p_{\boldsymbol{\theta}}(\mathbf{x}_0) &\leq \mathbb{E}_{q(\mathbf{x}_{1:T} \mid \mathbf{x}_0)}\!\left[ -\log \frac{p_{\boldsymbol{\theta}}(\mathbf{x}_{0:T})}{q(\mathbf{x}_{1:T} \mid \mathbf{x}_0)} \right] \\\\ &=: \mathcal{L}(\mathbf{x}_0). \end{align*} \] The inequality is the standard Jensen bound applied with the latent sequence \((\mathbf{x}_1, \ldots, \mathbf{x}_T)\) playing the role of the latents and \(q(\mathbf{x}_{1:T} \mid \mathbf{x}_0)\) playing the role of the variational distribution. Minimizing \(\mathcal{L}(\mathbf{x}_0)\) over \(\boldsymbol{\theta}\) therefore lowers a bound on \(-\log p_{\boldsymbol{\theta}}(\mathbf{x}_0)\), not that quantity itself.

The integrand factorizes along the chain. The reverse model \(p_{\boldsymbol{\theta}}(\mathbf{x}_{0:T}) = p(\mathbf{x}_T) \prod_{t=1}^T p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t)\) uses the standard Gaussian prior \(p(\mathbf{x}_T) = \mathcal{N}(\mathbf{0}, \mathbf{I})\), and the forward joint factorizes as \(q(\mathbf{x}_{1:T} \mid \mathbf{x}_0) = \prod_{t=1}^T q(\mathbf{x}_t \mid \mathbf{x}_{t-1})\). Substituting, \[ \mathcal{L}(\mathbf{x}_0) = \mathbb{E}_q\!\left[ -\log p(\mathbf{x}_T) - \sum_{t=1}^{T} \log \frac{p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t)}{q(\mathbf{x}_t \mid \mathbf{x}_{t-1})} \right]. \]

The summand mixes forward and reverse transitions, which makes direct identification with the reverse posterior of the previous subsection awkward. The crucial step is to rewrite each forward transition using Bayes' rule together with the Markov property, exactly as in the reverse-posterior derivation. For \(t \geq 2\), \[ \begin{align*} q(\mathbf{x}_t \mid \mathbf{x}_{t-1}) &= q(\mathbf{x}_t \mid \mathbf{x}_{t-1}, \mathbf{x}_0) \\\\ &= \frac{q(\mathbf{x}_{t-1} \mid \mathbf{x}_t, \mathbf{x}_0)\, q(\mathbf{x}_t \mid \mathbf{x}_0)}{q(\mathbf{x}_{t-1} \mid \mathbf{x}_0)}. \end{align*} \] The first equality is the Markov property. The second is Bayes' rule. The case \(t = 1\) is left untouched, since \(q(\mathbf{x}_1 \mid \mathbf{x}_0)\) already conditions on \(\mathbf{x}_0\) and no rewriting is needed.

Substituting this identity for every \(t \geq 2\) and separating the \(t = 1\) term, the ratio inside the sum becomes \[ \log \frac{p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t)}{q(\mathbf{x}_t \mid \mathbf{x}_{t-1})} = \log \frac{p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t)}{q(\mathbf{x}_{t-1} \mid \mathbf{x}_t, \mathbf{x}_0)} + \log \frac{q(\mathbf{x}_{t-1} \mid \mathbf{x}_0)}{q(\mathbf{x}_t \mid \mathbf{x}_0)}. \] The first term has the structure of a Kullback-Leibler integrand. It compares the learned reverse kernel to the reverse posterior derived above. The second term telescopes when summed over \(t = 2, \ldots, T\), collapsing to \(\log q(\mathbf{x}_1 \mid \mathbf{x}_0) - \log q(\mathbf{x}_T \mid \mathbf{x}_0)\).

The first part of this residue cancels against the \(t = 1\) term that was left untouched, and the second part combines with \(-\log p(\mathbf{x}_T)\) to form a Kullback-Leibler divergence comparing the forward terminal to the reverse prior. After collecting terms, the negative ELBO admits the following decomposition.

Theorem: Decomposition of the Diffusion Negative ELBO

The negative ELBO of the diffusion model decomposes as \[ \mathcal{L}(\mathbf{x}_0) = \mathcal{L}_T + \sum_{t=2}^{T} \mathcal{L}_{t-1} + \mathcal{L}_0, \] with \[ \begin{align*} \mathcal{L}_T &= D_{\mathbb{KL}}\!\left( q(\mathbf{x}_T \mid \mathbf{x}_0) \,\|\, p(\mathbf{x}_T) \right), \\\\[4pt] \mathcal{L}_{t-1} &= \mathbb{E}_{q(\mathbf{x}_t \mid \mathbf{x}_0)}\!\left[ D_{\mathbb{KL}}\!\left( q(\mathbf{x}_{t-1} \mid \mathbf{x}_t, \mathbf{x}_0) \,\|\, p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t) \right) \right], \\\\[4pt] \mathcal{L}_0 &= -\,\mathbb{E}_{q(\mathbf{x}_1 \mid \mathbf{x}_0)}\!\left[ \log p_{\boldsymbol{\theta}}(\mathbf{x}_0 \mid \mathbf{x}_1) \right]. \end{align*} \]

Three observations follow, one for each kind of term.

The key conceptual point is that the variational problem has reduced to a sequence of step-local matching problems. For each \(t\), we bring the learned reverse kernel close to the analytically known reverse posterior. We exploit this in the next subsection by giving \(p_{\boldsymbol{\theta}}\) and \(q\) the same Gaussian functional form, so that each KL term becomes a Gaussian-to-Gaussian divergence with a closed-form expression.

Noise Prediction and the DDPM Loss

The decomposition above leaves us with the diffusion terms \(\mathcal{L}_{t-1}\), each of which must be turned into something a network can be trained against. The key step is to choose the functional form of the learned reverse kernel so that the Kullback-Leibler divergence inside \(\mathcal{L}_{t-1}\) becomes a closed-form expression in the network outputs.

We take \(p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t) = \mathcal{N}(\mathbf{x}_{t-1};\, \boldsymbol{\mu}_{\boldsymbol{\theta}}(\mathbf{x}_t, t),\, \sigma_t^2 \mathbf{I})\), with the variance fixed by hand to a schedule-dependent constant \(\sigma_t^2\). Only the mean \(\boldsymbol{\mu}_{\boldsymbol{\theta}}\) is learned. Typical choices are \(\sigma_t^2 = \beta_t\) or the reverse-posterior variance \(\tilde{\beta}_t\), both used in the literature. We revisit the question when we discuss sampling.

With this choice, both kernels in the KL are isotropic Gaussians, and the divergence reduces to a squared distance between means plus a term \(C_t\) that does not depend on \(\boldsymbol{\theta}\): \[ D_{\mathbb{KL}}\!\left( q(\mathbf{x}_{t-1} \mid \mathbf{x}_t, \mathbf{x}_0) \,\|\, p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t) \right) = \frac{1}{2 \sigma_t^2} \left\| \tilde{\boldsymbol{\mu}}_t(\mathbf{x}_t, \mathbf{x}_0) - \boldsymbol{\mu}_{\boldsymbol{\theta}}(\mathbf{x}_t, t) \right\|^2 + C_t. \] The term \(C_t\) vanishes when \(\sigma_t^2 = \tilde{\beta}_t\).

A direct parameterization of \(\boldsymbol{\mu}_{\boldsymbol{\theta}}\) as a free neural network output is possible but was found to perform worse in practice. The cleaner approach is to exploit the structure of the reverse posterior mean. Recall the one-shot reparameterization \(\mathbf{x}_t = \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\varepsilon}\). Solving for \(\mathbf{x}_0\) and substituting into the expression for \(\tilde{\boldsymbol{\mu}}_t(\mathbf{x}_t, \mathbf{x}_0)\) derived above (the algebra simplifies neatly because the two coefficients combine through the same identity \(\beta_t + \alpha_t (1 - \bar{\alpha}_{t-1}) = 1 - \bar{\alpha}_t\)) yields \[ \tilde{\boldsymbol{\mu}}_t(\mathbf{x}_t, \mathbf{x}_0) = \frac{1}{\sqrt{\alpha_t}} \left( \mathbf{x}_t - \frac{\beta_t}{\sqrt{1 - \bar{\alpha}_t}}\,\boldsymbol{\varepsilon} \right). \]

The reverse posterior mean, viewed at the noisy sample \(\mathbf{x}_t\), is a deterministic function of the noise \(\boldsymbol{\varepsilon}\) that was used to produce \(\mathbf{x}_t\) from \(\mathbf{x}_0\). This suggests parameterizing \(\boldsymbol{\mu}_{\boldsymbol{\theta}}\) in the same form, with a learned noise predictor \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t)\) replacing the true noise: \[ \boldsymbol{\mu}_{\boldsymbol{\theta}}(\mathbf{x}_t, t) = \frac{1}{\sqrt{\alpha_t}} \left( \mathbf{x}_t - \frac{\beta_t}{\sqrt{1 - \bar{\alpha}_t}}\,\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t) \right). \]

The two means differ only through \(\boldsymbol{\varepsilon} - \boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\), and substituting into the squared-distance form of the KL gives the diffusion term \[ \mathcal{L}_{t-1} = \mathbb{E}_{\boldsymbol{\varepsilon}}\!\left[ \frac{\beta_t^2}{2 \sigma_t^2 \alpha_t (1 - \bar{\alpha}_t)} \,\left\| \boldsymbol{\varepsilon} - \boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\!\left(\sqrt{\bar{\alpha}_t}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\varepsilon},\, t \right) \right\|^2 \right] + C_t, \] with the expectation over \(\boldsymbol{\varepsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\) at fixed \(\mathbf{x}_0\), as in the decomposition. Training the diffusion model has reduced to a regression problem. Given a noisy sample and a time index, predict the noise.

The time-dependent weight in front of the squared error, call it \(\lambda_t\), is what a faithful derivation of the variational bound produces. The objective actually used in practice discards it, for reasons recorded in modification (iii) below.

Theorem: The DDPM Training Objective

The simplified DDPM loss is \[ \mathcal{L}_{\text{simple}}(\boldsymbol{\theta}) = \mathbb{E}_{t,\,\mathbf{x}_0,\,\boldsymbol{\varepsilon}}\!\left[ \left\| \boldsymbol{\varepsilon} - \boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\!\left( \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\varepsilon},\, t \right) \right\|^2 \right], \] where \(t \sim \mathrm{Unif}\{1, \ldots, T\}\), \(\mathbf{x}_0 \sim q_0\), and \(\boldsymbol{\varepsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\). This objective is obtained from the full negative ELBO, averaged over \(\mathbf{x}_0 \sim q_0\), by three modifications.

(i) Dropping the prior term. The prior term \(\mathcal{L}_T = D_{\mathbb{KL}}(q(\mathbf{x}_T \mid \mathbf{x}_0) \,\|\, p(\mathbf{x}_T))\) depends only on the schedule \(\{\beta_t\}\) and not on the network parameters \(\boldsymbol{\theta}\). Its gradient with respect to \(\boldsymbol{\theta}\) is therefore identically zero, and it can be removed without affecting training.

(ii) Folding the reconstruction term into the same form. The reconstruction term \(\mathcal{L}_0 = -\mathbb{E}_{q(\mathbf{x}_1 \mid \mathbf{x}_0)}[\log p_{\boldsymbol{\theta}}(\mathbf{x}_0 \mid \mathbf{x}_1)]\) has a different functional form from the diffusion terms \(\mathcal{L}_{t-1}\) for \(t \geq 2\). Treating \(\mathbf{x}_0\) as continuous and modeling \(p_{\boldsymbol{\theta}}(\mathbf{x}_0 \mid \mathbf{x}_1)\) as a Gaussian with the same noise-prediction parameterization evaluated at \(t = 1\), the reconstruction term reduces (up to a constant) to a \(t = 1\) instance of the squared-error term inside the expectation. The sum over diffusion terms can therefore be extended from \(t = 2, \ldots, T\) to \(t = 1, \ldots, T\), and the extended sum absorbs the reconstruction term. (For genuinely discrete data such as image pixels, a separate discretized decoder is needed at \(t = 1\). The loss above continues to give the correct gradient signal for the rest of the chain.)

(iii) Setting the time-dependent weight to one. The principled per-step weight \(\lambda_t = \beta_t^2 / (2 \sigma_t^2 \alpha_t (1 - \bar{\alpha}_t))\) is replaced by \(\lambda_t = 1\). This is an empirical choice. The \(\boldsymbol{\theta}\)-independent constants \(C_t\) are dropped as well, for the same reason as the prior term. The resulting objective is no longer, in general, a variational bound on \(-\log p_{\boldsymbol{\theta}}(\mathbf{x}_0)\), but it weights all noise levels equally and was found by Ho, Jain, and Abbeel to produce visibly better samples.

Algorithm: DDPM Training Input: Dataset \(\mathcal{D}\), schedule \(\{\beta_t\}_{t=1}^T\), network \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\) repeat    \(\mathbf{x}_0 \sim \mathcal{D}\);    \(t \sim \mathrm{Unif}\{1, \ldots, T\}\);    \(\boldsymbol{\varepsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\);    \(\mathbf{x}_t \leftarrow \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\varepsilon}\);    Take a gradient step on \(\nabla_{\boldsymbol{\theta}} \left\| \boldsymbol{\varepsilon} - \boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t) \right\|^2\); until converged

The architecture of \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\) depends on the data modality. For images, a common choice is a U-Net with skip connections between matching resolution levels and a small time embedding injected into each block. A transformer backbone (the diffusion transformer, or DiT) is another. The point is that \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\) is just a neural network that maps a tensor and a scalar time index to a tensor of the same shape as the input. From the perspective of training, all the structural insight of the diffusion model lives in the loss above, not in the network architecture.

The Score-Matching Perspective

The term score function has appeared before in a different sense. In the classical statistical setting that goes back to Fisher, the score function is the gradient of the log-likelihood with respect to the parameters of a model, \(s(\boldsymbol{\theta}) = \nabla_{\boldsymbol{\theta}} \log p(\mathbf{x} \mid \boldsymbol{\theta})\). This is the quantity that underlies the Fisher information matrix and natural gradient descent.

The diffusion-model literature, following the score-matching framework of Hyvärinen (2005), uses the same term for a formally different quantity, a gradient taken with respect to the data instead. Both are gradients of log-densities, but they live in different spaces and serve different purposes. The parameter gradient is for inference and optimization, while the data gradient describes the geometry of the distribution at a point in sample space. Throughout this subsection, score function refers exclusively to the data-gradient version.

Definition: Score Function (Data Gradient)

Let \(p\) be a strictly positive, differentiable probability density on \(\mathbb{R}^d\). The score function of \(p\) is the gradient of the log-density with respect to the data argument, \[ \mathbf{s}(\mathbf{x}) := \nabla_{\mathbf{x}} \log p(\mathbf{x}). \] Geometrically, \(\mathbf{s}(\mathbf{x})\) is a vector field on \(\mathbb{R}^d\) that points, at every \(\mathbf{x}\), in the direction of steepest ascent of the probability density.

We now connect this object to the diffusion model. The forward chain provides a family of marginal distributions \(q_t(\mathbf{x}_t) := \int q(\mathbf{x}_t \mid \mathbf{x}_0)\, q_0(\mathbf{x}_0)\, d\mathbf{x}_0\), one for each time index \(t\), that interpolate between the data distribution at \(t = 0\) and approximately standard Gaussian noise at \(t = T\). For \(t \geq 1\), each \(q_t\) is a Gaussian smoothing of the data distribution, hence strictly positive and differentiable, with its own score function \(\mathbf{s}_t(\mathbf{x}_t) = \nabla_{\mathbf{x}_t} \log q_t(\mathbf{x}_t)\), and the diffusion model turns out to be implicitly estimating all of them at once.

The mechanism is a direct calculation. The conditional forward kernel \(q(\mathbf{x}_t \mid \mathbf{x}_0) = \mathcal{N}(\mathbf{x}_t;\, \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0,\, (1 - \bar{\alpha}_t)\mathbf{I})\) has, as a Gaussian in \(\mathbf{x}_t\), the score \[ \nabla_{\mathbf{x}_t} \log q(\mathbf{x}_t \mid \mathbf{x}_0) = -\,\frac{\mathbf{x}_t - \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0}{1 - \bar{\alpha}_t}. \] Substituting the reparameterization \(\mathbf{x}_t = \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\varepsilon}\) gives the key identity \[ \nabla_{\mathbf{x}_t} \log q(\mathbf{x}_t \mid \mathbf{x}_0) = -\,\frac{\boldsymbol{\varepsilon}}{\sqrt{1 - \bar{\alpha}_t}}. \]

The conditional score is, up to a negative scaling, exactly the Gaussian noise that produced \(\mathbf{x}_t\) from \(\mathbf{x}_0\). A theorem due to Vincent (2011), known as denoising score matching, extends this from the conditional to the marginal. A network trained to predict the noise on samples drawn from the forward chain learns the score of the marginal \(q_t\). Writing \(\sigma_t := \sqrt{1 - \bar{\alpha}_t}\), the noise-prediction network \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\) and the score network \(\mathbf{s}_{\boldsymbol{\theta}}\) are linked by the parameterization \[ \mathbf{s}_{\boldsymbol{\theta}}(\mathbf{x}_t, t) = -\,\frac{\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t)}{\sigma_t}. \]

The argument takes two steps. Assume the data distribution permits differentiating \(q_t(\mathbf{x}_t) = \int q(\mathbf{x}_t \mid \mathbf{x}_0)\, q_0(\mathbf{x}_0)\, d\mathbf{x}_0\) under the integral sign. Dividing the derivative by \(q_t(\mathbf{x}_t)\) and writing \(q(\mathbf{x}_0 \mid \mathbf{x}_t)\) for the posterior of the clean data given the noisy sample, we obtain \[ \begin{align*} \nabla_{\mathbf{x}_t} \log q_t(\mathbf{x}_t) &= \int \nabla_{\mathbf{x}_t} \log q(\mathbf{x}_t \mid \mathbf{x}_0)\, q(\mathbf{x}_0 \mid \mathbf{x}_t)\, d\mathbf{x}_0 \\\\ &= -\,\frac{\mathbb{E}[\boldsymbol{\varepsilon} \mid \mathbf{x}_t]}{\sigma_t}, \end{align*} \] where the second line uses the key identity above. On the other hand, among all square-integrable functions of \(\mathbf{x}_t\), the squared error \(\mathbb{E}\,\|\boldsymbol{\varepsilon} - \mathbf{g}(\mathbf{x}_t)\|^2\) over the joint law of \((\mathbf{x}_0, \boldsymbol{\varepsilon})\) is minimized by \(\mathbf{g}(\mathbf{x}_t) = \mathbb{E}[\boldsymbol{\varepsilon} \mid \mathbf{x}_t]\), since conditional expectation is the orthogonal projection in \(L^2\). A noise predictor that attains the minimum at step \(t\) therefore satisfies \(-\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t) / \sigma_t = \nabla_{\mathbf{x}_t} \log q_t(\mathbf{x}_t)\).

Remark on notation. The symbol \(\sigma_t\) is used in two distinct senses on this page, and the literature is not always careful to distinguish them. In the previous subsection, \(\sigma_t^2\) denoted the variance of the learned reverse kernel \(p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t)\), a hyperparameter typically set to \(\beta_t\) or \(\tilde{\beta}_t\). Here \(\sigma_t = \sqrt{1 - \bar{\alpha}_t}\) is the standard deviation of the forward marginal \(q(\mathbf{x}_t \mid \mathbf{x}_0)\), determined entirely by the noise schedule. The two quantities coincide only when \(\sigma_t^2 = 1 - \bar{\alpha}_t\), which is not the conventional choice for the reverse-kernel variance. We retain the standard overloaded notation because it is universal in the diffusion literature, but the reader should keep the distinction in mind.

Under this identification, the noise-prediction objective of DDPM and the denoising score-matching objective of score-based models coincide up to the choice of per-step weighting. The two frameworks are not identical (DDPM is set up as a discrete-time variational model and score-based generative modeling as a continuous-time score-matching model), but their trained networks are interchangeable through the parameterization above.

With this identification in hand, we return to the thread opened in the introduction. The page on PCA and autoencoders took on faith that a denoising autoencoder trained at noise level \(\sigma\) implicitly learns the score \(\nabla_{\mathbf{x}} \log p(\mathbf{x})\) of the data distribution, via the asymptotic identity \(\mathbf{r}^*(\mathbf{x}) - \mathbf{x} = \sigma^2 \nabla_{\mathbf{x}} \log p(\mathbf{x}) + o(\sigma^2)\) of Vincent (2011) and Alain and Bengio (2014).

What we have just developed is the multi-scale version of that statement. Rather than fixing one noise level, the diffusion model trains a single network across the full schedule of noise levels \(\{\sigma_t\}_{t=1}^T\), and the resulting noise predictor doubles as a score estimator for the entire family of marginals \(\{q_t\}_{t=1}^T\). The promise made on the autoencoder page (that score-matching theory justifies the denoising construction) is what the diffusion framework redeems.

In the continuous-time formulation previewed in the introduction, the score function is the central object. Anderson's 1982 reversal theorem expresses the reverse drift in terms of \(\nabla \log q_t\), and the score-based generative modeling framework of Song and Ermon (2019) and Song et al. (2021) takes this viewpoint as primary. The discrete DDPM picture developed on this page corresponds to a particular time-discretization of the variance-preserving SDE.

Sampling: DDPM and DDIM

Ancestral Sampling

With \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\) trained, generation proceeds by simulating the reverse chain. Starting from \(\mathbf{x}_T \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\), we sample successively from \(p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t) = \mathcal{N}(\mathbf{x}_{t-1};\, \boldsymbol{\mu}_{\boldsymbol{\theta}}(\mathbf{x}_t, t),\, \sigma_t^2 \mathbf{I})\) for \(t = T, T-1, \ldots, 1\), and return \(\mathbf{x}_0\). Because each variable is drawn from its conditional distribution given its parent in the directed chain \(\mathbf{x}_T \to \cdots \to \mathbf{x}_0\), this procedure is called ancestral sampling. Substituting the noise-prediction form of \(\boldsymbol{\mu}_{\boldsymbol{\theta}}\) derived earlier gives the explicit step rule used in practice.

Algorithm: DDPM Ancestral Sampling Input: Trained network \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\), schedule \(\{\beta_t\}_{t=1}^T\), reverse-kernel variances \(\{\sigma_t^2\}_{t=1}^T\) \(\mathbf{x}_T \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\); for \(t = T, T-1, \ldots, 1\) do    if \(t \gt 1\): \(\mathbf{z} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\); else: \(\mathbf{z} = \mathbf{0}\);    \(\mathbf{x}_{t-1} \leftarrow \dfrac{1}{\sqrt{\alpha_t}} \left( \mathbf{x}_t - \dfrac{\beta_t}{\sqrt{1 - \bar{\alpha}_t}}\,\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t) \right) + \sigma_t \mathbf{z}\); Output: \(\mathbf{x}_0\)

Two design choices remain.

  1. The reverse-kernel variance \(\sigma_t^2\), which was fixed by hand during training rather than learned. Ho, Jain, and Abbeel observed that two natural choices give comparable sample quality. Setting \(\sigma_t^2 = \beta_t\) is optimal when \(\mathbf{x}_0 \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\). Setting \(\sigma_t^2 = \tilde{\beta}_t\), the reverse-posterior variance computed earlier, is optimal when \(\mathbf{x}_0\) is a single fixed point. For data with coordinatewise unit variance, the two are the extreme choices, giving upper and lower bounds on the entropy of the reverse process. In practice both work well, with the difference between them small compared with other sources of variation in the sampler.
  2. The suppression of noise at the final step. When \(t = 1\), the algorithm sets \(\mathbf{z} = \mathbf{0}\) and returns the deterministic mean. This avoids injecting Gaussian noise into the final decoded output, which would otherwise visibly degrade samples in image domains.

The principal limitation of ancestral sampling is its cost. With \(T = 1000\) (the value used in the original DDPM experiments), generating a single sample requires one thousand sequential network evaluations, and the steps cannot be parallelized because each \(\mathbf{x}_{t-1}\) depends on \(\mathbf{x}_t\). This is the bottleneck that the next subsection addresses.

Deterministic Sampling with DDIM

Song, Meng, and Ermon's denoising diffusion implicit models (DDIM, 2021) rest on one observation. One can sample from a DDPM in far fewer steps, and with a fully deterministic reverse process if desired, without retraining the network. The same \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\) that was trained against the DDPM objective is reused, and only the sampling procedure changes.

The mechanism is a family of reverse kernels indexed by stochasticity parameters \(\tilde{\sigma}_t \geq 0\) with \(\tilde{\sigma}_t^2 \leq 1 - \bar{\alpha}_{t-1}\). For each choice of \(\tilde{\sigma}_t\), Song, Meng, and Ermon construct a reverse distribution \[ q_{\tilde{\sigma}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t, \mathbf{x}_0) = \mathcal{N}\!\left( \mathbf{x}_{t-1}; \sqrt{\bar{\alpha}_{t-1}}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_{t-1} - \tilde{\sigma}_t^2}\, \cdot \frac{\mathbf{x}_t - \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0}{\sqrt{1 - \bar{\alpha}_t}}, \tilde{\sigma}_t^2 \mathbf{I} \right) \] that is consistent with the forward marginals \(q(\mathbf{x}_t \mid \mathbf{x}_0)\) used during training, although the forward process it induces is not in general Markovian.

Two members of this family are distinguished. Setting \(\tilde{\sigma}_t^2 = \tilde{\beta}_t\), the reverse-posterior variance, recovers the DDPM reverse posterior of the previous section. Setting \(\tilde{\sigma}_t = 0\) yields a fully deterministic update, the DDIM sampler.

Remark on notation. The symbol \(\tilde{\sigma}_t\) here is yet another \(\sigma\)-like quantity, distinct from the reverse-kernel variance \(\sigma_t^2\) of DDPM sampling and from the forward marginal standard deviation \(\sigma_t = \sqrt{1 - \bar{\alpha}_t}\) used in the score-matching subsection. We use the tilde to mark it as the DDIM stochasticity parameter, distinct from both.

With \(\mathbf{x}_0\) replaced by the network's prediction \(\hat{\mathbf{x}}_0(\mathbf{x}_t, t) := \frac{\mathbf{x}_t - \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t)}{\sqrt{\bar{\alpha}_t}}\), obtained by solving the forward reparameterization for \(\mathbf{x}_0\), and \(\tilde{\sigma}_t\) set to zero, the deterministic DDIM step takes a particularly transparent form.

Algorithm: DDIM Deterministic Sampling Input: Trained network \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\), schedule \(\{\beta_t\}_{t=1}^T\), step subsequence \(\tau_1 \lt \tau_2 \lt \cdots \lt \tau_S\) with \(\tau_S = T\) \(\mathbf{x}_{\tau_S} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\); for \(i = S, S-1, \ldots, 1\) do    \(t \leftarrow \tau_i\); \( s \leftarrow \tau_{i-1}\) (with \(\tau_0 := 0\), \(\bar{\alpha}_{\tau_0} := 1\));    \(\hat{\mathbf{x}}_0 \leftarrow \dfrac{\mathbf{x}_t - \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t)}{\sqrt{\bar{\alpha}_t}}\);    \(\mathbf{x}_s \leftarrow \sqrt{\bar{\alpha}_s}\,\hat{\mathbf{x}}_0 + \sqrt{1 - \bar{\alpha}_s}\,\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t)\); Output: \(\mathbf{x}_0\)

The step decomposes into two interpretable operations.

  1. The current noisy sample \(\mathbf{x}_t\) is mapped to a prediction of the clean data \(\hat{\mathbf{x}}_0\) by undoing the forward reparameterization with the network's noise estimate.
  2. This predicted clean sample is re-corrupted to the target time \(s\), using the same noise estimate \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t)\) rather than a fresh Gaussian draw. The coefficients \(\sqrt{\bar{\alpha}_s}\) and \(\sqrt{1 - \bar{\alpha}_s}\), whose squares sum to one, are those of the one-shot forward reparameterization at time \(s\), so the result has the form of a forward sample at time \(s\) built from \(\hat{\mathbf{x}}_0\) and the estimated noise.

Because no fresh randomness enters once \(\mathbf{x}_T\) has been drawn, the entire trajectory from \(\mathbf{x}_T\) to \(\mathbf{x}_0\) is a deterministic function of the initial noise, and seeds can be replayed exactly.

A second feature of the DDIM construction is that the step indices need not visit every \(t\) in \(\{1, \ldots, T\}\). Any subsequence \(\tau_1 \lt \cdots \lt \tau_S = T\) of the training schedule is valid for sampling, and choosing \(S\) much smaller than \(T\) yields a substantial speedup. Typical practice uses \(S\) on the order of 50 with \(T = 1000\) during training, a twenty-fold reduction in network evaluations per sample, with image quality that approximates the full-chain DDPM sampler.

Conditional Generation

Practical diffusion systems rarely sample from the unconditional data distribution alone. They generate conditioned on auxiliary information \(\mathbf{c}\) (a class label, a low-resolution image to upsample, a text prompt), and the user expects the sampled output to reflect that condition. The minimal modification is to make the network condition-aware. We replace \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t)\) with \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t, \mathbf{c})\) and train on \((\mathbf{x}_0, \mathbf{c})\) pairs drawn jointly from data.

The mechanism of injection depends on the modality of \(\mathbf{c}\). Class labels are typically embedded and added to the time embedding, conditioning images are concatenated to \(\mathbf{x}_t\) along the channel axis, and text prompts are encoded by a separate model and consumed through cross-attention at multiple resolutions inside the network. The loss is unchanged, the simple denoising objective applied to the conditional network.

Conditional training alone, however, tends to produce samples that are only weakly attached to \(\mathbf{c}\). The network happily falls back on the prior whenever the conditional signal is ambiguous. The standard remedy, introduced by Ho and Salimans (2021), is classifier-free guidance. During training, the conditioning input is replaced by a null token \(\varnothing\) with some fixed probability (typically 10 to 20 percent), so that the same network learns both the conditional predictor \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t, \mathbf{c})\) and the unconditional predictor \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t, \varnothing)\).

At sampling time, the two predictions are combined into a guided estimate \[ \tilde{\boldsymbol{\varepsilon}}_{\boldsymbol{\theta}}(\mathbf{x}_t, t, \mathbf{c}) = (1 + w)\,\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t, \mathbf{c}) - w\,\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t, \varnothing), \] which is then substituted for \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\) inside any of the sampling algorithms above.

The guidance weight \(w \geq 0\) governs the trade-off. Setting \(w = 0\) recovers the unguided conditional sampler, while larger \(w\) extrapolates beyond the conditional prediction, away from the unconditional one. The resulting samples adhere more tightly to \(\mathbf{c}\), at the cost of reduced diversity and occasional artifacts.

The score-matching perspective developed earlier illuminates the construction. Bayes' rule on log-densities gives the identity \(\nabla_{\mathbf{x}_t} \log p(\mathbf{x}_t \mid \mathbf{c}) = \nabla_{\mathbf{x}_t} \log p(\mathbf{x}_t) + \nabla_{\mathbf{x}_t} \log p(\mathbf{c} \mid \mathbf{x}_t)\), so the conditional score decomposes into an unconditional score and a classifier-like correction. Classifier-free guidance, viewed at the score level, replaces this exact decomposition with an extrapolation. The guided score \((1+w)\mathbf{s}_{\boldsymbol{\theta}}(\mathbf{x}_t, t, \mathbf{c}) - w\,\mathbf{s}_{\boldsymbol{\theta}}(\mathbf{x}_t, t, \varnothing)\) overweights the conditional direction relative to its true Bayesian value and so biases samples toward regions where the classifier-like likelihood \(p(\mathbf{c} \mid \mathbf{x}_t)\) is high.

Interactive Demo

The visualization below applies the full diffusion construction (forward chain, score function, and reverse sampling) to a point cloud built from the MATH-CS COMPASS logo. Each point carries five coordinates, two spatial \((x, y)\) and three color \((r, g, b)\), and the diffusion is run jointly on all five. The forward direction drives both the geometry and the color of the logo toward a pure Gaussian distribution. The reverse direction reconstructs points on the data manifold from noise, with the logo emerging as the support of the distribution being sampled.

Try the forward direction first to watch the logo dissolve, then switch to reverse and observe the same construction running backwards. The time slider scrubs through any intermediate state. Three overlays expose the machinery rather than just the result.

Finally, when the reverse process reaches \(t = 0\), a readout reports the mean distance from the generated points to the data manifold, turning the visual comparison between samplers into a number.

DDPM versus DDIM: Stochastic versus Deterministic Reverse Paths

The sampler toggle is the most pedagogically informative control. Both DDPM and DDIM start from the same initial draw \(\mathbf{x}_T\) (use the "regenerate noise" button to draw a new one), and both end on the data manifold.

What differs is the structure of the trajectory between them. The DDPM update is stochastic. Each step draws a fresh Gaussian, so repeated runs from the same \(\mathbf{x}_T\) would in principle produce different reconstructions (this demo caches a single such run per noise seed for visualization). The DDIM update is deterministic. The entire trajectory is a fixed function of \(\mathbf{x}_T\), and the demo's DDIM sampler reaches the data manifold in roughly five times fewer steps than DDPM (twenty against one hundred here).

Because DDIM compresses many small DDPM steps into a few large jumps, the reconstruction in this demo lands close to, but not exactly on, the original logo points, so the points look slightly hazier than under DDPM. The trails overlay shows the cause and the \(t = 0\) readout shows the effect. For the same starting noise, the mean distance to the data under DDPM is clearly smaller than under the twenty-step DDIM sampler.

This residual is the discretization error of the few-step DDIM sampler, the small accuracy cost that pays for the speedup. The number of steps therefore trades accuracy for speed, and more refined few-step samplers aim to shrink this error at the same cost.

The construction visualized here is a small instance of the mathematics that underlies diffusion models at scale. Three replacements turn it into such a model. The five-dimensional point cloud becomes a high-dimensional data tensor, the point distribution becomes the data distribution, and the analytical score becomes a trained neural network. Some models run the diffusion in a learned latent space rather than directly on the data. Larger models add several ingredients on top (classifier-free guidance, learned encoders for the conditioning input, alternative parameterizations and samplers), but the discrete-time DDPM and DDIM updates developed on this page remain the structural core.

What makes the reverse direction possible is the score function \(\nabla_{\mathbf{x}_t} \log q_t(\mathbf{x}_t)\). Given the gradient of the log-density of every intermediate marginal, walking back from noise to data is a matter of repeated denoising steps, each one guided by the score evaluated at the current state.

The score-field overlay shows this directly. Switch it on during the reverse process and the arrows point the way from noise back toward the logo, growing sharper as the marginal concentrates. In this demo the score is computed exactly from the empirical point cloud. In practice it is learned from data, which is what diffusion training is really for.