The Forward Diffusion Process
The forward process is the simpler of the two halves of a diffusion model. It is fixed, requires no training, and is specified
by a single design choice, the rate at which Gaussian noise is added to the data at each step. Let
\(\mathbf{x}_0 \in \mathbb{R}^d\) denote a data sample drawn from an unknown data distribution \(q_0(\mathbf{x}_0)\), and let
\(T\) denote the total number of diffusion steps. We choose a sequence of small positive numbers
\(\{\beta_t\}_{t=1}^T \subset (0, 1)\), called the noise schedule, and define a Markov chain
\(\mathbf{x}_0 \to \mathbf{x}_1 \to \cdots \to \mathbf{x}_T\) by the following transition.
Definition: Forward Diffusion Kernel
Given a noise schedule \(\{\beta_t\}_{t=1}^T \subset (0, 1)\), the forward diffusion kernel
is the Gaussian transition
\[
q(\mathbf{x}_t \mid \mathbf{x}_{t-1}) = \mathcal{N}\!\left(\mathbf{x}_t; \sqrt{1 - \beta_t}\,\mathbf{x}_{t-1}, \beta_t \mathbf{I}\right),
\quad t = 1, \ldots, T.
\]
Three properties of this kernel matter.
-
The chain is
first-order Markov.
Each \(\mathbf{x}_t\) depends only on its immediate predecessor \(\mathbf{x}_{t-1}\), so the joint conditional
distribution factorizes as
\[
q(\mathbf{x}_{1:T} \mid \mathbf{x}_0) = \prod_{t=1}^{T} q(\mathbf{x}_t \mid \mathbf{x}_{t-1}).
\]
-
The kernel contains no learned parameters.
Here \(\beta_t\) is fixed by the schedule, and the mean and covariance are
closed-form functions of \(\mathbf{x}_{t-1}\).
-
The scaling \(\sqrt{1 - \beta_t}\) on the mean is not arbitrary. It is chosen so that, if
\(\mathbf{x}_{t-1}\) has identity covariance, then \(\mathbf{x}_t\) does too:
\(\operatorname{Cov}(\mathbf{x}_t) = (1-\beta_t)\mathbf{I} + \beta_t \mathbf{I} = \mathbf{I}\). For general data,
independence of the injected noise gives
\(\operatorname{Cov}(\mathbf{x}_t) = \bar{\alpha}_t \operatorname{Cov}(\mathbf{x}_0) + (1 - \bar{\alpha}_t)\mathbf{I}\),
with \(\bar{\alpha}_t\) as defined below, so the marginal covariance moves from that of the data toward the identity
along the chain.
Forward kernels with this scaling are called variance-preserving, and the form above is the variance-preserving
construction of Ho, Jain, and Abbeel.
Closed-Form Marginal
Because the forward chain is a linear Gaussian Markov chain, the marginal distribution
\(q(\mathbf{x}_t \mid \mathbf{x}_0)\) at any time \(t\) can be computed in closed form. We introduce the abbreviations
\[
\alpha_t := 1 - \beta_t, \quad \bar{\alpha}_t := \prod_{s=1}^{t} \alpha_s,
\]
so that \(\bar{\alpha}_t\) is the cumulative product of the per-step retention factors.
Proof:
We argue by induction on \(t\). The base case \(t = 1\) is the kernel itself with
\(\bar{\alpha}_1 = \alpha_1 = 1 - \beta_1\). For the inductive step, suppose
\(q(\mathbf{x}_{t-1} \mid \mathbf{x}_0) = \mathcal{N}(\mathbf{x}_{t-1};\, \sqrt{\bar{\alpha}_{t-1}}\,\mathbf{x}_0,\, (1-\bar{\alpha}_{t-1})\mathbf{I})\).
By the reparameterization trick,
we can write
\[
\mathbf{x}_{t-1} = \sqrt{\bar{\alpha}_{t-1}}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_{t-1}}\,\boldsymbol{\varepsilon}_{t-1}, \quad \boldsymbol{\varepsilon}_{t-1} \sim \mathcal{N}(\mathbf{0}, \mathbf{I}),
\]
and similarly \(\mathbf{x}_t = \sqrt{\alpha_t}\,\mathbf{x}_{t-1} + \sqrt{\beta_t}\,\boldsymbol{\varepsilon}_t\)
with \(\boldsymbol{\varepsilon}_t \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\) independent of
\(\boldsymbol{\varepsilon}_{t-1}\). Substituting,
\[
\mathbf{x}_t = \sqrt{\alpha_t \bar{\alpha}_{t-1}}\,\mathbf{x}_0 + \sqrt{\alpha_t (1 - \bar{\alpha}_{t-1})}\,\boldsymbol{\varepsilon}_{t-1} + \sqrt{\beta_t}\,\boldsymbol{\varepsilon}_t.
\]
The last two terms are independent zero-mean Gaussians, so their sum is Gaussian. Using
\(\alpha_t \bar{\alpha}_{t-1} = \bar{\alpha}_t\), which holds by definition, its variance is
\[
\begin{align*}
\alpha_t(1 - \bar{\alpha}_{t-1}) + \beta_t
&= \alpha_t - \alpha_t \bar{\alpha}_{t-1} + (1 - \alpha_t) \\\\
&= 1 - \bar{\alpha}_t,
\end{align*}
\]
and the coefficient of \(\mathbf{x}_0\) is \(\sqrt{\alpha_t \bar{\alpha}_{t-1}} = \sqrt{\bar{\alpha}_t}\). Hence
\[
\mathbf{x}_t = \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\varepsilon}, \quad \boldsymbol{\varepsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I}),
\]
which is the reparameterization of the claimed Gaussian.
The identity at the end of the proof deserves its own line. Any noisy sample at time \(t\) is produced in one shot from the
clean data as
\[
\mathbf{x}_t = \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\varepsilon},
\quad \boldsymbol{\varepsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I}).
\]
This one-shot reparameterization is the central computational identity of training. Rather than simulating the chain step
by step, the trainer samples \(t\) uniformly and \(\boldsymbol{\varepsilon}\) from a standard Gaussian, and produces a
noisy sample directly.
The Endpoint and the Choice of Schedule
The schedule \(\{\beta_t\}\) is chosen so that \(\bar{\alpha}_T\) is very close to zero. From the closed-form marginal, this forces
\(q(\mathbf{x}_T \mid \mathbf{x}_0) \approx \mathcal{N}(\mathbf{0}, \mathbf{I})\), and the dependence on \(\mathbf{x}_0\) is washed
out. Two schedules are most common in practice. The linear schedule interpolates \(\beta_t\) linearly between
small endpoints (Ho, Jain, and Abbeel, 2020). The cosine schedule keeps \(\bar{\alpha}_t\) close to one for longer
at small \(t\) and so destroys information less abruptly, which improved results on lower-resolution images (Nichol and Dhariwal,
2021). Typical values of \(T\) range from a few hundred to a few thousand. The schedule and \(T\) are hyperparameters of the model,
not learned quantities.
The Reverse Process and Loss
The Reverse Posterior
To generate samples, we would like to run the chain backwards, starting from
\(\mathbf{x}_T \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\) and successively sampling from \(q(\mathbf{x}_{t-1} \mid \mathbf{x}_t)\)
until we recover \(\mathbf{x}_0\). The obstruction is that this reverse transition is not directly computable. Writing it as
\[
q(\mathbf{x}_{t-1} \mid \mathbf{x}_t) = \frac{q(\mathbf{x}_t \mid \mathbf{x}_{t-1})\, q(\mathbf{x}_{t-1})}{q(\mathbf{x}_t)},
\]
we see that both \(q(\mathbf{x}_{t-1})\) and \(q(\mathbf{x}_t)\) require integrating the forward chain against the unknown data
distribution \(q_0(\mathbf{x}_0)\), and so are inaccessible.
A different conditional, however, is tractable. It is the reverse process conditioned on the clean data \(\mathbf{x}_0\).
This auxiliary distribution will play the role of an oracle in the loss derivation of the next subsection. We cannot use it
for sampling, since \(\mathbf{x}_0\) is precisely what we are trying to generate, but we can use it during training, where
\(\mathbf{x}_0\) is drawn from the training set and is therefore available.
The derivation begins from Bayes' rule together with the Markov property of the forward chain. Because \(\mathbf{x}_t\) depends on
\(\mathbf{x}_{t-1}\) only through the forward kernel and not through \(\mathbf{x}_0\), we have
\(q(\mathbf{x}_t \mid \mathbf{x}_{t-1}, \mathbf{x}_0) = q(\mathbf{x}_t \mid \mathbf{x}_{t-1})\). Bayes' rule then gives
\[
q(\mathbf{x}_{t-1} \mid \mathbf{x}_t, \mathbf{x}_0)
= \frac{q(\mathbf{x}_t \mid \mathbf{x}_{t-1})\, q(\mathbf{x}_{t-1} \mid \mathbf{x}_0)}{q(\mathbf{x}_t \mid \mathbf{x}_0)}.
\]
Every quantity on the right is a Gaussian whose mean and variance we already know. The denominator is independent of
\(\mathbf{x}_{t-1}\) and serves only as a normalizing constant. The numerator is a product of two Gaussian densities in
\(\mathbf{x}_{t-1}\), so for \(t \geq 2\) its logarithm is, up to terms free of \(\mathbf{x}_{t-1}\),
\[
-\frac{1}{2} \left( \frac{\left\| \mathbf{x}_t - \sqrt{\alpha_t}\,\mathbf{x}_{t-1} \right\|^2}{\beta_t}
+ \frac{\left\| \mathbf{x}_{t-1} - \sqrt{\bar{\alpha}_{t-1}}\,\mathbf{x}_0 \right\|^2}{1 - \bar{\alpha}_{t-1}} \right).
\]
This is a quadratic in \(\mathbf{x}_{t-1}\) whose coefficient matrix is
\[
\left( \frac{\alpha_t}{\beta_t} + \frac{1}{1 - \bar{\alpha}_{t-1}} \right) \mathbf{I} = \frac{1 - \bar{\alpha}_t}{\beta_t (1 - \bar{\alpha}_{t-1})}\,\mathbf{I},
\]
where we used \(\alpha_t (1 - \bar{\alpha}_{t-1}) + \beta_t = 1 - \bar{\alpha}_t\). Completing the square therefore yields a
multivariate normal distribution
whose covariance is the inverse of this matrix and whose mean is that covariance applied to the linear coefficient
\(\sqrt{\alpha_t}\,\mathbf{x}_t / \beta_t + \sqrt{\bar{\alpha}_{t-1}}\,\mathbf{x}_0 / (1 - \bar{\alpha}_{t-1})\).
Definition: Reverse Posterior
For \(t \geq 2\), the reverse posterior conditioned on \(\mathbf{x}_0\) is the Gaussian
\[
q(\mathbf{x}_{t-1} \mid \mathbf{x}_t, \mathbf{x}_0)
= \mathcal{N}\!\left(\mathbf{x}_{t-1}; \tilde{\boldsymbol{\mu}}_t(\mathbf{x}_t, \mathbf{x}_0), \tilde{\beta}_t \mathbf{I}\right),
\]
with mean and variance
\[
\begin{align*}
\tilde{\boldsymbol{\mu}}_t(\mathbf{x}_t, \mathbf{x}_0)
&= \frac{\sqrt{\bar{\alpha}_{t-1}}\,\beta_t}{1 - \bar{\alpha}_t}\,\mathbf{x}_0
+ \frac{\sqrt{\alpha_t}\,(1 - \bar{\alpha}_{t-1})}{1 - \bar{\alpha}_t}\,\mathbf{x}_t, \\\\
\tilde{\beta}_t &= \frac{1 - \bar{\alpha}_{t-1}}{1 - \bar{\alpha}_t}\,\beta_t.
\end{align*}
\]
First, the mean \(\tilde{\boldsymbol{\mu}}_t\) is a linear combination of \(\mathbf{x}_0\) and \(\mathbf{x}_t\) with positive
coefficients. Both contributions are necessary because \(\mathbf{x}_{t-1}\) lies between them along the chain. Second, the variance
\(\tilde{\beta}_t\) depends only on the schedule, not on \(\mathbf{x}_0\) or \(\mathbf{x}_t\), a consequence of the linear-Gaussian
structure that we will exploit when matching this distribution against a learned reverse kernel. This distribution is the per-step
target that the learned model \(p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t)\) is asked to approximate, as the next
subsection makes precise.
The Evidence Lower Bound
With the reverse posterior in hand, we turn to the training objective. The
evidence lower bound
of the diffusion model takes a particularly clean form. It is a sum of Kullback-Leibler terms, one for each step of the chain.
Following the convention of the diffusion literature, we work with the negative ELBO, written as a variational upper bound on
the negative log-likelihood:
\[
\begin{align*}
-\log p_{\boldsymbol{\theta}}(\mathbf{x}_0)
&\leq \mathbb{E}_{q(\mathbf{x}_{1:T} \mid \mathbf{x}_0)}\!\left[ -\log \frac{p_{\boldsymbol{\theta}}(\mathbf{x}_{0:T})}{q(\mathbf{x}_{1:T} \mid \mathbf{x}_0)} \right] \\\\
&=: \mathcal{L}(\mathbf{x}_0).
\end{align*}
\]
The inequality is the standard Jensen bound applied with the latent sequence \((\mathbf{x}_1, \ldots, \mathbf{x}_T)\) playing the
role of the latents and \(q(\mathbf{x}_{1:T} \mid \mathbf{x}_0)\) playing the role of the variational distribution. Minimizing
\(\mathcal{L}(\mathbf{x}_0)\) over \(\boldsymbol{\theta}\) therefore lowers a bound on
\(-\log p_{\boldsymbol{\theta}}(\mathbf{x}_0)\), not that quantity itself.
The integrand factorizes along the chain. The reverse model
\(p_{\boldsymbol{\theta}}(\mathbf{x}_{0:T}) = p(\mathbf{x}_T) \prod_{t=1}^T p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t)\)
uses the standard Gaussian prior \(p(\mathbf{x}_T) = \mathcal{N}(\mathbf{0}, \mathbf{I})\), and the forward joint factorizes as
\(q(\mathbf{x}_{1:T} \mid \mathbf{x}_0) = \prod_{t=1}^T q(\mathbf{x}_t \mid \mathbf{x}_{t-1})\). Substituting,
\[
\mathcal{L}(\mathbf{x}_0) = \mathbb{E}_q\!\left[ -\log p(\mathbf{x}_T) - \sum_{t=1}^{T} \log \frac{p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t)}{q(\mathbf{x}_t \mid \mathbf{x}_{t-1})} \right].
\]
The summand mixes forward and reverse transitions, which makes direct identification with the reverse posterior of the previous
subsection awkward. The crucial step is to rewrite each forward transition using Bayes' rule together with the Markov property,
exactly as in the reverse-posterior derivation. For \(t \geq 2\),
\[
\begin{align*}
q(\mathbf{x}_t \mid \mathbf{x}_{t-1})
&= q(\mathbf{x}_t \mid \mathbf{x}_{t-1}, \mathbf{x}_0) \\\\
&= \frac{q(\mathbf{x}_{t-1} \mid \mathbf{x}_t, \mathbf{x}_0)\, q(\mathbf{x}_t \mid \mathbf{x}_0)}{q(\mathbf{x}_{t-1} \mid \mathbf{x}_0)}.
\end{align*}
\]
The first equality is the Markov property. The second is Bayes' rule. The case \(t = 1\) is left untouched, since
\(q(\mathbf{x}_1 \mid \mathbf{x}_0)\) already conditions on \(\mathbf{x}_0\) and no rewriting is needed.
Substituting this identity for every \(t \geq 2\) and separating the \(t = 1\) term, the ratio inside the sum becomes
\[
\log \frac{p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t)}{q(\mathbf{x}_t \mid \mathbf{x}_{t-1})}
= \log \frac{p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t)}{q(\mathbf{x}_{t-1} \mid \mathbf{x}_t, \mathbf{x}_0)} + \log \frac{q(\mathbf{x}_{t-1} \mid \mathbf{x}_0)}{q(\mathbf{x}_t \mid \mathbf{x}_0)}.
\]
The first term has the structure of a Kullback-Leibler integrand. It compares the learned reverse kernel to the reverse
posterior derived above. The second term telescopes when summed over \(t = 2, \ldots, T\), collapsing to
\(\log q(\mathbf{x}_1 \mid \mathbf{x}_0) - \log q(\mathbf{x}_T \mid \mathbf{x}_0)\).
The first part of this residue
cancels against the \(t = 1\) term that was left untouched, and the second part combines with \(-\log p(\mathbf{x}_T)\)
to form a Kullback-Leibler divergence comparing the forward terminal to the reverse prior. After collecting terms,
the negative ELBO admits the following decomposition.
Theorem: Decomposition of the Diffusion Negative ELBO
The negative ELBO of the diffusion model decomposes as
\[
\mathcal{L}(\mathbf{x}_0) = \mathcal{L}_T + \sum_{t=2}^{T} \mathcal{L}_{t-1} + \mathcal{L}_0,
\]
with
\[
\begin{align*}
\mathcal{L}_T &= D_{\mathbb{KL}}\!\left( q(\mathbf{x}_T \mid \mathbf{x}_0) \,\|\, p(\mathbf{x}_T) \right), \\\\[4pt]
\mathcal{L}_{t-1} &= \mathbb{E}_{q(\mathbf{x}_t \mid \mathbf{x}_0)}\!\left[ D_{\mathbb{KL}}\!\left( q(\mathbf{x}_{t-1} \mid \mathbf{x}_t, \mathbf{x}_0) \,\|\, p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t) \right) \right], \\\\[4pt]
\mathcal{L}_0 &= -\,\mathbb{E}_{q(\mathbf{x}_1 \mid \mathbf{x}_0)}\!\left[ \log p_{\boldsymbol{\theta}}(\mathbf{x}_0 \mid \mathbf{x}_1) \right].
\end{align*}
\]
Three observations follow, one for each kind of term.
-
The prior term \(\mathcal{L}_T\)
compares the forward terminal distribution to the fixed Gaussian prior of the reverse chain. Because the schedule is chosen
so that \(q(\mathbf{x}_T \mid \mathbf{x}_0) \approx \mathcal{N}(\mathbf{0}, \mathbf{I})\), this term is approximately zero
and contains no learnable parameters, so it is typically ignored during training.
-
The diffusion terms \(\mathcal{L}_{t-1}\) for \(t = 2, \ldots, T\) carry the entire training signal. Each one
asks the learned reverse kernel \(p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t)\) to match the reverse posterior
\(q(\mathbf{x}_{t-1} \mid \mathbf{x}_t, \mathbf{x}_0)\) in
Kullback-Leibler divergence, used
here in its continuous form, with the sum over outcomes replaced by an integral of densities.
-
The reconstruction term \(\mathcal{L}_0\) handles the final step \(\mathbf{x}_1 \to \mathbf{x}_0\),
which is treated separately because \(\mathbf{x}_0\) lies in the data space (typically discrete pixel values)
while \(\mathbf{x}_1\) does not.
The key conceptual point is that the variational problem has reduced to a sequence of step-local matching problems.
For each \(t\), we bring the learned reverse kernel close to the analytically known reverse posterior. We exploit this in the
next subsection by giving \(p_{\boldsymbol{\theta}}\) and \(q\) the same Gaussian functional form, so that each KL term becomes
a Gaussian-to-Gaussian divergence with a closed-form expression.
Noise Prediction and the DDPM Loss
The decomposition above leaves us with the diffusion terms \(\mathcal{L}_{t-1}\), each of which must be turned into something a
network can be trained against. The key step is to choose the functional form of the learned reverse kernel so that the
Kullback-Leibler divergence inside \(\mathcal{L}_{t-1}\) becomes a closed-form expression in the network outputs.
We take \(p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t) = \mathcal{N}(\mathbf{x}_{t-1};\, \boldsymbol{\mu}_{\boldsymbol{\theta}}(\mathbf{x}_t, t),\, \sigma_t^2 \mathbf{I})\),
with the variance fixed by hand to a schedule-dependent constant \(\sigma_t^2\). Only the mean
\(\boldsymbol{\mu}_{\boldsymbol{\theta}}\) is learned. Typical choices are \(\sigma_t^2 = \beta_t\) or the reverse-posterior
variance \(\tilde{\beta}_t\), both used in the literature. We revisit the question when we discuss sampling.
With this choice, both kernels in the KL are isotropic Gaussians, and the divergence reduces to a squared distance between means
plus a term \(C_t\) that does not depend on \(\boldsymbol{\theta}\):
\[
D_{\mathbb{KL}}\!\left( q(\mathbf{x}_{t-1} \mid \mathbf{x}_t, \mathbf{x}_0) \,\|\, p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t) \right)
= \frac{1}{2 \sigma_t^2} \left\| \tilde{\boldsymbol{\mu}}_t(\mathbf{x}_t, \mathbf{x}_0) - \boldsymbol{\mu}_{\boldsymbol{\theta}}(\mathbf{x}_t, t) \right\|^2 + C_t.
\]
The term \(C_t\) vanishes when \(\sigma_t^2 = \tilde{\beta}_t\).
A direct parameterization of \(\boldsymbol{\mu}_{\boldsymbol{\theta}}\) as a free neural network output is possible but was found
to perform worse in practice. The cleaner approach is to exploit the structure of the reverse posterior mean. Recall the one-shot
reparameterization \(\mathbf{x}_t = \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\varepsilon}\).
Solving for \(\mathbf{x}_0\) and substituting into the expression for \(\tilde{\boldsymbol{\mu}}_t(\mathbf{x}_t, \mathbf{x}_0)\)
derived above (the algebra simplifies neatly because the two coefficients combine through the same identity
\(\beta_t + \alpha_t (1 - \bar{\alpha}_{t-1}) = 1 - \bar{\alpha}_t\)) yields
\[
\tilde{\boldsymbol{\mu}}_t(\mathbf{x}_t, \mathbf{x}_0)
= \frac{1}{\sqrt{\alpha_t}} \left( \mathbf{x}_t - \frac{\beta_t}{\sqrt{1 - \bar{\alpha}_t}}\,\boldsymbol{\varepsilon} \right).
\]
The reverse posterior mean, viewed at the noisy sample \(\mathbf{x}_t\), is a deterministic function of the
noise \(\boldsymbol{\varepsilon}\) that was used to produce \(\mathbf{x}_t\) from \(\mathbf{x}_0\). This suggests
parameterizing \(\boldsymbol{\mu}_{\boldsymbol{\theta}}\) in the same form, with a learned noise predictor
\(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t)\) replacing the true noise:
\[
\boldsymbol{\mu}_{\boldsymbol{\theta}}(\mathbf{x}_t, t)
= \frac{1}{\sqrt{\alpha_t}} \left( \mathbf{x}_t - \frac{\beta_t}{\sqrt{1 - \bar{\alpha}_t}}\,\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t) \right).
\]
The two means differ only through \(\boldsymbol{\varepsilon} - \boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\), and substituting
into the squared-distance form of the KL gives the diffusion term
\[
\mathcal{L}_{t-1} = \mathbb{E}_{\boldsymbol{\varepsilon}}\!\left[ \frac{\beta_t^2}{2 \sigma_t^2 \alpha_t (1 - \bar{\alpha}_t)} \,\left\| \boldsymbol{\varepsilon} - \boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\!\left(\sqrt{\bar{\alpha}_t}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\varepsilon},\, t \right) \right\|^2 \right] + C_t,
\]
with the expectation over \(\boldsymbol{\varepsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\) at fixed \(\mathbf{x}_0\), as in
the decomposition. Training the diffusion model has reduced to a regression problem. Given a noisy sample and a time index,
predict the noise.
The time-dependent weight in front of the squared error, call it \(\lambda_t\), is what a faithful derivation of the variational
bound produces. The objective actually used in practice discards it, for reasons recorded in modification (iii) below.
Theorem: The DDPM Training Objective
The simplified DDPM loss is
\[
\mathcal{L}_{\text{simple}}(\boldsymbol{\theta})
= \mathbb{E}_{t,\,\mathbf{x}_0,\,\boldsymbol{\varepsilon}}\!\left[ \left\| \boldsymbol{\varepsilon} - \boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\!\left( \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\varepsilon},\, t \right) \right\|^2 \right],
\]
where \(t \sim \mathrm{Unif}\{1, \ldots, T\}\), \(\mathbf{x}_0 \sim q_0\), and
\(\boldsymbol{\varepsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\). This objective is obtained from the full negative ELBO,
averaged over \(\mathbf{x}_0 \sim q_0\), by three modifications.
(i) Dropping the prior term. The prior term \(\mathcal{L}_T = D_{\mathbb{KL}}(q(\mathbf{x}_T \mid \mathbf{x}_0) \,\|\, p(\mathbf{x}_T))\)
depends only on the schedule \(\{\beta_t\}\) and not on the network parameters \(\boldsymbol{\theta}\). Its gradient with respect to
\(\boldsymbol{\theta}\) is therefore identically zero, and it can be removed without affecting training.
(ii) Folding the reconstruction term into the same form. The reconstruction term
\(\mathcal{L}_0 = -\mathbb{E}_{q(\mathbf{x}_1 \mid \mathbf{x}_0)}[\log p_{\boldsymbol{\theta}}(\mathbf{x}_0 \mid \mathbf{x}_1)]\)
has a different functional form from the diffusion terms \(\mathcal{L}_{t-1}\) for \(t \geq 2\). Treating \(\mathbf{x}_0\) as
continuous and modeling \(p_{\boldsymbol{\theta}}(\mathbf{x}_0 \mid \mathbf{x}_1)\) as a Gaussian with the same
noise-prediction parameterization evaluated at \(t = 1\), the reconstruction term reduces (up to a constant) to a \(t = 1\)
instance of the squared-error term inside the expectation. The sum over diffusion terms can therefore be extended from
\(t = 2, \ldots, T\) to \(t = 1, \ldots, T\), and the extended sum absorbs the reconstruction term. (For genuinely discrete
data such as image pixels, a separate discretized decoder is needed at \(t = 1\). The loss above continues to give the correct
gradient signal for the rest of the chain.)
(iii) Setting the time-dependent weight to one. The principled per-step weight
\(\lambda_t = \beta_t^2 / (2 \sigma_t^2 \alpha_t (1 - \bar{\alpha}_t))\) is replaced by \(\lambda_t = 1\). This is an empirical
choice. The \(\boldsymbol{\theta}\)-independent constants \(C_t\) are dropped as well, for the same reason as the prior term.
The resulting objective is no longer, in general, a variational bound on \(-\log p_{\boldsymbol{\theta}}(\mathbf{x}_0)\), but
it weights all noise levels equally and was found by Ho, Jain, and Abbeel to produce visibly better samples.
Algorithm: DDPM Training
Input: Dataset \(\mathcal{D}\), schedule \(\{\beta_t\}_{t=1}^T\), network \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\)
repeat
\(\mathbf{x}_0 \sim \mathcal{D}\);
\(t \sim \mathrm{Unif}\{1, \ldots, T\}\);
\(\boldsymbol{\varepsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\);
\(\mathbf{x}_t \leftarrow \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\varepsilon}\);
Take a gradient step on \(\nabla_{\boldsymbol{\theta}} \left\| \boldsymbol{\varepsilon} - \boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t) \right\|^2\);
until converged
The architecture of \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\) depends on the data modality. For images, a common choice is
a U-Net with skip connections between matching resolution levels and a small time embedding injected into each block. A transformer
backbone (the diffusion transformer, or DiT) is another. The point is that \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\) is
just a neural network that maps a tensor and a scalar time index to a tensor of the same shape as the input. From the perspective
of training, all the structural insight of the diffusion model lives in the loss above, not in the network architecture.
The Score-Matching Perspective
The term score function has appeared before in a different sense. In the classical statistical setting that goes back to
Fisher, the score function is the
gradient of the log-likelihood with respect to the parameters of a model,
\(s(\boldsymbol{\theta}) = \nabla_{\boldsymbol{\theta}} \log p(\mathbf{x} \mid \boldsymbol{\theta})\). This is the quantity that
underlies the Fisher information matrix and natural gradient descent.
The diffusion-model literature, following the score-matching framework of Hyvärinen (2005), uses the same term for a formally
different quantity, a gradient taken with respect to the data instead. Both are gradients of log-densities, but they live
in different spaces and serve different purposes. The parameter gradient is for inference and optimization, while the data gradient
describes the geometry of the distribution at a point in sample space. Throughout this subsection, score function refers
exclusively to the data-gradient version.
Definition: Score Function (Data Gradient)
Let \(p\) be a strictly positive, differentiable probability density on \(\mathbb{R}^d\). The score function
of \(p\) is the gradient of the log-density with respect to the data argument,
\[
\mathbf{s}(\mathbf{x}) := \nabla_{\mathbf{x}} \log p(\mathbf{x}).
\]
Geometrically, \(\mathbf{s}(\mathbf{x})\) is a vector field on \(\mathbb{R}^d\) that points, at every \(\mathbf{x}\), in the
direction of steepest ascent of the probability density.
We now connect this object to the diffusion model. The forward chain provides a family of marginal distributions
\(q_t(\mathbf{x}_t) := \int q(\mathbf{x}_t \mid \mathbf{x}_0)\, q_0(\mathbf{x}_0)\, d\mathbf{x}_0\), one for each time index \(t\),
that interpolate between the data distribution at \(t = 0\) and approximately standard Gaussian noise at \(t = T\). For
\(t \geq 1\), each \(q_t\) is a Gaussian smoothing of the data distribution, hence strictly positive and differentiable, with its
own score function \(\mathbf{s}_t(\mathbf{x}_t) = \nabla_{\mathbf{x}_t} \log q_t(\mathbf{x}_t)\), and the diffusion model turns out
to be implicitly estimating all of them at once.
The mechanism is a direct calculation. The conditional forward kernel
\(q(\mathbf{x}_t \mid \mathbf{x}_0) = \mathcal{N}(\mathbf{x}_t;\, \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0,\, (1 - \bar{\alpha}_t)\mathbf{I})\)
has, as a Gaussian in \(\mathbf{x}_t\), the score
\[
\nabla_{\mathbf{x}_t} \log q(\mathbf{x}_t \mid \mathbf{x}_0)
= -\,\frac{\mathbf{x}_t - \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0}{1 - \bar{\alpha}_t}.
\]
Substituting the reparameterization
\(\mathbf{x}_t = \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\varepsilon}\) gives the key identity
\[
\nabla_{\mathbf{x}_t} \log q(\mathbf{x}_t \mid \mathbf{x}_0)
= -\,\frac{\boldsymbol{\varepsilon}}{\sqrt{1 - \bar{\alpha}_t}}.
\]
The conditional score is, up to a negative scaling, exactly the Gaussian noise that produced \(\mathbf{x}_t\) from
\(\mathbf{x}_0\). A theorem due to Vincent (2011), known as denoising score matching, extends this from
the conditional to the marginal. A network trained to predict the noise on samples drawn from the forward chain learns the
score of the marginal \(q_t\). Writing \(\sigma_t := \sqrt{1 - \bar{\alpha}_t}\), the noise-prediction network
\(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\) and the score network \(\mathbf{s}_{\boldsymbol{\theta}}\) are linked
by the parameterization
\[
\mathbf{s}_{\boldsymbol{\theta}}(\mathbf{x}_t, t) = -\,\frac{\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t)}{\sigma_t}.
\]
The argument takes two steps. Assume the data distribution permits differentiating
\(q_t(\mathbf{x}_t) = \int q(\mathbf{x}_t \mid \mathbf{x}_0)\, q_0(\mathbf{x}_0)\, d\mathbf{x}_0\) under the integral sign.
Dividing the derivative by \(q_t(\mathbf{x}_t)\) and writing \(q(\mathbf{x}_0 \mid \mathbf{x}_t)\) for the posterior of the clean
data given the noisy sample, we obtain
\[
\begin{align*}
\nabla_{\mathbf{x}_t} \log q_t(\mathbf{x}_t)
&= \int \nabla_{\mathbf{x}_t} \log q(\mathbf{x}_t \mid \mathbf{x}_0)\, q(\mathbf{x}_0 \mid \mathbf{x}_t)\, d\mathbf{x}_0 \\\\
&= -\,\frac{\mathbb{E}[\boldsymbol{\varepsilon} \mid \mathbf{x}_t]}{\sigma_t},
\end{align*}
\]
where the second line uses the key identity above. On the other hand, among all square-integrable functions of \(\mathbf{x}_t\),
the squared error \(\mathbb{E}\,\|\boldsymbol{\varepsilon} - \mathbf{g}(\mathbf{x}_t)\|^2\) over the joint law of
\((\mathbf{x}_0, \boldsymbol{\varepsilon})\) is minimized by
\(\mathbf{g}(\mathbf{x}_t) = \mathbb{E}[\boldsymbol{\varepsilon} \mid \mathbf{x}_t]\), since
conditional expectation is the orthogonal projection in \(L^2\).
A noise predictor that attains the minimum at step \(t\) therefore satisfies
\(-\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t) / \sigma_t = \nabla_{\mathbf{x}_t} \log q_t(\mathbf{x}_t)\).
Remark on notation. The symbol \(\sigma_t\) is used in two distinct senses on this page, and the literature
is not always careful to distinguish them. In the previous subsection, \(\sigma_t^2\) denoted the variance of the learned
reverse kernel \(p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t)\), a hyperparameter typically set to \(\beta_t\)
or \(\tilde{\beta}_t\). Here \(\sigma_t = \sqrt{1 - \bar{\alpha}_t}\) is the standard deviation of the forward marginal
\(q(\mathbf{x}_t \mid \mathbf{x}_0)\), determined entirely by the noise schedule. The two quantities coincide only when
\(\sigma_t^2 = 1 - \bar{\alpha}_t\), which is not the conventional choice for the reverse-kernel variance. We retain the standard
overloaded notation because it is universal in the diffusion literature, but the reader should keep the distinction in mind.
Under this identification, the noise-prediction objective of DDPM and the denoising score-matching objective of score-based
models coincide up to the choice of per-step weighting. The two frameworks are not identical (DDPM is set up as a discrete-time
variational model and score-based generative modeling as a continuous-time score-matching model), but their trained networks
are interchangeable through the parameterization above.
With this identification in hand, we return to the thread opened in the introduction. The
page on PCA and autoencoders took on faith
that a denoising autoencoder trained at noise level \(\sigma\) implicitly learns the score
\(\nabla_{\mathbf{x}} \log p(\mathbf{x})\) of the data distribution, via the asymptotic identity
\(\mathbf{r}^*(\mathbf{x}) - \mathbf{x} = \sigma^2 \nabla_{\mathbf{x}} \log p(\mathbf{x}) + o(\sigma^2)\) of Vincent (2011) and
Alain and Bengio (2014).
What we have just developed is the multi-scale version of that statement. Rather than fixing one noise level, the diffusion model
trains a single network across the full schedule of noise levels \(\{\sigma_t\}_{t=1}^T\), and the resulting noise predictor
doubles as a score estimator for the entire family of marginals \(\{q_t\}_{t=1}^T\). The promise made on the autoencoder page (that
score-matching theory justifies the denoising construction) is what the diffusion framework redeems.
In the continuous-time formulation previewed in the introduction, the score function is the central object. Anderson's 1982
reversal theorem expresses the reverse drift in terms of \(\nabla \log q_t\), and the score-based generative modeling framework of
Song and Ermon (2019) and Song et al. (2021) takes this viewpoint as primary. The discrete DDPM picture developed on this page
corresponds to a particular time-discretization of the variance-preserving SDE.
Sampling: DDPM and DDIM
Ancestral Sampling
With \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\) trained, generation proceeds by simulating the reverse chain. Starting from
\(\mathbf{x}_T \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\), we sample successively from
\(p_{\boldsymbol{\theta}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t) = \mathcal{N}(\mathbf{x}_{t-1};\, \boldsymbol{\mu}_{\boldsymbol{\theta}}(\mathbf{x}_t, t),\, \sigma_t^2 \mathbf{I})\)
for \(t = T, T-1, \ldots, 1\), and return \(\mathbf{x}_0\). Because each variable is drawn from its conditional distribution given
its parent in the directed chain \(\mathbf{x}_T \to \cdots \to \mathbf{x}_0\), this procedure is called
ancestral sampling. Substituting the noise-prediction form of \(\boldsymbol{\mu}_{\boldsymbol{\theta}}\) derived
earlier gives the explicit step rule used in practice.
Algorithm: DDPM Ancestral Sampling
Input: Trained network \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\), schedule \(\{\beta_t\}_{t=1}^T\), reverse-kernel variances \(\{\sigma_t^2\}_{t=1}^T\)
\(\mathbf{x}_T \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\);
for \(t = T, T-1, \ldots, 1\) do
if \(t \gt 1\): \(\mathbf{z} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\); else: \(\mathbf{z} = \mathbf{0}\);
\(\mathbf{x}_{t-1} \leftarrow \dfrac{1}{\sqrt{\alpha_t}} \left( \mathbf{x}_t - \dfrac{\beta_t}{\sqrt{1 - \bar{\alpha}_t}}\,\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t) \right) + \sigma_t \mathbf{z}\);
Output: \(\mathbf{x}_0\)
Two design choices remain.
-
The reverse-kernel variance \(\sigma_t^2\), which was fixed by hand during training rather than learned. Ho,
Jain, and Abbeel observed that two natural choices give comparable sample quality. Setting \(\sigma_t^2 = \beta_t\) is optimal
when \(\mathbf{x}_0 \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\). Setting \(\sigma_t^2 = \tilde{\beta}_t\), the reverse-posterior
variance computed earlier, is optimal when \(\mathbf{x}_0\) is a single fixed point. For data with coordinatewise unit
variance, the two are the extreme choices, giving upper and lower bounds on the entropy of the reverse process. In practice
both work well, with the difference between them small compared with other sources of variation in the sampler.
-
The suppression of noise at the final step. When \(t = 1\), the algorithm sets \(\mathbf{z} = \mathbf{0}\)
and returns the deterministic mean. This avoids injecting Gaussian noise into the final decoded output, which would otherwise
visibly degrade samples in image domains.
The principal limitation of ancestral sampling is its cost. With \(T = 1000\) (the value used in the original DDPM experiments),
generating a single sample requires one thousand sequential network evaluations, and the steps cannot be parallelized because each
\(\mathbf{x}_{t-1}\) depends on \(\mathbf{x}_t\). This is the bottleneck that the next subsection addresses.
Deterministic Sampling with DDIM
Song, Meng, and Ermon's denoising diffusion implicit models (DDIM, 2021) rest on one observation.
One can sample from a DDPM in far fewer steps, and with a fully deterministic reverse process if desired,
without retraining the network. The same \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\) that was trained against
the DDPM objective is reused, and only the sampling procedure changes.
The mechanism is a family of reverse kernels indexed by stochasticity parameters \(\tilde{\sigma}_t \geq 0\) with
\(\tilde{\sigma}_t^2 \leq 1 - \bar{\alpha}_{t-1}\). For each choice of \(\tilde{\sigma}_t\), Song, Meng, and Ermon construct
a reverse distribution
\[
q_{\tilde{\sigma}}(\mathbf{x}_{t-1} \mid \mathbf{x}_t, \mathbf{x}_0)
= \mathcal{N}\!\left( \mathbf{x}_{t-1}; \sqrt{\bar{\alpha}_{t-1}}\,\mathbf{x}_0 + \sqrt{1 - \bar{\alpha}_{t-1} - \tilde{\sigma}_t^2}\, \cdot \frac{\mathbf{x}_t - \sqrt{\bar{\alpha}_t}\,\mathbf{x}_0}{\sqrt{1 - \bar{\alpha}_t}}, \tilde{\sigma}_t^2 \mathbf{I} \right)
\]
that is consistent with the forward marginals \(q(\mathbf{x}_t \mid \mathbf{x}_0)\) used during training, although the forward
process it induces is not in general Markovian.
Two members of this family are distinguished. Setting \(\tilde{\sigma}_t^2 = \tilde{\beta}_t\), the reverse-posterior variance,
recovers the DDPM reverse posterior of the previous section. Setting \(\tilde{\sigma}_t = 0\) yields a fully deterministic update,
the DDIM sampler.
Remark on notation. The symbol \(\tilde{\sigma}_t\) here is yet another \(\sigma\)-like quantity, distinct from the
reverse-kernel variance \(\sigma_t^2\) of DDPM sampling and from the forward marginal standard deviation
\(\sigma_t = \sqrt{1 - \bar{\alpha}_t}\) used in the score-matching subsection. We use the tilde to mark it as the DDIM
stochasticity parameter, distinct from both.
With \(\mathbf{x}_0\) replaced by the network's prediction
\(\hat{\mathbf{x}}_0(\mathbf{x}_t, t) := \frac{\mathbf{x}_t - \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t)}{\sqrt{\bar{\alpha}_t}}\),
obtained by solving the forward reparameterization for \(\mathbf{x}_0\), and \(\tilde{\sigma}_t\) set to zero, the deterministic
DDIM step takes a particularly transparent form.
Algorithm: DDIM Deterministic Sampling
Input: Trained network \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\), schedule \(\{\beta_t\}_{t=1}^T\), step subsequence \(\tau_1 \lt \tau_2 \lt \cdots \lt \tau_S\) with \(\tau_S = T\)
\(\mathbf{x}_{\tau_S} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})\);
for \(i = S, S-1, \ldots, 1\) do
\(t \leftarrow \tau_i\); \( s \leftarrow \tau_{i-1}\) (with \(\tau_0 := 0\), \(\bar{\alpha}_{\tau_0} := 1\));
\(\hat{\mathbf{x}}_0 \leftarrow \dfrac{\mathbf{x}_t - \sqrt{1 - \bar{\alpha}_t}\,\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t)}{\sqrt{\bar{\alpha}_t}}\);
\(\mathbf{x}_s \leftarrow \sqrt{\bar{\alpha}_s}\,\hat{\mathbf{x}}_0 + \sqrt{1 - \bar{\alpha}_s}\,\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t)\);
Output: \(\mathbf{x}_0\)
The step decomposes into two interpretable operations.
-
The current noisy sample \(\mathbf{x}_t\) is mapped to a prediction
of the clean data \(\hat{\mathbf{x}}_0\) by undoing the forward reparameterization with the network's noise estimate.
-
This predicted clean sample is re-corrupted to the target time \(s\), using the same noise estimate
\(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t)\) rather than a fresh Gaussian draw. The coefficients
\(\sqrt{\bar{\alpha}_s}\) and \(\sqrt{1 - \bar{\alpha}_s}\), whose squares sum to one, are those of the one-shot forward
reparameterization at time \(s\), so the result has the form of a forward sample at time \(s\) built from
\(\hat{\mathbf{x}}_0\) and the estimated noise.
Because no fresh randomness enters once \(\mathbf{x}_T\) has been drawn, the entire trajectory from \(\mathbf{x}_T\) to \(\mathbf{x}_0\)
is a deterministic function of the initial noise, and seeds can be replayed exactly.
A second feature of the DDIM construction is that the step indices need not visit every \(t\) in \(\{1, \ldots, T\}\). Any subsequence
\(\tau_1 \lt \cdots \lt \tau_S = T\) of the training schedule is valid for sampling, and choosing \(S\) much smaller than \(T\)
yields a substantial speedup. Typical practice uses \(S\) on the order of 50 with \(T = 1000\) during training, a twenty-fold
reduction in network evaluations per sample, with image quality that approximates the full-chain DDPM sampler.
Conditional Generation
Practical diffusion systems rarely sample from the unconditional data distribution alone. They generate conditioned on auxiliary
information \(\mathbf{c}\) (a class label, a low-resolution image to upsample, a text prompt), and the user expects the sampled
output to reflect that condition. The minimal modification is to make the network condition-aware. We replace
\(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t)\) with
\(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t, \mathbf{c})\) and train on \((\mathbf{x}_0, \mathbf{c})\) pairs
drawn jointly from data.
The mechanism of injection depends on the modality of \(\mathbf{c}\). Class labels are typically embedded and added to the time
embedding, conditioning images are concatenated to \(\mathbf{x}_t\) along the channel axis, and text prompts are encoded by a
separate model and consumed through
cross-attention at multiple resolutions
inside the network. The loss is unchanged, the simple denoising objective applied to the conditional network.
Conditional training alone, however, tends to produce samples that are only weakly attached to \(\mathbf{c}\). The network
happily falls back on the prior whenever the conditional signal is ambiguous. The standard remedy, introduced by Ho and
Salimans (2021), is classifier-free guidance. During training, the conditioning input is replaced by
a null token \(\varnothing\) with some fixed probability (typically 10 to 20 percent), so that the same network learns
both the conditional predictor \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t, \mathbf{c})\) and the
unconditional predictor \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t, \varnothing)\).
At sampling time, the two predictions are combined into a guided estimate
\[
\tilde{\boldsymbol{\varepsilon}}_{\boldsymbol{\theta}}(\mathbf{x}_t, t, \mathbf{c})
= (1 + w)\,\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t, \mathbf{c})
- w\,\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}(\mathbf{x}_t, t, \varnothing),
\]
which is then substituted for \(\boldsymbol{\varepsilon}_{\boldsymbol{\theta}}\) inside any of the sampling algorithms above.
The guidance weight \(w \geq 0\) governs the trade-off. Setting \(w = 0\) recovers the unguided conditional
sampler, while larger \(w\) extrapolates beyond the conditional prediction, away from the unconditional one. The resulting samples
adhere more tightly to \(\mathbf{c}\), at the cost of reduced diversity and occasional artifacts.
The score-matching perspective developed earlier illuminates the construction. Bayes' rule on log-densities gives the identity
\(\nabla_{\mathbf{x}_t} \log p(\mathbf{x}_t \mid \mathbf{c}) = \nabla_{\mathbf{x}_t} \log p(\mathbf{x}_t) + \nabla_{\mathbf{x}_t} \log p(\mathbf{c} \mid \mathbf{x}_t)\),
so the conditional score decomposes into an unconditional score and a classifier-like correction. Classifier-free guidance, viewed
at the score level, replaces this exact decomposition with an extrapolation. The guided score
\((1+w)\mathbf{s}_{\boldsymbol{\theta}}(\mathbf{x}_t, t, \mathbf{c}) - w\,\mathbf{s}_{\boldsymbol{\theta}}(\mathbf{x}_t, t, \varnothing)\)
overweights the conditional direction relative to its true Bayesian value and so biases samples toward regions where the
classifier-like likelihood \(p(\mathbf{c} \mid \mathbf{x}_t)\) is high.
Interactive Demo
The visualization below applies the full diffusion construction (forward chain, score function, and reverse sampling) to a point
cloud built from the MATH-CS COMPASS logo. Each point carries five coordinates, two spatial \((x, y)\) and three color
\((r, g, b)\), and the diffusion is run jointly on all five. The forward direction drives both the geometry and the color of the
logo toward a pure Gaussian distribution. The reverse direction reconstructs points on the data manifold from noise, with the logo
emerging as the support of the distribution being sampled.
Try the forward direction first to watch the logo dissolve, then switch to reverse and observe the same construction running
backwards. The time slider scrubs through any intermediate state. Three overlays expose the machinery rather than just the result.
-
Trails follow twelve tracked points through recent steps. Under DDPM they wander in a
Brownian zigzag, while under DDIM the same points move along a few straight, deterministic jumps, the structural difference
between the two samplers, visible per particle.
-
Score field draws the \((x, y)\) components of the exact
score \(\nabla_{\mathbf{x}_t} \log q_t\) as arrows at a subsample of the points (arrow length saturates with magnitude so
the near-data regime, where the score blows up, stays readable).
-
x̂0 ghost shows, during the
reverse process, the model's current prediction \(\hat{\mathbf{x}}_0\) of where each point will land, the quantity the DDIM
update is built around.
Finally, when the reverse process reaches \(t = 0\), a readout reports the mean distance from the
generated points to the data manifold, turning the visual comparison between samplers into a number.
DDPM versus DDIM: Stochastic versus Deterministic Reverse Paths
The sampler toggle is the most pedagogically informative control. Both DDPM and DDIM start from the same initial draw
\(\mathbf{x}_T\) (use the "regenerate noise" button to draw a new one), and both end on the data manifold.
What differs is the structure of the trajectory between them. The DDPM update is stochastic. Each step draws a fresh Gaussian,
so repeated runs from the same \(\mathbf{x}_T\) would in principle produce different reconstructions (this demo caches a single
such run per noise seed for visualization). The DDIM update is deterministic. The entire trajectory is a fixed function of
\(\mathbf{x}_T\), and the demo's DDIM sampler reaches the data manifold in roughly five times fewer steps than DDPM (twenty
against one hundred here).
Because DDIM compresses many small DDPM steps into a few large jumps, the reconstruction in this demo lands close to, but not
exactly on, the original logo points, so the points look slightly hazier than under DDPM. The trails overlay shows the cause
and the \(t = 0\) readout shows the effect. For the same starting noise, the mean distance to the data under DDPM is clearly
smaller than under the twenty-step DDIM sampler.
This residual is the discretization error of the few-step DDIM sampler, the small accuracy cost that pays for the speedup.
The number of steps therefore trades accuracy for speed, and more refined few-step samplers aim to shrink this error at
the same cost.
The construction visualized here is a small instance of the mathematics that underlies diffusion models at scale. Three
replacements turn it into such a model. The five-dimensional point cloud becomes a high-dimensional data tensor, the point
distribution becomes the data distribution, and the analytical score becomes a trained neural network. Some models run the
diffusion in a learned latent space rather than directly on the data. Larger models add several ingredients on top (classifier-free
guidance, learned encoders for the conditioning input, alternative parameterizations and samplers), but the discrete-time DDPM and
DDIM updates developed on this page remain the structural core.
What makes the reverse direction possible is the score function \(\nabla_{\mathbf{x}_t} \log q_t(\mathbf{x}_t)\). Given the
gradient of the log-density of every intermediate marginal, walking back from noise to data is a matter of repeated denoising
steps, each one guided by the score evaluated at the current state.
The score-field overlay shows this directly. Switch it on during the reverse process and the arrows point the way from noise back
toward the logo, growing sharper as the marginal concentrates. In this demo the score is computed exactly from the empirical point
cloud. In practice it is learned from data, which is what diffusion training is really for.