Flow Matching

The Trajectory Problem Learning the Velocity Field Diffusion as a Special Case Interactive Demo

The Trajectory Problem

Flow matching is a principle for training generative models that is closely related to diffusion. The two differ in which trajectory from noise to data is integrated, and in why that trajectory is chosen.

The point of departure is our treatment of diffusion models. There, generation was framed as the reversal of a noising process. A sample is produced by starting from Gaussian noise and following a trajectory back toward the data distribution, one denoising step at a time. That page also noted, in passing, that the discrete denoising chain admits a continuous-time limit. In that limit the deterministic sampler traces out a smooth curve governed by an ordinary differential equation, a probability-flow ODE, rather than a noisy stochastic walk. Flow matching takes this deterministic, continuous-time picture as primary, and in doing so it defines the probability path directly instead of deriving it from a noising process.

The ancestral sampler of a diffusion model follows a stochastic trajectory. Fresh Gaussian noise is injected at every step, so the path between noise and data wanders. A wandering path is expensive. Faithfully integrating it demands many small steps, which is why classical diffusion sampling can require hundreds or thousands of network evaluations to produce a single sample.

Flow matching asks a more direct question. Rather than reversing a stochastic corruption, it seeks a deterministic map from a source distribution \(p_{\text{init}}\), typically but not necessarily Gaussian noise, to the data distribution \(p_{\text{data}}\), realized as the time-1 flow of an ODE. The object to be learned is the velocity field driving that ODE. Its central advantage, developed over the next two sections, is that the trajectories can be made nearly straight, and a straight path is cheap to integrate.

Two consequences of the deterministic ODE viewpoint deserve early mention, because they explain why flow matching is not merely a recasting of diffusion but a genuine generalization. First, the source distribution need not be Gaussian. Any distribution one can sample from may serve as \(p_{\text{init}}\), a freedom that diffusion's noise-centric construction does not naturally permit. Second, diffusion itself reappears inside this framework as one particular choice of trajectory, rather than as a separate theory. We develop the construction first, then return to both points once the machinery is in place.

Learning the Velocity Field

Fix a source distribution \(p_{\text{init}}\) and the data distribution \(p_{\text{data}}\), both on \(\mathbb{R}^d\). Flow matching seeks a time-dependent velocity field \(\mathbf{u}_t(\mathbf{x})\), defined for \(t \in [0, 1]\), such that points drawn from \(p_{\text{init}}\) and carried along the field arrive at time \(1\) distributed according to \(p_{\text{data}}\). Concretely, if \(\mathbf{x}_t\) solves the ordinary differential equation

\[ \frac{\mathrm{d}}{\mathrm{d}t}\,\mathbf{x}_t = \mathbf{u}_t(\mathbf{x}_t), \quad \mathbf{x}_0 \sim p_{\text{init}}, \]

then its distribution at each intermediate time should trace out a prescribed probability path \((p_t)_{0 \le t \le 1}\), a curve in the space of distributions running from \(p_0 = p_{\text{init}}\) to \(p_1 = p_{\text{data}}\). The flow of this ODE, the map carrying each starting point to its time-\(1\) endpoint, is the generator we want. Sampling means drawing \(\mathbf{x}_0\) from \(p_{\text{init}}\) and integrating forward to \(t = 1\). The flow is a well-defined map when the ODE has exactly one solution on \([0, 1]\) from every starting point. By the existence-and-uniqueness theory for ODEs (the Picard-Lindelöf theorem), this holds when \(\mathbf{u}_t(\mathbf{x})\) is continuous in \(t\) and satisfies a Lipschitz condition in \(\mathbf{x}\) with a constant independent of \(t\). Uniqueness also means that distinct trajectories never cross. We take this regularity for granted and turn to the real problem of finding the field.

A path of distributions, built one data point at a time

We are free to choose the probability path, and the choice is made by first prescribing it conditionally. For a fixed data point \(\mathbf{z} \sim p_{\text{data}}\), a conditional probability path \(p_t(\cdot \mid \mathbf{z})\) is a family of distributions interpolating from the source noise at \(t=0\) to a point mass concentrated on \(\mathbf{z}\) at \(t=1\). A common choice is the Gaussian conditional path

\[ p_t(\cdot \mid \mathbf{z}) = \mathcal{N}\!\big(\alpha_t \mathbf{z},\, \beta_t^2 I\big), \]

where the scalar schedules \(\alpha_t, \beta_t\) are differentiable and run monotonically from \(\alpha_0 = 0, \beta_0 = 1\) (pure noise) to \(\alpha_1 = 1, \beta_1 = 0\) (the variance collapses and the Gaussian concentrates on \(\mathbf{z}\)), with \(\alpha_t \gt 0\) for \(t \gt 0\) and \(\beta_t \gt 0\) for \(t \lt 1\). These are the coefficients of the mean and the standard deviation of the path, not the variance schedule \(\beta_t\) and its complement \(\alpha_t = 1 - \beta_t\) of the diffusion page. Time runs from noise at \(t = 0\) to data at \(t = 1\). Sampling from the path is immediate. Draw \(\boldsymbol{\varepsilon} \sim \mathcal{N}(\mathbf{0}, I)\) and set \(\mathbf{x}_t = \alpha_t \mathbf{z} + \beta_t \boldsymbol{\varepsilon}\). The simplest schedule of all, \(\alpha_t = t\) and \(\beta_t = 1 - t\), moves each conditional sample \(\mathbf{x}_t = t\mathbf{z} + (1 - t)\boldsymbol{\varepsilon}\) along the straight segment from \(\boldsymbol{\varepsilon}\) to \(\mathbf{z}\), a property we return to when we weigh the cost of sampling.

One point of generality deserves care. With \(\alpha_0 = 0\) and \(\beta_0 = 1\) the conditional path begins at \(p_0(\cdot \mid \mathbf{z}) = \mathcal{N}(\mathbf{0}, I)\) independently of \(\mathbf{z}\), so mixing over data points leaves the source pinned to \(p_{\text{init}} = \mathcal{N}(\mathbf{0}, I)\). This standard construction recovers a Gaussian source, the same starting point as diffusion. The freedom to take an arbitrary \(p_{\text{init}}\), promised earlier, is realized by a more general construction that conditions the path on a source-target pair \((\mathbf{x}_0, \mathbf{z})\) rather than on \(\mathbf{z}\) alone. The marginalization argument below goes through unchanged for that case. We develop the Gaussian-source version.

The path we actually care about is the marginal probability path, obtained by mixing the conditional paths over all data points,

\[ p_t(\mathbf{x}) = \int p_t(\mathbf{x} \mid \mathbf{z})\, p_{\text{data}}(\mathbf{z})\, \mathrm{d}\mathbf{z}. \]

One can sample from it (draw a data point, then draw from its conditional path), but its density is an intractable integral. This tractable-to-sample, intractable-to-evaluate split is the crux of everything that follows.

The continuity equation, and the trick it powers

What links a probability path to a velocity field is a conservation law. A field \(\mathbf{u}_t\) transports the distribution \(p_t\) correctly (meaning a particle started from \(p_0\) and obeying the ODE stays distributed according to \(p_t\)) if and only if the pair satisfies the continuity equation

\[ \partial_t\, p_t(\mathbf{x}) = -\,\operatorname{div}\!\big(p_t\, \mathbf{u}_t\big)(\mathbf{x}). \]

The reading is physical. The left side is the rate at which probability mass at \(\mathbf{x}\) changes in time. The divergence on the right measures net outflow of mass carried by the field, so its negative is net inflow. Probability mass moves only by being carried along the field, never created or destroyed at a point, so the two sides must agree. The same equation expresses conservation of mass in a fluid, compressible or not, and conservation of electric charge. We assume the regularity of \(p_t\) and \(\mathbf{u}_t\) under which this equivalence holds.

The continuity equation is what makes the conditional construction pay off. The marginal velocity field that transports \(p_t\) is recovered from the conditional fields by a weighted average. Each data point contributes the field aimed at it, weighted by how plausibly the current location \(\mathbf{x}\) arose from that point.

Theorem: The Marginalization Trick

Suppose that for each data point \(\mathbf{z}\) the conditional field \(\mathbf{u}_t(\cdot \mid \mathbf{z})\) transports the conditional path, so that \[ \partial_t\, p_t(\mathbf{x} \mid \mathbf{z}) = -\operatorname{div}\big(p_t(\cdot \mid \mathbf{z})\, \mathbf{u}_t(\cdot \mid \mathbf{z})\big)(\mathbf{x}), \] and assume that \(p_t\) is continuous and that \(\partial_t\) and \(\operatorname{div}\) may be exchanged with integration over \(\mathbf{z}\). Define the marginal field, wherever \(p_t(\mathbf{x}) \gt 0\), by \[ \mathbf{u}_t(\mathbf{x}) = \int \mathbf{u}_t(\mathbf{x} \mid \mathbf{z})\, \frac{p_t(\mathbf{x} \mid \mathbf{z})\, p_{\text{data}}(\mathbf{z})}{p_t(\mathbf{x})}\, \mathrm{d}\mathbf{z}. \] Then the pair \((p_t, \mathbf{u}_t)\) satisfies the continuity equation \(\partial_t\, p_t = -\operatorname{div}(p_t\, \mathbf{u}_t)\) on the open set where \(p_t \gt 0\).

This formula, the marginalization trick, is the theoretical heart of flow matching. Its weights are the posterior density of \(\mathbf{z}\) given \(\mathbf{x}_t = \mathbf{x}\), which is defined wherever \(p_t(\mathbf{x}) \gt 0\).

Proof:

Differentiating the mixture under the integral sign and applying the conditional continuity equation, we obtain \[ \begin{align*} \partial_t\, p_t(\mathbf{x}) &= -\int \operatorname{div}\big(p_t(\cdot \mid \mathbf{z})\, \mathbf{u}_t(\cdot \mid \mathbf{z})\big)(\mathbf{x})\, p_{\text{data}}(\mathbf{z})\, \mathrm{d}\mathbf{z} \\\\ &= -\operatorname{div}\big(p_t\, \mathbf{u}_t\big)(\mathbf{x}), \end{align*} \] since the integral of \(p_t(\mathbf{x} \mid \mathbf{z})\, \mathbf{u}_t(\mathbf{x} \mid \mathbf{z})\, p_{\text{data}}(\mathbf{z})\) over \(\mathbf{z}\) is \(p_t(\mathbf{x})\, \mathbf{u}_t(\mathbf{x})\) by the definition of the marginal field.

The marginal field is intractable, just like the marginal density. The conditional field, by contrast, is explicit. Differentiating \(\mathbf{x}_t = \alpha_t \mathbf{z} + \beta_t \boldsymbol{\varepsilon}\) in \(t\) gives the velocity \(\dot\alpha_t \mathbf{z} + \dot\beta_t \boldsymbol{\varepsilon}\), and substituting \(\boldsymbol{\varepsilon} = (\mathbf{x}_t - \alpha_t \mathbf{z})/\beta_t\) expresses it as a function of the current position,

\[ \mathbf{u}_t(\mathbf{x} \mid \mathbf{z}) = \dot\alpha_t\, \mathbf{z} + \frac{\dot\beta_t}{\beta_t}\big(\mathbf{x} - \alpha_t \mathbf{z}\big), \quad 0 \le t \lt 1. \]

A direct computation confirms that this field transports the Gaussian conditional path. The remaining question is whether we can train against the easy conditional field and still recover the hard marginal one.

Conditional flow matching

We can, and the reason is a clean identity between two losses. The loss we would like to minimize regresses a network \(\mathbf{u}_t^{\boldsymbol{\theta}}\) against the marginal field,

\[ \mathcal{L}_{\text{FM}}(\boldsymbol{\theta}) = \mathbb{E}\Big[\big\|\,\mathbf{u}_t^{\boldsymbol{\theta}}(\mathbf{x}_t) - \mathbf{u}_t(\mathbf{x}_t)\,\big\|^2\Big], \]

with \(\mathbf{x}_t \sim p_t\). The loss we can minimize, the conditional flow matching objective, regresses the network against the explicit conditional field,

\[ \mathcal{L}_{\text{CFM}}(\boldsymbol{\theta}) = \mathbb{E}\Big[\big\|\,\mathbf{u}_t^{\boldsymbol{\theta}}(\mathbf{x}_t) - \mathbf{u}_t(\mathbf{x}_t \mid \mathbf{z})\,\big\|^2\Big], \]

with the expectation over a random time \(t \sim \mathrm{Unif}[0,1]\), a data point \(\mathbf{z} \sim p_{\text{data}}\), and a point \(\mathbf{x}_t \sim p_t(\cdot \mid \mathbf{z})\) on the conditional path. In both losses \(t\) is uniform on \([0,1]\).

Theorem: Conditional and Marginal Flow Matching Differ by a Constant

Let \(\mathbf{u}_t\) be the marginal field of the marginalization trick, and suppose that \(\mathbb{E}\big[\|\mathbf{u}_t(\mathbf{x}_t \mid \mathbf{z})\|^2\big]\) and \(\mathbb{E}\big[\|\mathbf{u}_t(\mathbf{x}_t)\|^2\big]\) are finite. Then for every \(\boldsymbol{\theta}\) with \(\mathbb{E}\big[\|\mathbf{u}_t^{\boldsymbol{\theta}}(\mathbf{x}_t)\|^2\big]\) finite, the flow matching and conditional flow matching losses satisfy \[ \begin{align*} \mathcal{L}_{\text{CFM}}(\boldsymbol{\theta}) &= \mathcal{L}_{\text{FM}}(\boldsymbol{\theta}) \\\\ &\quad + \mathbb{E}\big[\|\mathbf{u}_t(\mathbf{x}_t \mid \mathbf{z})\|^2\big] - \mathbb{E}\big[\|\mathbf{u}_t(\mathbf{x}_t)\|^2\big], \end{align*} \] and the last two terms do not depend on \(\boldsymbol{\theta}\). In particular the two losses have the same minimizers, and the same gradient wherever either is differentiable.

Proof:

Drawing \(\mathbf{z}\) and then \(\mathbf{x}_t\) produces \(\mathbf{x}_t \sim p_t\), so the two expectations of \(\|\mathbf{u}_t^{\boldsymbol{\theta}}(\mathbf{x}_t)\|^2\) coincide. The cross terms are integrable because \(|\langle \mathbf{a}, \mathbf{b} \rangle| \le \frac{1}{2}(\|\mathbf{a}\|^2 + \|\mathbf{b}\|^2)\), and for each \(t\) the definition of the marginal field gives \[ \begin{align*} \iint \big\langle \mathbf{u}_t^{\boldsymbol{\theta}}(\mathbf{x}), \mathbf{u}_t(\mathbf{x} \mid \mathbf{z}) \big\rangle\, p_t(\mathbf{x} \mid \mathbf{z})\, p_{\text{data}}(\mathbf{z})\, \mathrm{d}\mathbf{z}\, \mathrm{d}\mathbf{x} &= \int \big\langle \mathbf{u}_t^{\boldsymbol{\theta}}(\mathbf{x}), \mathbf{u}_t(\mathbf{x}) \big\rangle\, p_t(\mathbf{x})\, \mathrm{d}\mathbf{x}. \end{align*} \] Averaging over \(t\), the cross terms of the two losses agree as well. Expanding both squares, every term that depends on \(\boldsymbol{\theta}\) cancels in the difference, which leaves the stated constant.

Minimizing the conditional loss is therefore exactly minimizing the marginal one. The network is only ever shown conditional targets, yet over a sufficiently expressive family the minimizer is their posterior average, the marginal field that transports \(p_{\text{init}}\) into \(p_{\text{data}}\).

Regressing against the conditional field is what makes training simulation-free. No ODE is solved during training. Each gradient step needs only a noise sample, a data sample, a random time, and the conditional velocity between them, a closed-form target evaluated without any integration through the field being learned. Contrast this with the alternatives the trick avoids. The marginal loss cannot be evaluated, since its target is the intractable integral above, and fitting the flow by maximum likelihood would require solving the ODE through the network being trained at every step. The marginalization trick converts these into ordinary regression.

Diffusion as a Special Case

We can now make good on a claim from the opening. Diffusion is not a rival to flow matching but an instance of it. The flow-matching construction transports a distribution along a probability path using a deterministic ODE driven by the velocity field \(\mathbf{u}_t\). The stochastic samplers of diffusion models correspond, in the continuous-time limit, to transporting a distribution along a probability path with a stochastic differential equation instead. The two are connected by a single knob.

One knob: from deterministic flow to stochastic diffusion

Given a velocity field \(\mathbf{u}_t\) that transports the path \(p_t\), one may add a noise term and obtain a stochastic differential equation that transports the very same path. For any diffusion coefficient \(\sigma_t \ge 0\), the stochastic differential equation

\[ \mathrm{d}\mathbf{x}_t = \Big[\mathbf{u}_t(\mathbf{x}_t) + \tfrac{\sigma_t^2}{2}\, \mathbf{s}_t(\mathbf{x}_t)\Big]\mathrm{d}t + \sigma_t\, \mathrm{d}\mathbf{w}_t, \]

started from \(\mathbf{x}_0 \sim p_{\text{init}}\), produces trajectories distributed according to \(p_t\) at every time. Here \(\mathbf{s}_t = \nabla \log p_t\) is the score function of the path and \(\mathbf{w}_t\) is Brownian motion. The score is defined because, for \(t \lt 1\), the Gaussian path is a mixture of nondegenerate Gaussians and so has \(p_t \gt 0\) everywhere. The deterministic flow is the endpoint \(\sigma_t = 0\). The noise vanishes, the score term vanishes with it, and the ODE of the previous section is recovered. Stochastic diffusion samplers live at \(\sigma_t \gt 0\), where the trajectories acquire the wandering character of Brownian motion while the distribution they carry is unchanged. The variance-preserving noising of classical denoising diffusion corresponds to paths with \(\alpha_t^2 + \beta_t^2 = 1\), read with time running from noise to data and with pure noise reached only approximately. Its ancestral sampler corresponds to one particular \(\sigma_t \gt 0\), and the deterministic DDIM sampler to \(\sigma_t = 0\).

The bookkeeping that guarantees the path is preserved is, once again, a conservation law. A stochastic differential equation with drift \(\boldsymbol{\mu}_t\) and diffusion coefficient \(\sigma_t\) evolves its density according to the Fokker-Planck equation

\[ \partial_t\, p_t = -\operatorname{div}(p_t\, \boldsymbol{\mu}_t) + \tfrac{\sigma_t^2}{2}\,\Delta p_t. \]

The first term is the familiar transport of the continuity equation. The Laplacian \(\Delta p_t\) is the diffusion term that also appears in the heat equation. It is the mathematical signature of mass spreading stochastically. Substituting the drift of our equation, \(\boldsymbol{\mu}_t = \mathbf{u}_t + \tfrac{\sigma_t^2}{2}\,\mathbf{s}_t\), makes the two parts conspire. The score contributes the term

\[ -\tfrac{\sigma_t^2}{2}\operatorname{div}(p_t \nabla \log p_t) = -\tfrac{\sigma_t^2}{2}\,\Delta p_t, \]

since \(p_t \nabla \log p_t = \nabla p_t\), and this cancels the diffusion term identically. What remains is

\[ \partial_t\, p_t = -\operatorname{div}(p_t\, \mathbf{u}_t), \]

the continuity equation of the previous section, for every \(\sigma_t\). This is the precise reason the whole family carries the same path. The score term is calibrated to absorb the diffusion the Brownian motion injects, and the cancellation is proved, with its hypotheses made explicit, as the Fokker-Planck invariant family theorem. Concluding that the stochastic differential equation has marginals \(p_t\) also uses uniqueness of solutions to the Fokker-Planck equation from the initial density \(p_{\text{init}}\), which we assume without proof.

The score and the velocity field carry the same information

The appearance of the score function above is not a coincidence of notation. For the Gaussian path, both fields are affine functions of the posterior mean \(\hat{\mathbf{z}}_t(\mathbf{x})\), the average of \(\mathbf{z}\) under the weights of the marginalization trick. Averaging the conditional field gives the first line below. Differentiating the Gaussian inside the mixture integral, again under the integral sign, gives the second.

\[ \begin{align*} \mathbf{u}_t(\mathbf{x}) &= \dot\alpha_t\, \hat{\mathbf{z}}_t(\mathbf{x}) + \frac{\dot\beta_t}{\beta_t}\big(\mathbf{x} - \alpha_t \hat{\mathbf{z}}_t(\mathbf{x})\big), \\\\ \mathbf{s}_t(\mathbf{x}) &= \frac{\alpha_t \hat{\mathbf{z}}_t(\mathbf{x}) - \mathbf{x}}{\beta_t^2}. \end{align*} \]

For \(0 \lt t \lt 1\), eliminating \(\hat{\mathbf{z}}_t\) leaves

\[ \mathbf{u}_t(\mathbf{x}) = \frac{\dot\alpha_t}{\alpha_t}\, \mathbf{x} + \frac{\beta_t\,(\dot\alpha_t \beta_t - \alpha_t \dot\beta_t)}{\alpha_t}\, \mathbf{s}_t(\mathbf{x}). \]

The coefficient of the score is nonzero whenever \(\dot\alpha_t \gt 0\) or \(\dot\beta_t \lt 0\), and then each field is recoverable from the other. What the diffusion literature calls learning the score function, and what flow matching calls learning the velocity field, are two parametrizations of one underlying quantity. The diffusion development reached it through a discrete noising chain, the evidence lower bound, and denoising score matching. The flow-matching development reached it through deterministic transport and the marginalization trick. They arrive at the same place.

Why the deterministic end is attractive

If every \(\sigma_t\) carries the same distribution, why prefer the deterministic end? Because of sampling cost. A stochastic trajectory wanders and typically needs many small steps. A deterministic trajectory, especially one whose path has been chosen to be nearly straight, can be integrated in few, and in the ideal limit a single step. Each step is one evaluation of the network, so straighter and deterministic means faster generation. The schedule whose conditional trajectories are straight segments from noise to data is the basis of rectified flow. Genuine straightness of the marginal trajectories connects to optimal transport. Under the quadratic cost, the optimal coupling between two distributions moves every particle along a straight line at constant speed.

In this sense flow matching generalizes diffusion. It keeps what made diffusion work (a tractable training signal, a smooth path from noise to data), exposes the stochasticity as a tunable coefficient rather than a built-in commitment, and frees the path itself to be chosen. Diffusion occupies the well-studied variance-preserving region of that design space, sampled at \(\sigma_t \gt 0\) by its ancestral sampler and at \(\sigma_t = 0\) by DDIM, and flow matching opens the rest of it.

Interactive Demo

Both panels integrate the same seeded batch of particles through the exact marginal velocity field of a five-mode Gaussian-mixture target. The field is available in closed form here (mixture responsibilities and a per-mode posterior mean), so nothing is trained or approximated, and only the probability path differs between the panels. Blue terminals land within a mode, and orange ones miss.

The slider sets the number of Euler steps \(N\), and each panel reports two numbers. The first is \(E(N)\), the mean distance between its terminals and a fine reference solution of the same ODE. This is the pure discretization error, which falls toward zero as \(N\) grows. The second is a straightness ratio, the trajectory path length divided by the displacement, which equals \(1\) exactly when the marginal trajectories are straight lines.

The ratio is where the demo earns its keep, because it comes out backwards from what a first reading of flow matching might suggest. By construction, the conditional trajectories of the linear path are straight segments, yet its marginal trajectories, the ones actually drawn in the left panel, are visibly bent (ratio around \(1.6\), against roughly \(1.15\) for the cosine path). The mechanism is in the field itself. The term \((\dot\beta_t/\beta_t)\,\mathbf{x}\) of the marginal field, with \(\dot\beta_t/\beta_t = -1/(1-t)\) on the linear path, contracts every particle toward the origin while the posterior-mean component pulls outward, so particles dip inward before swinging out to their mode. Conditional straightness does not survive marginalization. The marginal field is a posterior-weighted average of conditional fields, and averaging bends.

This gap is exactly what the rectification (Reflow) step of rectified flow exists to close. Re-coupling noise samples to the points the flow actually transports them to, then re-fitting with straight conditional segments between the coupled pairs, makes the marginal trajectories themselves straighter, and repeating it drives them toward straight lines. One round of flow matching with straight conditional paths, as shown here, does not.

The \(E(N)\) readout tells the complementary story. For both schedules the Euler error decays as \(N\) grows. At this problem scale the difference between the two paths is modest, since a few steps suffice for either once the field is exact. The large step-count savings reported for rectified flow belong to the rectified regime, where the marginal paths have been straightened. The toy here shows where the baseline stands before that step.