Heat Equations

Introduction Heat Equation on a Bounded Interval Heat Equation on the Real Line Properties of the Heat Kernel

Introduction

Three pages have built the toolkit of Fourier analysis: the Fourier series, the Fourier transform, and their Hilbert-space formulation in Fourier analysis in Hilbert spaces. From a computer science perspective, it can appear essentially complete. Orthogonal decomposition, Plancherel's identity, convolution as pointwise multiplication, and the FFT cover what most CS curricula treat as the entirety of the subject. Fourier is the engine of signal processing, image compression, and spectral algorithms. Yet a question implicit in our earlier development of Fourier series remains unanswered: what was Fourier analysis invented for? The present page returns to that original purpose. We use the Fourier machinery already built to solve a partial differential equation.

Until recently, partial differential equations sat outside the CS mainstream. PDEs were the province of physics and applied mathematics. The heat equation belonged to thermodynamics, the wave equation to acoustics and electromagnetism, the Laplace equation to electrostatics and fluid flow. A CS student might encounter a discretized PDE in a numerical methods elective, but the continuous theory itself rarely entered the curriculum.

This perception has been overturned by developments of the deep-learning era. The forward process of a diffusion generative model, the mechanism behind much of modern image and video synthesis, is in continuous time a stochastic diffusion. Its density evolution is governed by the Fokker-Planck equation, a parabolic PDE that reduces to the heat equation in the simplest case of vanishing drift and constant diffusion coefficient. The continuous domain, long dismissed as a physicist's concern, has become impossible for the ML researcher to ignore.

The heat equation \[ \partial_t u(x, t) = k\, \partial_{xx} u(x, t) \] is the canonical example of a parabolic partial differential equation, where \(u(x, t)\) represents temperature (or, abstractly, the density of a diffusing quantity) at position \(x\) and time \(t\), and \(k \gt 0\) is a positive constant called the diffusivity.

Two distinct senses motivate its study here. Mathematically, the heat equation is the parabolic representative of the classical trichotomy of parabolic, hyperbolic, and elliptic equations. The other two representatives, the wave equation and the Laplace equation, are developed in subsequent pages. In application, it underlies the Gaussian forward corruption of diffusion models, the graph heat kernel of spectral graph neural networks, and the Gaussian scale-space of classical image processing. The treatment is one-dimensional and classical: separation of variables on a bounded interval, the Fourier transform on the real line, the heat kernel and its qualitative properties. Higher dimensions, weak solutions in Sobolev spaces, and semigroup theory belong to more advanced treatments. The stochastic interpretation via Brownian motion is taken up in the pages on Itô diffusions.

Heat Equation on a Bounded Interval

Heat conduction was the very problem that motivated Joseph Fourier to introduce trigonometric series in his 1822 Théorie analytique de la chaleur. We begin with the prototype on which the entire theory rests: temperature evolution in a one-dimensional rod of length \(L\) whose two ends are held at zero temperature. The mathematical formulation is the initial-boundary value problem \[ \begin{cases} \partial_t u = k\, \partial_{xx} u, & 0 \lt x \lt L,\ t \gt 0, \\\\ u(0, t) = u(L, t) = 0, & t \geq 0, \\\\ u(x, 0) = f(x), & 0 \leq x \leq L, \end{cases} \] where \(f\) is a prescribed initial temperature profile. At the two corners \((0, 0)\) and \((L, 0)\) the boundary and initial conditions both apply, and they agree only if \(f(0) = f(L) = 0\). The theorem below fixes the precise sense in which each condition is imposed.

The boundary conditions \(u(0,t) = u(L,t) = 0\) are the homogeneous Dirichlet conditions, corresponding physically to the rod ends being held in contact with a thermal reservoir at temperature zero. Our goal is an explicit solution formula that exposes both how the temperature evolves and how that evolution reflects the spectral structure of the spatial differential operator \(\partial_{xx}\) on \([0, L]\).

The strategy proceeds in three stages. Separation of variables reduces the PDE to two ordinary differential equations linked by a separation constant. The spatial ODE together with the boundary conditions becomes an eigenvalue problem whose solutions are the standing waves of the rod. Superposition of these standing waves, weighted by the Fourier sine coefficients of \(f\), gives the complete solution. The idea of superposing standing modes goes back to Daniel Bernoulli, and Fourier made it systematic. The method reaches far beyond the heat equation and underpins much of the linear theory of partial differential equations on bounded domains.

Separation of Variables

We seek solutions in the product form \[ u(x, t) = \phi(x)\, G(t), \] setting aside the initial condition for the moment and demanding only that \(u\) satisfy the PDE and the homogeneous boundary conditions. Substituting into the heat equation and computing the relevant partial derivatives gives \(\phi(x) G'(t) = k\, \phi''(x) G(t)\). Dividing both sides by \(k\, \phi(x)\, G(t)\), which is legitimate wherever neither factor vanishes, separates the variables: \[ \frac{G'(t)}{k\, G(t)} = \frac{\phi''(x)}{\phi(x)}. \] The left-hand side depends only on \(t\), the right-hand side only on \(x\). Two functions of independent variables can agree everywhere only if both are equal to a common constant, which we denote \(-\lambda\): \[ \frac{G'(t)}{k\, G(t)} = \frac{\phi''(x)}{\phi(x)} = -\lambda. \]

The minus sign is a convention chosen with foresight. We shall find that the physically relevant values of \(\lambda\) are positive, in which case \(G(t)\) decays rather than grows. The separation equation splits into two ordinary differential equations, \[ \begin{align*} \phi''(x) + \lambda\, \phi(x) &= 0, \\\\ G'(t) + \lambda k\, G(t) &= 0, \end{align*} \] coupled only through the shared constant \(\lambda\).

The homogeneous boundary conditions \(u(0, t) = u(L, t) = 0\) translate into conditions on \(\phi\) alone. Requiring \(\phi(0)\, G(t) = \phi(L)\, G(t) = 0\) for all \(t\) and rejecting the trivial possibility \(G \equiv 0\), which gives \(u \equiv 0\) and is of no interest, we obtain \[ \phi(0) = \phi(L) = 0. \]

The Eigenvalue Problem

The spatial problem \[ \begin{align*} \phi''(x) + \lambda\, \phi(x) &= 0, \\\\ \phi(0) = \phi(L) &= 0, \end{align*} \] is a boundary value problem. Two conditions are imposed at two distinct points of the interval, not at a single point as for an initial value problem. It is fundamentally different in character from the initial value problems familiar from a first course in ordinary differential equations. The trivial function \(\phi \equiv 0\) satisfies the ODE and both boundary conditions regardless of \(\lambda\). The question is whether nontrivial solutions exist, and if so for which values of \(\lambda\).

Definition: Eigenvalues and Eigenfunctions of the Dirichlet Laplacian on \([0, L]\)

A number \(\lambda \in \mathbb{C}\) is an eigenvalue of the boundary value problem \[ \begin{align*} \phi''(x) + \lambda\, \phi(x) &= 0, \\\\ \phi(0) = \phi(L) &= 0 \end{align*} \] if there exists a twice-differentiable solution \(\phi : [0, L] \to \mathbb{C}\) with \(\phi \not\equiv 0\). Any such \(\phi\) is called an eigenfunction corresponding to \(\lambda\). The terminology arises because the problem can be rewritten as \(-\phi'' = \lambda\, \phi\), expressing \(\phi\) as an eigenvector of the differential operator \(-d^2/dx^2\) acting on twice-differentiable functions on \([0, L]\) that vanish at both endpoints.

Every eigenvalue is real and positive. Let \(\phi\) be an eigenfunction for \(\lambda\). Since \(\phi'' = -\lambda\, \phi\) is continuous, we may multiply \(-\phi'' = \lambda\, \phi\) by the complex conjugate \(\overline{\phi}\) and integrate by parts over \([0, L]\). The boundary term \([\phi'\, \overline{\phi}]_0^L\) vanishes because \(\phi(0) = \phi(L) = 0\), which gives the energy identity \[ \begin{align*} \lambda \int_0^L |\phi(x)|^2\, dx &= -\int_0^L \phi''(x)\, \overline{\phi(x)}\, dx \\\\ &= \int_0^L |\phi'(x)|^2\, dx. \end{align*} \] The integral on the left is positive, because \(\phi\) is continuous and not identically zero. Hence \(\lambda\) is real and \(\lambda \geq 0\). If \(\lambda = 0\), the identity forces \(\phi' \equiv 0\), so \(\phi\) is constant, and the boundary conditions make it zero. Therefore \(\lambda \gt 0\).

To find the eigenvalues themselves, we solve the ODE directly. The characteristic equation \(r^2 + \lambda = 0\) has roots \(r = \pm\sqrt{-\lambda}\), and for real \(\lambda\) the sign decides whether these roots are real, zero, or purely imaginary. The identity above has already excluded \(\lambda \leq 0\). The direct computation confirms this independently, and it is what identifies the positive values that occur. We therefore consider three cases. Throughout, the constants \(c_1\) and \(c_2\) are complex numbers.

Case 1: \(\lambda \lt 0\).
Writing \(\lambda = -s\) with \(s \gt 0\), the characteristic roots are \(r = \pm\sqrt{s}\), and the general solution is \[ \phi(x) = c_1 \cosh\sqrt{s}\, x + c_2 \sinh\sqrt{s}\, x. \] The condition \(\phi(0) = 0\) forces \(c_1 = 0\), leaving \(\phi(x) = c_2 \sinh\sqrt{s}\, x\). The condition \(\phi(L) = 0\) requires \(c_2 \sinh\sqrt{s}\, L = 0\). Since \(\sinh\) vanishes only at zero and \(\sqrt{s}\, L \gt 0\), we must have \(c_2 = 0\), giving only the trivial solution. No negative eigenvalue exists.

Case 2: \(\lambda = 0\).
The ODE reduces to \(\phi'' = 0\) with general solution \[ \phi(x) = c_1 + c_2 x. \] The boundary conditions \(\phi(0) = c_1 = 0\) and \(\phi(L) = c_2 L = 0\) (with \(L \gt 0\)) force both constants to vanish. \(\lambda = 0\) is not an eigenvalue.

Case 3: \(\lambda \gt 0\).
The characteristic roots are purely imaginary, \(r = \pm i\sqrt{\lambda}\), and the general solution in trigonometric form is \[ \phi(x) = c_1 \cos\sqrt{\lambda}\, x + c_2 \sin\sqrt{\lambda}\, x. \] The condition \(\phi(0) = 0\) again forces \(c_1 = 0\). The remaining condition \(\phi(L) = c_2 \sin\sqrt{\lambda}\, L = 0\) admits nontrivial solutions \((c_2 \neq 0)\) precisely when \(\sin\sqrt{\lambda}\, L = 0\), that is, when \(\sqrt{\lambda}\, L = n\pi\) for some positive integer \(n\). (The value \(n = 0\) is excluded because it gives \(\lambda = 0\), already handled. A negative integer \(n = -m\) with \(m \geq 1\) yields the same eigenvalue \(\lambda_m\) and, since \(\sin(-m\pi x/L) = -\sin(m\pi x/L)\), the same eigenfunction up to sign.)

Theorem: Eigenvalues and Eigenfunctions of the Dirichlet Laplacian on \([0, L]\)

The eigenvalues and corresponding eigenfunctions of the boundary value problem \[ \begin{align*} \phi''(x) + \lambda\, \phi(x) &= 0, \\\\ \phi(0) = \phi(L) &= 0 \end{align*} \] are \[ \begin{align*} \lambda_n &= \left(\frac{n\pi}{L}\right)^{\!2}, \\\\ \phi_n(x) &= \sin\frac{n\pi x}{L}, \end{align*} \] for \(n = 1, 2, 3, \ldots\). The eigenvalues form a strictly increasing sequence \(0 \lt \lambda_1 \lt \lambda_2 \lt \cdots\) tending to infinity. Each eigenvalue is simple. Its eigenspace is one-dimensional, so every eigenfunction for \(\lambda_n\) is a complex multiple of \(\phi_n\).

Proof:

The case analysis above establishes that nontrivial solutions exist only when \(\lambda \gt 0\) and \(\sqrt{\lambda}\, L = n\pi\) for some positive integer \(n\), that is, when \(\lambda = \lambda_n = (n\pi/L)^2\). The corresponding eigenfunctions \(\phi_n(x) = \sin(n\pi x/L)\) span a one-dimensional space. Any other solution of the ODE with eigenvalue \(\lambda_n\) is a complex linear combination of \(\sin\sqrt{\lambda_n}\, x\) and \(\cos\sqrt{\lambda_n}\, x\), and the boundary condition at \(x = 0\) eliminates the cosine component while \(\phi(L) = 0\) is automatically satisfied by \(\sin\sqrt{\lambda_n}\, x\).

The sequence is strictly increasing because \(\lambda_{n+1} - \lambda_n = (2n + 1)\pi^2/L^2 \gt 0\). No other eigenvalues exist, because the energy identity preceding the case analysis shows that all eigenvalues are real, and the three cases cover every real value.

Insight: Spectral Geometry of an Interval

The eigenfunctions \(\phi_n(x) = \sin(n\pi x/L)\) are the standing waves of a rod fixed at both ends. They are the same modes that govern a vibrating string of length \(L\) with fixed endpoints. The \(n\)th mode has \(n - 1\) internal nodes (zeros) and oscillates \(n/2\) times along the rod. The eigenvalue \(\lambda_n = (n\pi/L)^2\) measures the "spatial frequency squared" of the mode. Higher modes oscillate more rapidly and, as we shall see, decay more rapidly in time.

The quadratic growth \(\lambda_n = (\pi/L)^2 n^2\) is no accident. It reflects the fact that the Laplacian on a one-dimensional interval is a second-order operator. Weyl's law in higher dimensions generalizes this scaling. On a bounded \(d\)-dimensional domain of volume \(V\), the \(n\)th Dirichlet eigenvalue satisfies \(\lambda_n \sim C_d\, (n/V)^{2/d}\) as \(n \to \infty\), where \(C_d = 4\pi^2 \omega_d^{-2/d}\) and \(\omega_d\) is the volume of the unit ball in \(\mathbb{R}^d\). Here \(\sim\) means that the ratio of the two sides tends to \(1\). For \(d = 1\), where \(V = L\) and \(\omega_1 = 2\), this gives \(C_1 = \pi^2\), in agreement with the exact formula above. We state Weyl's law without proof. The spectral structure of a domain encodes deep geometric information. This theme connects to the spectral analysis of graph Laplacians in spectral graph theory and to the manifold theory that lies ahead in subsequent pages.

Time Evolution and Product Solutions

With the eigenvalues in hand, the time-dependent ODE \(G'(t) + \lambda_n k\, G(t) = 0\) associated with the \(n\)th mode is a first-order linear equation with constant coefficient \(\lambda_n k\), whose general solution is the exponential \[ \begin{align*} G_n(t) &= e^{-\lambda_n k\, t} \\\\ &= \exp\!\left(-k\!\left(\tfrac{n\pi}{L}\right)^{\!2} t\right), \end{align*} \] chosen with the normalization \(G_n(0) = 1\) so that the multiplicative constant is absorbed into the spatial part at the time of superposition. Multiplying the spatial eigenfunction by its time evolution gives the \(n\)th product solution or normal mode: \[ u_n(x, t) = \sin\frac{n\pi x}{L}\, \exp\!\left(-k\!\left(\tfrac{n\pi}{L}\right)^{\!2} t\right). \]

Each \(u_n\) is a solution of the PDE \(\partial_t u = k\, \partial_{xx} u\) satisfying the homogeneous Dirichlet boundary conditions. Three features stand out. First, every mode decays exponentially in time. This is the hallmark of parabolic behavior, and it distinguishes the heat equation sharply from the wave equation, whose modes oscillate without decay. Second, higher modes decay faster. The decay rate \(k\lambda_n = k(n\pi/L)^2\) grows quadratically in \(n\), so spatial features of high frequency are damped on a much shorter timescale than coarse features. The thermal diffusion time of the \(n\)th mode is \(\tau_n = 1/(k\lambda_n) = L^2/(k n^2 \pi^2)\). Third, the spatial profile is preserved. Each \(u_n\) factors as a pure sine wave times a time-dependent amplitude, so the shape \(\sin(n\pi x/L)\) is fixed and only its amplitude shrinks.

Superposition and the Series Solution

The heat equation is linear and homogeneous, and the boundary conditions are also homogeneous. Consequently any finite linear combination of product solutions is again a solution: \[ \sum_{n=1}^{N} B_n\, u_n(x, t) = \sum_{n=1}^{N} B_n \sin\frac{n\pi x}{L}\, \exp\!\left(-k\!\left(\tfrac{n\pi}{L}\right)^{\!2} t\right) \] satisfies the PDE and the Dirichlet boundary conditions for any choice of coefficients \(B_1, \ldots, B_N\). The principle of superposition extends this to infinite series, provided the series converges in an appropriate sense. Termwise differentiation in both \(x\) and \(t\) is justified for \(t \gt 0\) whenever the coefficients \(B_n\) are bounded, because the exponential factors \(\exp(-k(n\pi/L)^2 t)\) decay so rapidly in \(n\) that the differentiated series converges uniformly on every closed half-strip \(0 \leq x \leq L\), \(t \geq t_0 \gt 0\). Here we use a standard fact of real analysis without proof. A convergent series of differentiable functions may be differentiated term by term on any interval where the series of derivatives converges uniformly.

To match the initial condition \(u(x, 0) = f(x)\), evaluate the formal infinite series at \(t = 0\). The exponential factors collapse to unity, leaving \[ f(x) = \sum_{n=1}^{\infty} B_n \sin\frac{n\pi x}{L}, \quad 0 \leq x \leq L. \] This is precisely a Fourier sine series expansion of \(f\) on \([0, L]\). The orthogonality of the trigonometric system on \([-L, L]\), restricted to the half-interval \([0, L]\) by the evenness of the product \(\sin(n\pi x/L)\sin(m\pi x/L)\), gives the coefficient formula immediately. Multiplying both sides by \(\sin(m\pi x/L)\) and integrating over \([0, L]\) selects the term with \(n = m\), using \[ \int_0^L \sin\frac{n\pi x}{L} \sin\frac{m\pi x}{L}\, dx = \begin{cases} L/2, & m = n, \\\\ 0, & m \neq n. \end{cases} \] The resulting coefficients are the Fourier sine coefficients of \(f\). We collect these observations into a single statement.

Theorem: Series Solution of the Heat Equation on a Bounded Interval

Let \(f \in L^2([0, L])\) be a real-valued initial temperature profile. We call \(u\) a solution of the initial-boundary value problem \[ \begin{cases} \partial_t u = k\, \partial_{xx} u, & 0 \lt x \lt L,\ t \gt 0, \\\\ u(0, t) = u(L, t) = 0, & t \gt 0, \\\\ u(\cdot, t) \to f \text{ in } L^2([0, L]), & t \to 0^+ \end{cases} \] if \(u\) is real-valued, satisfies the three conditions above, and has the following regularity. The functions \(u\), \(\partial_x u\), \(\partial_{xx} u\), and \(\partial_t u\) exist and are continuous on \([0, L] \times (0, \infty)\), with one-sided \(x\)-derivatives at \(x = 0\) and \(x = L\). This problem has exactly one solution, namely \[ u(x, t) = \sum_{n=1}^{\infty} B_n \sin\frac{n\pi x}{L}\, \exp\!\left(-k\!\left(\tfrac{n\pi}{L}\right)^{\!2} t\right), \] where the coefficients are the Fourier sine coefficients of the initial datum, \[ B_n = \frac{2}{L}\int_0^L f(x) \sin\frac{n\pi x}{L}\, dx. \] The series converges in \(L^2([0, L])\) for every \(t \geq 0\) and, for \(t \gt 0\), converges absolutely and uniformly on \([0, L]\) together with all of its \(x\)-derivatives and \(t\)-derivatives.

Proof:

Step 1: the series is a solution.
Termwise the series satisfies the PDE and the Dirichlet boundary conditions, as each product solution \(u_n\) does. The convergence claims are a consequence of the rapid decay of the exponential factors. The bound \(|B_n| \leq \|f\|_{L^2}\sqrt{2/L}\) follows from Bessel's inequality for the orthonormal system \(\{\sqrt{2/L}\sin(n\pi x/L)\}_{n \geq 1}\) in \(L^2([0, L])\). Fix \(t_0 \gt 0\). Every termwise derivative of the series is bounded on \([0, L] \times [t_0, \infty)\) by a constant multiple of \(n^j \exp(-k(n\pi/L)^2 t_0)\) for some integer \(j \geq 0\), and these bounds decay faster than any inverse power of \(n\). The series and all of its termwise derivatives therefore converge absolutely and uniformly on \([0, L] \times [t_0, \infty)\). By the termwise-differentiation fact recalled above, \(u\) is \(C^\infty\) there, up to \(x = 0\) and \(x = L\), and the termwise identities pass to the sum. Since \(t_0 \gt 0\) is arbitrary, \(u\) has the regularity required of a solution. It is real-valued because \(f\) is.

Step 2: the initial datum.
The system \(\{\sqrt{2/L}\sin(n\pi x/L)\}_{n \geq 1}\) is an orthonormal basis of \(L^2([0, L])\). To see this, let \(g \in L^2([0, L])\) be orthogonal to every \(\sin(n\pi x/L)\), and let \(G\) be its odd extension to \([-L, L]\). For every integer \(n \geq 0\), the product of \(G\) with \(\cos(n\pi x/L)\) is odd, so its integral over \([-L, L]\) vanishes. For \(n \geq 1\), the product with \(\sin(n\pi x/L)\) is even, so its integral is twice the integral of \(g \sin(n\pi x/L)\) over \([0, L]\), which is zero. Hence \(G\) is orthogonal to every \(e^{in\pi x/L}\) with \(n \in \mathbb{Z}\), and the completeness of the full trigonometric system on \(L^2([-L, L])\) gives \(G = 0\), so \(g = 0\). An orthonormal system to which only the zero vector is orthogonal is an orthonormal basis, by the characterization recalled in the completeness argument for Fourier series. In particular, at \(t = 0\) the series is the Fourier sine series of \(f\), which converges to \(f\) in \(L^2([0, L])\).

For \(t \gt 0\) the series for \(u(\cdot, t)\) converges uniformly on \([0, L]\), hence in \(L^2([0, L])\), and the inner product is continuous, so its coefficients in this basis may be computed term by term. The coefficients of \(u(\cdot, t) - f\) are therefore \(\sqrt{L/2}\, B_n (e^{-k\lambda_n t} - 1)\), with \(B_n\) real because \(f\) is real-valued. The \(L^2\)-continuity \(u(\cdot, t) \to f\) as \(t \to 0^+\) then follows from Parseval's identity: \[ \|u(\cdot, t) - f\|_{L^2}^2 = \frac{L}{2}\sum_{n=1}^{\infty} B_n^2 \left(e^{-k\lambda_n t} - 1\right)^2, \] which tends to \(0\) as \(t \to 0^+\) by dominated convergence, applied to the counting measure on the positive integers along an arbitrary sequence \(t_j \to 0^+\). Each term vanishes pointwise and is bounded uniformly in \(t \geq 0\) by \(B_n^2\), since \(0 \leq e^{-k\lambda_n t} \leq 1\), while \((L/2)\sum_n B_n^2 = \|f\|_{L^2}^2 \lt \infty\).

Step 3: uniqueness.
Let \(u_1\) and \(u_2\) be two solutions, and set \(w = u_1 - u_2\). Then \(w\) has the same regularity, satisfies \(\partial_t w = k\, \partial_{xx} w\) for \(0 \lt x \lt L\) and \(t \gt 0\), and vanishes at \(x = 0\) and \(x = L\) for \(t \gt 0\). Moreover, \[ \|w(\cdot, t)\|_{L^2} \leq \|u_1(\cdot, t) - f\|_{L^2} + \|u_2(\cdot, t) - f\|_{L^2} \] for every \(t \gt 0\), so \(\|w(\cdot, t)\|_{L^2} \to 0\) as \(t \to 0^+\). Define \(E(t) = \int_0^L w(x, t)^2\, dx\) for \(t \gt 0\). On each compact rectangle \([0, L] \times [t_0, t_1]\) with \(t_0 \gt 0\), the functions \(w\) and \(\partial_t w\) are continuous and hence bounded. By the mean value theorem, the difference quotients of \(w^2\) in \(t\) are bounded there by \(2 \sup|w| \sup|\partial_t w|\). Dominated convergence, which applies along every sequence of increments with limit zero, therefore justifies differentiation under the integral sign. Using \(\partial_t w = k\, \partial_{xx} w\), \[ \begin{align*} \frac{dE}{dt} &= 2\!\int_0^L w\, \partial_t w\, dx \\\\ &= 2k\!\int_0^L w\, \partial_{xx} w\, dx \\\\ &= -2k\!\int_0^L (\partial_x w)^2\, dx \\\\ &\leq 0. \end{align*} \] The integration by parts in the third equality is valid because \(w(\cdot, t)\) is twice continuously differentiable on \([0, L]\), and the boundary term \(2k\,[w\, \partial_x w]_0^L\) drops out by the Dirichlet condition. Since \(dE/dt \leq 0\), \(E\) is non-increasing on \((0, \infty)\). For \(0 \lt s \lt t\) this gives \(E(t) \leq E(s) = \|w(\cdot, s)\|_{L^2}^2\), and letting \(s \to 0^+\) yields \(E(t) = 0\). Since \(w(\cdot, t)\) is continuous, it vanishes identically, and \(u_1 = u_2\).

The argument in Step 3 is the energy method. The class of solutions in the theorem contains every function that has the stated regularity for \(t \gt 0\), satisfies the heat equation with the Dirichlet condition, and extends continuously to \([0, L] \times [0, \infty)\) with \(u(\cdot, 0) = f\). Continuity on the compact rectangle \([0, L] \times [0, 1]\) makes \(u(\cdot, t) \to f\) uniformly, hence in \(L^2\), so the uniqueness statement applies to them.

Attaining the initial datum continuously asks more of \(f\) than square-integrability. If a solution \(u\), extended by \(u(\cdot, 0) = f\), is continuous on \([0, L] \times [0, \infty)\), then \(f\) is continuous and \(f(0) = \lim_{t \to 0^+} u(0, t) = 0\), and likewise \(f(L) = 0\). These endpoint conditions are therefore necessary, and together with a mild regularity assumption they are also sufficient. We call \(f\) piecewise \(C^1\) on \([0, L]\) if there are points \(0 = s_0 \lt s_1 \lt \cdots \lt s_m = L\) such that the restriction of \(f\) to each \([s_{j-1}, s_j]\) is continuously differentiable, with one-sided derivatives at the endpoints.

Lemma: Absolute Summability of Sine Coefficients

Let \(f\) be continuous and piecewise \(C^1\) on \([0, L]\) with \(f(0) = f(L) = 0\), and let \[ B_n = \frac{2}{L}\int_0^L f(x) \sin\frac{n\pi x}{L}\, dx. \] Then \(\sum_{n=1}^{\infty} |B_n| \lt \infty\), and the Fourier sine series \(\sum_{n} B_n \sin(n\pi x/L)\) converges to \(f(x)\) absolutely and uniformly on \([0, L]\).

Proof:

Write \(I_j\) for the integral of \(f(x) \sin(n\pi x/L)\) over \([s_{j-1}, s_j]\). Integration by parts gives \[ \begin{align*} I_j &= \Bigl[-\frac{L}{n\pi}\, f(x) \cos\frac{n\pi x}{L}\Bigr]_{s_{j-1}}^{s_j} \\\\ &\quad + \frac{L}{n\pi}\int_{s_{j-1}}^{s_j} f'(x) \cos\frac{n\pi x}{L}\, dx. \end{align*} \] Summing over \(j\), the boundary terms telescope because \(f\) is continuous, and the two surviving terms vanish because \(f(0) = f(L) = 0\). Hence \[ \begin{align*} B_n &= \frac{L}{n\pi}\, a_n, \\\\ a_n &= \frac{2}{L}\int_0^L f'(x) \cos\frac{n\pi x}{L}\, dx, \end{align*} \] where \(f'\) is bounded and continuous except at the finitely many points \(s_j\), so it is square-integrable. The orthogonality relations, restricted to \([0, L]\) because \(\cos(n\pi x/L)\cos(m\pi x/L)\) is an even function, show that the functions \(\sqrt{2/L}\cos(n\pi x/L)\) with \(n \geq 1\) form an orthonormal system in \(L^2([0, L])\). The inner product of \(f'\) with the \(n\)th of them is \(\sqrt{L/2}\, a_n\), so Bessel's inequality gives \(\frac{L}{2}\sum_n |a_n|^2 \leq \|f'\|_{L^2}^2 \lt \infty\). The Cauchy-Schwarz inequality for sums then yields \[ \begin{align*} \sum_{n=1}^{\infty} |B_n| &= \frac{L}{\pi} \sum_{n=1}^{\infty} \frac{|a_n|}{n} \\\\ &\leq \frac{L}{\pi} \Bigl(\sum_{n=1}^{\infty} \frac{1}{n^2}\Bigr)^{1/2} \Bigl(\sum_{n=1}^{\infty} |a_n|^2\Bigr)^{1/2} \\\\ &\lt \infty. \end{align*} \]

Since \(|B_n \sin(n\pi x/L)| \leq |B_n|\), the Weierstrass M-test shows that the sine series converges uniformly on \([0, L]\), and absolutely at every point. Its sum \(S\) is continuous, as a uniform limit of continuous functions. The M-test and the continuity of uniform limits are standard facts that we take for granted. Uniform convergence on the bounded interval implies convergence in \(L^2([0, L])\), so the partial sums tend to \(S\) in \(L^2\). They also tend to \(f\) in \(L^2\), because the functions \(\sqrt{2/L}\sin(n\pi x/L)\) form an orthonormal basis of \(L^2([0, L])\), as Step 2 of the proof of the series solution theorem shows. Limits in \(L^2\) are unique, so \(S = f\) almost everywhere. Since \(f\) is continuous by hypothesis, \(S - f\) is continuous, and a continuous function that vanishes almost everywhere vanishes everywhere, because otherwise it would be nonzero on an interval of positive length. Hence \(S = f\) on \([0, L]\).

Theorem: Continuous Attainment of the Initial Datum

Let \(f\) be real-valued, continuous, and piecewise \(C^1\) on \([0, L]\) with \(f(0) = f(L) = 0\), and let \(u\) be the solution given by the series solution theorem. Then the series defining \(u\) converges absolutely and uniformly on \([0, L] \times [0, \infty)\). With \(u(x, 0) = f(x)\), the function \(u\) is continuous on \([0, L] \times [0, \infty)\), so \(u(x, t) \to f(x_0)\) as \((x, t) \to (x_0, 0)\) for every \(x_0 \in [0, L]\).

Proof:

On \([0, L] \times [0, \infty)\), each term has absolute value at most \(|B_n|\), and \(\sum_n |B_n| \lt \infty\) by the summability lemma. The M-test gives absolute and uniform convergence there, so the sum is continuous. At \(t = 0\) it reduces to the sine series of \(f\), whose sum is \(f\) by the same lemma.

The worked example below satisfies these hypotheses. For a merely continuous \(f\) with \(f(0) = f(L) = 0\), the series defining \(u\) need not converge pointwise at \(t = 0\), but the function \(u\), with \(u(\cdot, 0) = f\), is still continuous on \([0, L] \times [0, \infty)\). The proof of that statement needs an approximation argument that we do not give here.

The series solution formula has a clean structure. Each Fourier mode of the initial datum evolves independently, weighted by an exponential decay whose rate is the eigenvalue times the diffusivity. The Laplacian \(\partial_{xx}\) on \([0, L]\) with Dirichlet boundary conditions is diagonalized by the orthonormal basis \(\{\sqrt{2/L}\sin(n\pi x/L)\}\), and the heat semigroup \(e^{tk\partial_{xx}}\) is defined on this basis as multiplication by \(e^{-tk\lambda_n}\). This is diagonalization by a unitary change of basis, the viewpoint in which the Hilbert-space treatment of Fourier analysis culminates. Functions of an unbounded self-adjoint operator, such as the Dirichlet Laplacian, are treated in general by the functional calculus of such operators, which lies beyond the scope of this page. The same picture recurs for elliptic operators on manifolds and for graph Laplacians in spectral graph theory.

A Worked Example

Example: Parabolic Initial Profile

Take \(L = 1\) and the initial datum \(f(x) = x(1 - x)\), a parabolic bump vanishing at both endpoints and reaching its maximum \(1/4\) at the midpoint \(x = 1/2\). The Fourier sine coefficients are computed by direct integration: \[ B_n = 2\int_0^1 x(1 - x) \sin(n\pi x)\, dx. \] Two integrations by parts, or the standard tabulated identities, give \[ \begin{align*} \int_0^1 x \sin(n\pi x)\, dx &= \frac{(-1)^{n+1}}{n\pi}, \\\\ \int_0^1 x^2 \sin(n\pi x)\, dx &= \frac{(-1)^{n+1}}{n\pi} - \frac{2[1 - (-1)^n]}{(n\pi)^3}. \end{align*} \] Subtracting yields \[ \begin{align*} B_n &= 2\!\left(\frac{(-1)^{n+1}}{n\pi} - \left[\frac{(-1)^{n+1}}{n\pi} - \frac{2[1-(-1)^n]}{(n\pi)^3}\right]\right) \\\\ &= \frac{4[1 - (-1)^n]}{(n\pi)^3}. \end{align*} \] Since \(1 - (-1)^n\) equals \(2\) when \(n\) is odd and \(0\) when \(n\) is even, only odd modes contribute, and writing \(n = 2j + 1\), \[ \begin{align*} B_{2j+1} &= \frac{8}{(2j+1)^3 \pi^3}, \\\\ B_{2j} &= 0. \end{align*} \] The solution is therefore \[ u(x, t) = \frac{8}{\pi^3} \sum_{j=0}^{\infty} \frac{1}{(2j+1)^3}\, \sin((2j+1)\pi x)\, \exp\!\left(-k (2j+1)^2 \pi^2 t\right). \]

Three qualitative features illustrate the general structure. The vanishing of even-mode coefficients reflects the spatial symmetry of \(f\) about \(x = 1/2\). The initial profile is even with respect to that midpoint, and the even-indexed eigenfunctions \(\sin(2j\pi x)\) are odd about it, hence orthogonal. The coefficients decay as \(1/(2j+1)^3\), an algebraic rate set by the smoothness of \(f\) and its compatibility with the boundary conditions. Finally, the time exponential \(\exp(-k(2j+1)^2 \pi^2 t)\) damps high modes far faster than the fundamental \((j = 0)\) mode \(\sin(\pi x) e^{-k\pi^2 t}\). The shape relaxes toward a sine arch on the timescale \(1/(8k\pi^2)\), since the ratio of the \(j = 1\) term to the fundamental carries the factor \(e^{-8k\pi^2 t}\). After a few multiples of that time the solution is well approximated by its first term, a sine arch whose amplitude decays to the zero equilibrium on the longer timescale \(1/(k\pi^2)\). Because \(f\) is continuously differentiable with \(f(0) = f(1) = 0\), the continuous attainment theorem applies, and \(u(x, t) \to f(x_0)\) as \((x, t) \to (x_0, 0)\) for every \(x_0 \in [0, 1]\).

Insight: From Bounded Interval to Real Line

The series formula above is the complete solution of the heat equation on \([0, L]\) with Dirichlet boundary conditions. Every initial profile decomposes into Dirichlet eigenmodes, each evolving independently by exponential damping with rate \(k(n\pi/L)^2\). The structure is rigid: a countable basis, a discrete spectrum, a sum.

The next section removes the bounded interval and asks what happens on the entire real line, where the boundary is sent to infinity. The Dirichlet eigenmodes \(\sin(n\pi x/L)\) lose their normalization as \(L \to \infty\), and the discrete spectrum thickens into a continuum. The replacement of the Fourier series by the Fourier transform is precisely what is needed. On \(\mathbb{R}\), the heat equation will be solved by a single integral kernel.

Heat Equation on the Real Line

Replacing the rod \([0, L]\) by the full line \(\mathbb{R}\) changes the analytic situation in a structurally important way. On a bounded interval, the normalized Dirichlet eigenfunctions \(\sqrt{2/L}\sin(n\pi x/L)\) provide a countable orthonormal basis and the solution is a discrete sum over normal modes. On the real line, no such basis of eigenfunctions of \(-\partial_{xx}\) exists in \(L^2(\mathbb{R})\). The formal eigenfunctions \(e^{i\xi x}\) are bounded but not square-integrable, and the "expansion in eigenmodes" must be replaced by an integral over a continuous frequency variable.

That replacement is precisely the passage from Fourier series to the Fourier transform, carried out on the page devoted to it. Here we put that machinery to its first serious use, solving an initial-value problem on an unbounded domain.

The problem we now consider is: \[ \begin{cases} \partial_t u = k\, \partial_{xx} u, & x \in \mathbb{R},\ t \gt 0, \\\\ u(x, 0) = f(x), & x \in \mathbb{R}, \end{cases} \] with \(k \gt 0\) and \(f \in L^1(\mathbb{R})\). We additionally require three things of \(u\). First, the partial derivatives \(\partial_x u\), \(\partial_{xx} u\), and \(\partial_t u\) exist on \(\mathbb{R} \times (0, \infty)\), and for each fixed \(t \gt 0\) the four functions \(u(\cdot, t)\), \(\partial_x u(\cdot, t)\), \(\partial_{xx} u(\cdot, t)\), and \(\partial_t u(\cdot, t)\) are continuous. Second, these four functions tend to \(0\) as \(|x| \to \infty\) for each \(t \gt 0\). Third, for every compact \(J \subset (0, \infty)\) there is an integrable function \(g_J\) on \(\mathbb{R}\) with \(|u|, |\partial_x u|, |\partial_{xx} u|, |\partial_t u| \leq g_J\) on \(\mathbb{R} \times J\). We call such a \(u\) a decaying solution. This restriction picks out the natural function class for the problem and rules out, for instance, the addition of a constant temperature or a polynomially growing background field. The initial condition is imposed as a limit, \(u(\cdot, t) \to f\) as \(t \to 0^+\), in a sense made precise in the main theorem, and the uniqueness of the solution within this class is treated alongside that theorem.

Fourier-Transforming the Equation

Define the spatial Fourier transform of \(u(\cdot, t)\) for each fixed \(t\): \[ \hat{u}(\xi, t) = \int_{-\infty}^{\infty} u(x, t)\, e^{ix\xi}\, dx. \] Under the transform, a double integration by parts with respect to \(x\), the argument that proves the differentiation property, turns \(\partial_{xx}\) into multiplication by \((-i\xi)^2 = -\xi^2\). Applying the transform to both sides of the heat equation therefore converts the partial differential equation in \((x, t)\) into an ordinary differential equation in \(t\) alone, with \(\xi\) entering only as a parameter:

Theorem: Fourier-Transformed Heat Equation

Let \(u\) be a decaying solution of the heat equation on \(\mathbb{R} \times (0, \infty)\) with \(u(\cdot, t) \to f\) in \(L^1(\mathbb{R})\) as \(t \to 0^+\), and let \(\hat{u}(\xi, t)\) denote its spatial Fourier transform. Then, for each \(\xi \in \mathbb{R}\), the function \(t \mapsto \hat{u}(\xi, t)\) satisfies, for \(t \gt 0\), \[ \begin{align*} \partial_t \hat{u}(\xi, t) &= -k \xi^2\, \hat{u}(\xi, t), \\\\ \lim_{t \to 0^+} \hat{u}(\xi, t) &= \hat{f}(\xi). \end{align*} \] The unique such function is \[ \hat{u}(\xi, t) = \hat{f}(\xi)\, e^{-k \xi^2 t}. \]

Proof:

Differentiation under the integral sign in \(t\) is justified by the dominated convergence theorem, applied to difference quotients, since \(|\partial_t u(x, t)| \leq g_J(x)\) for \(t\) in any compact \(J \subset (0, \infty)\) and \(g_J\) is integrable. Then \[ \begin{align*} \partial_t \hat{u}(\xi, t) &= \int_{-\infty}^{\infty} \partial_t u(x, t)\, e^{ix\xi}\, dx \\\\ &= \int_{-\infty}^{\infty} k\, \partial_{xx} u(x, t)\, e^{ix\xi}\, dx \\\\ &= k \cdot (-i\xi)^2\, \hat{u}(\xi, t) \\\\ &= -k\xi^2\, \hat{u}(\xi, t), \end{align*} \] where the third equality is the computation behind the differentiation property of the Fourier transform. That property is stated for Schwartz functions, which \(u(\cdot, t)\) need not be, so we carry out the computation directly. Since \(u(\cdot, t)\) is twice continuously differentiable, two integrations by parts on \([-R, R]\) give \[ \begin{align*} \int_{-R}^{R} \partial_{xx} u\, e^{ix\xi}\, dx &= \Bigl[\partial_x u\, e^{ix\xi}\Bigr]_{-R}^{R} \\\\ &\quad - i\xi \Bigl[u\, e^{ix\xi}\Bigr]_{-R}^{R} \\\\ &\quad + (i\xi)^2 \int_{-R}^{R} u\, e^{ix\xi}\, dx. \end{align*} \] As \(R \to \infty\), the boundary terms vanish by the assumed decay of \(u\) and \(\partial_x u\) at infinity, and the two integrals converge to integrals over \(\mathbb{R}\) because \(u(\cdot, t)\) and \(\partial_{xx} u(\cdot, t)\) are integrable. Since \((i\xi)^2 = (-i\xi)^2 = -\xi^2\), this is the third equality.

For each fixed \(\xi\), the equation \(\partial_t \hat{u}(\xi, t) = -k\xi^2\, \hat{u}(\xi, t)\) is a scalar linear ODE in \(t\) on \((0, \infty)\) with constant coefficient \(-k\xi^2\) (\(\xi\) is held fixed). Hence \(\hat{u}(\xi, t)\, e^{k\xi^2 t}\) has zero derivative, and \(\hat{u}(\xi, t) = C(\xi)\, e^{-k\xi^2 t}\) for some constant \(C(\xi)\). The limit at \(t = 0\) holds because \[ \begin{align*} |\hat{u}(\xi, t) - \hat{f}(\xi)| &\leq \int_{-\infty}^{\infty} |u(x, t) - f(x)|\, dx \\\\ &= \|u(\cdot, t) - f\|_{L^1} \longrightarrow 0, \end{align*} \] and letting \(t \to 0^+\) in \(\hat{u}(\xi, t) = C(\xi)\, e^{-k\xi^2 t}\) gives \(C(\xi) = \hat{f}(\xi)\).

This reduction is the central simplification. The Fourier transform diagonalizes the spatial derivative operator, turning a partial differential equation into a family of decoupled scalar ODEs indexed by frequency. Each frequency \(\xi\) decays independently at rate \(k\xi^2\), proportional to the square of the frequency. High-frequency components are therefore suppressed exponentially faster than low-frequency ones. The smoothing phenomenon is, in the frequency domain, simply multiplication by a rapidly decaying Gaussian envelope.

The Heat Kernel

To recover \(u(x, t)\) we must invert the Fourier transform of the right-hand side, that is, identify the function whose Fourier transform is \(e^{-k\xi^2 t}\). This is the moment at which the Gaussian calculation done once on the Fourier-transform page pays for itself across all of parabolic theory. There it was established (as a worked example) that \[ \widehat{e^{-x^2/(2\sigma^2)}}(\xi) = \sigma \sqrt{2\pi}\, e^{-\sigma^2 \xi^2 / 2} \] for any \(\sigma \gt 0\). Choosing \(\sigma^2 = 2kt\) makes the right-hand side equal \(\sigma\sqrt{2\pi}\, e^{-k\xi^2 t}\), which differs from the target \(e^{-k\xi^2 t}\) only by the normalization factor \(\sigma\sqrt{2\pi} = \sqrt{4\pi kt}\). Dividing through gives the function whose Fourier transform is \(e^{-k\xi^2 t}\). This function is fundamental enough to deserve a name.

Definition: Heat Kernel on \(\mathbb{R}\)

The heat kernel (or fundamental solution of the heat equation) on \(\mathbb{R}\) with diffusivity \(k \gt 0\) is \[ K_t(x) = \frac{1}{\sqrt{4\pi k t}}\, \exp\!\left(-\frac{x^2}{4kt}\right), \quad t \gt 0,\ x \in \mathbb{R}. \] Its spatial Fourier transform is \[ \widehat{K_t}(\xi) = e^{-k \xi^2 t}. \]

The Fourier identity in the definition is the substitution \(\sigma^2 = 2kt\) applied to the Gaussian Fourier transform displayed above, divided by \(\sqrt{4\pi kt}\). No fresh integration is required. In the spatial variable, \(K_t\) is a centered Gaussian bell whose width \(\sqrt{2kt}\) grows like \(\sqrt{t}\), the characteristic length scale of diffusive spreading.

Solution by Convolution

With \(\hat{u}(\xi, t) = \hat{f}(\xi)\, \widehat{K_t}(\xi)\), the convolution theorem identifies the right-hand side at once. A pointwise product of Fourier transforms is the Fourier transform of a convolution, so \(\hat{u}(\cdot, t)\) is the transform of \(f * K_t\). The theorem below confirms that this convolution is a solution, and the only one in the decaying class.

Theorem: Heat Kernel Solution on \(\mathbb{R}\)

Let \(f \in L^1(\mathbb{R})\). The convolution of \(f\) with the heat kernel, \[ \begin{align*} u(x, t) &= (f * K_t)(x) \\\\ &= \int_{-\infty}^{\infty} K_t(x - y)\, f(y)\, dy, \end{align*} \] is a decaying solution of the heat equation \(\partial_t u = k\, \partial_{xx} u\) on \(\mathbb{R} \times (0, \infty)\) with \(u(\cdot, t) \to f\) in \(L^1(\mathbb{R})\) as \(t \to 0^+\). It is the only decaying solution with this property. If in addition \(f\) is bounded and continuous, then \(u(x, t) \to f(x)\) as \(t \to 0^+\) at every \(x \in \mathbb{R}\). Explicitly, \[ u(x, t) = \frac{1}{\sqrt{4\pi kt}} \int_{-\infty}^{\infty} \exp\!\left(-\frac{(x - y)^2}{4kt}\right) f(y)\, dy. \]

Proof:

Step 1: a representation of the convolution.
The Gaussian transform displayed before the definition of the heat kernel, read with the roles of the space and frequency variables exchanged and with \(\sigma^2 = 1/(2kt)\), gives \(\int e^{iz\xi}\, e^{-k\xi^2 t}\, d\xi = \sqrt{\pi/(kt)}\, e^{-z^2/(4kt)}\). Dividing by \(2\pi\) and using that the Gaussian is even, \[ K_t(z) = \frac{1}{2\pi} \int_{-\infty}^{\infty} e^{-k\xi^2 t}\, e^{-iz\xi}\, d\xi. \] Substituting this into \(u(x, t) = \int f(y)\, K_t(x - y)\, dy\) produces a double integral whose integrand has absolute value \(|f(y)|\, e^{-k\xi^2 t}\), an integrable function on \(\mathbb{R}^2\) because \(f \in L^1(\mathbb{R})\). By Fubini's theorem, applied to the real and imaginary parts separately, the order of integration can be exchanged, and the inner integral in \(y\) is \(\hat{f}(\xi)\). Hence \[ u(x, t) = \frac{1}{2\pi}\int_{-\infty}^{\infty} \hat{f}(\xi)\, e^{-k\xi^2 t}\, e^{-ix\xi}\, d\xi. \] No inversion theorem is used here. The identity is a consequence of the Gaussian computation and Fubini's theorem alone. The integral form of \(u\) is the definition of convolution, and the symmetry \(K_t(x - y) = K_t(y - x)\) makes the two orders of the arguments interchangeable.

Step 2: the convolution is a decaying solution.
Fix \(t_0 \gt 0\) and work on \(t \geq t_0\) with increments in \(t\) that stay in \([t_0, \infty)\). For \(t \geq t_0\) and every integer \(n \geq 0\), \[ |\xi|^n\, |\hat{f}(\xi)|\, e^{-k\xi^2 t} \leq \|f\|_{L^1}\, |\xi|^n\, e^{-k\xi^2 t_0}, \] and the right-hand side is integrable in \(\xi\). By the dominated convergence theorem, applied to difference quotients along any sequence of increments tending to zero, the representation of Step 1 may be differentiated under the integral sign for \(t \gt 0\). This applies any number of times in \(x\) and once in \(t\). Each derivative in \(x\) multiplies the integrand by \(-i\xi\), and the derivative in \(t\) multiplies it by \(-k\xi^2\). The operators \(\partial_t\) and \(k\, \partial_{xx}\) therefore act on the integrand in the same way, so \(\partial_t u = k\, \partial_{xx} u\) holds at every point of \(\mathbb{R} \times (0, \infty)\).

Each of \(u\), \(\partial_x u\), \(\partial_{xx} u\), and \(\partial_t u\) at time \(t\) is, up to the factor \(1/(2\pi)\) and the reflection \(x \mapsto -x\), the Fourier transform of an integrable function of \(\xi\). By the Riemann-Lebesgue lemma, all four are continuous in \(x\) and tend to \(0\) as \(|x| \to \infty\). For the integrable bounds, let \(J = [t_0, t_1] \subset (0, \infty)\). Each \(\partial_x^n K_t(z)\) with \(n \leq 2\) is a polynomial in \(z\) and \(t^{-1/2}\) times \(e^{-z^2/(4kt)}\), so there is a constant \(C_J\) with \[ \begin{align*} |\partial_x^n K_t(z)| &\leq h_J(z) \quad \text{for all } t \in J \text{ and } n \leq 2, \\\\ h_J(z) &= C_J\, (1 + |z|)^2\, e^{-z^2/(4kt_1)}, \end{align*} \] and \(h_J\) is integrable. The convolution integral \(u = f * K_t\) may also be differentiated under the integral sign, since the difference quotients in \(x\) of \(f(y)\, \partial_x^n K_t(x - y)\) are bounded by \(|f(y)|\, \sup_{t \in J} \|\partial_x^{n+1} K_t\|_\infty\), an integrable function of \(y\). This gives \(\partial_x^n u(\cdot, t) = f * \partial_x^n K_t\), hence \(|\partial_x^n u(x, t)| \leq (|f| * h_J)(x)\) for \(t \in J\). The function \(g_J = |f| * h_J\) is integrable with \(\|g_J\|_{L^1} \leq \|f\|_{L^1} \|h_J\|_{L^1}\) by Tonelli's theorem, and \(|\partial_t u| = k\, |\partial_{xx} u| \leq k\, g_J\). So \(\max(1, k)\, g_J\) is an integrable bound for all four functions on \(\mathbb{R} \times J\), and \(u\) is a decaying solution.

Step 3: the initial datum.
That the convolution attains the initial datum in the two senses stated, in \(L^1(\mathbb{R})\) for \(f \in L^1(\mathbb{R})\) and pointwise when \(f\) is bounded and continuous, is the approximate-identity property of the Gaussian family \(K_t\) as \(t \to 0^+\). The verification of this property is omitted here.

Step 4: uniqueness.
Let \(u\) be any decaying solution with \(u(\cdot, t) \to f\) in \(L^1(\mathbb{R})\) as \(t \to 0^+\). By the Fourier-transformed heat equation, \(\hat{u}(\xi, t) = \hat{f}(\xi)\, e^{-k\xi^2 t}\) for every \(\xi\) and \(t \gt 0\), and by the convolution theorem the same function is \(\widehat{f * K_t}(\xi)\). Fix \(t \gt 0\) and set \(w = u(\cdot, t) - f * K_t\). Then \(w \in L^1(\mathbb{R})\) and \(\hat{w} \equiv 0\). Moreover \(w\) is continuous and tends to \(0\) at infinity, by the hypotheses on \(u\) and by Step 2, so \(w\) is bounded, and therefore \(w \in L^2(\mathbb{R})\) as well.

Let \(\psi\) be a Schwartz function. The Fourier inversion formula, whose proof covers the Schwartz class, gives \(\psi(x) = \frac{1}{2\pi} \int \hat{\psi}(\xi)\, e^{-ix\xi}\, d\xi\), so \[ \begin{align*} \int_{-\infty}^{\infty} w(x)\, \overline{\psi(x)}\, dx &= \frac{1}{2\pi} \int_{-\infty}^{\infty} \int_{-\infty}^{\infty} w(x)\, \overline{\hat{\psi}(\xi)}\, e^{ix\xi}\, d\xi\, dx \\\\ &= \frac{1}{2\pi} \int_{-\infty}^{\infty} \overline{\hat{\psi}(\xi)}\, \hat{w}(\xi)\, d\xi \\\\ &= 0, \end{align*} \] where the exchange of the two integrals is justified by Fubini's theorem, the integrand being bounded in absolute value by \(|w(x)|\, |\hat{\psi}(\xi)|\), which is integrable on \(\mathbb{R}^2\) because \(w \in L^1\) and \(\hat{\psi} \in \mathcal{S}(\mathbb{R})\). By the density of the Schwartz class in \(L^2(\mathbb{R})\), there are Schwartz functions \(\psi_n \to w\) in \(L^2\). The Cauchy-Schwarz inequality gives \(|\langle w, w - \psi_n \rangle| \leq \|w\|_{L^2}\, \|w - \psi_n\|_{L^2} \to 0\), so \[ \|w\|_{L^2}^2 = \langle w, w \rangle = \lim_{n \to \infty} \langle w, \psi_n \rangle = 0. \] Hence \(w = 0\) almost everywhere, and since \(w\) is continuous, \(w \equiv 0\). That is, \(u(\cdot, t) = f * K_t\) for every \(t \gt 0\).

Remark on the class.
The uniqueness statement depends on the decay restriction in an essential way. Without a growth bound, non-trivial solutions with the zero initial datum exist. Tychonoff's classical counterexample exhibits a non-zero \(u\) with \(u(x, t) \to 0\) pointwise as \(t \to 0^+\). This solution violates \(|u(x, t)| \leq C e^{a x^2}\) for every choice of the constants \(C\) and \(a\), so the decay qualifier in the theorem is doing genuine work. We omit the construction. Uniqueness also holds among solutions that attain the initial datum only pointwise, provided \(u\) is continuous up to \(t = 0\) on a finite time interval \([0, T]\) and satisfies a growth bound there, such as boundedness or \(|u(x, t)| \leq C e^{a x^2}\). That result rests on the maximum principle rather than on the Fourier transform, and we do not prove it here.

The convolution form admits a transparent physical reading. The temperature at \(x\) at time \(t\) is a weighted average of the initial temperature, with weights given by the heat kernel centered at \(x\). The kernel is sharply peaked at \(y = x\) for small \(t\) and broadens with width \(\sqrt{2kt}\), so that nearby initial values contribute heavily at first and progressively more distant values contribute as time advances.

In the limit \(t \to 0^+\), \(K_t\) concentrates into a Dirac mass at \(0\), and the convolution recovers \(f\) in the two senses recorded in the theorem.

Insight: The Heat Kernel as a Gaussian Density

The expression for \(K_t\) is not a new special function: setting \(\sigma^2 = 2kt\), \[ K_t(x) = \frac{1}{\sqrt{2\pi \sigma^2}}\, \exp\!\left(-\frac{x^2}{2\sigma^2}\right), \] which is the probability density function of the normal distribution with mean \(0\) and variance \(2kt\). Viewed in this light, the identity \(\int K_t(x)\, dx = 1\) is the Gaussian integral. We will exploit it in the next section as conservation of total heat.

The agreement is not coincidence. The heat equation governs the time evolution of the probability density of Brownian motion. If \(B_t\) is a Brownian motion with variance \(2kt\) at time \(t\) (a constant rescaling of standard Brownian motion), then its density satisfies \(\partial_t p = k\, \partial_{xx} p\) with initial mass concentrated at the origin. The solution is \(p(x, t) = K_t(x)\).

The convolution \(u(x, t) = (f * K_t)(x)\) admits the probabilistic reading \(u(x, t) = \mathbb{E}[f(x + B_t)]\). The temperature at \((x, t)\) is the expected value of the initial profile sampled along a random walk emanating from \(x\). The width \(\sqrt{2kt}\) is then nothing but the standard deviation of the random walk's displacement, growing as \(\sqrt{t}\). This is the hallmark of diffusive scaling, in contrast to the linear-in-\(t\) growth of ballistic motion. The full development of Brownian motion, the Itô calculus that formalizes this connection, and the Fokker-Planck equation that generalizes it to drift-diffusion processes are carried out in the stochastic analysis pages. What is established here is the bridge from a purely analytic computation to a probabilistic object. The probability pages cross it from the other side.

Properties of the Heat Kernel

The convolution solution \(u(x, t) = (f * K_t)(x)\) places the qualitative behavior of every initial-value problem on the same Gaussian footing. All features of the solution can be read off from a single object, the heat kernel \(K_t\). Four properties capture the essential character of parabolic evolution and, taken together, distinguish the heat equation from the other two members of the parabolic-hyperbolic-elliptic trichotomy. They are mass conservation, instantaneous smoothing, the maximum principle, and the ill-posedness of the backward problem.

Conservation of Total Heat

The total integral of the temperature is preserved in time. For the heat kernel itself, this is the statement that \(K_t\) is a probability density. The Gaussian integral identity gives \[ \begin{align*} \int_{-\infty}^{\infty} K_t(x)\, dx &= \frac{1}{\sqrt{4\pi kt}} \int_{-\infty}^{\infty} e^{-x^2/(4kt)}\, dx \\\\ &= \frac{1}{\sqrt{4\pi kt}} \cdot \sqrt{4\pi kt} \\\\ &= 1 \end{align*} \] for every \(t \gt 0\).

For a general initial datum \(f \in L^1(\mathbb{R})\), Fubini's theorem applied to the convolution \(u = f * K_t\) gives \[ \begin{align*} \int_{-\infty}^{\infty} u(x, t)\, dx &= \int_{-\infty}^{\infty} \int_{-\infty}^{\infty} K_t(x - y)\, f(y)\, dy\, dx \\\\ &= \int_{-\infty}^{\infty} f(y) \left( \int_{-\infty}^{\infty} K_t(x - y)\, dx \right) dy \\\\ &= \int_{-\infty}^{\infty} f(y)\, dy, \end{align*} \] so \(\int u(\cdot, t)\, dx = \int f\, dx\) for all \(t \gt 0\). Physically, heat is neither created nor destroyed. It merely redistributes itself across the line. The same identity, viewed probabilistically, says that the diffusing density remains a probability density at every time, a fact that lies at the foundation of the Fokker-Planck equation.

Instantaneous Smoothing

The most characteristic feature of parabolic evolution is that the solution becomes smooth instantly, no matter how rough the initial datum is. The mechanism is most transparent in the frequency domain. From \(\hat{u}(\xi, t) = \hat{f}(\xi)\, e^{-k\xi^2 t}\), the high-frequency content of \(u(\cdot, t)\) is suppressed by a Gaussian factor that decays super-exponentially in \(|\xi|\). Concretely, for any \(t \gt 0\) and any \(n \in \mathbb{N}\), \[ |\xi|^n\, |\hat{u}(\xi, t)| \leq \|f\|_{L^1} \cdot |\xi|^n\, e^{-k\xi^2 t} \] (using the elementary bound \(|\hat{f}(\xi)| \leq \|f\|_{L^1}\) valid for any \(f \in L^1\)). The right-hand side is integrable in \(\xi\) for every \(n\), since polynomial growth is beaten by Gaussian decay.

To conclude smoothness from this integrability, recall the representation \[ u(x, t) = \frac{1}{2\pi}\int_{-\infty}^{\infty} \hat{u}(\xi, t)\, e^{-ix\xi}\, d\xi \] established in Step 1 of the proof of the heat kernel solution, with \(\hat{u}(\xi, t) = \hat{f}(\xi)\, e^{-k\xi^2 t}\). Differentiating \(n\) times in \(x\) under the integral sign brings down a factor \((-i\xi)^n\): \[ \partial_x^n u(x, t) = \frac{1}{2\pi}\int_{-\infty}^{\infty} (-i\xi)^n\, \hat{u}(\xi, t)\, e^{-ix\xi}\, d\xi. \]

Differentiation under the integral is justified precisely because the bound above makes the integrand absolutely integrable in \(\xi\), uniformly in \(x\). The same bound also makes the right-hand side a continuous function of \(x\) by the dominated convergence theorem. Hence \(\partial_x^n u(\cdot, t)\) exists and is continuous for every \(n\), so \(u(\cdot, t)\) is of class \(C^n\) for every \(n\), and therefore of class \(C^\infty\), even if \(f\) is merely \(L^1\) with jump discontinuities or worse. The smoothing acts in an arbitrarily short time. Any \(t \gt 0\), however small, produces an infinitely differentiable function.

The same conclusion can be reached from the convolution side. The kernel \(K_t\) is \(C^\infty\) and decays together with all its derivatives, and convolution with a smooth rapidly decaying kernel inherits the kernel's smoothness regardless of the regularity of the other factor. The frequency-domain argument is preferred here because it identifies the precise quantitative mechanism, Gaussian damping of high frequencies, and makes the rate of smoothing explicit.

The Maximum Principle

For this subsection we strengthen our hypotheses and assume that the initial datum \(f\) is bounded on \(\mathbb{R}\), so that \(\inf_{y \in \mathbb{R}} f(y)\) and \(\sup_{y \in \mathbb{R}} f(y)\) are both finite real numbers. Boundedness is what makes the two bounds below informative. The argument interprets \(u(x, t)\) as a weighted average of values of \(f\), and for an unbounded \(f\) at least one of the bounds is infinite and says nothing.

Because \(K_t\) is non-negative and integrates to one, the convolution \(u(x, t) = \int K_t(x - y)\, f(y)\, dy\) is a weighted average of the values of \(f\), with weights given by the heat kernel. A weighted average cannot exceed the supremum of its inputs nor fall below their infimum: \[ \inf_{y \in \mathbb{R}} f(y) \leq u(x, t) \leq \sup_{y \in \mathbb{R}} f(y) \quad \text{for all } x \in \mathbb{R},\ t \gt 0. \]

The two-sided bound is the maximum principle on the real line. It states that the temperature at any point at any time is bounded above by the maximum initial temperature and below by the minimum. Heat does not spontaneously concentrate to produce hot spots hotter than what was initially present, nor cold spots colder. The result is genuinely a statement about parabolic evolution. Unlike the heat equation, the wave equation generally lacks an analogous pointwise maximum principle. In higher dimensions, focusing can amplify solutions beyond the initial range, and even in one dimension a non-zero initial velocity can drive the solution outside the bounds of \(f\). The conserved quantity of the wave equation is the global mechanical energy rather than a pointwise bound. The Laplace equation has a related but distinct boundary-data form.

The Backward Problem is Ill-Posed

Smoothing is irreversible. The same Gaussian factor \(e^{-k\xi^2 t}\) that damps high frequencies in forward time amplifies them exponentially when run backwards. To attempt to recover \(u(\cdot, 0) = f\) from knowledge of \(u(\cdot, T)\) at some later time \(T \gt 0\), one would need to multiply \(\hat{u}(\xi, T)\) by \(e^{+k\xi^2 T}\). At frequency \(\xi = 10\) with \(kT = 0.14\), the amplification factor is \(e^{14} \approx 1.2 \times 10^6\). At \(\xi = 20\) it exceeds \(10^{24}\).

Any high-frequency noise present in \(u(\cdot, T)\), whether from measurement error, numerical round-off, or any other source, is amplified beyond all bounds. The map \(f \mapsto u(\cdot, T)\) is well-defined and even smoothing, but its inverse is not continuous. On \(L^1 \cap L^2(\mathbb{R})\) the map is injective, since \(\hat{u}(\cdot, T) = \hat{f}\, e^{-k\xi^2 T}\) determines \(\hat{f}\), and by Plancherel's theorem \(\hat{f} = 0\) forces \(\|f\|_{L^2} = 0\). The failure of continuity can be exhibited explicitly in the \(L^2\) norm. Let \(f_m(x) = e^{-x^2/2}\cos(mx)\) for \(m = 1, 2, \ldots\). Each \(f_m\) lies in \(L^1 \cap L^2(\mathbb{R})\), and the same Gaussian transform gives \[ \begin{align*} \|f_m\|_{L^2}^2 &= \tfrac{\sqrt{\pi}}{2}\bigl(1 + e^{-m^2}\bigr), \\\\ \widehat{f_m}(\xi) &= \sqrt{\tfrac{\pi}{2}}\,\bigl(e^{-(\xi + m)^2/2} + e^{-(\xi - m)^2/2}\bigr). \end{align*} \] The solution \(u_m(\cdot, T) = f_m * K_T\) is integrable by Tonelli's theorem and bounded by \(\|f_m\|_{L^1} \sup K_T\), hence in \(L^1 \cap L^2(\mathbb{R})\), and its transform is \(\widehat{f_m}\, e^{-k\xi^2 T}\) by the convolution theorem. Plancherel's theorem, the inequality \((\alpha + \beta)^2 \leq 2\alpha^2 + 2\beta^2\), and the symmetry \(\xi \mapsto -\xi\) give \[ \begin{align*} \|u_m(\cdot, T)\|_{L^2}^2 &= \frac{1}{2\pi}\int_{-\infty}^{\infty} |\widehat{f_m}(\xi)|^2\, e^{-2k\xi^2 T}\, d\xi \\\\ &\leq \int_{-\infty}^{\infty} e^{-(\xi - m)^2}\, e^{-2kT\xi^2}\, d\xi \\\\ &= \sqrt{\frac{\pi}{1 + 2kT}}\, \exp\!\Bigl(-\frac{2kT\, m^2}{1 + 2kT}\Bigr), \end{align*} \] where the last step completes the square in the exponent.

As \(m \to \infty\), \(\|f_m\|_{L^2}^2\) stays at least \(\sqrt{\pi}/2\) while \(\|u_m(\cdot, T)\|_{L^2} \to 0\). No bound of the form \(\|f\|_{L^2} \leq C\, \|u(\cdot, T)\|_{L^2}\) can therefore hold, and the inverse map is not continuous in the \(L^2\) norm. This is what is meant by saying the backward heat equation is ill-posed in the sense of Hadamard.

The forward irreversibility just established stands in sharp contrast to the time-reversibility of the wave equation, where the backward problem is as well-posed as the forward one and no information is lost. Diffusion generative models run a noising process of this forward, smoothing kind, but they do not attempt the ill-posed inversion. Their aim is not to recover the particular sample that was noised, only to draw new samples from the data distribution. The reverse dynamics that achieve this can be written in terms of the score of each intermediate marginal distribution. The page on diffusion models shows that the noise prediction of the network is, up to scaling, an estimate of that score.