Heat Equation on a Bounded Interval
Heat conduction was the very problem that motivated Joseph Fourier to introduce trigonometric
series in his 1822 Théorie analytique de la chaleur. We begin with the prototype
on which the entire theory rests: temperature evolution in a one-dimensional rod of length \(L\)
whose two ends are held at zero temperature. The mathematical formulation is the
initial-boundary value problem
\[
\begin{cases}
\partial_t u = k\, \partial_{xx} u, & 0 \lt x \lt L,\ t \gt 0, \\\\
u(0, t) = u(L, t) = 0, & t \geq 0, \\\\
u(x, 0) = f(x), & 0 \leq x \leq L,
\end{cases}
\]
where \(f\) is a prescribed initial temperature profile. At the two corners \((0, 0)\) and
\((L, 0)\) the boundary and initial conditions both apply, and they agree only if
\(f(0) = f(L) = 0\). The theorem below fixes the precise sense in which each condition is
imposed.
The boundary conditions \(u(0,t) = u(L,t) = 0\) are the homogeneous Dirichlet
conditions, corresponding physically to the rod ends being held in contact with a thermal
reservoir at temperature zero. Our goal is an explicit solution formula that exposes both how
the temperature evolves and how that evolution reflects the spectral structure of the spatial
differential operator \(\partial_{xx}\) on \([0, L]\).
The strategy proceeds in three stages. Separation of variables reduces the PDE to two
ordinary differential equations linked by a separation constant. The spatial ODE together with
the boundary conditions becomes an eigenvalue problem whose solutions are the standing
waves of the rod. Superposition of these standing waves, weighted by the Fourier sine
coefficients of \(f\), gives the complete solution. The idea of superposing standing modes goes
back to Daniel Bernoulli, and Fourier made it systematic. The method reaches far beyond the heat
equation and underpins much of the linear theory of partial differential equations on bounded
domains.
Separation of Variables
We seek solutions in the product form
\[
u(x, t) = \phi(x)\, G(t),
\]
setting aside the initial condition for the moment and demanding only that \(u\) satisfy the PDE
and the homogeneous boundary conditions. Substituting into the heat equation and computing the
relevant partial derivatives gives \(\phi(x) G'(t) = k\, \phi''(x) G(t)\). Dividing both sides
by \(k\, \phi(x)\, G(t)\), which is legitimate wherever neither factor vanishes, separates the
variables:
\[
\frac{G'(t)}{k\, G(t)} = \frac{\phi''(x)}{\phi(x)}.
\]
The left-hand side depends only on \(t\), the right-hand side only on \(x\). Two functions of
independent variables can agree everywhere only if both are equal to a common constant, which we
denote \(-\lambda\):
\[
\frac{G'(t)}{k\, G(t)} = \frac{\phi''(x)}{\phi(x)} = -\lambda.
\]
The minus sign is a convention chosen with foresight. We shall find that the physically relevant
values of \(\lambda\) are positive, in which case \(G(t)\) decays rather than grows. The
separation equation splits into two ordinary differential equations,
\[
\begin{align*}
\phi''(x) + \lambda\, \phi(x) &= 0, \\\\
G'(t) + \lambda k\, G(t) &= 0,
\end{align*}
\]
coupled only through the shared constant \(\lambda\).
The homogeneous boundary conditions \(u(0, t) = u(L, t) = 0\) translate into conditions on
\(\phi\) alone. Requiring \(\phi(0)\, G(t) = \phi(L)\, G(t) = 0\) for all \(t\) and rejecting
the trivial possibility \(G \equiv 0\), which gives \(u \equiv 0\) and is of no interest, we
obtain
\[
\phi(0) = \phi(L) = 0.
\]
The Eigenvalue Problem
The spatial problem
\[
\begin{align*}
\phi''(x) + \lambda\, \phi(x) &= 0, \\\\
\phi(0) = \phi(L) &= 0,
\end{align*}
\]
is a boundary value problem. Two conditions are imposed at two distinct points
of the interval, not at a single point as for an initial value problem. It is fundamentally
different in character from the initial value problems familiar from a first course in ordinary
differential equations. The trivial function \(\phi \equiv 0\) satisfies the ODE and both
boundary conditions regardless of \(\lambda\). The question is whether nontrivial solutions
exist, and if so for which values of \(\lambda\).
Definition: Eigenvalues and Eigenfunctions of the Dirichlet Laplacian on \([0, L]\)
A number \(\lambda \in \mathbb{C}\) is an eigenvalue of the boundary value
problem
\[
\begin{align*}
\phi''(x) + \lambda\, \phi(x) &= 0, \\\\
\phi(0) = \phi(L) &= 0
\end{align*}
\]
if there exists a twice-differentiable solution \(\phi : [0, L] \to \mathbb{C}\) with
\(\phi \not\equiv 0\). Any such \(\phi\) is called an eigenfunction
corresponding to \(\lambda\). The terminology arises because the problem can be rewritten as
\(-\phi'' = \lambda\, \phi\), expressing \(\phi\) as an eigenvector of the differential operator
\(-d^2/dx^2\) acting on twice-differentiable functions on \([0, L]\) that vanish at both
endpoints.
Every eigenvalue is real and positive. Let \(\phi\) be an eigenfunction for \(\lambda\). Since
\(\phi'' = -\lambda\, \phi\) is continuous, we may multiply \(-\phi'' = \lambda\, \phi\) by the
complex conjugate \(\overline{\phi}\) and integrate by parts over \([0, L]\). The boundary term
\([\phi'\, \overline{\phi}]_0^L\) vanishes because \(\phi(0) = \phi(L) = 0\), which gives the
energy identity
\[
\begin{align*}
\lambda \int_0^L |\phi(x)|^2\, dx
&= -\int_0^L \phi''(x)\, \overline{\phi(x)}\, dx \\\\
&= \int_0^L |\phi'(x)|^2\, dx.
\end{align*}
\]
The integral on the left is positive, because \(\phi\) is continuous and not identically zero.
Hence \(\lambda\) is real and \(\lambda \geq 0\). If \(\lambda = 0\), the identity forces
\(\phi' \equiv 0\), so \(\phi\) is constant, and the boundary conditions make it zero.
Therefore \(\lambda \gt 0\).
To find the eigenvalues themselves, we solve the ODE directly. The characteristic equation
\(r^2 + \lambda = 0\) has roots \(r = \pm\sqrt{-\lambda}\), and for real \(\lambda\) the sign
decides whether these roots are real, zero, or purely imaginary. The identity above has already
excluded \(\lambda \leq 0\). The direct computation confirms this independently, and it is what
identifies the positive values that occur. We therefore consider three cases. Throughout, the
constants \(c_1\) and \(c_2\) are complex numbers.
Case 1: \(\lambda \lt 0\).
Writing \(\lambda = -s\) with \(s \gt 0\), the characteristic roots are \(r = \pm\sqrt{s}\), and
the general solution is
\[
\phi(x) = c_1 \cosh\sqrt{s}\, x + c_2 \sinh\sqrt{s}\, x.
\]
The condition \(\phi(0) = 0\) forces \(c_1 = 0\), leaving \(\phi(x) = c_2 \sinh\sqrt{s}\, x\).
The condition \(\phi(L) = 0\) requires \(c_2 \sinh\sqrt{s}\, L = 0\). Since \(\sinh\) vanishes
only at zero and \(\sqrt{s}\, L \gt 0\), we must have \(c_2 = 0\), giving only the trivial solution.
No negative eigenvalue exists.
Case 2: \(\lambda = 0\).
The ODE reduces to \(\phi'' = 0\) with general solution
\[
\phi(x) = c_1 + c_2 x.
\]
The boundary conditions \(\phi(0) = c_1 = 0\) and \(\phi(L) = c_2 L = 0\) (with \(L \gt 0\))
force both constants to vanish. \(\lambda = 0\) is not an eigenvalue.
Case 3: \(\lambda \gt 0\).
The characteristic roots are purely imaginary,
\(r = \pm i\sqrt{\lambda}\), and the general solution in trigonometric form is
\[
\phi(x) = c_1 \cos\sqrt{\lambda}\, x + c_2 \sin\sqrt{\lambda}\, x.
\]
The condition \(\phi(0) = 0\) again forces \(c_1 = 0\). The remaining condition
\(\phi(L) = c_2 \sin\sqrt{\lambda}\, L = 0\) admits nontrivial solutions \((c_2 \neq 0)\)
precisely when \(\sin\sqrt{\lambda}\, L = 0\), that is, when \(\sqrt{\lambda}\, L = n\pi\) for
some positive integer \(n\). (The value \(n = 0\) is excluded because it gives \(\lambda = 0\),
already handled. A negative integer \(n = -m\) with \(m \geq 1\) yields the same eigenvalue
\(\lambda_m\) and, since \(\sin(-m\pi x/L) = -\sin(m\pi x/L)\), the same eigenfunction up to
sign.)
Theorem: Eigenvalues and Eigenfunctions of the Dirichlet Laplacian on \([0, L]\)
The eigenvalues and corresponding eigenfunctions of the boundary value problem
\[
\begin{align*}
\phi''(x) + \lambda\, \phi(x) &= 0, \\\\
\phi(0) = \phi(L) &= 0
\end{align*}
\]
are
\[
\begin{align*}
\lambda_n &= \left(\frac{n\pi}{L}\right)^{\!2}, \\\\
\phi_n(x) &= \sin\frac{n\pi x}{L},
\end{align*}
\]
for \(n = 1, 2, 3, \ldots\). The eigenvalues form a strictly increasing sequence
\(0 \lt \lambda_1 \lt \lambda_2 \lt \cdots\) tending to infinity. Each eigenvalue is simple. Its
eigenspace is one-dimensional, so every eigenfunction for \(\lambda_n\) is a complex multiple
of \(\phi_n\).
Proof:
The case analysis above establishes that nontrivial solutions exist only when
\(\lambda \gt 0\) and \(\sqrt{\lambda}\, L = n\pi\) for some positive integer \(n\), that is,
when \(\lambda = \lambda_n = (n\pi/L)^2\). The corresponding eigenfunctions
\(\phi_n(x) = \sin(n\pi x/L)\) span a one-dimensional space. Any other solution of the ODE
with eigenvalue \(\lambda_n\) is a complex linear combination of \(\sin\sqrt{\lambda_n}\, x\)
and \(\cos\sqrt{\lambda_n}\, x\), and the boundary condition at \(x = 0\) eliminates the cosine
component while \(\phi(L) = 0\) is automatically satisfied by \(\sin\sqrt{\lambda_n}\, x\).
The sequence is strictly increasing because
\(\lambda_{n+1} - \lambda_n = (2n + 1)\pi^2/L^2 \gt 0\). No other eigenvalues exist, because
the energy identity preceding the case analysis shows that all eigenvalues are real, and the
three cases cover every real value.
Insight: Spectral Geometry of an Interval
The eigenfunctions \(\phi_n(x) = \sin(n\pi x/L)\) are the standing waves of
a rod fixed at both ends. They are the same modes that govern a vibrating string of length
\(L\) with fixed endpoints. The \(n\)th mode has \(n - 1\) internal nodes (zeros) and
oscillates \(n/2\) times along the rod. The eigenvalue \(\lambda_n = (n\pi/L)^2\) measures
the "spatial frequency squared" of the mode. Higher modes oscillate more rapidly and, as we
shall see, decay more rapidly in time.
The quadratic growth \(\lambda_n = (\pi/L)^2 n^2\) is no accident. It reflects the fact that
the Laplacian on a one-dimensional interval is a second-order operator. Weyl's law in higher
dimensions generalizes this scaling. On a bounded \(d\)-dimensional domain of volume \(V\),
the \(n\)th Dirichlet eigenvalue satisfies \(\lambda_n \sim C_d\, (n/V)^{2/d}\) as
\(n \to \infty\), where \(C_d = 4\pi^2 \omega_d^{-2/d}\) and \(\omega_d\) is the volume of the
unit ball in \(\mathbb{R}^d\). Here \(\sim\) means that the ratio of the two sides tends to
\(1\). For \(d = 1\), where \(V = L\) and \(\omega_1 = 2\), this gives \(C_1 = \pi^2\), in
agreement with the exact formula above. We state Weyl's law without proof. The spectral
structure of a domain encodes deep geometric information. This theme connects to the
spectral analysis of graph Laplacians in
spectral graph theory
and to the manifold theory that lies ahead in subsequent pages.
Time Evolution and Product Solutions
With the eigenvalues in hand, the time-dependent ODE \(G'(t) + \lambda_n k\, G(t) = 0\)
associated with the \(n\)th mode is a first-order linear equation with constant coefficient
\(\lambda_n k\), whose general solution is the exponential
\[
\begin{align*}
G_n(t) &= e^{-\lambda_n k\, t} \\\\
&= \exp\!\left(-k\!\left(\tfrac{n\pi}{L}\right)^{\!2} t\right),
\end{align*}
\]
chosen with the normalization \(G_n(0) = 1\) so that the multiplicative constant is absorbed
into the spatial part at the time of superposition. Multiplying the spatial eigenfunction by
its time evolution gives the \(n\)th product solution or
normal mode:
\[
u_n(x, t) = \sin\frac{n\pi x}{L}\, \exp\!\left(-k\!\left(\tfrac{n\pi}{L}\right)^{\!2} t\right).
\]
Each \(u_n\) is a solution of the PDE \(\partial_t u = k\, \partial_{xx} u\) satisfying the
homogeneous Dirichlet boundary conditions. Three features stand out. First, every mode
decays exponentially in time. This is the hallmark of parabolic behavior, and it
distinguishes the heat equation sharply from the wave equation, whose modes oscillate without
decay. Second, higher modes decay faster. The decay rate \(k\lambda_n = k(n\pi/L)^2\)
grows quadratically in \(n\), so spatial features of high frequency are damped on a much shorter
timescale than coarse features. The thermal diffusion time of the \(n\)th mode is
\(\tau_n = 1/(k\lambda_n) = L^2/(k n^2 \pi^2)\). Third, the spatial profile is
preserved. Each \(u_n\) factors as a pure sine wave times a time-dependent amplitude,
so the shape \(\sin(n\pi x/L)\) is fixed and only its amplitude shrinks.
Superposition and the Series Solution
The heat equation is linear and homogeneous, and the boundary conditions are also homogeneous.
Consequently any finite linear combination of product solutions is again a solution:
\[
\sum_{n=1}^{N} B_n\, u_n(x, t)
= \sum_{n=1}^{N} B_n \sin\frac{n\pi x}{L}\, \exp\!\left(-k\!\left(\tfrac{n\pi}{L}\right)^{\!2} t\right)
\]
satisfies the PDE and the Dirichlet boundary conditions for any choice of coefficients
\(B_1, \ldots, B_N\). The principle of superposition extends this to infinite
series, provided the series converges in an appropriate sense. Termwise differentiation in both
\(x\) and \(t\) is justified for \(t \gt 0\) whenever the coefficients \(B_n\) are bounded,
because the exponential factors \(\exp(-k(n\pi/L)^2 t)\) decay so rapidly in \(n\) that the
differentiated series converges uniformly on every closed half-strip \(0 \leq x \leq L\),
\(t \geq t_0 \gt 0\). Here we use a standard fact of real analysis without proof. A convergent
series of differentiable functions may be differentiated term by term on any interval where the
series of derivatives converges uniformly.
To match the initial condition \(u(x, 0) = f(x)\), evaluate the formal infinite series at
\(t = 0\). The exponential factors collapse to unity, leaving
\[
f(x) = \sum_{n=1}^{\infty} B_n \sin\frac{n\pi x}{L}, \quad 0 \leq x \leq L.
\]
This is precisely a Fourier sine series expansion of \(f\) on \([0, L]\). The
orthogonality of the
trigonometric system on \([-L, L]\), restricted to the half-interval \([0, L]\) by
the evenness of the product \(\sin(n\pi x/L)\sin(m\pi x/L)\), gives the coefficient formula
immediately. Multiplying both sides by \(\sin(m\pi x/L)\) and integrating over \([0, L]\)
selects the term with \(n = m\), using
\[
\int_0^L \sin\frac{n\pi x}{L} \sin\frac{m\pi x}{L}\, dx = \begin{cases} L/2, & m = n, \\\\ 0, & m \neq n. \end{cases}
\]
The resulting coefficients are the Fourier sine coefficients of \(f\). We collect these
observations into a single statement.
Theorem: Series Solution of the Heat Equation on a Bounded Interval
Let \(f \in L^2([0, L])\) be a real-valued initial temperature profile. We call \(u\) a
solution of the initial-boundary value problem
\[
\begin{cases}
\partial_t u = k\, \partial_{xx} u, & 0 \lt x \lt L,\ t \gt 0, \\\\
u(0, t) = u(L, t) = 0, & t \gt 0, \\\\
u(\cdot, t) \to f \text{ in } L^2([0, L]), & t \to 0^+
\end{cases}
\]
if \(u\) is real-valued, satisfies the three conditions above, and has the following
regularity. The functions \(u\), \(\partial_x u\), \(\partial_{xx} u\), and \(\partial_t u\)
exist and are continuous on \([0, L] \times (0, \infty)\), with one-sided \(x\)-derivatives
at \(x = 0\) and \(x = L\). This problem has exactly one solution, namely
\[
u(x, t) = \sum_{n=1}^{\infty} B_n \sin\frac{n\pi x}{L}\, \exp\!\left(-k\!\left(\tfrac{n\pi}{L}\right)^{\!2} t\right),
\]
where the coefficients are the Fourier sine coefficients of the initial datum,
\[
B_n = \frac{2}{L}\int_0^L f(x) \sin\frac{n\pi x}{L}\, dx.
\]
The series converges in \(L^2([0, L])\) for every \(t \geq 0\) and, for \(t \gt 0\),
converges absolutely and uniformly on \([0, L]\) together with all of its
\(x\)-derivatives and \(t\)-derivatives.
Proof:
Step 1: the series is a solution.
Termwise the series satisfies the PDE and the Dirichlet boundary conditions, as each product
solution \(u_n\) does. The convergence claims are a consequence of the rapid decay of the
exponential factors. The bound \(|B_n| \leq \|f\|_{L^2}\sqrt{2/L}\) follows from
Bessel's inequality
for the orthonormal system \(\{\sqrt{2/L}\sin(n\pi x/L)\}_{n \geq 1}\) in \(L^2([0, L])\). Fix
\(t_0 \gt 0\). Every termwise derivative of the series is bounded on
\([0, L] \times [t_0, \infty)\) by a constant multiple of \(n^j \exp(-k(n\pi/L)^2 t_0)\) for
some integer \(j \geq 0\), and these bounds decay faster than any inverse power of \(n\). The
series and all of its termwise derivatives therefore converge absolutely and uniformly on
\([0, L] \times [t_0, \infty)\). By the termwise-differentiation fact recalled above, \(u\) is
\(C^\infty\) there, up to \(x = 0\) and \(x = L\), and the termwise identities pass to the sum.
Since \(t_0 \gt 0\) is arbitrary, \(u\) has the regularity required of a solution. It is
real-valued because \(f\) is.
Step 2: the initial datum.
The system \(\{\sqrt{2/L}\sin(n\pi x/L)\}_{n \geq 1}\) is an orthonormal basis of
\(L^2([0, L])\). To see this, let \(g \in L^2([0, L])\) be orthogonal to every
\(\sin(n\pi x/L)\), and let \(G\) be its odd extension to \([-L, L]\). For every integer
\(n \geq 0\), the product of \(G\) with \(\cos(n\pi x/L)\) is odd, so its integral over
\([-L, L]\) vanishes. For \(n \geq 1\), the product with \(\sin(n\pi x/L)\) is even, so its
integral is twice the integral of \(g \sin(n\pi x/L)\) over \([0, L]\), which is zero. Hence
\(G\) is orthogonal to every \(e^{in\pi x/L}\) with \(n \in \mathbb{Z}\), and the
completeness
of the full trigonometric system on \(L^2([-L, L])\) gives \(G = 0\), so
\(g = 0\). An orthonormal system to which only the zero vector is orthogonal is an
orthonormal basis, by the characterization recalled in the
completeness argument for Fourier series.
In particular, at \(t = 0\) the series is the Fourier sine series of \(f\), which converges to
\(f\) in \(L^2([0, L])\).
For \(t \gt 0\) the series for \(u(\cdot, t)\) converges uniformly on \([0, L]\), hence in
\(L^2([0, L])\), and the inner product is continuous, so its coefficients in this basis may be
computed term by term. The coefficients of \(u(\cdot, t) - f\) are therefore
\(\sqrt{L/2}\, B_n (e^{-k\lambda_n t} - 1)\), with \(B_n\) real because \(f\) is real-valued.
The \(L^2\)-continuity \(u(\cdot, t) \to f\) as \(t \to 0^+\) then follows from
Parseval's identity:
\[
\|u(\cdot, t) - f\|_{L^2}^2 = \frac{L}{2}\sum_{n=1}^{\infty} B_n^2
\left(e^{-k\lambda_n t} - 1\right)^2,
\]
which tends to \(0\) as \(t \to 0^+\) by
dominated convergence,
applied to the counting measure on the positive integers along an arbitrary sequence
\(t_j \to 0^+\). Each term vanishes pointwise and is bounded uniformly in \(t \geq 0\) by
\(B_n^2\), since \(0 \leq e^{-k\lambda_n t} \leq 1\), while
\((L/2)\sum_n B_n^2 = \|f\|_{L^2}^2 \lt \infty\).
Step 3: uniqueness.
Let \(u_1\) and \(u_2\) be two solutions, and set \(w = u_1 - u_2\). Then \(w\) has the same
regularity, satisfies \(\partial_t w = k\, \partial_{xx} w\) for \(0 \lt x \lt L\) and
\(t \gt 0\), and vanishes at \(x = 0\) and \(x = L\) for \(t \gt 0\). Moreover,
\[
\|w(\cdot, t)\|_{L^2} \leq \|u_1(\cdot, t) - f\|_{L^2} + \|u_2(\cdot, t) - f\|_{L^2}
\]
for every \(t \gt 0\), so \(\|w(\cdot, t)\|_{L^2} \to 0\) as \(t \to 0^+\). Define
\(E(t) = \int_0^L w(x, t)^2\, dx\) for \(t \gt 0\). On each compact rectangle
\([0, L] \times [t_0, t_1]\) with \(t_0 \gt 0\), the functions \(w\) and \(\partial_t w\) are
continuous and hence bounded. By the mean value theorem, the difference quotients of \(w^2\)
in \(t\) are bounded there by \(2 \sup|w| \sup|\partial_t w|\). Dominated convergence, which
applies along every sequence of increments with limit zero, therefore justifies
differentiation under the integral sign. Using \(\partial_t w = k\, \partial_{xx} w\),
\[
\begin{align*}
\frac{dE}{dt}
&= 2\!\int_0^L w\, \partial_t w\, dx \\\\
&= 2k\!\int_0^L w\, \partial_{xx} w\, dx \\\\
&= -2k\!\int_0^L (\partial_x w)^2\, dx \\\\
&\leq 0.
\end{align*}
\]
The integration by parts in the third equality is valid because \(w(\cdot, t)\) is twice
continuously differentiable on \([0, L]\), and the boundary term
\(2k\,[w\, \partial_x w]_0^L\) drops out by the Dirichlet condition. Since \(dE/dt \leq 0\),
\(E\) is non-increasing on \((0, \infty)\). For \(0 \lt s \lt t\) this gives
\(E(t) \leq E(s) = \|w(\cdot, s)\|_{L^2}^2\), and letting \(s \to 0^+\) yields \(E(t) = 0\).
Since \(w(\cdot, t)\) is continuous, it vanishes identically, and \(u_1 = u_2\).
The argument in Step 3 is the energy method. The class of solutions in the
theorem contains every function that has the stated regularity for \(t \gt 0\), satisfies the
heat equation with the Dirichlet condition, and extends continuously to
\([0, L] \times [0, \infty)\) with \(u(\cdot, 0) = f\). Continuity on the compact rectangle
\([0, L] \times [0, 1]\) makes \(u(\cdot, t) \to f\) uniformly, hence in \(L^2\), so the
uniqueness statement applies to them.
Attaining the initial datum continuously asks more of \(f\) than square-integrability. If a
solution \(u\), extended by \(u(\cdot, 0) = f\), is continuous on
\([0, L] \times [0, \infty)\), then \(f\) is continuous and
\(f(0) = \lim_{t \to 0^+} u(0, t) = 0\), and likewise \(f(L) = 0\). These endpoint conditions
are therefore necessary, and together with a mild regularity assumption they are also
sufficient. We call \(f\) piecewise \(C^1\) on \([0, L]\) if there are points
\(0 = s_0 \lt s_1 \lt \cdots \lt s_m = L\) such that the restriction of \(f\) to each
\([s_{j-1}, s_j]\) is continuously differentiable, with one-sided derivatives at the endpoints.
Lemma: Absolute Summability of Sine Coefficients
Let \(f\) be continuous and piecewise \(C^1\) on \([0, L]\) with \(f(0) = f(L) = 0\), and let
\[
B_n = \frac{2}{L}\int_0^L f(x) \sin\frac{n\pi x}{L}\, dx.
\]
Then \(\sum_{n=1}^{\infty} |B_n| \lt \infty\), and the Fourier sine series
\(\sum_{n} B_n \sin(n\pi x/L)\) converges to \(f(x)\) absolutely and uniformly on
\([0, L]\).
Proof:
Write \(I_j\) for the integral of \(f(x) \sin(n\pi x/L)\) over \([s_{j-1}, s_j]\). Integration
by parts gives
\[
\begin{align*}
I_j
&= \Bigl[-\frac{L}{n\pi}\, f(x) \cos\frac{n\pi x}{L}\Bigr]_{s_{j-1}}^{s_j} \\\\
&\quad + \frac{L}{n\pi}\int_{s_{j-1}}^{s_j} f'(x) \cos\frac{n\pi x}{L}\, dx.
\end{align*}
\]
Summing over \(j\), the boundary terms telescope because \(f\) is continuous, and the two
surviving terms vanish because \(f(0) = f(L) = 0\). Hence
\[
\begin{align*}
B_n &= \frac{L}{n\pi}\, a_n, \\\\
a_n &= \frac{2}{L}\int_0^L f'(x) \cos\frac{n\pi x}{L}\, dx,
\end{align*}
\]
where \(f'\) is bounded and continuous except at the finitely many points \(s_j\), so it is
square-integrable. The
orthogonality
relations, restricted to \([0, L]\) because \(\cos(n\pi x/L)\cos(m\pi x/L)\) is an
even function, show that the functions \(\sqrt{2/L}\cos(n\pi x/L)\) with \(n \geq 1\) form an
orthonormal system in \(L^2([0, L])\). The inner product of \(f'\) with the \(n\)th of them is
\(\sqrt{L/2}\, a_n\), so
Bessel's inequality
gives \(\frac{L}{2}\sum_n |a_n|^2 \leq \|f'\|_{L^2}^2 \lt \infty\). The Cauchy-Schwarz
inequality for sums then yields
\[
\begin{align*}
\sum_{n=1}^{\infty} |B_n|
&= \frac{L}{\pi} \sum_{n=1}^{\infty} \frac{|a_n|}{n} \\\\
&\leq \frac{L}{\pi} \Bigl(\sum_{n=1}^{\infty} \frac{1}{n^2}\Bigr)^{1/2}
\Bigl(\sum_{n=1}^{\infty} |a_n|^2\Bigr)^{1/2} \\\\
&\lt \infty.
\end{align*}
\]
Since \(|B_n \sin(n\pi x/L)| \leq |B_n|\), the Weierstrass M-test shows that the sine series
converges uniformly on \([0, L]\), and absolutely at every point. Its sum \(S\) is continuous,
as a uniform limit of continuous functions. The M-test and the continuity of uniform limits
are standard facts that we take for granted. Uniform convergence on the bounded interval
implies convergence in \(L^2([0, L])\), so the partial sums tend to \(S\) in \(L^2\). They
also tend to \(f\) in \(L^2\), because the functions \(\sqrt{2/L}\sin(n\pi x/L)\) form an
orthonormal basis of \(L^2([0, L])\), as Step 2 of the proof of the
series solution theorem
shows. Limits in \(L^2\) are unique, so \(S = f\) almost everywhere. Since \(f\) is
continuous by hypothesis, \(S - f\) is continuous, and a continuous function that vanishes
almost everywhere vanishes everywhere, because otherwise it would be nonzero on an interval
of positive length. Hence \(S = f\) on \([0, L]\).
Theorem: Continuous Attainment of the Initial Datum
Let \(f\) be real-valued, continuous, and piecewise \(C^1\) on \([0, L]\) with
\(f(0) = f(L) = 0\), and let \(u\) be the solution given by the
series solution
theorem. Then the series defining \(u\) converges absolutely and uniformly on
\([0, L] \times [0, \infty)\). With \(u(x, 0) = f(x)\), the function \(u\) is continuous on
\([0, L] \times [0, \infty)\), so \(u(x, t) \to f(x_0)\) as \((x, t) \to (x_0, 0)\) for every
\(x_0 \in [0, L]\).
Proof:
On \([0, L] \times [0, \infty)\), each term has absolute value at most \(|B_n|\), and
\(\sum_n |B_n| \lt \infty\) by the
summability lemma.
The M-test gives absolute and uniform convergence there, so the sum is continuous. At
\(t = 0\) it reduces to the sine series of \(f\), whose sum is \(f\) by the same lemma.
The worked example below satisfies these hypotheses. For a merely continuous \(f\) with
\(f(0) = f(L) = 0\), the series defining \(u\) need not converge pointwise at \(t = 0\), but
the function \(u\), with \(u(\cdot, 0) = f\), is still continuous on
\([0, L] \times [0, \infty)\). The proof of that statement needs an approximation argument
that we do not give here.
The series solution formula has a clean structure. Each Fourier mode of the initial datum
evolves independently, weighted by an exponential decay whose rate is the eigenvalue times the
diffusivity. The Laplacian \(\partial_{xx}\) on \([0, L]\) with Dirichlet boundary
conditions is diagonalized by the orthonormal basis \(\{\sqrt{2/L}\sin(n\pi x/L)\}\), and the
heat semigroup \(e^{tk\partial_{xx}}\) is defined on this basis as multiplication by
\(e^{-tk\lambda_n}\). This is diagonalization by a unitary change of basis, the viewpoint in
which the Hilbert-space treatment of Fourier
analysis culminates. Functions of an unbounded self-adjoint operator, such as the
Dirichlet Laplacian, are treated in general by the functional calculus of such operators, which
lies beyond the scope of this page. The same picture recurs for elliptic operators on manifolds
and for graph Laplacians in spectral graph theory.
A Worked Example
Example: Parabolic Initial Profile
Take \(L = 1\) and the initial datum \(f(x) = x(1 - x)\), a parabolic bump vanishing at both
endpoints and reaching its maximum \(1/4\) at the midpoint \(x = 1/2\). The Fourier sine
coefficients are computed by direct integration:
\[
B_n = 2\int_0^1 x(1 - x) \sin(n\pi x)\, dx.
\]
Two integrations by parts, or the standard tabulated identities, give
\[
\begin{align*}
\int_0^1 x \sin(n\pi x)\, dx &= \frac{(-1)^{n+1}}{n\pi}, \\\\
\int_0^1 x^2 \sin(n\pi x)\, dx &= \frac{(-1)^{n+1}}{n\pi} - \frac{2[1 - (-1)^n]}{(n\pi)^3}.
\end{align*}
\]
Subtracting yields
\[
\begin{align*}
B_n &= 2\!\left(\frac{(-1)^{n+1}}{n\pi} - \left[\frac{(-1)^{n+1}}{n\pi} - \frac{2[1-(-1)^n]}{(n\pi)^3}\right]\right) \\\\
&= \frac{4[1 - (-1)^n]}{(n\pi)^3}.
\end{align*}
\]
Since \(1 - (-1)^n\) equals \(2\) when \(n\) is odd and \(0\) when \(n\) is even, only odd
modes contribute, and writing \(n = 2j + 1\),
\[
\begin{align*}
B_{2j+1} &= \frac{8}{(2j+1)^3 \pi^3}, \\\\
B_{2j} &= 0.
\end{align*}
\]
The solution is therefore
\[
u(x, t) = \frac{8}{\pi^3} \sum_{j=0}^{\infty} \frac{1}{(2j+1)^3}\, \sin((2j+1)\pi x)\,
\exp\!\left(-k (2j+1)^2 \pi^2 t\right).
\]
Three qualitative features illustrate the general structure. The vanishing of even-mode
coefficients reflects the spatial symmetry of \(f\) about \(x = 1/2\). The initial profile
is even with respect to that midpoint, and the even-indexed eigenfunctions \(\sin(2j\pi x)\)
are odd about it, hence orthogonal. The coefficients decay as \(1/(2j+1)^3\), an algebraic
rate set by the smoothness of \(f\) and its compatibility with the boundary conditions.
Finally, the time exponential \(\exp(-k(2j+1)^2 \pi^2 t)\) damps high modes far faster than
the fundamental \((j = 0)\) mode \(\sin(\pi x) e^{-k\pi^2 t}\). The shape relaxes toward a
sine arch on the timescale \(1/(8k\pi^2)\), since the ratio of the \(j = 1\) term to the
fundamental carries the factor \(e^{-8k\pi^2 t}\). After a few multiples of that time the
solution is well approximated by its first term, a sine arch whose amplitude decays to the
zero equilibrium on the longer timescale \(1/(k\pi^2)\). Because \(f\) is continuously
differentiable with \(f(0) = f(1) = 0\), the
continuous attainment theorem
applies, and \(u(x, t) \to f(x_0)\) as \((x, t) \to (x_0, 0)\) for every \(x_0 \in [0, 1]\).
Insight: From Bounded Interval to Real Line
The series formula above is the complete solution of the heat equation on \([0, L]\) with
Dirichlet boundary conditions. Every initial profile decomposes into Dirichlet eigenmodes,
each evolving independently by exponential damping with rate \(k(n\pi/L)^2\). The structure
is rigid: a countable basis, a discrete spectrum, a sum.
The next section removes the bounded interval and asks what happens on the entire real line,
where the boundary is sent to infinity. The Dirichlet eigenmodes \(\sin(n\pi x/L)\) lose
their normalization as \(L \to \infty\), and the discrete spectrum thickens into a
continuum. The replacement of the Fourier series by the Fourier transform is precisely what
is needed. On \(\mathbb{R}\), the heat equation will be solved by a single integral kernel.
Heat Equation on the Real Line
Replacing the rod \([0, L]\) by the full line \(\mathbb{R}\) changes the analytic situation in a
structurally important way. On a bounded interval, the normalized Dirichlet eigenfunctions
\(\sqrt{2/L}\sin(n\pi x/L)\) provide a countable orthonormal basis and the solution is a
discrete sum over normal modes. On the real line, no such basis of eigenfunctions of
\(-\partial_{xx}\) exists in \(L^2(\mathbb{R})\). The formal eigenfunctions \(e^{i\xi x}\)
are bounded but not square-integrable, and the "expansion in eigenmodes" must be replaced by an
integral over a continuous frequency variable.
That replacement is precisely the passage from Fourier series to the
Fourier transform,
carried out on the page devoted to it. Here we put that machinery to its first serious use,
solving an initial-value problem on an unbounded domain.
The problem we now consider is:
\[
\begin{cases}
\partial_t u = k\, \partial_{xx} u, & x \in \mathbb{R},\ t \gt 0, \\\\
u(x, 0) = f(x), & x \in \mathbb{R},
\end{cases}
\]
with \(k \gt 0\) and \(f \in L^1(\mathbb{R})\). We additionally require three things of \(u\).
First, the partial derivatives \(\partial_x u\), \(\partial_{xx} u\), and \(\partial_t u\) exist
on \(\mathbb{R} \times (0, \infty)\), and for each fixed \(t \gt 0\) the four functions
\(u(\cdot, t)\), \(\partial_x u(\cdot, t)\), \(\partial_{xx} u(\cdot, t)\), and
\(\partial_t u(\cdot, t)\) are continuous. Second, these four functions tend to \(0\) as
\(|x| \to \infty\) for each \(t \gt 0\). Third, for every compact \(J \subset (0, \infty)\) there
is an integrable function \(g_J\) on \(\mathbb{R}\) with
\(|u|, |\partial_x u|, |\partial_{xx} u|, |\partial_t u| \leq g_J\) on \(\mathbb{R} \times J\).
We call such a \(u\) a decaying solution. This restriction picks out the natural
function class for the problem and rules out, for instance, the addition of a constant
temperature or a polynomially growing background field. The initial condition is imposed as a
limit, \(u(\cdot, t) \to f\) as \(t \to 0^+\), in a sense made precise in the main theorem, and
the uniqueness of the solution within this class is treated alongside that theorem.
Fourier-Transforming the Equation
Define the spatial Fourier transform of \(u(\cdot, t)\) for each fixed \(t\):
\[
\hat{u}(\xi, t) = \int_{-\infty}^{\infty} u(x, t)\, e^{ix\xi}\, dx.
\]
Under the transform, a double integration by parts with respect to \(x\), the argument that
proves the
differentiation
property, turns \(\partial_{xx}\) into multiplication by \((-i\xi)^2 = -\xi^2\).
Applying the transform to both sides of the heat equation therefore converts the partial
differential equation in \((x, t)\) into an ordinary differential equation in \(t\) alone, with
\(\xi\) entering only as a parameter:
Theorem: Fourier-Transformed Heat Equation
Let \(u\) be a decaying solution of the heat equation on \(\mathbb{R} \times (0, \infty)\)
with \(u(\cdot, t) \to f\) in \(L^1(\mathbb{R})\) as \(t \to 0^+\), and let \(\hat{u}(\xi, t)\)
denote its spatial Fourier transform. Then, for each \(\xi \in \mathbb{R}\), the function
\(t \mapsto \hat{u}(\xi, t)\) satisfies, for \(t \gt 0\),
\[
\begin{align*}
\partial_t \hat{u}(\xi, t) &= -k \xi^2\, \hat{u}(\xi, t), \\\\
\lim_{t \to 0^+} \hat{u}(\xi, t) &= \hat{f}(\xi).
\end{align*}
\]
The unique such function is
\[
\hat{u}(\xi, t) = \hat{f}(\xi)\, e^{-k \xi^2 t}.
\]
Proof:
Differentiation under the integral sign in \(t\) is justified by the
dominated convergence
theorem, applied to difference quotients, since \(|\partial_t u(x, t)| \leq g_J(x)\)
for \(t\) in any compact \(J \subset (0, \infty)\) and \(g_J\) is integrable. Then
\[
\begin{align*}
\partial_t \hat{u}(\xi, t) &= \int_{-\infty}^{\infty} \partial_t u(x, t)\, e^{ix\xi}\, dx \\\\
&= \int_{-\infty}^{\infty} k\, \partial_{xx} u(x, t)\, e^{ix\xi}\, dx \\\\
&= k \cdot (-i\xi)^2\, \hat{u}(\xi, t) \\\\
&= -k\xi^2\, \hat{u}(\xi, t),
\end{align*}
\]
where the third equality is the computation behind the differentiation property of the Fourier
transform. That property is stated for Schwartz functions, which \(u(\cdot, t)\) need not be,
so we carry out the computation directly. Since \(u(\cdot, t)\) is twice continuously
differentiable, two integrations by parts on \([-R, R]\) give
\[
\begin{align*}
\int_{-R}^{R} \partial_{xx} u\, e^{ix\xi}\, dx
&= \Bigl[\partial_x u\, e^{ix\xi}\Bigr]_{-R}^{R} \\\\
&\quad - i\xi \Bigl[u\, e^{ix\xi}\Bigr]_{-R}^{R} \\\\
&\quad + (i\xi)^2 \int_{-R}^{R} u\, e^{ix\xi}\, dx.
\end{align*}
\]
As \(R \to \infty\), the boundary terms vanish by the assumed decay of \(u\) and
\(\partial_x u\) at infinity, and the two integrals converge to integrals over \(\mathbb{R}\)
because \(u(\cdot, t)\) and \(\partial_{xx} u(\cdot, t)\) are integrable. Since
\((i\xi)^2 = (-i\xi)^2 = -\xi^2\), this is the third equality.
For each fixed \(\xi\), the equation
\(\partial_t \hat{u}(\xi, t) = -k\xi^2\, \hat{u}(\xi, t)\) is a scalar linear
ODE in \(t\) on \((0, \infty)\) with constant coefficient \(-k\xi^2\) (\(\xi\) is held
fixed). Hence \(\hat{u}(\xi, t)\, e^{k\xi^2 t}\) has zero derivative, and
\(\hat{u}(\xi, t) = C(\xi)\, e^{-k\xi^2 t}\) for some constant \(C(\xi)\). The limit at
\(t = 0\) holds because
\[
\begin{align*}
|\hat{u}(\xi, t) - \hat{f}(\xi)|
&\leq \int_{-\infty}^{\infty} |u(x, t) - f(x)|\, dx \\\\
&= \|u(\cdot, t) - f\|_{L^1} \longrightarrow 0,
\end{align*}
\]
and letting \(t \to 0^+\) in \(\hat{u}(\xi, t) = C(\xi)\, e^{-k\xi^2 t}\) gives
\(C(\xi) = \hat{f}(\xi)\).
This reduction is the central simplification. The Fourier transform diagonalizes the spatial derivative
operator, turning a partial differential equation into a family of decoupled scalar ODEs indexed
by frequency. Each frequency \(\xi\) decays independently at rate \(k\xi^2\), proportional to
the square of the frequency. High-frequency components are therefore suppressed exponentially
faster than low-frequency ones. The smoothing phenomenon is, in the frequency domain, simply
multiplication by a rapidly decaying Gaussian envelope.
The Heat Kernel
To recover \(u(x, t)\) we must invert the Fourier transform of the right-hand side,
that is, identify the function whose Fourier transform is \(e^{-k\xi^2 t}\). This is
the moment at which the Gaussian calculation done once on the Fourier-transform page
pays for itself across all of parabolic theory. There it was established (as a
worked example) that
\[
\widehat{e^{-x^2/(2\sigma^2)}}(\xi) = \sigma \sqrt{2\pi}\, e^{-\sigma^2 \xi^2 / 2}
\]
for any \(\sigma \gt 0\). Choosing \(\sigma^2 = 2kt\) makes the right-hand side equal
\(\sigma\sqrt{2\pi}\, e^{-k\xi^2 t}\), which differs from the target \(e^{-k\xi^2 t}\)
only by the normalization factor \(\sigma\sqrt{2\pi} = \sqrt{4\pi kt}\). Dividing
through gives the function whose Fourier transform is \(e^{-k\xi^2 t}\). This
function is fundamental enough to deserve a name.
Definition: Heat Kernel on \(\mathbb{R}\)
The heat kernel (or fundamental solution of the heat
equation) on \(\mathbb{R}\) with diffusivity \(k \gt 0\) is
\[
K_t(x) = \frac{1}{\sqrt{4\pi k t}}\, \exp\!\left(-\frac{x^2}{4kt}\right),
\quad t \gt 0,\ x \in \mathbb{R}.
\]
Its spatial Fourier transform is
\[
\widehat{K_t}(\xi) = e^{-k \xi^2 t}.
\]
The Fourier identity in the definition is the substitution \(\sigma^2 = 2kt\) applied to the
Gaussian Fourier transform displayed above, divided by \(\sqrt{4\pi kt}\). No fresh integration
is required. In the spatial variable, \(K_t\) is a centered Gaussian bell whose width
\(\sqrt{2kt}\) grows like \(\sqrt{t}\), the characteristic length scale of diffusive spreading.
Solution by Convolution
With \(\hat{u}(\xi, t) = \hat{f}(\xi)\, \widehat{K_t}(\xi)\), the
convolution theorem
identifies the right-hand side at once. A pointwise product of Fourier transforms is the
Fourier transform of a convolution, so \(\hat{u}(\cdot, t)\) is the transform of \(f * K_t\).
The theorem below confirms that this convolution is a solution, and the only one in the
decaying class.
Theorem: Heat Kernel Solution on \(\mathbb{R}\)
Let \(f \in L^1(\mathbb{R})\). The convolution of \(f\) with the heat kernel,
\[
\begin{align*}
u(x, t) &= (f * K_t)(x) \\\\
&= \int_{-\infty}^{\infty} K_t(x - y)\, f(y)\, dy,
\end{align*}
\]
is a decaying solution of the heat equation \(\partial_t u = k\, \partial_{xx} u\) on
\(\mathbb{R} \times (0, \infty)\) with \(u(\cdot, t) \to f\) in \(L^1(\mathbb{R})\) as
\(t \to 0^+\). It is the only decaying solution with this property. If in addition
\(f\) is bounded and continuous, then \(u(x, t) \to f(x)\) as \(t \to 0^+\) at every
\(x \in \mathbb{R}\). Explicitly,
\[
u(x, t) = \frac{1}{\sqrt{4\pi kt}} \int_{-\infty}^{\infty}
\exp\!\left(-\frac{(x - y)^2}{4kt}\right) f(y)\, dy.
\]
Proof:
Step 1: a representation of the convolution.
The Gaussian transform displayed before the definition of the heat kernel, read with the
roles of the space and frequency variables exchanged and with \(\sigma^2 = 1/(2kt)\), gives
\(\int e^{iz\xi}\, e^{-k\xi^2 t}\, d\xi = \sqrt{\pi/(kt)}\, e^{-z^2/(4kt)}\). Dividing by
\(2\pi\) and using that the Gaussian is even,
\[
K_t(z) = \frac{1}{2\pi} \int_{-\infty}^{\infty} e^{-k\xi^2 t}\, e^{-iz\xi}\, d\xi.
\]
Substituting this into \(u(x, t) = \int f(y)\, K_t(x - y)\, dy\) produces a double integral
whose integrand has absolute value \(|f(y)|\, e^{-k\xi^2 t}\), an integrable function on
\(\mathbb{R}^2\) because \(f \in L^1(\mathbb{R})\). By
Fubini's
theorem, applied to the real and imaginary parts separately, the order of
integration can be exchanged, and the inner integral in \(y\) is \(\hat{f}(\xi)\). Hence
\[
u(x, t) = \frac{1}{2\pi}\int_{-\infty}^{\infty}
\hat{f}(\xi)\, e^{-k\xi^2 t}\, e^{-ix\xi}\, d\xi.
\]
No inversion theorem is used here. The identity is a consequence of the Gaussian
computation and Fubini's theorem alone. The integral form of \(u\) is the definition of
convolution, and the symmetry \(K_t(x - y) = K_t(y - x)\) makes the two orders of the
arguments interchangeable.
Step 2: the convolution is a decaying solution.
Fix \(t_0 \gt 0\) and work on \(t \geq t_0\) with increments in \(t\) that stay in
\([t_0, \infty)\). For \(t \geq t_0\) and every integer \(n \geq 0\),
\[
|\xi|^n\, |\hat{f}(\xi)|\, e^{-k\xi^2 t} \leq \|f\|_{L^1}\, |\xi|^n\, e^{-k\xi^2 t_0},
\]
and the right-hand side is integrable in \(\xi\). By the
dominated convergence
theorem, applied to difference quotients along any sequence of increments
tending to zero, the representation of Step 1 may be differentiated under the integral
sign for \(t \gt 0\). This applies any number of times in \(x\) and once in \(t\). Each
derivative in
\(x\) multiplies the integrand by \(-i\xi\), and the derivative in \(t\) multiplies it by
\(-k\xi^2\). The operators \(\partial_t\) and \(k\, \partial_{xx}\) therefore act on the
integrand in the same way, so \(\partial_t u = k\, \partial_{xx} u\) holds at every point
of \(\mathbb{R} \times (0, \infty)\).
Each of \(u\), \(\partial_x u\), \(\partial_{xx} u\), and \(\partial_t u\) at time \(t\) is,
up to the factor \(1/(2\pi)\) and the reflection \(x \mapsto -x\), the Fourier transform of an
integrable function of \(\xi\). By the
Riemann-Lebesgue
lemma, all four are continuous in \(x\) and tend to \(0\) as \(|x| \to \infty\).
For the integrable bounds, let \(J = [t_0, t_1] \subset (0, \infty)\).
Each \(\partial_x^n K_t(z)\) with \(n \leq 2\) is a polynomial in \(z\) and \(t^{-1/2}\)
times \(e^{-z^2/(4kt)}\), so there is a constant \(C_J\) with
\[
\begin{align*}
|\partial_x^n K_t(z)| &\leq h_J(z) \quad \text{for all } t \in J \text{ and } n \leq 2, \\\\
h_J(z) &= C_J\, (1 + |z|)^2\, e^{-z^2/(4kt_1)},
\end{align*}
\]
and \(h_J\) is integrable. The convolution integral \(u = f * K_t\) may also be
differentiated under the integral sign, since the difference quotients in \(x\) of
\(f(y)\, \partial_x^n K_t(x - y)\) are bounded by
\(|f(y)|\, \sup_{t \in J} \|\partial_x^{n+1} K_t\|_\infty\), an integrable function of
\(y\). This gives \(\partial_x^n u(\cdot, t) = f * \partial_x^n K_t\), hence
\(|\partial_x^n u(x, t)| \leq (|f| * h_J)(x)\) for \(t \in J\). The function
\(g_J = |f| * h_J\) is integrable with \(\|g_J\|_{L^1} \leq \|f\|_{L^1} \|h_J\|_{L^1}\) by
Tonelli's
theorem, and \(|\partial_t u| = k\, |\partial_{xx} u| \leq k\, g_J\). So
\(\max(1, k)\, g_J\) is an integrable bound for all four functions on \(\mathbb{R} \times J\),
and \(u\) is a decaying solution.
Step 3: the initial datum.
That the convolution attains the initial datum in the two senses stated, in
\(L^1(\mathbb{R})\) for \(f \in L^1(\mathbb{R})\) and pointwise when \(f\) is bounded and
continuous, is the approximate-identity property of the Gaussian family \(K_t\) as
\(t \to 0^+\). The verification of this property is omitted here.
Step 4: uniqueness.
Let \(u\) be any decaying solution with \(u(\cdot, t) \to f\) in \(L^1(\mathbb{R})\) as
\(t \to 0^+\). By the
Fourier-transformed heat
equation, \(\hat{u}(\xi, t) = \hat{f}(\xi)\, e^{-k\xi^2 t}\) for every \(\xi\)
and \(t \gt 0\), and by the
convolution
theorem the same function is \(\widehat{f * K_t}(\xi)\). Fix \(t \gt 0\) and
set \(w = u(\cdot, t) - f * K_t\). Then \(w \in L^1(\mathbb{R})\) and \(\hat{w} \equiv 0\).
Moreover \(w\) is continuous and tends to \(0\) at infinity, by the hypotheses on \(u\) and
by Step 2, so \(w\) is bounded, and therefore \(w \in L^2(\mathbb{R})\) as well.
Let \(\psi\) be a Schwartz function. The
Fourier
inversion formula, whose proof covers the Schwartz class, gives
\(\psi(x) = \frac{1}{2\pi} \int \hat{\psi}(\xi)\, e^{-ix\xi}\, d\xi\), so
\[
\begin{align*}
\int_{-\infty}^{\infty} w(x)\, \overline{\psi(x)}\, dx
&= \frac{1}{2\pi} \int_{-\infty}^{\infty} \int_{-\infty}^{\infty}
w(x)\, \overline{\hat{\psi}(\xi)}\, e^{ix\xi}\, d\xi\, dx \\\\
&= \frac{1}{2\pi} \int_{-\infty}^{\infty} \overline{\hat{\psi}(\xi)}\, \hat{w}(\xi)\, d\xi \\\\
&= 0,
\end{align*}
\]
where the exchange of the two integrals is justified by Fubini's theorem, the integrand
being bounded in absolute value by \(|w(x)|\, |\hat{\psi}(\xi)|\), which is integrable on
\(\mathbb{R}^2\) because \(w \in L^1\) and \(\hat{\psi} \in \mathcal{S}(\mathbb{R})\). By the
density of the
Schwartz class in \(L^2(\mathbb{R})\), there are Schwartz functions
\(\psi_n \to w\) in \(L^2\). The Cauchy-Schwarz inequality gives
\(|\langle w, w - \psi_n \rangle| \leq \|w\|_{L^2}\, \|w - \psi_n\|_{L^2} \to 0\), so
\[
\|w\|_{L^2}^2 = \langle w, w \rangle = \lim_{n \to \infty} \langle w, \psi_n \rangle = 0.
\]
Hence \(w = 0\) almost everywhere, and since \(w\) is continuous, \(w \equiv 0\). That is,
\(u(\cdot, t) = f * K_t\) for every \(t \gt 0\).
Remark on the class.
The uniqueness statement depends on the decay restriction in an essential way. Without a
growth bound, non-trivial solutions with the zero initial datum exist. Tychonoff's classical
counterexample exhibits a non-zero \(u\) with \(u(x, t) \to 0\) pointwise as \(t \to 0^+\).
This solution violates \(|u(x, t)| \leq C e^{a x^2}\) for every choice of the constants \(C\)
and \(a\), so the decay qualifier in the theorem is doing genuine work. We omit the
construction. Uniqueness also holds among solutions that attain the initial datum only
pointwise, provided \(u\) is continuous up to \(t = 0\) on a finite time interval
\([0, T]\) and satisfies a growth bound there, such as boundedness or
\(|u(x, t)| \leq C e^{a x^2}\). That result rests on the maximum principle rather than on the
Fourier transform, and we do not prove it here.
The convolution form admits a transparent physical reading. The temperature at \(x\) at time
\(t\) is a weighted average of the initial temperature, with weights given by the heat kernel
centered at \(x\). The kernel is sharply peaked at \(y = x\) for small \(t\) and broadens with
width \(\sqrt{2kt}\), so that nearby initial values contribute heavily at first and
progressively more distant values contribute as time advances.
In the limit \(t \to 0^+\), \(K_t\) concentrates into a Dirac mass at \(0\), and the convolution
recovers \(f\) in the two senses recorded in the theorem.
Insight: The Heat Kernel as a Gaussian Density
The expression for \(K_t\) is not a new special function: setting \(\sigma^2 = 2kt\),
\[
K_t(x) = \frac{1}{\sqrt{2\pi \sigma^2}}\, \exp\!\left(-\frac{x^2}{2\sigma^2}\right),
\]
which is the probability density function of the
normal
distribution with mean \(0\) and variance \(2kt\). Viewed in this light, the
identity \(\int K_t(x)\, dx = 1\) is the
Gaussian
integral. We will exploit it in the next section as conservation of total heat.
The agreement is not coincidence. The heat equation governs the time evolution of the
probability density of Brownian motion. If \(B_t\) is a Brownian motion
with variance \(2kt\) at time \(t\) (a constant rescaling of standard Brownian motion), then
its density satisfies \(\partial_t p = k\, \partial_{xx} p\) with initial mass concentrated
at the origin. The solution is \(p(x, t) = K_t(x)\).
The convolution \(u(x, t) = (f * K_t)(x)\) admits the probabilistic reading
\(u(x, t) = \mathbb{E}[f(x + B_t)]\). The temperature at \((x, t)\) is the expected value of
the initial profile sampled along a random walk emanating from \(x\). The width
\(\sqrt{2kt}\) is then nothing but the standard deviation of the random walk's displacement,
growing as \(\sqrt{t}\). This is the hallmark of diffusive scaling, in contrast to the
linear-in-\(t\) growth of ballistic motion. The full development of
Brownian motion, the
Itô calculus that
formalizes this connection, and the
Fokker-Planck
equation that generalizes it to drift-diffusion processes are carried out in the
stochastic analysis pages. What is established here is the bridge from a purely analytic
computation to a probabilistic object. The probability pages cross it from the other side.
Properties of the Heat Kernel
The convolution solution \(u(x, t) = (f * K_t)(x)\) places the qualitative behavior of every
initial-value problem on the same Gaussian footing. All features of the solution can be
read off from a single object, the heat kernel \(K_t\). Four properties capture the essential
character of parabolic evolution and, taken together, distinguish the heat equation from the
other two members of the parabolic-hyperbolic-elliptic trichotomy. They are mass conservation,
instantaneous smoothing, the maximum principle, and the ill-posedness of the backward problem.
Conservation of Total Heat
The total integral of the temperature is preserved in time. For the heat kernel itself, this is
the statement that \(K_t\) is a probability density. The
Gaussian integral
identity gives
\[
\begin{align*}
\int_{-\infty}^{\infty} K_t(x)\, dx
&= \frac{1}{\sqrt{4\pi kt}} \int_{-\infty}^{\infty} e^{-x^2/(4kt)}\, dx \\\\
&= \frac{1}{\sqrt{4\pi kt}} \cdot \sqrt{4\pi kt} \\\\
&= 1
\end{align*}
\]
for every \(t \gt 0\).
For a general initial datum \(f \in L^1(\mathbb{R})\),
Fubini's
theorem applied to the convolution \(u = f * K_t\) gives
\[
\begin{align*}
\int_{-\infty}^{\infty} u(x, t)\, dx
&= \int_{-\infty}^{\infty} \int_{-\infty}^{\infty} K_t(x - y)\, f(y)\, dy\, dx \\\\
&= \int_{-\infty}^{\infty} f(y) \left( \int_{-\infty}^{\infty} K_t(x - y)\, dx \right) dy \\\\
&= \int_{-\infty}^{\infty} f(y)\, dy,
\end{align*}
\]
so \(\int u(\cdot, t)\, dx = \int f\, dx\) for all \(t \gt 0\). Physically, heat is neither
created nor destroyed. It merely redistributes itself across the line. The same identity, viewed
probabilistically, says that the diffusing density remains a probability density at every time,
a fact that lies at the foundation of the
Fokker-Planck
equation.
Instantaneous Smoothing
The most characteristic feature of parabolic evolution is that the solution becomes
smooth instantly, no matter how rough the initial datum is. The mechanism is most
transparent in the frequency domain. From
\(\hat{u}(\xi, t) = \hat{f}(\xi)\, e^{-k\xi^2 t}\), the high-frequency content of
\(u(\cdot, t)\) is suppressed by a Gaussian factor that decays super-exponentially
in \(|\xi|\). Concretely, for any \(t \gt 0\) and any \(n \in \mathbb{N}\),
\[
|\xi|^n\, |\hat{u}(\xi, t)| \leq \|f\|_{L^1} \cdot |\xi|^n\, e^{-k\xi^2 t}
\]
(using the elementary bound \(|\hat{f}(\xi)| \leq \|f\|_{L^1}\) valid for any
\(f \in L^1\)). The right-hand side is integrable in \(\xi\) for every \(n\), since
polynomial growth is beaten by Gaussian decay.
To conclude smoothness from this integrability, recall the representation
\[
u(x, t) = \frac{1}{2\pi}\int_{-\infty}^{\infty} \hat{u}(\xi, t)\, e^{-ix\xi}\, d\xi
\]
established in Step 1 of the proof of the
heat kernel solution,
with \(\hat{u}(\xi, t) = \hat{f}(\xi)\, e^{-k\xi^2 t}\).
Differentiating \(n\) times in \(x\) under the integral sign brings down a factor \((-i\xi)^n\):
\[
\partial_x^n u(x, t) = \frac{1}{2\pi}\int_{-\infty}^{\infty} (-i\xi)^n\, \hat{u}(\xi, t)\, e^{-ix\xi}\, d\xi.
\]
Differentiation under the integral is justified precisely because the bound above makes the
integrand absolutely integrable in \(\xi\), uniformly in \(x\). The same bound also makes the
right-hand side a continuous function of \(x\) by the dominated convergence theorem. Hence
\(\partial_x^n u(\cdot, t)\) exists and is continuous for every \(n\), so \(u(\cdot, t)\) is of
class \(C^n\) for every \(n\), and therefore of class \(C^\infty\), even if \(f\) is merely
\(L^1\) with jump discontinuities or worse. The smoothing acts in an arbitrarily short time. Any
\(t \gt 0\), however small, produces an infinitely differentiable function.
The same conclusion can be reached from the convolution side. The kernel \(K_t\) is \(C^\infty\)
and decays together with all its derivatives, and convolution with a smooth rapidly decaying
kernel inherits the kernel's smoothness regardless of the regularity of the other factor. The
frequency-domain argument is preferred here because it identifies the precise quantitative
mechanism, Gaussian damping of high frequencies, and makes the rate of smoothing explicit.
The Maximum Principle
For this subsection we strengthen our hypotheses and assume that the initial datum \(f\) is
bounded on \(\mathbb{R}\), so that \(\inf_{y \in \mathbb{R}} f(y)\) and
\(\sup_{y \in \mathbb{R}} f(y)\) are both finite real numbers. Boundedness is what makes the two
bounds below informative. The argument interprets \(u(x, t)\) as a weighted average of values of
\(f\), and for an unbounded \(f\) at least one of the bounds is infinite and says nothing.
Because \(K_t\) is non-negative and integrates to one, the convolution
\(u(x, t) = \int K_t(x - y)\, f(y)\, dy\) is a weighted average of the values of \(f\),
with weights given by the heat kernel. A weighted average cannot exceed the supremum of its
inputs nor fall below their infimum:
\[
\inf_{y \in \mathbb{R}} f(y) \leq u(x, t) \leq \sup_{y \in \mathbb{R}} f(y)
\quad \text{for all } x \in \mathbb{R},\ t \gt 0.
\]
The two-sided bound is the maximum principle on the real line. It states that
the temperature at any point at any time is bounded above by the maximum initial temperature and
below by the minimum. Heat does not spontaneously concentrate to produce hot spots hotter than
what was initially present, nor cold spots colder. The result is genuinely a statement about
parabolic evolution. Unlike the heat equation, the wave equation generally lacks an analogous
pointwise maximum principle. In higher dimensions, focusing can amplify solutions beyond the
initial range, and even in one dimension a non-zero initial velocity can drive the solution
outside the bounds of \(f\). The conserved quantity of the wave equation is the global
mechanical energy rather than a pointwise bound. The
Laplace equation has a related but
distinct boundary-data form.
The Backward Problem is Ill-Posed
Smoothing is irreversible. The same Gaussian factor \(e^{-k\xi^2 t}\) that damps high
frequencies in forward time amplifies them exponentially when run backwards. To attempt to
recover \(u(\cdot, 0) = f\) from knowledge of \(u(\cdot, T)\) at some later time \(T \gt 0\),
one would need to multiply \(\hat{u}(\xi, T)\) by \(e^{+k\xi^2 T}\). At frequency \(\xi = 10\)
with \(kT = 0.14\), the amplification factor is \(e^{14} \approx 1.2 \times 10^6\). At
\(\xi = 20\) it exceeds \(10^{24}\).
Any high-frequency noise present in \(u(\cdot, T)\), whether from measurement error, numerical
round-off, or any other source, is amplified beyond all bounds. The map
\(f \mapsto u(\cdot, T)\) is well-defined and even smoothing, but its inverse is not
continuous. On \(L^1 \cap L^2(\mathbb{R})\) the map is injective, since
\(\hat{u}(\cdot, T) = \hat{f}\, e^{-k\xi^2 T}\) determines \(\hat{f}\), and by
Plancherel's
theorem \(\hat{f} = 0\) forces \(\|f\|_{L^2} = 0\). The failure of continuity can
be exhibited explicitly in the \(L^2\) norm. Let
\(f_m(x) = e^{-x^2/2}\cos(mx)\) for \(m = 1, 2, \ldots\). Each \(f_m\) lies in
\(L^1 \cap L^2(\mathbb{R})\), and the same Gaussian transform gives
\[
\begin{align*}
\|f_m\|_{L^2}^2 &= \tfrac{\sqrt{\pi}}{2}\bigl(1 + e^{-m^2}\bigr), \\\\
\widehat{f_m}(\xi) &= \sqrt{\tfrac{\pi}{2}}\,\bigl(e^{-(\xi + m)^2/2} + e^{-(\xi - m)^2/2}\bigr).
\end{align*}
\]
The solution \(u_m(\cdot, T) = f_m * K_T\) is integrable by Tonelli's theorem and bounded by
\(\|f_m\|_{L^1} \sup K_T\), hence in \(L^1 \cap L^2(\mathbb{R})\), and its transform is
\(\widehat{f_m}\, e^{-k\xi^2 T}\) by the convolution theorem. Plancherel's theorem, the
inequality \((\alpha + \beta)^2 \leq 2\alpha^2 + 2\beta^2\), and the symmetry
\(\xi \mapsto -\xi\) give
\[
\begin{align*}
\|u_m(\cdot, T)\|_{L^2}^2
&= \frac{1}{2\pi}\int_{-\infty}^{\infty} |\widehat{f_m}(\xi)|^2\, e^{-2k\xi^2 T}\, d\xi \\\\
&\leq \int_{-\infty}^{\infty} e^{-(\xi - m)^2}\, e^{-2kT\xi^2}\, d\xi \\\\
&= \sqrt{\frac{\pi}{1 + 2kT}}\, \exp\!\Bigl(-\frac{2kT\, m^2}{1 + 2kT}\Bigr),
\end{align*}
\]
where the last step completes the square in the exponent.
As \(m \to \infty\), \(\|f_m\|_{L^2}^2\) stays at least \(\sqrt{\pi}/2\) while
\(\|u_m(\cdot, T)\|_{L^2} \to 0\). No bound of the form
\(\|f\|_{L^2} \leq C\, \|u(\cdot, T)\|_{L^2}\) can therefore hold, and the inverse map is not
continuous in the \(L^2\) norm. This is what is meant by saying the
backward heat equation is ill-posed in the sense of Hadamard.
The forward irreversibility just established stands in sharp contrast to the time-reversibility
of the wave equation, where the backward
problem is as well-posed as the forward one and no information is lost. Diffusion generative
models run a noising process of this forward, smoothing kind, but they do not attempt the
ill-posed inversion. Their aim is not to recover the particular sample that was noised, only to
draw new samples from the data distribution. The reverse dynamics that achieve this can be
written in terms of the score of each intermediate marginal distribution. The page on
diffusion models
shows that the noise prediction of the network is, up to scaling, an estimate of that score.