Two Density Lemmas
The Itô integral was built as a machine that turns integrands into random variables. This
page asks the machine's inverse question. Which random variables come out? The answer is one
of the structural surprises of the subject. Up to an additive constant, every
square-integrable variable that depends only on the Brownian path up to time \(T\) is an Itô
integral, and every square-integrable martingale of the Brownian filtration is a constant
plus a running Itô integral. Noise is not merely one source of randomness among many. Over
its own filtration, it is the only one.
Throughout, \(w_t\) is the standard one-dimensional Brownian motion of the preceding pages,
started at \(w_0 = 0\), with Brownian filtration \(\{\mathcal{F}_t\}_{t \geq 0}\) containing
all null sets, and \(T \gt 0\) is a fixed horizon. The stage is the
space
\(L^2(\mathcal{F}_T) = L^2(\Omega, \mathcal{F}_T, \mathbb{P})\) of square-integrable
\(\mathcal{F}_T\)-measurable variables, a
Hilbert space
under \(\langle X, Y \rangle = \mathbb{E}[XY]\). The strategy of the whole page is a density
ladder. We show that variables of an extremely special form, exponentials of Itô integrals of
deterministic integrands, already fill \(L^2(\mathcal{F}_T)\) densely. Representation is then
proved for the special form by a computation and extended to everything by taking limits.
This section builds the ladder's two rungs.
What a Measurable Variable Must Look Like
The first tool answers a basic question. If a random variable is measurable with respect to
the \(\sigma\)-algebra generated by finitely many others, how must it depend on them? The
answer is: as a Borel function, and in no other way. For a random vector
\(\mathbf{W} : \Omega \to \mathbb{R}^d\), write
\(\sigma(\mathbf{W}) = \{ \mathbf{W}^{-1}(B) : B \in \mathcal{B}(\mathbb{R}^d) \}\) for the
\(\sigma\)-algebra it generates, the preimage structure making this a \(\sigma\)-algebra
directly.
Theorem: Factorization of Measurable Variables
Let \(\mathbf{W} : \Omega \to \mathbb{R}^d\) be a random vector and
\(Y : \Omega \to \mathbb{R}\) a function. Then \(Y\) is
\(\sigma(\mathbf{W})\)-measurable if and only if there exists a Borel measurable
\(g : \mathbb{R}^d \to \mathbb{R}\) with \(Y = g(\mathbf{W})\).
Proof.
If \(Y = g(\mathbf{W})\) with \(g\) Borel, then for Borel \(E \subseteq \mathbb{R}\) the
preimage \(Y^{-1}(E) = \mathbf{W}^{-1}(g^{-1}(E))\) lies in \(\sigma(\mathbf{W})\), which
is the easy direction. For the converse, climb from indicators. If \(Y = \mathbf{1}_A\)
with \(A \in \sigma(\mathbf{W})\), then \(A = \mathbf{W}^{-1}(B)\) for some Borel
\(B\), and \(Y = \mathbf{1}_B(\mathbf{W})\). If \(Y\) is a simple function, a finite
linear combination of such indicators, take the same combination of the corresponding
\(\mathbf{1}_B\). For general \(\sigma(\mathbf{W})\)-measurable \(Y\), the standard
staircase truncations
\[
Y_n = \Bigl( \bigl( 2^{-n} \lfloor 2^{n} Y \rfloor \bigr) \vee (-n) \Bigr) \wedge n
\]
are simple, \(\sigma(\mathbf{W})\)-measurable, and converge to \(Y\) pointwise, since
\(|Y_n - Y| \leq 2^{-n}\) once \(|Y| \lt n\). Write \(Y_n = g_n(\mathbf{W})\) with
\(g_n\) Borel by the simple case, and set
\(g(\mathbf{x}) = \limsup_n g_n(\mathbf{x})\) where the limit superior is finite and
\(g(\mathbf{x}) = 0\) elsewhere, a Borel function. At every
\(\omega\), \(g_n(\mathbf{W}(\omega)) = Y_n(\omega) \to Y(\omega)\), so the limit
superior at \(\mathbf{W}(\omega)\) is finite and equals \(Y(\omega)\). Hence
\(Y = g(\mathbf{W})\).
Cylinder Functions Are Dense
Call a random variable a cylinder function of the path if it has the form
\(\varphi(w_{t_1}, \ldots, w_{t_k})\) for some \(k\), some times
\(t_1, \ldots, t_k \in [0, T]\), and some \(\varphi \in C_0^\infty(\mathbb{R}^k)\), a smooth
function vanishing outside a bounded set. A cylinder function reads the path through
finitely many samples and processes them smoothly. The first rung of the ladder says that
nothing in \(L^2(\mathcal{F}_T)\) is out of their reach.
Theorem: Density of Cylinder Functions
The set of cylinder functions is dense in \(L^2(\mathcal{F}_T)\). Every
\(g \in L^2(\mathcal{F}_T)\) is the \(L^2\)-limit of a sequence of cylinder functions.
Proof.
Step 1 (Countably many samples suffice).
Let \(t_1, t_2, \ldots\) enumerate a dense subset of \([0, T]\) and let
\(\mathcal{H}_k = \sigma(w_{t_1}, \ldots, w_{t_k})\), an increasing family with union
generating a \(\sigma\)-algebra \(\mathcal{H}_\infty = \sigma\bigl( \bigcup_k \mathcal{H}_k \bigr)\).
We claim \(\mathcal{F}_T\) exceeds \(\mathcal{H}_\infty\) only by null sets. For any
\(s \in [0, T]\), pick \(t_{k_i} \to s\) from the dense set. Paths are continuous outside
a null set, so \(w_s = \lim_i w_{t_{k_i}}\) almost surely, and a pointwise limit of
\(\mathcal{H}_\infty\)-measurable functions, adjusted on the exceptional null set, is
measurable with respect to \(\mathcal{H}_\infty \vee \mathcal{N}\), the augmentation by
null sets. Since \(\mathcal{F}_T\) is generated by exactly these variables \(w_s\) and
the null sets, \(\mathcal{F}_T \subseteq \mathcal{H}_\infty \vee \mathcal{N}\).
Step 2 (Augmentation is invisible to \(L^2\)).
Consider the family
\(\mathcal{G}\) of sets \(A \subseteq \Omega\) differing from some
\(A' \in \mathcal{H}_\infty\) by a null set, meaning
\(A \,\triangle\, A' \in \mathcal{N}\). It contains
\(\mathcal{H}_\infty\) and every null set, and it is a \(\sigma\)-algebra. Complements
preserve the symmetric difference, and for countable unions
\(\bigl( \bigcup A_i \bigr) \triangle \bigl( \bigcup A_i' \bigr)
\subseteq \bigcup ( A_i \triangle A_i' )\), a countable union of null sets. Hence
\(\mathcal{H}_\infty \vee \mathcal{N} \subseteq \mathcal{G}\), and every
\(\mathcal{F}_T\)-measurable indicator equals an \(\mathcal{H}_\infty\)-measurable
indicator almost surely. Climbing through simple functions and pointwise limits as in the
factorization proof, every \(\mathcal{F}_T\)-measurable variable agrees almost surely
with an \(\mathcal{H}_\infty\)-measurable one, so it suffices to approximate
\(g \in L^2(\mathcal{H}_\infty)\).
Step 3 (Down to finitely many samples).
Let \(\mathcal{S}\) be the closure in \(L^2(\mathbb{P})\) of
\(\bigcup_k L^2(\mathcal{H}_k)\), a closed subspace. The family
\(\{ A \in \mathcal{H}_\infty : \mathbf{1}_A \in \mathcal{S} \}\) contains the algebra
\(\bigcup_k \mathcal{H}_k\), and it is stable under complements, by linearity and
\(\mathbf{1}_{A^c} = 1 - \mathbf{1}_A\), and under increasing countable unions, since
\(\mathbf{1}_{A_i} \to \mathbf{1}_{\cup A_i}\) pointwise with the constant dominator one,
so in \(L^2\) by the
dominated convergence theorem.
The passage from an algebra with these closure properties to the full generated
\(\sigma\)-algebra is the monotone-class step this track has taken on faith since the
construction of the integral, and we extend the same trust to this one application.
Simple functions are then in \(\mathcal{S}\) by linearity, and general
\(g \in L^2(\mathcal{H}_\infty)\) by the staircase truncations of the factorization
proof, the squared differences being dominated by the integrable \(4 g^2\).
So \(\mathcal{S} = L^2(\mathcal{H}_\infty)\), and it suffices to approximate
\(g \in L^2(\mathcal{H}_k)\) for fixed \(k\).
Step 4 (Smoothing the dependence).
Fix \(k\) and \(g \in L^2(\mathcal{H}_k)\). By the factorization theorem,
\(g = g_k(\mathbf{W})\) for a Borel \(g_k : \mathbb{R}^k \to \mathbb{R}\), where
\(\mathbf{W} = (w_{t_1}, \ldots, w_{t_k})\), and writing \(\mu\) for the law of
\(\mathbf{W}\) on \(\mathbb{R}^k\),
\[
\mathbb{E}\bigl[ \bigl( g_k(\mathbf{W}) - \varphi(\mathbf{W}) \bigr)^2 \bigr]
= \int_{\mathbb{R}^k} ( g_k - \varphi )^2\, d\mu ,
\]
so the problem becomes approximating \(g_k \in L^2(\mu)\) by
\(\varphi \in C_0^\infty(\mathbb{R}^k)\) in \(L^2(\mu)\). That smooth compactly
supported functions are dense in \(L^2\) of a finite Borel measure on
\(\mathbb{R}^k\) is a standard fact of real analysis, by truncation, regularity of the
measure, and mollification, and we record it as trusted rather than build the mollifier
machinery here. It is the single analytic input of this lemma. Chaining the four steps
completes the proof.
Exponentials Are Dense
The second rung replaces smooth cylinder functions by a family with far more structure, the
exponentials of Itô integrals with deterministic integrands. For
\(h \in L^2([0, T])\) deterministic, the integral \(\int_0^T h\, dw_s\) is defined, the
integrand being trivially adapted and square integrable, and we consider
\[
\mathcal{E}(h)
= \exp \Bigl( \int_0^T h(s)\, dw_s - \tfrac{1}{2} \int_0^T h(s)^2\, ds \Bigr) .
\]
The deterministic correction \(-\tfrac{1}{2} \int h^2\) is a constant factor and does not
affect spans or density. It is included now because the next section will reveal it as the
Itô correction in disguise, the exact discount that makes \(\mathcal{E}(h)\) a martingale.
Theorem: Density of Exponential Functionals
The linear span of \(\{ \mathcal{E}(h) \}\), with \(h\) ranging over the deterministic
step functions on \([0, T]\), is dense in \(L^2(\mathcal{F}_T)\). A fortiori, the
span over all deterministic \(h \in L^2([0, T])\) is dense.
Proof.
Step 1 (Reduction to a vanishing pairing).
If the closed span were a proper subspace of the Hilbert space
\(L^2(\mathcal{F}_T)\), the
projection theorem
would produce a nonzero \(g\) orthogonal to every \(\mathcal{E}(h)\). It therefore
suffices to show that such orthogonality forces \(g = 0\). We will show it forces
\(\mathbb{E}[\varphi(\mathbf{W})\, g] = 0\) for every cylinder function. Then,
choosing cylinder functions \(\varphi_n(\mathbf{W}_n) \to g\) in \(L^2\) by the
density theorem above,
\(0 = \mathbb{E}[\varphi_n(\mathbf{W}_n)\, g] \to \mathbb{E}[g^2]\), and \(g = 0\).
Step 2 (From exponentials of integrals to exponentials of samples).
Fix times \(0 \leq t_1 \lt \cdots \lt t_k \leq T\) and \(\lambda \in \mathbb{R}^k\). The
step function \(h = \sum_i c_i \mathbf{1}_{[0, t_i]}\) has
\(\int_0^T h\, dw = \sum_i c_i\, w_{t_i}\), by
linearity of the integral
and the identity \(\int_0^T \mathbf{1}_{[0, t_i]}\, dw = w_{t_i}\), and the coefficients
\(c_i\) can be chosen to produce any prescribed \(\lambda_i\) as the total weight on
\(w_{t_i}\). Orthogonality to \(\mathcal{E}(h)\), with the deterministic factor
\(\exp(\tfrac{1}{2} \int h^2)\) multiplied back in, therefore gives
\[
\begin{align*}
G(\lambda)
&= \mathbb{E}\Bigl[ \exp\bigl( \lambda_1 w_{t_1} + \cdots + \lambda_k w_{t_k} \bigr)\, g \Bigr] \\\\
&= 0
\quad \forall\, \lambda \in \mathbb{R}^k .
\end{align*}
\]
Every expectation here is finite. The exponent \(\lambda \cdot \mathbf{W}\) is a linear
combination of the Gaussian process, hence a centered Gaussian variable with some
variance \(\sigma^2\), and completing the square against the density gives the Gaussian
moment generating function,
\[
\begin{align*}
\mathbb{E}\bigl[ e^{s Z} \bigr]
&= \int_{-\infty}^{\infty} e^{s z}\, \frac{1}{\sqrt{2 \pi \sigma^2}}
e^{-z^2 / 2 \sigma^2}\, dz \\\\
&= e^{s^2 \sigma^2 / 2}
\int_{-\infty}^{\infty} \frac{1}{\sqrt{2 \pi \sigma^2}}
e^{-(z - s \sigma^2)^2 / 2 \sigma^2}\, dz \\\\
&= e^{s^2 \sigma^2 / 2}
\end{align*}
\]
for \(Z\) centered Gaussian with variance \(\sigma^2 \gt 0\), the degenerate case
\(\sigma^2 = 0\) being trivial, the shifted density integrating
to one by the
Gaussian integral.
With the
Cauchy-Schwarz inequality,
\(\mathbb{E}[e^{\lambda \cdot \mathbf{W}} |g|]
\leq \mathbb{E}[e^{2 \lambda \cdot \mathbf{W}}]^{1/2}\, \|g\|_{L^2} \lt \infty\), so
\(G\) is well defined, and the display will be recycled twice more on this page.
Step 3 (Two measures with one generating function).
Split \(g = g^+ - g^-\) and define two finite Borel measures on \(\mathbb{R}^k\) as the
\(g^\pm\)-weighted laws of the sample vector,
\[
\nu^\pm(B) = \mathbb{E}\bigl[ \mathbf{1}_B(\mathbf{W})\, g^\pm \bigr] ,
\quad B \in \mathcal{B}(\mathbb{R}^k) ,
\]
countably additive by the
monotone convergence theorem
and finite since \(\mathbb{E}[g^\pm] \leq \mathbb{E}[|g|] \leq \|g\|_{L^2}\). The
vanishing of \(G\) says exactly that the two measures have the same moment generating
function everywhere,
\[
\int_{\mathbb{R}^k} e^{\lambda \cdot x}\, d\nu^+(x)
= \int_{\mathbb{R}^k} e^{\lambda \cdot x}\, d\nu^-(x)
\quad \forall\, \lambda \in \mathbb{R}^k ,
\]
the passage from expectations of \(e^{\lambda \cdot \mathbf{W}} g^\pm\) to integrals
against \(\nu^\pm\) being the image-measure substitution, run through indicators, simple
functions, and monotone limits. We now invoke one fact at statement level. Two finite
Borel measures on \(\mathbb{R}^k\) whose
moment generating functions
exist and agree on all of \(\mathbb{R}^k\) are equal. This uniqueness theorem sits at the
same epistemic tier as the continuity theorem already recorded without proof on the
convergence page, its standard proofs running through analytic continuation or Fourier
methods that this track has not built, and it is the single trusted input of this lemma.
Granting it, \(\nu^+ = \nu^-\).
Step 4 (Back up the ladder).
Equal measures integrate every bounded Borel function equally, so for every
\(\varphi \in C_0^\infty(\mathbb{R}^k)\),
\[
\begin{align*}
\mathbb{E}\bigl[ \varphi(\mathbf{W})\, g \bigr]
&= \int \varphi\, d\nu^+ - \int \varphi\, d\nu^- \\\\
&= 0 .
\end{align*}
\]
The times and \(k\) were arbitrary, so \(g\) is orthogonal to every cylinder function,
and Step 1 closes the argument.
The ladder is built. Anything in \(L^2(\mathcal{F}_T)\) can be reached by linear
combinations of the exponentials \(\mathcal{E}(h)\), and the entire representation problem
now reduces to a single question. Can one exponential be written as a constant plus an Itô
integral? The next section answers it with a computation the calculus pages have made
routine.
The Exponential Martingale
The density theorem hands the representation problem a dense family, and the family was
chosen because its members can be represented by a single computation. Two facts are needed.
The law of \(\int_0^t h\, dw\) for deterministic \(h\), which controls every integrability
question, and the Itô differential of \(\mathcal{E}(h)\), which is the representation
itself. The Gaussian law below holds for every deterministic \(h \in L^2([0, T])\). The
exponential computation that follows it is run for step functions, with values \(c_i\) on
the intervals \([r_i, r_{i+1})\) of a partition \(0 = r_0 \lt \cdots \lt r_m = T\),
which is all the density theorem requires.
Theorem: Integrals of Deterministic Integrands Are Gaussian
Let \(h \in L^2([0, T])\) be deterministic and \(t \in [0, T]\). Then
\(I_t = \int_0^t h(s)\, dw_s\) is a centered Gaussian random variable with variance
\(\int_0^t h(s)^2\, ds\), and for every \(s \in \mathbb{R}\),
\[
\mathbb{E}\bigl[ e^{s I_t} \bigr]
= \exp \Bigl( \tfrac{s^2}{2} \int_0^t h(r)^2\, dr \Bigr) .
\]
Proof.
Step 1 (Step integrands, exactly).
Suppose first that \(h\) is a step function. Then
\(I_t = \sum_i c_i (w_{r_{i+1} \wedge t} - w_{r_i \wedge t})\) is a finite linear
combination of increments over disjoint intervals. Its moment generating function
factors over the increments, which are
independent of the past
so the expectation of the product peels one factor at a time from the last increment
backwards, by the
product rule,
each peeled factor being a Gaussian moment generating function computed by the completed square
of the previous section. Multiplying,
\[
\begin{align*}
\mathbb{E}\bigl[ e^{s I_t} \bigr]
&= \prod_i \exp \Bigl( \tfrac{s^2 c_i^2}{2}
\bigl( r_{i+1} \wedge t - r_i \wedge t \bigr) \Bigr) \\\\
&= \exp \Bigl( \tfrac{s^2}{2} \int_0^t h(r)^2\, dr \Bigr) .
\end{align*}
\]
Step 2 (General integrands by \(L^1\) transfer).
For general deterministic \(h \in L^2([0, T])\), the density input already trusted in
the cylinder lemma provides smooth compactly supported approximants of \(h\) in
\(L^2([0, t])\), and left-endpoint discretization, exact in \(L^2\) up to the modulus
of uniform continuity, turns those into step functions, so there are step functions
\(h_n \to h\) in \(L^2([0, t])\). The
Itô isometry
gives \(I^{(n)} = \int_0^t h_n\, dw \to I_t\) in \(L^2(\mathbb{P})\). Fix
\(s \in \mathbb{R}\). The elementary bound
\(|e^{a} - e^{b}| \leq |a - b|\, (e^{a} + e^{b})\), from the mean value of the
exponential between \(a\) and \(b\), and the Cauchy-Schwarz inequality give
\[
\Bigl| \mathbb{E}\bigl[ e^{s I^{(n)}} \bigr] - \mathbb{E}\bigl[ e^{s I_t} \bigr] \Bigr|
\leq |s|\, \bigl\| I^{(n)} - I_t \bigr\|_{L^2}
\Bigl( \bigl\| e^{s I^{(n)}} \bigr\|_{L^2} + \bigl\| e^{s I_t} \bigr\|_{L^2} \Bigr) .
\]
The first factor tends to zero. The norms are bounded along the sequence, since
\(\mathbb{E}[ e^{2 s I^{(n)}} ] = \exp( 2 s^2 \int_0^t h_n^2 )\) converges by Step 1, and
\(\mathbb{E}[ e^{2 s I_t} ] \lt \infty\) because along an almost surely convergent
subsequence
Fatou's
lemma
bounds it by the limit of the exact values. So
\[
\begin{align*}
\mathbb{E}[ e^{s I_t} ] &= \lim_n \exp( \tfrac{s^2}{2} \int_0^t h_n^2 ) \\\\
&= \exp( \tfrac{s^2}{2} \int_0^t h^2 ) .
\end{align*}
\]
Step 3 (Identifying the law).
The law of \(I_t\) and the centered Gaussian law with variance
\(\sigma^2 = \int_0^t h^2\, dr\) are two finite Borel measures whose moment generating
functions exist and agree everywhere, the Gaussian side by the completed square. The
uniqueness of finite measures with a common moment generating function, the input
already trusted in the density lemma, identifies the two laws, and \(I_t\) is centered
Gaussian with variance \(\sigma^2\).
Now the main computation. For a deterministic step function \(h\), define the process
Definition: Exponential Martingale
For deterministic \(h \in L^2([0, T])\), the exponential martingale of
\(h\) is the process
\[
Z_t = \exp \Bigl( \int_0^t h(s)\, dw_s
- \tfrac{1}{2} \int_0^t h(s)^2\, ds \Bigr) ,
\quad 0 \leq t \leq T ,
\]
so that \(Z_0 = 1\) and \(Z_T = \mathcal{E}(h)\), the \(dw\)-term taken in its
continuous realization.
The name is a claim, and the claim is the theorem. The deterministic discount
\(-\tfrac{1}{2} \int h^2\), carried unexplained through the density section, is exactly the
Itô correction of the exponential, engineered so that the drift of \(Z\) cancels to zero.
Theorem: The Exponential Martingale Identity
Let \(h\) be a deterministic step function on \([0, T]\) and \(Z\) its exponential
martingale. Then \(h(s) Z_s \in \mathcal{V}(0, T)\), and for every \(t \in [0, T]\),
almost surely,
\[
\begin{align*}
Z_t &= 1 + \int_0^t h(s)\, Z_s\, dw_s , \\\\
dZ_t &= h(t)\, Z_t\, dw_t .
\end{align*}
\]
In particular \(Z\) is a
martingale
with \(\mathbb{E}[Z_t] = 1\) for all \(t\).
Proof.
Step 1 (Membership).
The process \(h(s) Z_s\) is progressively measurable, being a product of a deterministic
step function and a continuous adapted process, and by the Gaussian theorem, with
\(I_s = \int_0^s h\, dw\),
\[
\begin{align*}
\mathbb{E}\bigl[ Z_s^2 \bigr]
&= e^{- \int_0^s h^2}\, \mathbb{E}\bigl[ e^{2 I_s} \bigr] \\\\
&= e^{- \int_0^s h^2}\, e^{2 \int_0^s h^2} \\\\
&= e^{\int_0^s h^2} \\\\
&\leq e^{\|h\|_{L^2}^2} ,
\end{align*}
\]
so \(\mathbb{E}\bigl[ \int_0^T h^2 Z_s^2\, ds \bigr]
\leq \|h\|_{L^2}^2\, e^{\|h\|_{L^2}^2} \lt \infty\), and
\(h Z \in \mathcal{V}(0, T)\).
Step 2 (The identity, one piece at a time).
The process \(X_t = I_t\) is an Itô process with \(u = 0\) and \(v = h\). The natural
map \(g(t, x) = \exp( x - \tfrac{1}{2} \int_0^t h^2 )\) is not \(C^2\) in \(t\), its
time derivative jumping where \(h\) jumps, so we work one interval of the partition at a
time with a smooth stand-in. Fix a piece \([r_i, r_{i+1}]\), on which \(h \equiv c_i\),
and define
\[
\begin{align*}
g_i(t, x) &= \exp \Bigl( x - A(r_i) - \tfrac{c_i^2}{2} ( t - r_i ) \Bigr) , \\\\
A(r) &= \tfrac{1}{2} \int_0^{r} h^2\, ds ,
\end{align*}
\]
a \(C^2\) function on all of \([0, \infty) \times \mathbb{R}\) that agrees with the
natural map for \(t \in [r_i, r_{i+1}]\), where \(g_i(t, I_t) = Z_t\). Its membership
condition holds, the diffusion entry \(h(s)\, g_i(s, I_s)\) differing from \(h Z\) by
the bounded deterministic factor \(e^{A(s) - A(r_i) - c_i^2 (s - r_i)/2}\), so the
Itô formula
applies to \(g_i\). Its drift integrand is
\[
\frac{\partial g_i}{\partial t} + \tfrac{1}{2}\, h(s)^2\,
\frac{\partial^2 g_i}{\partial x^2}
= \Bigl( - \tfrac{c_i^2}{2} + \tfrac{h(s)^2}{2} \Bigr)\, g_i(s, I_s) ,
\]
which vanishes identically for \(s \in [r_i, r_{i+1})\), where \(h(s) = c_i\). This is
the discount doing its work. The half squared integrand subtracted in the exponent is
consumed, to the last drop, by the half second derivative the chain rule inserts.
Applying the formula at times \(r_i\) and \(t \in [r_i, r_{i+1}]\) and subtracting, the
\(ds\)-integrals over \([r_i, t]\) carry only the vanishing drift, and
\[
\begin{align*}
Z_t - Z_{r_i}
&= \int_{r_i}^{t} h(s)\, g_i(s, I_s)\, dw_s \\\\
&= \int_{r_i}^{t} h(s)\, Z_s\, dw_s ,
\end{align*}
\]
the integrands agreeing on the interval of integration and the interval splitting being
the
additivity of the integral.
Telescoping over the pieces up to \(t\), with \(Z_0 = 1\), yields the identity, one null
set per piece and finitely many pieces.
Step 3 (Martingale).
With \(h Z \in \mathcal{V}(0, T)\), the process \(Z_t - 1\) is an Itô integral process
and hence a
martingale
with zero mean, so \(Z\) is a martingale with \(\mathbb{E}[Z_t] = 1\).
This is the first representation. The exponential \(\mathcal{E}(h) = Z_T\) equals its mean,
one, plus the Itô integral of the explicit integrand \(h Z \in \mathcal{V}(0, T)\). The
density theorem says such exponentials span everything, and the next section turns the span
and a limit into the general statement.
The Itô Representation Theorem
The pieces assemble. Exponentials are representable by an explicit computation, exponentials
span a dense subspace, and the Itô isometry makes the passage to limits not merely possible
but automatic, because it converts convergence of the represented variables into convergence
of the representing integrands. One point deserves care before the proof, and it is where
the progressive machinery of the formula page pays an unadvertised dividend. Limits of
integrands must again be legitimate integrands, and the class \(\mathcal{V}(0, T)\) must be
closed under \(L^2\)-limits. The membership conditions of \(\mathcal{V}\) constrain
measurability, not just size, and an \(L^2\)-limit is only an equivalence class, so
closedness needs an argument, supplied inside the proof by a limit superior of progressively
measurable representatives.
Theorem: The Itô Representation Theorem
Let \(F \in L^2(\mathcal{F}_T)\). Then there exists a progressively measurable
\(f \in \mathcal{V}(0, T)\), unique up to equality
\(\lambda \otimes \mathbb{P}\)-almost everywhere, such that, almost surely,
\[
F = \mathbb{E}[F] + \int_0^T f(s, \omega)\, dw_s
\]
Proof.
Step 1 (The dense subspace is represented).
For a single exponential, the previous section gives
\(\mathcal{E}(h) = 1 + \int_0^T h\, Z^{h}_s\, dw_s\) with
\(h Z^{h} \in \mathcal{V}(0, T)\) and \(\mathbb{E}[\mathcal{E}(h)] = 1\), writing
\(Z^h\) for the exponential martingale of \(h\). For a finite linear combination
\(F_0 = \sum_j a_j\, \mathcal{E}(h_j)\) with step functions \(h_j\), linearity of the
integral gives
\[
\begin{align*}
F_0
&= \sum_j a_j + \int_0^T \Bigl( \sum_j a_j\, h_j(s)\, Z^{h_j}_s \Bigr) dw_s \\\\
&= \mathbb{E}[F_0] + \int_0^T f_0(s, \omega)\, dw_s ,
\end{align*}
\]
the mean identification using \(\mathbb{E}[\mathcal{E}(h_j)] = 1\), and the integrand
\(f_0\), a finite combination of progressively measurable members of
\(\mathcal{V}(0, T)\), lies in \(\mathcal{V}(0, T)\) and is progressively measurable.
Step 2 (The isometry transports the Cauchy property).
Let \(F \in L^2(\mathcal{F}_T)\) and, by the density theorem, choose span elements
\(F_n \to F\) in \(L^2(\mathbb{P})\), each represented as in Step 1 by a progressively
measurable \(f_n \in \mathcal{V}(0, T)\). Means converge,
\(|\mathbb{E}[F_n] - \mathbb{E}[F]| \leq \mathbb{E}[|F_n - F|] \leq \|F_n - F\|_{L^2}\)
by the Cauchy-Schwarz inequality against the constant one. Subtracting two
representations and applying the
Itô isometry,
with a Tonelli exchange writing the squared norm as an iterated integral,
\[
\begin{align*}
\int_0^T \mathbb{E}\bigl[ ( f_n - f_m )^2 \bigr]\, ds
&= \mathbb{E}\Bigl[ \Bigl( \int_0^T ( f_n - f_m )\, dw_s \Bigr)^{2} \Bigr] \\\\
&= \mathbb{E}\Bigl[ \bigl( F_n - F_m - \mathbb{E}[ F_n - F_m ] \bigr)^{2} \Bigr] \\\\
&\leq \mathbb{E}\bigl[ ( F_n - F_m )^2 \bigr] ,
\end{align*}
\]
the last inequality because subtracting the mean can only decrease the second moment,
\(\mathbb{E}[(X - \mathbb{E}X)^2] = \mathbb{E}[X^2] - (\mathbb{E}X)^2\). The right side
tends to zero as \(n, m \to \infty\), so \((f_n)\) is a Cauchy sequence in
\(L^2([0, T] \times \Omega,\ \lambda \otimes \mathbb{P})\), and by
completeness
it converges to some \(f\) in that space.
Step 3 (The limit is a legitimate integrand).
Along a
subsequence,
\(f_{n_k} \to f\) at \(\lambda \otimes \mathbb{P}\)-almost every point of
\([0, T] \times \Omega\). Define
\[
\tilde{f}(s, \omega)
= \limsup_{k}\, f_{n_k}(s, \omega) ,
\]
set to zero where the limit superior fails to be finite. Each \(f_{n_k}\) restricted to
\([0, t] \times \Omega\) is \(\mathcal{B}([0, t]) \otimes \mathcal{F}_t\)-measurable,
limit superiors preserve measurability, and the finiteness set is measurable, so
\(\tilde{f}\) is
progressively measurable.
At almost every point \(\tilde{f} = f\), so \(\tilde{f}\) lies in the same
\(L^2\)-class, is jointly measurable and adapted through progressivity, and is square
integrable. Hence \(\tilde{f} \in \mathcal{V}(0, T)\), and we rename it \(f\). The class
\(\mathcal{V}(0, T)\) is, in this sense, closed in
\(L^2(\lambda \otimes \mathbb{P})\).
Step 4 (Passing to the limit).
By the isometry once more, \(\int_0^T f_n\, dw \to \int_0^T f\, dw\) in
\(L^2(\mathbb{P})\), and combining the three convergences,
\[
\begin{align*}
F &= \lim_n F_n \\\\
&= \lim_n \Bigl( \mathbb{E}[F_n] + \int_0^T f_n\, dw_s \Bigr) \\\\
&= \mathbb{E}[F] + \int_0^T f\, dw_s ,
\end{align*}
\]
the limits taken in \(L^2(\mathbb{P})\) and the identity holding almost surely.
Step 5 (Uniqueness).
If \(f_1, f_2 \in \mathcal{V}(0, T)\) both represent \(F\), subtracting the two
identities gives \(\int_0^T (f_1 - f_2)\, dw_s = 0\) almost surely, and the isometry,
with the same Tonelli exchange, turns the vanishing of the integral into the vanishing of
the integrand,
\[
\begin{align*}
0 &= \mathbb{E}\Bigl[ \Bigl( \int_0^T ( f_1 - f_2 )\, dw_s \Bigr)^{2} \Bigr] \\\\
&= \int_0^T \mathbb{E}\bigl[ ( f_1 - f_2 )^2 \bigr]\, ds ,
\end{align*}
\]
so \(f_1 = f_2\) at \(\lambda \otimes \mathbb{P}\)-almost every point.
The theorem repays a moment of astonishment before it is put to work. Nothing about
\(F\) was assumed beyond square integrability and measurability with respect to the
Brownian history. Any such quantity, the maximum of the path, the time it spends above a
level, an arbitrarily contorted functional of the trajectory, is a constant plus a
stochastic integral, synthesized from the same noise it observes by some adapted,
square-integrable recipe \(f\). The theorem asserts the recipe exists and is essentially
unique, and it does not say how to find it. Producing \(f\) explicitly is its own
industry, and the door it opens for mathematical finance, where \(f\) is a hedging
strategy, is a signpost the closing section records.
The Martingale Representation Theorem
The representation theorem freezes time at a single horizon \(T\). The final step lets the
horizon run, and represents entire martingales rather than single variables. Two of the
earlier pages meet here. The representation supplies an integrand for each fixed time, and
the
martingale property of the Itô integral
forces the integrands at different times to agree, so that one recipe serves all horizons
at once.
Theorem: The Martingale Representation Theorem
Let \(\{M_t\}_{t \geq 0}\) be a
martingale
with respect to the Brownian filtration \(\{\mathcal{F}_t\}\) with
\(M_t \in L^2(\mathbb{P})\) for every \(t\). Then there exists a progressively
measurable process \(g\) with \(g \in \mathcal{V}(0, t)\) for every \(t \gt 0\), unique
up to \(\lambda \otimes \mathbb{P}\)-null sets, such that for every \(t \geq 0\), almost
surely,
\[
M_t = \mathbb{E}[M_0] + \int_0^t g(s, \omega)\, dw_s .
\]
Proof.
Step 1 (A representation at every horizon).
Fix \(t \gt 0\). The variable \(M_t\) lies in \(L^2(\mathcal{F}_t)\), so the
representation theorem with horizon \(t\) supplies a progressively measurable
\(f^{(t)} \in \mathcal{V}(0, t)\), unique almost everywhere, with
\[
M_t = \mathbb{E}[M_t] + \int_0^t f^{(t)}(s, \omega)\, dw_s ,
\]
The mean is constant in \(t\). By the
tower property
and the martingale identity,
\(\mathbb{E}[M_t] = \mathbb{E}\bigl[ \mathbb{E}[M_t \mid \mathcal{F}_0] \bigr]
= \mathbb{E}[M_0]\).
Step 2 (Consistency of the recipes).
Let \(0 \lt t_1 \lt t_2\). Conditioning the horizon-\(t_2\) representation on
\(\mathcal{F}_{t_1}\), the martingale identity for \(M\) on the left and the martingale
property of the Itô integral on the right give, almost surely,
\[
\begin{align*}
M_{t_1}
&= \mathbb{E}\bigl[ M_{t_2} \mid \mathcal{F}_{t_1} \bigr] \\\\
&= \mathbb{E}[M_0]
+ \mathbb{E}\Bigl[ \int_0^{t_2} f^{(t_2)}\, dw_s \,\Bigm|\, \mathcal{F}_{t_1} \Bigr] \\\\
&= \mathbb{E}[M_0] + \int_0^{t_1} f^{(t_2)}\, dw_s .
\end{align*}
\]
The restriction of \(f^{(t_2)}\) to \([0, t_1]\) lies in \(\mathcal{V}(0, t_1)\) and
represents \(M_{t_1}\), as does \(f^{(t_1)}\), so the uniqueness clause of the
representation theorem forces, at \(\lambda \otimes \mathbb{P}\)-almost every point
of \([0, t_1] \times \Omega\),
\[
f^{(t_2)} = f^{(t_1)}
\]
Recipes for longer horizons extend recipes for shorter ones. Nothing needs to be
recomputed as time grows.
Step 3 (One process for all horizons).
Define \(g\) by gluing along integer horizons,
\[
g(s, \omega) = f^{(n)}(s, \omega) ,
\quad
s \in (n - 1, n] ,\ n \in \mathbb{N} ,
\]
and \(g(0, \cdot) = f^{(1)}(0, \cdot)\). Each piece is the product of a progressively
measurable process and a deterministic indicator, and the pieces live on disjoint time
intervals, so \(g\) is progressively measurable. Fix \(t \gt 0\) and let
\(N = \lceil t \rceil\). By Step 2, chained from each integer \(n \leq N\) up to \(N\),
the process \(g\) agrees with \(f^{(N)}\) at almost every point of
\([0, t] \times \Omega\), so \(g \in \mathcal{V}(0, t)\) and, integrands equal almost
everywhere having equal integrals by the isometry,
\[
\begin{align*}
\int_0^t g\, dw_s
&= \int_0^t f^{(N)}\, dw_s \\\\
&= M_t - \mathbb{E}[M_t] \\\\
&= M_t - \mathbb{E}[M_0]
\end{align*}
\]
almost surely, where the middle equality is Step 2 applied to the pair \((t, N)\). This
is the claimed identity for \(t \gt 0\). At \(t = 0\) the identity reads
\(M_0 = \mathbb{E}[M_0]\), which holds almost surely because \(\mathcal{F}_0\), being
generated by \(w_0 = 0\) and the null sets, consists of null sets and their
complements, so \(\mathcal{F}_0\)-measurable variables are almost surely constant.
Step 4 (Uniqueness).
If \(g\) and \(g'\) both work, then for every \(t\) they both represent \(M_t\) over
\([0, t]\), so they agree \(\lambda \otimes \mathbb{P}\)-almost everywhere on
\([0, t] \times \Omega\) by the representation theorem's uniqueness, and letting
\(t \to \infty\) through the integers, almost everywhere on
\([0, \infty) \times \Omega\).
The time-zero case that closed the proof deserves its own sentence, because it is a
statement about information, not about integrals. A martingale of the Brownian filtration
cannot even randomize its starting value. Every square-integrable fair game riding on
Brownian information is, in full, a constant plus a running Itô integral.
Two remarks close the chapter's accounts. First, dimensions. The same ladder, with sample
vectors drawn from all components, componentwise exponentials, and the matrix-valued
integrands of the multi-dimensional integral, delivers the \(n\)-dimensional statements,
every \(L^2\) functional of an \(n\)-dimensional Brownian path being a constant plus
\(\sum_j \int f_j\, dw^{(j)}\). We record the vector version at statement level, the
argument repeating rather than deepening. Second, what was trusted. This page consumed
exactly two statement-level inputs beyond the track's standing conventions, the density of
smooth compactly supported functions in \(L^2\) of a finite measure, and the uniqueness of
finite measures through their moment generating functions. Everything else, including the
convergence theorem for conditional expectations that the classical route runs through, was
either proved or bypassed.
The Recipe Is the Product
The representation theorem is an existence statement with an industry attached to its
constructive gap. In mathematical finance the theorem is the reason Brownian market
models are complete. A claim paying \(F\) at time \(T\) is a constant plus
\(\int_0^T f\, dw\), the constant is the price and the integrand \(f\) is the hedging
strategy, the recipe for replicating the claim by trading continuously against the
noise. The theorem guarantees the hedge exists and says nothing about its formula, and
computing or approximating \(f\), by partial differential equations, by Malliavin-type
calculus, or by training a neural network to output the integrand directly, is a live
boundary between stochastic analysis and machine learning.
The chapter that began with a chain rule ends with an inventory. Over its own filtration,
Brownian motion is not one source of randomness among many but the only one, every
observable a constant plus an integral against it, every fair game a running such integral.
The calculus, the conversion, and now the representation are the complete toolkit for the
next destination, equations that define a process implicitly through its own noise, where
the coefficients of the unknown appear inside the integrals that build it.