The Martingale Representation Theorem

Two Density Lemmas The Exponential Martingale The Itô Representation Theorem The Martingale Representation Theorem

Two Density Lemmas

The Itô integral was built as a machine that turns integrands into random variables. This page asks the machine's inverse question. Which random variables come out? The answer is one of the structural surprises of the subject. Up to an additive constant, every square-integrable variable that depends only on the Brownian path up to time \(T\) is an Itô integral, and every square-integrable martingale of the Brownian filtration is a constant plus a running Itô integral. Noise is not merely one source of randomness among many. Over its own filtration, it is the only one.

Throughout, \(w_t\) is the standard one-dimensional Brownian motion of the preceding pages, started at \(w_0 = 0\), with Brownian filtration \(\{\mathcal{F}_t\}_{t \geq 0}\) containing all null sets, and \(T \gt 0\) is a fixed horizon. The stage is the space \(L^2(\mathcal{F}_T) = L^2(\Omega, \mathcal{F}_T, \mathbb{P})\) of square-integrable \(\mathcal{F}_T\)-measurable variables, a Hilbert space under \(\langle X, Y \rangle = \mathbb{E}[XY]\). The strategy of the whole page is a density ladder. We show that variables of an extremely special form, exponentials of Itô integrals of deterministic integrands, already fill \(L^2(\mathcal{F}_T)\) densely. Representation is then proved for the special form by a computation and extended to everything by taking limits. This section builds the ladder's two rungs.

What a Measurable Variable Must Look Like

The first tool answers a basic question. If a random variable is measurable with respect to the \(\sigma\)-algebra generated by finitely many others, how must it depend on them? The answer is: as a Borel function, and in no other way. For a random vector \(\mathbf{W} : \Omega \to \mathbb{R}^d\), write \(\sigma(\mathbf{W}) = \{ \mathbf{W}^{-1}(B) : B \in \mathcal{B}(\mathbb{R}^d) \}\) for the \(\sigma\)-algebra it generates, the preimage structure making this a \(\sigma\)-algebra directly.

Theorem: Factorization of Measurable Variables

Let \(\mathbf{W} : \Omega \to \mathbb{R}^d\) be a random vector and \(Y : \Omega \to \mathbb{R}\) a function. Then \(Y\) is \(\sigma(\mathbf{W})\)-measurable if and only if there exists a Borel measurable \(g : \mathbb{R}^d \to \mathbb{R}\) with \(Y = g(\mathbf{W})\).

Proof.

If \(Y = g(\mathbf{W})\) with \(g\) Borel, then for Borel \(E \subseteq \mathbb{R}\) the preimage \(Y^{-1}(E) = \mathbf{W}^{-1}(g^{-1}(E))\) lies in \(\sigma(\mathbf{W})\), which is the easy direction. For the converse, climb from indicators. If \(Y = \mathbf{1}_A\) with \(A \in \sigma(\mathbf{W})\), then \(A = \mathbf{W}^{-1}(B)\) for some Borel \(B\), and \(Y = \mathbf{1}_B(\mathbf{W})\). If \(Y\) is a simple function, a finite linear combination of such indicators, take the same combination of the corresponding \(\mathbf{1}_B\). For general \(\sigma(\mathbf{W})\)-measurable \(Y\), the standard staircase truncations

\[ Y_n = \Bigl( \bigl( 2^{-n} \lfloor 2^{n} Y \rfloor \bigr) \vee (-n) \Bigr) \wedge n \]

are simple, \(\sigma(\mathbf{W})\)-measurable, and converge to \(Y\) pointwise, since \(|Y_n - Y| \leq 2^{-n}\) once \(|Y| \lt n\). Write \(Y_n = g_n(\mathbf{W})\) with \(g_n\) Borel by the simple case, and set \(g(\mathbf{x}) = \limsup_n g_n(\mathbf{x})\) where the limit superior is finite and \(g(\mathbf{x}) = 0\) elsewhere, a Borel function. At every \(\omega\), \(g_n(\mathbf{W}(\omega)) = Y_n(\omega) \to Y(\omega)\), so the limit superior at \(\mathbf{W}(\omega)\) is finite and equals \(Y(\omega)\). Hence \(Y = g(\mathbf{W})\).

Cylinder Functions Are Dense

Call a random variable a cylinder function of the path if it has the form \(\varphi(w_{t_1}, \ldots, w_{t_k})\) for some \(k\), some times \(t_1, \ldots, t_k \in [0, T]\), and some \(\varphi \in C_0^\infty(\mathbb{R}^k)\), a smooth function vanishing outside a bounded set. A cylinder function reads the path through finitely many samples and processes them smoothly. The first rung of the ladder says that nothing in \(L^2(\mathcal{F}_T)\) is out of their reach.

Theorem: Density of Cylinder Functions

The set of cylinder functions is dense in \(L^2(\mathcal{F}_T)\). Every \(g \in L^2(\mathcal{F}_T)\) is the \(L^2\)-limit of a sequence of cylinder functions.

Proof.

Step 1 (Countably many samples suffice).
Let \(t_1, t_2, \ldots\) enumerate a dense subset of \([0, T]\) and let \(\mathcal{H}_k = \sigma(w_{t_1}, \ldots, w_{t_k})\), an increasing family with union generating a \(\sigma\)-algebra \(\mathcal{H}_\infty = \sigma\bigl( \bigcup_k \mathcal{H}_k \bigr)\). We claim \(\mathcal{F}_T\) exceeds \(\mathcal{H}_\infty\) only by null sets. For any \(s \in [0, T]\), pick \(t_{k_i} \to s\) from the dense set. Paths are continuous outside a null set, so \(w_s = \lim_i w_{t_{k_i}}\) almost surely, and a pointwise limit of \(\mathcal{H}_\infty\)-measurable functions, adjusted on the exceptional null set, is measurable with respect to \(\mathcal{H}_\infty \vee \mathcal{N}\), the augmentation by null sets. Since \(\mathcal{F}_T\) is generated by exactly these variables \(w_s\) and the null sets, \(\mathcal{F}_T \subseteq \mathcal{H}_\infty \vee \mathcal{N}\).

Step 2 (Augmentation is invisible to \(L^2\)).
Consider the family \(\mathcal{G}\) of sets \(A \subseteq \Omega\) differing from some \(A' \in \mathcal{H}_\infty\) by a null set, meaning \(A \,\triangle\, A' \in \mathcal{N}\). It contains \(\mathcal{H}_\infty\) and every null set, and it is a \(\sigma\)-algebra. Complements preserve the symmetric difference, and for countable unions \(\bigl( \bigcup A_i \bigr) \triangle \bigl( \bigcup A_i' \bigr) \subseteq \bigcup ( A_i \triangle A_i' )\), a countable union of null sets. Hence \(\mathcal{H}_\infty \vee \mathcal{N} \subseteq \mathcal{G}\), and every \(\mathcal{F}_T\)-measurable indicator equals an \(\mathcal{H}_\infty\)-measurable indicator almost surely. Climbing through simple functions and pointwise limits as in the factorization proof, every \(\mathcal{F}_T\)-measurable variable agrees almost surely with an \(\mathcal{H}_\infty\)-measurable one, so it suffices to approximate \(g \in L^2(\mathcal{H}_\infty)\).

Step 3 (Down to finitely many samples).
Let \(\mathcal{S}\) be the closure in \(L^2(\mathbb{P})\) of \(\bigcup_k L^2(\mathcal{H}_k)\), a closed subspace. The family \(\{ A \in \mathcal{H}_\infty : \mathbf{1}_A \in \mathcal{S} \}\) contains the algebra \(\bigcup_k \mathcal{H}_k\), and it is stable under complements, by linearity and \(\mathbf{1}_{A^c} = 1 - \mathbf{1}_A\), and under increasing countable unions, since \(\mathbf{1}_{A_i} \to \mathbf{1}_{\cup A_i}\) pointwise with the constant dominator one, so in \(L^2\) by the dominated convergence theorem. The passage from an algebra with these closure properties to the full generated \(\sigma\)-algebra is the monotone-class step this track has taken on faith since the construction of the integral, and we extend the same trust to this one application. Simple functions are then in \(\mathcal{S}\) by linearity, and general \(g \in L^2(\mathcal{H}_\infty)\) by the staircase truncations of the factorization proof, the squared differences being dominated by the integrable \(4 g^2\). So \(\mathcal{S} = L^2(\mathcal{H}_\infty)\), and it suffices to approximate \(g \in L^2(\mathcal{H}_k)\) for fixed \(k\).

Step 4 (Smoothing the dependence).
Fix \(k\) and \(g \in L^2(\mathcal{H}_k)\). By the factorization theorem, \(g = g_k(\mathbf{W})\) for a Borel \(g_k : \mathbb{R}^k \to \mathbb{R}\), where \(\mathbf{W} = (w_{t_1}, \ldots, w_{t_k})\), and writing \(\mu\) for the law of \(\mathbf{W}\) on \(\mathbb{R}^k\),

\[ \mathbb{E}\bigl[ \bigl( g_k(\mathbf{W}) - \varphi(\mathbf{W}) \bigr)^2 \bigr] = \int_{\mathbb{R}^k} ( g_k - \varphi )^2\, d\mu , \]

so the problem becomes approximating \(g_k \in L^2(\mu)\) by \(\varphi \in C_0^\infty(\mathbb{R}^k)\) in \(L^2(\mu)\). That smooth compactly supported functions are dense in \(L^2\) of a finite Borel measure on \(\mathbb{R}^k\) is a standard fact of real analysis, by truncation, regularity of the measure, and mollification, and we record it as trusted rather than build the mollifier machinery here. It is the single analytic input of this lemma. Chaining the four steps completes the proof.

Exponentials Are Dense

The second rung replaces smooth cylinder functions by a family with far more structure, the exponentials of Itô integrals with deterministic integrands. For \(h \in L^2([0, T])\) deterministic, the integral \(\int_0^T h\, dw_s\) is defined, the integrand being trivially adapted and square integrable, and we consider

\[ \mathcal{E}(h) = \exp \Bigl( \int_0^T h(s)\, dw_s - \tfrac{1}{2} \int_0^T h(s)^2\, ds \Bigr) . \]

The deterministic correction \(-\tfrac{1}{2} \int h^2\) is a constant factor and does not affect spans or density. It is included now because the next section will reveal it as the Itô correction in disguise, the exact discount that makes \(\mathcal{E}(h)\) a martingale.

Theorem: Density of Exponential Functionals

The linear span of \(\{ \mathcal{E}(h) \}\), with \(h\) ranging over the deterministic step functions on \([0, T]\), is dense in \(L^2(\mathcal{F}_T)\). A fortiori, the span over all deterministic \(h \in L^2([0, T])\) is dense.

Proof.

Step 1 (Reduction to a vanishing pairing).
If the closed span were a proper subspace of the Hilbert space \(L^2(\mathcal{F}_T)\), the projection theorem would produce a nonzero \(g\) orthogonal to every \(\mathcal{E}(h)\). It therefore suffices to show that such orthogonality forces \(g = 0\). We will show it forces \(\mathbb{E}[\varphi(\mathbf{W})\, g] = 0\) for every cylinder function. Then, choosing cylinder functions \(\varphi_n(\mathbf{W}_n) \to g\) in \(L^2\) by the density theorem above, \(0 = \mathbb{E}[\varphi_n(\mathbf{W}_n)\, g] \to \mathbb{E}[g^2]\), and \(g = 0\).

Step 2 (From exponentials of integrals to exponentials of samples).
Fix times \(0 \leq t_1 \lt \cdots \lt t_k \leq T\) and \(\lambda \in \mathbb{R}^k\). The step function \(h = \sum_i c_i \mathbf{1}_{[0, t_i]}\) has \(\int_0^T h\, dw = \sum_i c_i\, w_{t_i}\), by linearity of the integral and the identity \(\int_0^T \mathbf{1}_{[0, t_i]}\, dw = w_{t_i}\), and the coefficients \(c_i\) can be chosen to produce any prescribed \(\lambda_i\) as the total weight on \(w_{t_i}\). Orthogonality to \(\mathcal{E}(h)\), with the deterministic factor \(\exp(\tfrac{1}{2} \int h^2)\) multiplied back in, therefore gives

\[ \begin{align*} G(\lambda) &= \mathbb{E}\Bigl[ \exp\bigl( \lambda_1 w_{t_1} + \cdots + \lambda_k w_{t_k} \bigr)\, g \Bigr] \\\\ &= 0 \quad \forall\, \lambda \in \mathbb{R}^k . \end{align*} \]

Every expectation here is finite. The exponent \(\lambda \cdot \mathbf{W}\) is a linear combination of the Gaussian process, hence a centered Gaussian variable with some variance \(\sigma^2\), and completing the square against the density gives the Gaussian moment generating function,

\[ \begin{align*} \mathbb{E}\bigl[ e^{s Z} \bigr] &= \int_{-\infty}^{\infty} e^{s z}\, \frac{1}{\sqrt{2 \pi \sigma^2}} e^{-z^2 / 2 \sigma^2}\, dz \\\\ &= e^{s^2 \sigma^2 / 2} \int_{-\infty}^{\infty} \frac{1}{\sqrt{2 \pi \sigma^2}} e^{-(z - s \sigma^2)^2 / 2 \sigma^2}\, dz \\\\ &= e^{s^2 \sigma^2 / 2} \end{align*} \]

for \(Z\) centered Gaussian with variance \(\sigma^2 \gt 0\), the degenerate case \(\sigma^2 = 0\) being trivial, the shifted density integrating to one by the Gaussian integral. With the Cauchy-Schwarz inequality, \(\mathbb{E}[e^{\lambda \cdot \mathbf{W}} |g|] \leq \mathbb{E}[e^{2 \lambda \cdot \mathbf{W}}]^{1/2}\, \|g\|_{L^2} \lt \infty\), so \(G\) is well defined, and the display will be recycled twice more on this page.

Step 3 (Two measures with one generating function).
Split \(g = g^+ - g^-\) and define two finite Borel measures on \(\mathbb{R}^k\) as the \(g^\pm\)-weighted laws of the sample vector,

\[ \nu^\pm(B) = \mathbb{E}\bigl[ \mathbf{1}_B(\mathbf{W})\, g^\pm \bigr] , \quad B \in \mathcal{B}(\mathbb{R}^k) , \]

countably additive by the monotone convergence theorem and finite since \(\mathbb{E}[g^\pm] \leq \mathbb{E}[|g|] \leq \|g\|_{L^2}\). The vanishing of \(G\) says exactly that the two measures have the same moment generating function everywhere,

\[ \int_{\mathbb{R}^k} e^{\lambda \cdot x}\, d\nu^+(x) = \int_{\mathbb{R}^k} e^{\lambda \cdot x}\, d\nu^-(x) \quad \forall\, \lambda \in \mathbb{R}^k , \]

the passage from expectations of \(e^{\lambda \cdot \mathbf{W}} g^\pm\) to integrals against \(\nu^\pm\) being the image-measure substitution, run through indicators, simple functions, and monotone limits. We now invoke one fact at statement level. Two finite Borel measures on \(\mathbb{R}^k\) whose moment generating functions exist and agree on all of \(\mathbb{R}^k\) are equal. This uniqueness theorem sits at the same epistemic tier as the continuity theorem already recorded without proof on the convergence page, its standard proofs running through analytic continuation or Fourier methods that this track has not built, and it is the single trusted input of this lemma. Granting it, \(\nu^+ = \nu^-\).

Step 4 (Back up the ladder).
Equal measures integrate every bounded Borel function equally, so for every \(\varphi \in C_0^\infty(\mathbb{R}^k)\),

\[ \begin{align*} \mathbb{E}\bigl[ \varphi(\mathbf{W})\, g \bigr] &= \int \varphi\, d\nu^+ - \int \varphi\, d\nu^- \\\\ &= 0 . \end{align*} \]

The times and \(k\) were arbitrary, so \(g\) is orthogonal to every cylinder function, and Step 1 closes the argument.

The ladder is built. Anything in \(L^2(\mathcal{F}_T)\) can be reached by linear combinations of the exponentials \(\mathcal{E}(h)\), and the entire representation problem now reduces to a single question. Can one exponential be written as a constant plus an Itô integral? The next section answers it with a computation the calculus pages have made routine.

The Exponential Martingale

The density theorem hands the representation problem a dense family, and the family was chosen because its members can be represented by a single computation. Two facts are needed. The law of \(\int_0^t h\, dw\) for deterministic \(h\), which controls every integrability question, and the Itô differential of \(\mathcal{E}(h)\), which is the representation itself. The Gaussian law below holds for every deterministic \(h \in L^2([0, T])\). The exponential computation that follows it is run for step functions, with values \(c_i\) on the intervals \([r_i, r_{i+1})\) of a partition \(0 = r_0 \lt \cdots \lt r_m = T\), which is all the density theorem requires.

Theorem: Integrals of Deterministic Integrands Are Gaussian

Let \(h \in L^2([0, T])\) be deterministic and \(t \in [0, T]\). Then \(I_t = \int_0^t h(s)\, dw_s\) is a centered Gaussian random variable with variance \(\int_0^t h(s)^2\, ds\), and for every \(s \in \mathbb{R}\),

\[ \mathbb{E}\bigl[ e^{s I_t} \bigr] = \exp \Bigl( \tfrac{s^2}{2} \int_0^t h(r)^2\, dr \Bigr) . \]

Proof.

Step 1 (Step integrands, exactly).
Suppose first that \(h\) is a step function. Then \(I_t = \sum_i c_i (w_{r_{i+1} \wedge t} - w_{r_i \wedge t})\) is a finite linear combination of increments over disjoint intervals. Its moment generating function factors over the increments, which are independent of the past so the expectation of the product peels one factor at a time from the last increment backwards, by the product rule, each peeled factor being a Gaussian moment generating function computed by the completed square of the previous section. Multiplying,

\[ \begin{align*} \mathbb{E}\bigl[ e^{s I_t} \bigr] &= \prod_i \exp \Bigl( \tfrac{s^2 c_i^2}{2} \bigl( r_{i+1} \wedge t - r_i \wedge t \bigr) \Bigr) \\\\ &= \exp \Bigl( \tfrac{s^2}{2} \int_0^t h(r)^2\, dr \Bigr) . \end{align*} \]

Step 2 (General integrands by \(L^1\) transfer).
For general deterministic \(h \in L^2([0, T])\), the density input already trusted in the cylinder lemma provides smooth compactly supported approximants of \(h\) in \(L^2([0, t])\), and left-endpoint discretization, exact in \(L^2\) up to the modulus of uniform continuity, turns those into step functions, so there are step functions \(h_n \to h\) in \(L^2([0, t])\). The Itô isometry gives \(I^{(n)} = \int_0^t h_n\, dw \to I_t\) in \(L^2(\mathbb{P})\). Fix \(s \in \mathbb{R}\). The elementary bound \(|e^{a} - e^{b}| \leq |a - b|\, (e^{a} + e^{b})\), from the mean value of the exponential between \(a\) and \(b\), and the Cauchy-Schwarz inequality give

\[ \Bigl| \mathbb{E}\bigl[ e^{s I^{(n)}} \bigr] - \mathbb{E}\bigl[ e^{s I_t} \bigr] \Bigr| \leq |s|\, \bigl\| I^{(n)} - I_t \bigr\|_{L^2} \Bigl( \bigl\| e^{s I^{(n)}} \bigr\|_{L^2} + \bigl\| e^{s I_t} \bigr\|_{L^2} \Bigr) . \]

The first factor tends to zero. The norms are bounded along the sequence, since \(\mathbb{E}[ e^{2 s I^{(n)}} ] = \exp( 2 s^2 \int_0^t h_n^2 )\) converges by Step 1, and \(\mathbb{E}[ e^{2 s I_t} ] \lt \infty\) because along an almost surely convergent subsequence Fatou's lemma bounds it by the limit of the exact values. So \[ \begin{align*} \mathbb{E}[ e^{s I_t} ] &= \lim_n \exp( \tfrac{s^2}{2} \int_0^t h_n^2 ) \\\\ &= \exp( \tfrac{s^2}{2} \int_0^t h^2 ) . \end{align*} \]

Step 3 (Identifying the law).
The law of \(I_t\) and the centered Gaussian law with variance \(\sigma^2 = \int_0^t h^2\, dr\) are two finite Borel measures whose moment generating functions exist and agree everywhere, the Gaussian side by the completed square. The uniqueness of finite measures with a common moment generating function, the input already trusted in the density lemma, identifies the two laws, and \(I_t\) is centered Gaussian with variance \(\sigma^2\).

Now the main computation. For a deterministic step function \(h\), define the process

Definition: Exponential Martingale

For deterministic \(h \in L^2([0, T])\), the exponential martingale of \(h\) is the process

\[ Z_t = \exp \Bigl( \int_0^t h(s)\, dw_s - \tfrac{1}{2} \int_0^t h(s)^2\, ds \Bigr) , \quad 0 \leq t \leq T , \]

so that \(Z_0 = 1\) and \(Z_T = \mathcal{E}(h)\), the \(dw\)-term taken in its continuous realization.

The name is a claim, and the claim is the theorem. The deterministic discount \(-\tfrac{1}{2} \int h^2\), carried unexplained through the density section, is exactly the Itô correction of the exponential, engineered so that the drift of \(Z\) cancels to zero.

Theorem: The Exponential Martingale Identity

Let \(h\) be a deterministic step function on \([0, T]\) and \(Z\) its exponential martingale. Then \(h(s) Z_s \in \mathcal{V}(0, T)\), and for every \(t \in [0, T]\), almost surely,

\[ \begin{align*} Z_t &= 1 + \int_0^t h(s)\, Z_s\, dw_s , \\\\ dZ_t &= h(t)\, Z_t\, dw_t . \end{align*} \]

In particular \(Z\) is a martingale with \(\mathbb{E}[Z_t] = 1\) for all \(t\).

Proof.

Step 1 (Membership).
The process \(h(s) Z_s\) is progressively measurable, being a product of a deterministic step function and a continuous adapted process, and by the Gaussian theorem, with \(I_s = \int_0^s h\, dw\),

\[ \begin{align*} \mathbb{E}\bigl[ Z_s^2 \bigr] &= e^{- \int_0^s h^2}\, \mathbb{E}\bigl[ e^{2 I_s} \bigr] \\\\ &= e^{- \int_0^s h^2}\, e^{2 \int_0^s h^2} \\\\ &= e^{\int_0^s h^2} \\\\ &\leq e^{\|h\|_{L^2}^2} , \end{align*} \]

so \(\mathbb{E}\bigl[ \int_0^T h^2 Z_s^2\, ds \bigr] \leq \|h\|_{L^2}^2\, e^{\|h\|_{L^2}^2} \lt \infty\), and \(h Z \in \mathcal{V}(0, T)\).

Step 2 (The identity, one piece at a time).
The process \(X_t = I_t\) is an Itô process with \(u = 0\) and \(v = h\). The natural map \(g(t, x) = \exp( x - \tfrac{1}{2} \int_0^t h^2 )\) is not \(C^2\) in \(t\), its time derivative jumping where \(h\) jumps, so we work one interval of the partition at a time with a smooth stand-in. Fix a piece \([r_i, r_{i+1}]\), on which \(h \equiv c_i\), and define

\[ \begin{align*} g_i(t, x) &= \exp \Bigl( x - A(r_i) - \tfrac{c_i^2}{2} ( t - r_i ) \Bigr) , \\\\ A(r) &= \tfrac{1}{2} \int_0^{r} h^2\, ds , \end{align*} \]

a \(C^2\) function on all of \([0, \infty) \times \mathbb{R}\) that agrees with the natural map for \(t \in [r_i, r_{i+1}]\), where \(g_i(t, I_t) = Z_t\). Its membership condition holds, the diffusion entry \(h(s)\, g_i(s, I_s)\) differing from \(h Z\) by the bounded deterministic factor \(e^{A(s) - A(r_i) - c_i^2 (s - r_i)/2}\), so the Itô formula applies to \(g_i\). Its drift integrand is

\[ \frac{\partial g_i}{\partial t} + \tfrac{1}{2}\, h(s)^2\, \frac{\partial^2 g_i}{\partial x^2} = \Bigl( - \tfrac{c_i^2}{2} + \tfrac{h(s)^2}{2} \Bigr)\, g_i(s, I_s) , \]

which vanishes identically for \(s \in [r_i, r_{i+1})\), where \(h(s) = c_i\). This is the discount doing its work. The half squared integrand subtracted in the exponent is consumed, to the last drop, by the half second derivative the chain rule inserts. Applying the formula at times \(r_i\) and \(t \in [r_i, r_{i+1}]\) and subtracting, the \(ds\)-integrals over \([r_i, t]\) carry only the vanishing drift, and

\[ \begin{align*} Z_t - Z_{r_i} &= \int_{r_i}^{t} h(s)\, g_i(s, I_s)\, dw_s \\\\ &= \int_{r_i}^{t} h(s)\, Z_s\, dw_s , \end{align*} \]

the integrands agreeing on the interval of integration and the interval splitting being the additivity of the integral. Telescoping over the pieces up to \(t\), with \(Z_0 = 1\), yields the identity, one null set per piece and finitely many pieces.

Step 3 (Martingale).
With \(h Z \in \mathcal{V}(0, T)\), the process \(Z_t - 1\) is an Itô integral process and hence a martingale with zero mean, so \(Z\) is a martingale with \(\mathbb{E}[Z_t] = 1\).

This is the first representation. The exponential \(\mathcal{E}(h) = Z_T\) equals its mean, one, plus the Itô integral of the explicit integrand \(h Z \in \mathcal{V}(0, T)\). The density theorem says such exponentials span everything, and the next section turns the span and a limit into the general statement.

The Itô Representation Theorem

The pieces assemble. Exponentials are representable by an explicit computation, exponentials span a dense subspace, and the Itô isometry makes the passage to limits not merely possible but automatic, because it converts convergence of the represented variables into convergence of the representing integrands. One point deserves care before the proof, and it is where the progressive machinery of the formula page pays an unadvertised dividend. Limits of integrands must again be legitimate integrands, and the class \(\mathcal{V}(0, T)\) must be closed under \(L^2\)-limits. The membership conditions of \(\mathcal{V}\) constrain measurability, not just size, and an \(L^2\)-limit is only an equivalence class, so closedness needs an argument, supplied inside the proof by a limit superior of progressively measurable representatives.

Theorem: The Itô Representation Theorem

Let \(F \in L^2(\mathcal{F}_T)\). Then there exists a progressively measurable \(f \in \mathcal{V}(0, T)\), unique up to equality \(\lambda \otimes \mathbb{P}\)-almost everywhere, such that, almost surely,

\[ F = \mathbb{E}[F] + \int_0^T f(s, \omega)\, dw_s \]

Proof.

Step 1 (The dense subspace is represented).
For a single exponential, the previous section gives \(\mathcal{E}(h) = 1 + \int_0^T h\, Z^{h}_s\, dw_s\) with \(h Z^{h} \in \mathcal{V}(0, T)\) and \(\mathbb{E}[\mathcal{E}(h)] = 1\), writing \(Z^h\) for the exponential martingale of \(h\). For a finite linear combination \(F_0 = \sum_j a_j\, \mathcal{E}(h_j)\) with step functions \(h_j\), linearity of the integral gives

\[ \begin{align*} F_0 &= \sum_j a_j + \int_0^T \Bigl( \sum_j a_j\, h_j(s)\, Z^{h_j}_s \Bigr) dw_s \\\\ &= \mathbb{E}[F_0] + \int_0^T f_0(s, \omega)\, dw_s , \end{align*} \]

the mean identification using \(\mathbb{E}[\mathcal{E}(h_j)] = 1\), and the integrand \(f_0\), a finite combination of progressively measurable members of \(\mathcal{V}(0, T)\), lies in \(\mathcal{V}(0, T)\) and is progressively measurable.

Step 2 (The isometry transports the Cauchy property).
Let \(F \in L^2(\mathcal{F}_T)\) and, by the density theorem, choose span elements \(F_n \to F\) in \(L^2(\mathbb{P})\), each represented as in Step 1 by a progressively measurable \(f_n \in \mathcal{V}(0, T)\). Means converge, \(|\mathbb{E}[F_n] - \mathbb{E}[F]| \leq \mathbb{E}[|F_n - F|] \leq \|F_n - F\|_{L^2}\) by the Cauchy-Schwarz inequality against the constant one. Subtracting two representations and applying the Itô isometry, with a Tonelli exchange writing the squared norm as an iterated integral,

\[ \begin{align*} \int_0^T \mathbb{E}\bigl[ ( f_n - f_m )^2 \bigr]\, ds &= \mathbb{E}\Bigl[ \Bigl( \int_0^T ( f_n - f_m )\, dw_s \Bigr)^{2} \Bigr] \\\\ &= \mathbb{E}\Bigl[ \bigl( F_n - F_m - \mathbb{E}[ F_n - F_m ] \bigr)^{2} \Bigr] \\\\ &\leq \mathbb{E}\bigl[ ( F_n - F_m )^2 \bigr] , \end{align*} \]

the last inequality because subtracting the mean can only decrease the second moment, \(\mathbb{E}[(X - \mathbb{E}X)^2] = \mathbb{E}[X^2] - (\mathbb{E}X)^2\). The right side tends to zero as \(n, m \to \infty\), so \((f_n)\) is a Cauchy sequence in \(L^2([0, T] \times \Omega,\ \lambda \otimes \mathbb{P})\), and by completeness it converges to some \(f\) in that space.

Step 3 (The limit is a legitimate integrand).
Along a subsequence, \(f_{n_k} \to f\) at \(\lambda \otimes \mathbb{P}\)-almost every point of \([0, T] \times \Omega\). Define

\[ \tilde{f}(s, \omega) = \limsup_{k}\, f_{n_k}(s, \omega) , \]

set to zero where the limit superior fails to be finite. Each \(f_{n_k}\) restricted to \([0, t] \times \Omega\) is \(\mathcal{B}([0, t]) \otimes \mathcal{F}_t\)-measurable, limit superiors preserve measurability, and the finiteness set is measurable, so \(\tilde{f}\) is progressively measurable. At almost every point \(\tilde{f} = f\), so \(\tilde{f}\) lies in the same \(L^2\)-class, is jointly measurable and adapted through progressivity, and is square integrable. Hence \(\tilde{f} \in \mathcal{V}(0, T)\), and we rename it \(f\). The class \(\mathcal{V}(0, T)\) is, in this sense, closed in \(L^2(\lambda \otimes \mathbb{P})\).

Step 4 (Passing to the limit).
By the isometry once more, \(\int_0^T f_n\, dw \to \int_0^T f\, dw\) in \(L^2(\mathbb{P})\), and combining the three convergences,

\[ \begin{align*} F &= \lim_n F_n \\\\ &= \lim_n \Bigl( \mathbb{E}[F_n] + \int_0^T f_n\, dw_s \Bigr) \\\\ &= \mathbb{E}[F] + \int_0^T f\, dw_s , \end{align*} \]

the limits taken in \(L^2(\mathbb{P})\) and the identity holding almost surely.

Step 5 (Uniqueness).
If \(f_1, f_2 \in \mathcal{V}(0, T)\) both represent \(F\), subtracting the two identities gives \(\int_0^T (f_1 - f_2)\, dw_s = 0\) almost surely, and the isometry, with the same Tonelli exchange, turns the vanishing of the integral into the vanishing of the integrand,

\[ \begin{align*} 0 &= \mathbb{E}\Bigl[ \Bigl( \int_0^T ( f_1 - f_2 )\, dw_s \Bigr)^{2} \Bigr] \\\\ &= \int_0^T \mathbb{E}\bigl[ ( f_1 - f_2 )^2 \bigr]\, ds , \end{align*} \]

so \(f_1 = f_2\) at \(\lambda \otimes \mathbb{P}\)-almost every point.

The theorem repays a moment of astonishment before it is put to work. Nothing about \(F\) was assumed beyond square integrability and measurability with respect to the Brownian history. Any such quantity, the maximum of the path, the time it spends above a level, an arbitrarily contorted functional of the trajectory, is a constant plus a stochastic integral, synthesized from the same noise it observes by some adapted, square-integrable recipe \(f\). The theorem asserts the recipe exists and is essentially unique, and it does not say how to find it. Producing \(f\) explicitly is its own industry, and the door it opens for mathematical finance, where \(f\) is a hedging strategy, is a signpost the closing section records.

The Martingale Representation Theorem

The representation theorem freezes time at a single horizon \(T\). The final step lets the horizon run, and represents entire martingales rather than single variables. Two of the earlier pages meet here. The representation supplies an integrand for each fixed time, and the martingale property of the Itô integral forces the integrands at different times to agree, so that one recipe serves all horizons at once.

Theorem: The Martingale Representation Theorem

Let \(\{M_t\}_{t \geq 0}\) be a martingale with respect to the Brownian filtration \(\{\mathcal{F}_t\}\) with \(M_t \in L^2(\mathbb{P})\) for every \(t\). Then there exists a progressively measurable process \(g\) with \(g \in \mathcal{V}(0, t)\) for every \(t \gt 0\), unique up to \(\lambda \otimes \mathbb{P}\)-null sets, such that for every \(t \geq 0\), almost surely,

\[ M_t = \mathbb{E}[M_0] + \int_0^t g(s, \omega)\, dw_s . \]

Proof.

Step 1 (A representation at every horizon).
Fix \(t \gt 0\). The variable \(M_t\) lies in \(L^2(\mathcal{F}_t)\), so the representation theorem with horizon \(t\) supplies a progressively measurable \(f^{(t)} \in \mathcal{V}(0, t)\), unique almost everywhere, with

\[ M_t = \mathbb{E}[M_t] + \int_0^t f^{(t)}(s, \omega)\, dw_s , \]

The mean is constant in \(t\). By the tower property and the martingale identity, \(\mathbb{E}[M_t] = \mathbb{E}\bigl[ \mathbb{E}[M_t \mid \mathcal{F}_0] \bigr] = \mathbb{E}[M_0]\).

Step 2 (Consistency of the recipes).
Let \(0 \lt t_1 \lt t_2\). Conditioning the horizon-\(t_2\) representation on \(\mathcal{F}_{t_1}\), the martingale identity for \(M\) on the left and the martingale property of the Itô integral on the right give, almost surely,

\[ \begin{align*} M_{t_1} &= \mathbb{E}\bigl[ M_{t_2} \mid \mathcal{F}_{t_1} \bigr] \\\\ &= \mathbb{E}[M_0] + \mathbb{E}\Bigl[ \int_0^{t_2} f^{(t_2)}\, dw_s \,\Bigm|\, \mathcal{F}_{t_1} \Bigr] \\\\ &= \mathbb{E}[M_0] + \int_0^{t_1} f^{(t_2)}\, dw_s . \end{align*} \]

The restriction of \(f^{(t_2)}\) to \([0, t_1]\) lies in \(\mathcal{V}(0, t_1)\) and represents \(M_{t_1}\), as does \(f^{(t_1)}\), so the uniqueness clause of the representation theorem forces, at \(\lambda \otimes \mathbb{P}\)-almost every point of \([0, t_1] \times \Omega\),

\[ f^{(t_2)} = f^{(t_1)} \]

Recipes for longer horizons extend recipes for shorter ones. Nothing needs to be recomputed as time grows.

Step 3 (One process for all horizons).
Define \(g\) by gluing along integer horizons,

\[ g(s, \omega) = f^{(n)}(s, \omega) , \quad s \in (n - 1, n] ,\ n \in \mathbb{N} , \]

and \(g(0, \cdot) = f^{(1)}(0, \cdot)\). Each piece is the product of a progressively measurable process and a deterministic indicator, and the pieces live on disjoint time intervals, so \(g\) is progressively measurable. Fix \(t \gt 0\) and let \(N = \lceil t \rceil\). By Step 2, chained from each integer \(n \leq N\) up to \(N\), the process \(g\) agrees with \(f^{(N)}\) at almost every point of \([0, t] \times \Omega\), so \(g \in \mathcal{V}(0, t)\) and, integrands equal almost everywhere having equal integrals by the isometry,

\[ \begin{align*} \int_0^t g\, dw_s &= \int_0^t f^{(N)}\, dw_s \\\\ &= M_t - \mathbb{E}[M_t] \\\\ &= M_t - \mathbb{E}[M_0] \end{align*} \]

almost surely, where the middle equality is Step 2 applied to the pair \((t, N)\). This is the claimed identity for \(t \gt 0\). At \(t = 0\) the identity reads \(M_0 = \mathbb{E}[M_0]\), which holds almost surely because \(\mathcal{F}_0\), being generated by \(w_0 = 0\) and the null sets, consists of null sets and their complements, so \(\mathcal{F}_0\)-measurable variables are almost surely constant.

Step 4 (Uniqueness).
If \(g\) and \(g'\) both work, then for every \(t\) they both represent \(M_t\) over \([0, t]\), so they agree \(\lambda \otimes \mathbb{P}\)-almost everywhere on \([0, t] \times \Omega\) by the representation theorem's uniqueness, and letting \(t \to \infty\) through the integers, almost everywhere on \([0, \infty) \times \Omega\).

The time-zero case that closed the proof deserves its own sentence, because it is a statement about information, not about integrals. A martingale of the Brownian filtration cannot even randomize its starting value. Every square-integrable fair game riding on Brownian information is, in full, a constant plus a running Itô integral.

Two remarks close the chapter's accounts. First, dimensions. The same ladder, with sample vectors drawn from all components, componentwise exponentials, and the matrix-valued integrands of the multi-dimensional integral, delivers the \(n\)-dimensional statements, every \(L^2\) functional of an \(n\)-dimensional Brownian path being a constant plus \(\sum_j \int f_j\, dw^{(j)}\). We record the vector version at statement level, the argument repeating rather than deepening. Second, what was trusted. This page consumed exactly two statement-level inputs beyond the track's standing conventions, the density of smooth compactly supported functions in \(L^2\) of a finite measure, and the uniqueness of finite measures through their moment generating functions. Everything else, including the convergence theorem for conditional expectations that the classical route runs through, was either proved or bypassed.

The Recipe Is the Product

The representation theorem is an existence statement with an industry attached to its constructive gap. In mathematical finance the theorem is the reason Brownian market models are complete. A claim paying \(F\) at time \(T\) is a constant plus \(\int_0^T f\, dw\), the constant is the price and the integrand \(f\) is the hedging strategy, the recipe for replicating the claim by trading continuously against the noise. The theorem guarantees the hedge exists and says nothing about its formula, and computing or approximating \(f\), by partial differential equations, by Malliavin-type calculus, or by training a neural network to output the integrand directly, is a live boundary between stochastic analysis and machine learning.

The chapter that began with a chain rule ends with an inventory. Over its own filtration, Brownian motion is not one source of randomness among many but the only one, every observable a constant plus an integral against it, every fair game a running such integral. The calculus, the conversion, and now the representation are the complete toolkit for the next destination, equations that define a process implicitly through its own noise, where the coefficients of the unknown appear inside the integrals that build it.