What the Construction Already Knows
The previous page built the
Itô integral
\(\int_S^T f(t, \omega)\, dw_t\) for integrands \(f\) in the class
\(\mathcal{V}(S, T)\),
as the \(L^2(\mathbb{P})\) limit of integrals of bounded
elementary processes,
and it closed with a promise: viewed as a process in its upper limit of integration, the integral is
a martingale and admits a continuous version. This page keeps that promise. Throughout, \(w_t\) is
the standard one-dimensional Brownian motion of the last two pages, carried on
\((\Omega, \mathcal{F}, \mathbb{P})\), and \(\{\mathcal{F}_t\}_{t \geq 0}\) is the
Brownian filtration,
with its convention that every set of \(\mathbb{P}\)-measure zero belongs to each
\(\mathcal{F}_t\); the norm \(\|f\|_{\mathcal{V}}^2 = \mathbb{E}\bigl[\int_S^T f^2\, dt\bigr]\) is
the one along which elementary processes approximate general integrands.
Before the process view, we collect what the fixed-interval construction already implies. The four
properties below are routinely summarized as "they hold for elementary processes, hence in the
limit" — but each one in fact crosses the limit by a different mechanism: the first by an exact
identity between approximating sums, the second by the linear structure of \(L^2\) limits, the
third by convergence of expectations, and the fourth by the measurability of almost-everywhere
limits. We cross each bridge in full, because two of them collect a toll: the third needs a norm
comparison between \(L^1\) and \(L^2\), and the fourth needs precisely the null-set convention
quoted above, adopted on the previous page and cashed in here for the first time.
Theorem: Properties of the Itô Integral
Let \(0 \leq S \lt U \lt T\), let \(f, g \in \mathcal{V}(S, T)\), and let \(a, b \in \mathbb{R}\).
Then \(f \in \mathcal{V}(S, U)\) and \(f \in \mathcal{V}(U, T)\), so that all integrals below
are defined, and the following hold almost surely.
(i) Additivity over intervals.
\(\displaystyle \int_S^T f\, dw_t = \int_S^U f\, dw_t + \int_U^T f\, dw_t\).
(ii) Linearity.
\(\displaystyle \int_S^T (a f + b g)\, dw_t = a \int_S^T f\, dw_t + b \int_S^T g\, dw_t\).
(iii) Zero mean.
\(\displaystyle \mathbb{E}\Bigl[ \int_S^T f\, dw_t \Bigr] = 0\).
(iv) Measurability. Every representative of the class
\(\int_S^T f\, dw_t \in L^2(\mathbb{P})\) is \(\mathcal{F}_T\)-measurable.
Proof
First the opening claim. Conditions (V1) and (V2) defining
\(\mathcal{V}\)
constrain \(f\) as a function on all of \([0, \infty) \times \Omega\) and do not mention the
interval; only (V3) does, and since \(f^2 \geq 0\), the integral of \(f^2\) over
\([S, U]\) or \([U, T]\) is bounded by the integral over \([S, T]\). Hence membership in
\(\mathcal{V}(S, T)\) implies membership in \(\mathcal{V}(S, U)\) and \(\mathcal{V}(U, T)\).
Throughout the proof, fix a sequence \(\{\phi_n\}\) of bounded elementary processes with
\(\|f - \phi_n\|_{\mathcal{V}} \to 0\) on \([S, T]\); such a sequence exists by the three
approximation steps of the previous page.
(i). Refine the partition of each \(\phi_n\) so that it contains the point
\(U\); as noted with the definition of the elementary integral, refinement changes neither the
function nor its integral. Now split each \(\phi_n\) at \(U\):
\[
\psi_n = \phi_n\, \mathbf{1}_{[S, U)} ,
\qquad
\psi_n' = \phi_n\, \mathbf{1}_{[U, T)} .
\]
The process \(\psi_n\) is bounded elementary on \([S, U]\): its partition is the part of the
refined partition up to \(U\), its levels are those of \(\phi_n\), and multiplication by a
deterministic indicator disturbs neither joint measurability nor adaptedness. The same holds
for \(\psi_n'\) on \([U, T]\). Splitting the defining sum of the elementary integral at \(U\)
gives, for every \(\omega\),
\[
\int_S^T \phi_n\, dw_t = \int_S^U \psi_n\, dw_t + \int_U^T \psi_n'\, dw_t .
\]
The sequences are admissible for \(f\) on the subintervals: since \(\psi_n = \phi_n\) on
\([S, U)\) and the endpoint \(\{U\}\) has Lebesgue measure zero,
\[
\mathbb{E}\Bigl[ \int_S^U (f - \psi_n)^2\, dt \Bigr]
= \mathbb{E}\Bigl[ \int_S^U (f - \phi_n)^2\, dt \Bigr]
\leq \|f - \phi_n\|_{\mathcal{V}}^2 \longrightarrow 0 ,
\]
and likewise on \([U, T]\). By the definition of the integral, the three integrals of
\(\phi_n\), \(\psi_n\), \(\psi_n'\) converge in \(L^2(\mathbb{P})\) to the three integrals of
\(f\) over \([S, T]\), \([S, U]\), \([U, T]\) respectively. Limits in \(L^2\) respect sums,
and two \(L^2\) limits of the same sequence agree almost surely, so the displayed identity
survives the passage: (i) holds almost surely.
(ii). This was established in the course of proving that the
definition of the integral
is legitimate: approximating \(f\) and \(g\) separately, combining the approximants linearly,
and using linearity of the elementary integral together with linearity of \(L^2\) limits. We
record the statement here for reference and add nothing to that argument.
(iii). For a bounded elementary process
\(\phi_n = \sum_j e_j\, \mathbf{1}_{[t_j, t_{j+1})}\), expectation is linear over the finite
defining sum by the
properties of expectation,
so
\[
\mathbb{E}\Bigl[ \int_S^T \phi_n\, dw_t \Bigr]
= \sum_{j} \mathbb{E}\bigl[ e_j\, \Delta w_j \bigr] ,
\qquad \Delta w_j = w_{t_{j+1}} - w_{t_j} .
\]
Each term vanishes by the argument that powered the elementary isometry: \(e_j\) is
\(\mathcal{F}_{t_j}\)-measurable and bounded, the increment \(\Delta w_j\) is
independent of \(\mathcal{F}_{t_j}\),
the
expectation of the product factors,
and increments are centered by (W1) of the
definition of Brownian motion.
Hence \(\mathbb{E}\bigl[ \int_S^T \phi_n\, dw_t \bigr] = 0\) for every \(n\). For the limit,
compare the norms of \(L^1\) and \(L^2\): writing
\(X_n = \int_S^T f\, dw_t - \int_S^T \phi_n\, dw_t\), the triangle inequality for expectation
and the
Cauchy-Schwarz inequality
in the inner product space \(L^2(\mathbb{P})\), applied against the constant function
\(\mathbf{1}\) with \(\|\mathbf{1}\|_{L^2(\mathbb{P})} = 1\), give, the subtracted term
having expectation zero by the above,
\[
\Bigl| \mathbb{E}\Bigl[ \int_S^T f\, dw_t \Bigr] \Bigr|
= \bigl| \mathbb{E}[X_n] \bigr|
\leq \mathbb{E}\bigl[ |X_n| \bigr]
\leq \| X_n \|_{L^2(\mathbb{P})}
\longrightarrow 0 ,
\]
the convergence being the definition of the integral. The left side does not depend on \(n\),
so it is zero.
(iv). Two steps: exhibit one \(\mathcal{F}_T\)-measurable representative, then
extend to all of them. Each elementary integral \(\sum_j e_j\, \Delta w_j\) is a genuine
function on \(\Omega\), and it is \(\mathcal{F}_T\)-measurable: every level \(e_j\) is
\(\mathcal{F}_{t_j}\)-measurable with \(\mathcal{F}_{t_j} \subseteq \mathcal{F}_T\), every
increment is a difference of values of the path at times at most \(T\), each
\(\mathcal{F}_T\)-measurable by the definition of the
Brownian filtration,
and sums and products of \(\mathcal{F}_T\)-measurable functions are
\(\mathcal{F}_T\)-measurable, the closure already used on the previous page. Since
\(\int_S^T \phi_n\, dw_t \to \int_S^T f\, dw_t\) in \(L^2(\mathbb{P})\), a
subsequence converges almost surely;
write \(g_k = \int_S^T \phi_{n_k}\, dw_t\) for it, and let \(\widetilde{X}\) equal
\(\lim_k g_k\) where that limit exists in \(\mathbb{R}\), and \(0\) elsewhere. Measurability
survives countable suprema and infima — for instance
\(\{\sup_k g_k \leq a\} = \bigcap_k \{g_k \leq a\}\) — so \(\limsup_k g_k\) and
\(\liminf_k g_k\) are \(\mathcal{F}_T\)-measurable maps into \([-\infty, \infty]\), the set
where they agree and are finite lies in \(\mathcal{F}_T\), and \(\widetilde{X}\) is
\(\mathcal{F}_T\)-measurable. Being the almost-sure limit of a subsequence converging in
\(L^2\), \(\widetilde{X}\) is a representative of \(\int_S^T f\, dw_t\).
Now let \(Y\) be any representative. As an element of \(L^2(\mathbb{P})\) it is
\(\mathcal{F}\)-measurable, and \(Y = \widetilde{X}\) off a \(\mathbb{P}\)-null set. For each
\(a \in \mathbb{R}\), the symmetric difference
\(D_a = \{Y \leq a\} \,\triangle\, \{\widetilde{X} \leq a\}\) is an
\(\mathcal{F}\)-measurable subset of \(\{Y \neq \widetilde{X}\}\), hence a set of
\(\mathbb{P}\)-measure zero, hence a member of \(\mathcal{F}_T\) — this is the moment the
null-set convention adopted with the Brownian filtration earns its keep. Since
\(A \,\triangle\, (A \,\triangle\, B) = B\) for any sets \(A, B\), we conclude
\(\{Y \leq a\} = \{\widetilde{X} \leq a\} \,\triangle\, D_a \in \mathcal{F}_T\), and \(Y\) is
\(\mathcal{F}_T\)-measurable.
Property (i) is the license to let the upper endpoint move. For \(f \in \mathcal{V}(0, T)\) and
\(t \in [0, T]\), the integral \(\int_0^t f\, dw_s\) is defined, and additivity makes the family
of integrals over nested intervals consistent: the integral up to \(t\) plus the increment from
\(t\) to \(t'\) is the integral up to \(t'\). Writing
\[
M_t(\omega) = \int_0^t f(s, \omega)\, dw_s ,
\qquad 0 \leq t \leq T ,
\]
we thus obtain not a single random variable but a stochastic process. Property (iv), applied on
the interval \([0, t]\), says that \(M_t\) is \(\mathcal{F}_t\)-measurable for every \(t\) — the
process \(\{M_t\}\) is \(\mathcal{F}_t\)-adapted, inheriting the causality of its integrand — and
property (iii) says that each \(M_t\) is centered. A centered, adapted process whose increments
are built from noise that no one can foresee: the next section gives this combination its name,
and proves that \(\{M_t\}\) deserves it.
Processes That Forecast Themselves
The process \(\{M_t\}\) assembled at the end of the last section is adapted, centered, and driven
by increments that the present cannot foresee. There is a classical name for processes with this
character. Its defining property is a statement about prediction: given everything observable now,
the best estimate of the process at any future time is its value now. Gains and losses are
possible, but neither is favored — the process models a fair game. The definition is stated
relative to an arbitrary
filtration,
not only the Brownian one, because the notion belongs to any setting where information accumulates
over time; the prediction itself is the
conditional expectation
given the current \(\sigma\)-algebra.
Definition: Martingale
Let \(\{\mathcal{M}_t\}_{t \geq 0}\) be a filtration on \((\Omega, \mathcal{F}, \mathbb{P})\).
A real-valued stochastic process \(\{X_t\}_{t \geq 0}\) is a martingale with
respect to \(\{\mathcal{M}_t\}\) (and with respect to \(\mathbb{P}\)) if:
(M1) the process \(\{X_t\}\) is
\(\mathcal{M}_t\)-adapted;
(M2) \(\mathbb{E}\bigl[ |X_t| \bigr] \lt \infty\) for every \(t \geq 0\);
(M3) \(\mathbb{E}\bigl[ X_s \mid \mathcal{M}_t \bigr] = X_t\) almost surely,
for all \(0 \leq t \leq s\).
Condition (M2) makes the conditional expectation in (M3) meaningful. The definition applies
verbatim when the time index ranges over a bounded interval \([0, T]\), the case our integrals
inhabit.
Two immediate readings. First, applying the averaging identity that defines conditional
expectation with \(A = \Omega\), condition (M3) gives
\(\mathbb{E}[X_s] = \mathbb{E}\bigl[ \mathbb{E}[X_s \mid \mathcal{M}_t] \bigr] = \mathbb{E}[X_t]\):
a martingale has constant expectation. Second, and stronger, (M3) pins down not just the average
over all of \(\Omega\) but the average over every event observable at time \(t\) — on each such
event, the future is expected to go nowhere. Constancy of expectation is what remains of the
martingale property when the information structure is thrown away.
The prototype is Brownian motion itself, with respect to its own filtration. The proof is a
dividend of machinery already in place: every step is a citation.
Theorem: Brownian Motion Is a Martingale
The process \(\{w_t\}_{t \geq 0}\) is a martingale with respect to the Brownian filtration
\(\{\mathcal{F}_t\}_{t \geq 0}\).
Proof
(M1). The event \(\{w_t \in F\}\), for \(F\) Borel, is one of the generating
sets of the
Brownian filtration
at time \(t\), so \(w_t\) is \(\mathcal{F}_t\)-measurable.
(M2). By the
Cauchy-Schwarz inequality
in \(L^2(\mathbb{P})\) against the constant function \(\mathbf{1}\), the same \(L^1\)-\(L^2\)
comparison as in the previous section,
\(\mathbb{E}[|w_t|] \leq \bigl( \mathbb{E}[w_t^2] \bigr)^{1/2} = \sqrt{t} \lt \infty\), the
variance being \(t\) by (W1) of the
definition of Brownian motion.
(M3). Let \(0 \leq t \leq s\) and write \(w_s = (w_s - w_t) + w_t\), both
summands integrable — \(w_t\) by the bound above, the increment by the triangle inequality. By
linearity of conditional expectation,
\[
\mathbb{E}[w_s \mid \mathcal{F}_t]
= \mathbb{E}[w_s - w_t \mid \mathcal{F}_t] + \mathbb{E}[w_t \mid \mathcal{F}_t] .
\]
For the first term, the increment is
independent of \(\mathcal{F}_t\)
— the lemma states precisely that every event \(\{w_s - w_t \in B\}\), and these events
constitute \(\sigma(w_s - w_t)\), decouples from every event in \(\mathcal{F}_t\) — so the
independence collapse
applies and yields \(\mathbb{E}[w_s - w_t \mid \mathcal{F}_t] = \mathbb{E}[w_s - w_t] = 0\),
increments being centered by (W1). For the second term, \(w_t\) is
\(\mathcal{F}_t\)-measurable and integrable, so the
take-out property
with the constant factor \(X = \mathbf{1}\) gives
\(\mathbb{E}[w_t \mid \mathcal{F}_t] = w_t \cdot \mathbb{E}[\mathbf{1} \mid \mathcal{F}_t] = w_t\).
Hence \(\mathbb{E}[w_s \mid \mathcal{F}_t] = w_t\) almost surely.
Brownian motion is the raw material; the Itô integral is machined from it. The next theorem says
the machining preserves the fair-game property — and the reason is the arrow of time built into
the class \(\mathcal{V}\): because the integrand is adapted, every increment of the integral pairs
a level readable from the past with a piece of noise invisible to it, and the conditional
expectation of such a pairing vanishes. Causality in, martingale out.
Theorem: The Itô Integral Is a Martingale
Let \(f \in \mathcal{V}(0, T)\). Then the process
\[
M_t(\omega) = \int_0^t f(s, \omega)\, dw_s ,
\qquad 0 \leq t \leq T ,
\]
with the convention \(M_0 = 0\), is a martingale with respect to
\(\{\mathcal{F}_t\}_{0 \leq t \leq T}\).
Proof
(M1) is part (iv) of the
properties theorem
applied on the interval \([0, t]\), and is trivial at \(t = 0\). (M2) is
quantitative: by the
Itô isometry
on \([0, t]\) and the \(L^1\)-\(L^2\) comparison,
\[
\mathbb{E}\bigl[ |M_t| \bigr]
\leq \bigl( \mathbb{E}[M_t^2] \bigr)^{1/2}
= \Bigl( \mathbb{E}\Bigl[ \int_0^t f^2\, ds \Bigr] \Bigr)^{1/2}
\leq \|f\|_{\mathcal{V}} \lt \infty ,
\]
the last bound because \(f^2 \geq 0\) makes the integral monotone in the interval.
(M3). Fix \(0 \leq t \leq s \leq T\). If \(t = s\) the claim is
\(\mathbb{E}[M_t \mid \mathcal{F}_t] = M_t\), which is the take-out property as in the
previous proof; assume \(t \lt s\). For \(t > 0\), additivity — part (i)
of the properties theorem, applied with the triple \((0, t, s)\), the membership
\(f \in \mathcal{V}(0, s)\) holding by the same monotonicity of (V3) — gives, almost surely,
\[
M_s = M_t + \int_t^s f\, dw_u ,
\]
and this identity holds trivially at \(t = 0\), where \(M_t = 0\) by convention and the
remaining integral is \(M_s\) itself. Conditional expectations of almost-surely equal
variables agree almost surely, by the uniqueness clause in their definition. By linearity of
conditional expectation and the take-out property as in the previous proof,
\(\mathbb{E}[M_t \mid \mathcal{F}_t] = M_t\), so everything reduces to a
conditional zero mean: almost surely,
\[
\mathbb{E}\Bigl[ \int_t^s f\, dw_u \,\Bigm|\, \mathcal{F}_t \Bigr] = 0 .
\]
Step 1: bounded elementary integrands. Membership
\(f \in \mathcal{V}(t, s)\) holds because (V1) and (V2) are interval-free and (V3) is
monotone in the interval — the argument that opened the properties theorem — so the three
approximation steps of the previous page,
applied on \([t, s]\), supply bounded
elementary processes
\(\phi_n\) with \(\mathbb{E}\bigl[ \int_t^s (f - \phi_n)^2\, du \bigr] \to 0\). Fix one such
\(\phi = \sum_j e_j \mathbf{1}_{[u_j, u_{j+1})}\) with partition
\(t = u_0 \lt u_1 \lt \cdots \lt u_m = s\) and bounded levels \(e_j\), each
\(\mathcal{F}_{u_j}\)-measurable by adaptedness; write
\(\Delta w_j = w_{u_{j+1}} - w_{u_j}\) for the matching increments, so that
\(\int_t^s \phi\, dw_u = \sum_j e_j\, \Delta w_j\). For a single term, since
\(\mathcal{F}_t \subseteq \mathcal{F}_{u_j}\) and \(e_j\, \Delta w_j \in L^1(\mathbb{P})\),
the
tower property
lets us condition first on the larger \(\sigma\)-algebra:
\[
\mathbb{E}\bigl[ e_j\, \Delta w_j \mid \mathcal{F}_t \bigr]
= \mathbb{E}\Bigl[\, \mathbb{E}\bigl[ e_j\, \Delta w_j \mid \mathcal{F}_{u_j} \bigr]
\,\Bigm|\, \mathcal{F}_t \Bigr]
= \mathbb{E}\Bigl[\, e_j\, \mathbb{E}\bigl[ \Delta w_j \mid \mathcal{F}_{u_j} \bigr]
\,\Bigm|\, \mathcal{F}_t \Bigr]
= 0 ,
\]
where the middle equality is the take-out property (the factor \(e_j\) is
\(\mathcal{F}_{u_j}\)-measurable and bounded, the increment integrable), and the last holds
because \(\mathbb{E}[\Delta w_j \mid \mathcal{F}_{u_j}] = \mathbb{E}[\Delta w_j] = 0\) by
increment independence, the independence collapse, and centeredness — the same chain as for
Brownian motion above. Summing over \(j\) by linearity,
\(\mathbb{E}\bigl[ \int_t^s \phi\, dw_u \mid \mathcal{F}_t \bigr] = 0\): the elementary
integral over the future is invisible to the present.
Step 2: passage to the limit. By the definition of the integral,
\(\int_t^s \phi_n\, dw_u \to \int_t^s f\, dw_u\) in \(L^2(\mathbb{P})\). Conditional
expectation given \(\mathcal{F}_t\) is linear and, by the
\(L^p\) contraction
with \(p = 2\), a contraction of \(L^2(\mathbb{P})\). Hence, using Step 1 for each \(n\),
\[
\Bigl\| \mathbb{E}\Bigl[ \int_t^s f\, dw_u \Bigm| \mathcal{F}_t \Bigr] \Bigr\|_{L^2(\mathbb{P})}
= \Bigl\| \mathbb{E}\Bigl[ \int_t^s (f - \phi_n)\, dw_u \Bigm| \mathcal{F}_t \Bigr] \Bigr\|_{L^2(\mathbb{P})}
\leq \Bigl\| \int_t^s (f - \phi_n)\, dw_u \Bigr\|_{L^2(\mathbb{P})}
\longrightarrow 0 ,
\]
the middle step also using linearity of the integral. A vector of norm zero in
\(L^2(\mathbb{P})\) vanishes almost surely, which is the conditional zero mean, and with it
(M3).
Fair Noise Is What Stochastic Gradient Descent Actually Needs
A stochastic optimization step
\(\boldsymbol{\theta}_{k+1} = \boldsymbol{\theta}_k - \eta \bigl( \nabla L(\boldsymbol{\theta}_k) + \boldsymbol{\xi}_k \bigr)\)
is driven by gradient noise \(\boldsymbol{\xi}_k\), and convergence analyses assume not that
the \(\boldsymbol{\xi}_k\) are independent — the noise at step \(k\) depends on the current
iterate \(\boldsymbol{\theta}_k\), itself a function of all earlier noise, so independence
across steps fails — but that
\(\mathbb{E}[\boldsymbol{\xi}_k \mid \mathcal{M}_k] = 0\), where \(\mathcal{M}_k\) is the
\(\sigma\)-algebra generated by the run so far. That is exactly the discrete-time shadow of
the conditional zero mean proved in Step 1: each noise term, conditioned on the accumulated
history, contributes nothing. Sequences with this property are called martingale differences,
because their partial sums form a discrete-time martingale by the same tower-property
computation as above, and
the maximal inequalities of the next section are among the standard tools for controlling
their cumulative effect over an entire run rather than one step at a time.
The martingale property, for all its strength, is a statement about two time points. The promises
outstanding from the previous page — a version of \(t \mapsto M_t\) that is continuous, and
control of the whole path — are statements about uncountably many time points at once, and
passing from two to uncountably many requires an inequality of a different caliber: one that
bounds the running maximum of a martingale by its value at the final time alone. That inequality
is due to Doob, and we prove it next.
Doob's Maximal Inequality
Everything proved so far constrains the integral at fixed times, or at pairs of times. The
promises still outstanding — a continuous version, control of a whole trajectory — are statements
about uncountably many times at once, and no union bound survives that passage: summing a
per-time estimate over even countably many times destroys it. What saves the situation is the
fair-game structure itself. If a martingale crosses a high level \(\lambda\) at some moment and
yet ends small, the path must lose ground systematically after the crossing — and (M3) says the
process expects to lose no ground on any event observable at the crossing time. Made precise,
this reasoning bounds the probability that the running maximum ever reaches \(\lambda\) by an
expectation involving the final value alone. The argument runs most naturally not for
martingales but for processes allowed to drift upward, and the inequality is inherited by
\(|M_t|^p\) from the martingale \(M_t\) precisely through that one-sided slack.
Definition: Submartingale
Let \(\{\mathcal{M}_t\}_{t \geq 0}\) be a filtration on \((\Omega, \mathcal{F}, \mathbb{P})\).
A real-valued process \(\{X_t\}_{t \geq 0}\) is a submartingale with respect
to \(\{\mathcal{M}_t\}\) if it satisfies (M1) and (M2) of the definition of a
martingale, together with
(M3') \(\mathbb{E}\bigl[ X_s \mid \mathcal{M}_t \bigr] \geq X_t\) almost
surely, for all \(0 \leq t \leq s\).
A submartingale is a game tilted in the player's favor: the forecast of the future, given the
present, never falls below the present. Reversing the inequality defines a
supermartingale, which we will not need. Averaging (M3') over \(A = \Omega\) shows that
the expectation \(t \mapsto \mathbb{E}[X_t]\) of a submartingale is nondecreasing, and every
martingale is in particular a submartingale. The maximal inequality is proved first for finitely
many time points — this is where all the probabilistic content lives; the passage to the
continuum afterwards is pure topology.
Lemma: Maximal Inequality over Finitely Many Times
Let \(\{X_t\}\) be a nonnegative submartingale with respect to \(\{\mathcal{M}_t\}\), let
\(0 \leq t_0 \lt t_1 \lt \cdots \lt t_n\) be times in its index set, and let \(\lambda > 0\). Then
\[
\lambda\, \mathbb{P}\Bigl[ \max_{0 \leq k \leq n} X_{t_k} \geq \lambda \Bigr]
\leq \mathbb{E}\Bigl[ X_{t_n}\, \mathbf{1}_{\{\max_k X_{t_k} \geq \lambda\}} \Bigr]
\leq \mathbb{E}\bigl[ X_{t_n} \bigr] .
\]
Proof
Decompose the event \(A = \{\max_k X_{t_k} \geq \lambda\}\) by the first time the
level is reached: for \(0 \leq k \leq n\), let
\[
A_k = \bigl\{ X_{t_0} \lt \lambda,\; \ldots,\; X_{t_{k-1}} \lt \lambda,\; X_{t_k} \geq \lambda \bigr\} .
\]
The \(A_k\) are pairwise disjoint with union \(A\), and \(A_k \in \mathcal{M}_{t_k}\): each
event \(\{X_{t_j} \lt \lambda\}\) with \(j \lt k\) lies in
\(\mathcal{M}_{t_j} \subseteq \mathcal{M}_{t_k}\), and \(\{X_{t_k} \geq \lambda\}\) lies in
\(\mathcal{M}_{t_k}\), both by adaptedness (M1).
On \(A_k\) the process has reached the level, so
\(\lambda\, \mathbf{1}_{A_k} \leq X_{t_k}\, \mathbf{1}_{A_k}\) pointwise, and by monotonicity
of expectation, among the
properties of expectation,
\[
\lambda\, \mathbb{P}(A_k) \leq \mathbb{E}\bigl[ X_{t_k}\, \mathbf{1}_{A_k} \bigr] .
\]
Now trade the time-\(t_k\) value for the terminal one. The submartingale inequality (M3')
gives \(X_{t_k} \leq \mathbb{E}[X_{t_n} \mid \mathcal{M}_{t_k}]\) almost surely; multiplying
by the nonnegative \(\mathbf{1}_{A_k}\) and taking expectations,
\[
\mathbb{E}\bigl[ X_{t_k}\, \mathbf{1}_{A_k} \bigr]
\leq \mathbb{E}\bigl[\, \mathbb{E}[X_{t_n} \mid \mathcal{M}_{t_k}]\, \mathbf{1}_{A_k} \bigr]
= \int_{A_k} \mathbb{E}[X_{t_n} \mid \mathcal{M}_{t_k}]\, d\mathbb{P}
= \int_{A_k} X_{t_n}\, d\mathbb{P} ,
\]
the last step being the averaging identity that
defines conditional expectation,
legitimate because \(A_k \in \mathcal{M}_{t_k}\). This is the fair-game mechanism in one
line: on an event decided by time \(t_k\), the terminal value integrates to at least as much
as the value at the crossing.
Summing over \(k\) and using disjointness,
\[
\lambda\, \mathbb{P}(A)
= \sum_{k=0}^{n} \lambda\, \mathbb{P}(A_k)
\leq \sum_{k=0}^{n} \mathbb{E}\bigl[ X_{t_n}\, \mathbf{1}_{A_k} \bigr]
= \mathbb{E}\bigl[ X_{t_n}\, \mathbf{1}_{A} \bigr]
\leq \mathbb{E}\bigl[ X_{t_n} \bigr] ,
\]
the final inequality because \(X_{t_n} \geq 0\).
To apply the lemma to a martingale \(M_t\), we feed it \(|M_t|^p\). The observation that convex
images of martingales drift upward is worth isolating: it is the standard device for converting
two-sided processes into nonnegative ones without losing the filtration structure.
Lemma: Convex Images of Martingales Are Submartingales
Let \(\{X_t\}\) be a martingale with respect to \(\{\mathcal{M}_t\}\), and let
\(\varphi : \mathbb{R} \to \mathbb{R}\) be a
convex function
with \(\varphi(X_t) \in L^1(\mathbb{P})\) for every \(t\). Then \(\{\varphi(X_t)\}\) is a
submartingale with respect to \(\{\mathcal{M}_t\}\).
Proof
A convex function on \(\mathbb{R}\) has finite left and right derivatives at every point —
the fact underlying the supporting-line proof of conditional Jensen — and is therefore
continuous, hence Borel measurable; so \(\varphi(X_t)\) is \(\mathcal{M}_t\)-measurable and
(M1) holds. (M2) is the integrability hypothesis. For (M3'), let \(0 \leq t \leq s\). Since
\(X_s \in L^1(\mathbb{P})\) by (M2) for the martingale and
\(\varphi(X_s) \in L^1(\mathbb{P})\) by hypothesis,
Jensen's inequality for conditional expectation
applies, and almost surely
\[
\varphi(X_t)
= \varphi\bigl( \mathbb{E}[X_s \mid \mathcal{M}_t] \bigr)
\leq \mathbb{E}\bigl[ \varphi(X_s) \mid \mathcal{M}_t \bigr] ,
\]
the first equality being the martingale property (M3) inside the continuous \(\varphi\).
One further reading convention, and the main theorem can be stated. A supremum of \(|M_t|\) over
the uncountable set \([0, T]\) is a supremum of uncountably many random variables and need not be
measurable, so the probability appearing below must be read with care: given the horizon \(T\),
we take the supremum along the countable set \(D = \bigcup_N D_N\) of dyadic times
\(D_N = \{ k T 2^{-N} : k = 0, 1, \ldots, 2^N \}\), which is dense in \([0, T]\) and contains
\(0\) and \(T\). A countable supremum of measurable functions is measurable, as in Section 1. For
almost every \(\omega\) the path \(t \mapsto M_t(\omega)\) is continuous by hypothesis, so
\(t \mapsto |M_t(\omega)|\) is continuous and its supremum over the dense set \(D\) equals its
supremum over all of \([0, T]\): the reading convention costs nothing where it matters.
Theorem: Doob's Martingale Inequality
Let \(T \geq 0\) and let \(\{M_t\}_{0 \leq t \leq T}\) be a martingale with respect to a filtration
\(\{\mathcal{M}_t\}\), such that \(t \mapsto M_t(\omega)\) is continuous for almost every
\(\omega\). Then for every \(p \geq 1\) and \(\lambda > 0\),
\[
\mathbb{P}\Bigl[ \sup_{0 \leq t \leq T} |M_t| \geq \lambda \Bigr]
\leq \frac{1}{\lambda^p}\, \mathbb{E}\bigl[ |M_T|^p \bigr] ,
\]
the supremum read along dyadic times as above.
Proof
If \(\mathbb{E}[|M_T|^p] = \infty\) there is nothing to prove, so assume it finite. The map
\(x \mapsto |x|^p\) is convex for \(p \geq 1\): on \([0, \infty)\) the power
\(u \mapsto u^p\) is convex and nondecreasing, its derivative \(p\, u^{p-1}\) being
nonnegative and nondecreasing, and composing a nondecreasing convex function with the convex
\(x \mapsto |x|\) preserves convexity. Moreover \(|M_t|^p\) is integrable for every
\(t \in [0, T]\): the martingale property gives \(M_t = \mathbb{E}[M_T \mid \mathcal{M}_t]\)
almost surely, so the
\(L^p\) contraction
yields \(\|M_t\|_{L^p(\mathbb{P})} \leq \|M_T\|_{L^p(\mathbb{P})} \lt \infty\). By the
convex-image lemma,
\(X_t = |M_t|^p\) is therefore a nonnegative submartingale on \([0, T]\).
Fix \(0 \lt \mu \lt \lambda\) and \(N \in \mathbb{N}\), and apply the
maximal inequality
to \(X\) at the finitely many times of \(D_N\), whose largest element is \(T\), with
threshold \(\mu^p\). Since \(u \mapsto u^p\) is strictly increasing on \([0, \infty)\), the
events \(\{\max_{t \in D_N} |M_t| \geq \mu\}\) and \(\{\max_{t \in D_N} X_t \geq \mu^p\}\)
coincide, so
\[
\mathbb{P}\Bigl[ \max_{t \in D_N} |M_t| > \mu \Bigr]
\leq \mathbb{P}\Bigl[ \max_{t \in D_N} |M_t| \geq \mu \Bigr]
\leq \frac{1}{\mu^p}\, \mathbb{E}\bigl[ |M_T|^p \bigr] .
\]
The sets \(D_N\) increase with \(N\), so the events
\(\{\max_{t \in D_N} |M_t| > \mu\}\) increase, and their union over \(N\) is
\(\{\sup_{t \in D} |M_t| > \mu\}\): a supremum exceeds \(\mu\) strictly exactly when some
member does. By
continuity of measure
from below,
\[
\mathbb{P}\Bigl[ \sup_{t \in D} |M_t| > \mu \Bigr]
= \lim_{N \to \infty} \mathbb{P}\Bigl[ \max_{t \in D_N} |M_t| > \mu \Bigr]
\leq \frac{1}{\mu^p}\, \mathbb{E}\bigl[ |M_T|^p \bigr] .
\]
Finally, \(\{\sup_{t \in D} |M_t| \geq \lambda\} \subseteq \{\sup_{t \in D} |M_t| > \mu\}\)
because \(\mu \lt \lambda\), so the left side of the theorem is bounded by
\(\mu^{-p}\, \mathbb{E}[|M_T|^p]\) for every \(\mu \in (0, \lambda)\); letting
\(\mu \uparrow \lambda\) gives the claim.
Note where each hypothesis worked. The strict inequality \(> \mu\) is what makes the events
increase to the union — with \(\geq\) the identity fails, since a supremum can reach a level no
single term attains — and the reserve \(\mu \lt \lambda\) repays that debt at the end. Path
continuity entered only through the reading convention: it is what entitles a countable set of
times to speak for the continuum.
Watching the Whole Run at Once
A confidence bound valid at one prespecified time becomes invalid the moment one monitors
continuously and stops on a favorable fluctuation — the peeking problem of sequential
testing, and the statistical face of taking a supremum over times. Doob's inequality is the
template for the repair: it prices the running maximum of an entire trajectory at the cost of
a single terminal moment, with no union bound over times and no widening factor growing with
the horizon. Modern anytime-valid inference — confidence sequences that a practitioner may
check after every observation, stopping whenever they like — is built on maximal inequalities
for exactly the martingale structures of this section, applied to the likelihood ratios and
error processes of the experiment.
Doob's inequality demands continuous paths, and there lies the last gap. The martingale
\(M_t = \int_0^t f\, dw_s\) of the
previous section
satisfies (M1)-(M3), but each \(M_t\) was constructed as an \(L^2\) limit, one \(t\) at a time,
determined only up to a null set — and nothing so far ties those uncountably many choices into a
single continuous path. Producing one process that is continuous in \(t\) and represents every
\(M_t\) at once is the task of the next section, and Doob's inequality, applied to the
differences of approximating integrals, is the tool that accomplishes it.
A Continuous Version
Here is the gap left by the construction, stated plainly. For each fixed \(t\), the integral
\(\int_0^t f\, dw_s\) is an element of \(L^2(\mathbb{P})\) — an equivalence class, pinned down
only up to a null set. A path \(t \mapsto M_t(\omega)\) requires choosing one representative for
every \(t\), and there are uncountably many \(t\): the null sets on which the choices misbehave
can accumulate into a set of full measure, and no property of the individual classes prevents it.
Continuity is a property of the joint selection, not of the separate values, and it must be
engineered. The plan: realize the approximating elementary integrals as genuinely continuous
paths, force a subsequence of them to converge uniformly in \(t\) on a single almost-sure
event — this is where Doob's inequality carries the load — and take the limit path by path. Two
pieces of vocabulary and one classical lemma first.
Definition: Version of a Process
Let \(\{X_t\}_{t \in I}\) and \(\{Y_t\}_{t \in I}\) be stochastic processes on
\((\Omega, \mathcal{F}, \mathbb{P})\) with a common index set \(I\). The process \(Y\) is a
version (or modification) of \(X\) if, for every
\(t \in I\),
\[
\mathbb{P}\bigl[ X_t = Y_t \bigr] = 1 .
\]
The relation is symmetric, and versions agree almost surely at each finite collection of times,
so every statement of the preceding sections — martingale property, isometry, zero mean —
transfers freely between versions. Path properties do not transfer: modifying a process at a
single, randomly located time produces a version with jumps. The saving grace is that continuity,
once achieved, is essentially unique: if two versions of the same process, indexed by an
interval, both have almost surely continuous paths, their difference vanishes almost surely at
each time in a countable dense subset of the interval, hence — one null set for countably many
times — at all of them at once, and continuity propagates the equality to every \(t\). Two
continuous versions are thus equal for all \(t\) simultaneously,
almost surely, and one may speak of the continuous version.
Lemma: Borel-Cantelli
Let \(\{A_k\}_{k \in \mathbb{N}} \subseteq \mathcal{F}\) satisfy
\(\sum_{k=1}^{\infty} \mathbb{P}(A_k) \lt \infty\). Then
\[
\mathbb{P}\Bigl[ \bigcap_{N \geq 1} \bigcup_{k \geq N} A_k \Bigr] = 0 ,
\]
that is, almost every \(\omega\) belongs to only finitely many of the \(A_k\).
Proof
First, countable subadditivity. For any sequence \(\{B_k\} \subseteq \mathcal{F}\), the sets
\(C_k = B_k \setminus \bigcup_{j \lt k} B_j\) are disjoint, lie in \(\mathcal{F}\), satisfy
\(C_k \subseteq B_k\), and have the same union as the \(B_k\); by countable additivity, which
is part of the
definition of a measure,
and by monotonicity — itself immediate from additivity applied to \(B_k = C_k \cup (B_k \setminus C_k)\) —
\[
\mathbb{P}\Bigl[ \bigcup_{k} B_k \Bigr]
= \sum_{k} \mathbb{P}(C_k)
\leq \sum_{k} \mathbb{P}(B_k) .
\]
Now let \(A\) denote the intersection in the statement. For every \(N\), monotonicity and
subadditivity give
\[
\mathbb{P}(A)
\leq \mathbb{P}\Bigl[ \bigcup_{k \geq N} A_k \Bigr]
\leq \sum_{k \geq N} \mathbb{P}(A_k) ,
\]
and the right side is the tail of a convergent series, hence tends to \(0\) as
\(N \to \infty\). The left side does not depend on \(N\), so \(\mathbb{P}(A) = 0\).
The lemma converts a summable sequence of failure probabilities into almost-sure eventual
success, and it is exactly the shape of conclusion we need: Doob's inequality will price each
failure event, the isometry will make the prices summable, and Borel-Cantelli will collect the
winnings. Here is the theorem that keeps the previous page's promise.
Theorem: The Itô Integral Admits a Continuous Version
Let \(f \in \mathcal{V}(0, T)\). There exists a stochastic process \(\{J_t\}_{0 \leq t \leq T}\)
on \((\Omega, \mathcal{F}, \mathbb{P})\) such that \(t \mapsto J_t(\omega)\) is continuous
for almost every \(\omega\), and, for every \(t \in [0, T]\),
\[
\mathbb{P}\Bigl[ J_t = \int_0^t f\, dw_s \Bigr] = 1 .
\]
Proof
Step 1: elementary integrals have continuous realizations.
Let \(\phi = \sum_{j} e_j\, \mathbf{1}_{[t_j, t_{j+1})}\) be bounded elementary on \([0, T]\) and
define, pointwise in \(\omega\) and for every \(t \in [0, T]\),
\[
I^{\phi}(t, \omega)
= \sum_{j} e_j(\omega) \bigl( w_{t \wedge t_{j+1}}(\omega) - w_{t \wedge t_j}(\omega) \bigr) ,
\]
where \(a \wedge b\) denotes \(\min(a, b)\).
For \(t \in [t_k, t_{k+1})\) the terms with \(j \lt k\) contribute full increments, the term
\(j = k\) contributes \(e_k (w_t - w_{t_k})\), and later terms vanish, while at \(t = T\)
every increment is complete — so \(I^{\phi}(t, \cdot)\)
is precisely the elementary integral of \(\phi\, \mathbf{1}_{[0, t)}\), the representative of
\(\int_0^t \phi\, dw_s\) used throughout Section 1. The formula is unchanged under refinement
of the partition, an inserted point splitting one term into two that telescope, and over a
common partition it is linear in \(\phi\). For almost every \(\omega\) the path
\(s \mapsto w_s(\omega)\) is continuous by (W3), and then each map
\(t \mapsto w_{t \wedge c}(\omega)\) is continuous, being a composition with the continuous
\(t \mapsto t \wedge c\); hence \(t \mapsto I^{\phi}(t, \omega)\), a finite linear
combination of such maps, is continuous.
Step 2: the maximal estimate.
Fix bounded elementary \(\phi_n\) with
\(\|f - \phi_n\|_{\mathcal{V}} \to 0\) and write \(I_n(t, \omega) = I^{\phi_n}(t, \omega)\).
For \(n, m\), pass to a common refinement: \(\phi_n - \phi_m\) is bounded elementary and, by
the linearity and refinement-invariance of Step 1,
\(I_n - I_m = I^{\phi_n - \phi_m}\) pointwise. For each fixed \(t\) this is a representative
of \(\int_0^t (\phi_n - \phi_m)\, dw_s\), so by the
martingale theorem
applied to \(\phi_n - \phi_m \in \mathcal{V}(0, T)\), the process
\(\{I_n(t) - I_m(t)\}_{0 \leq t \leq T}\) is a martingale with respect to
\(\{\mathcal{F}_t\}\) — conditions (M2) and (M3) concern only almost-sure classes, and (M1)
holds for every choice of representatives by part (iv) of the properties theorem — and by
Step 1 its paths are continuous almost surely.
Doob's inequality
with \(p = 2\), followed by the
Itô isometry
at the terminal time, gives for every \(\varepsilon > 0\)
\[
\mathbb{P}\Bigl[ \sup_{0 \leq t \leq T} |I_n(t) - I_m(t)| \geq \varepsilon \Bigr]
\leq \frac{1}{\varepsilon^2}\, \mathbb{E}\bigl[ (I_n(T) - I_m(T))^2 \bigr]
= \frac{1}{\varepsilon^2}\, \|\phi_n - \phi_m\|_{\mathcal{V}}^2 ,
\]
second moments of almost-surely equal variables being equal. The supremum is read along
dyadic times, as in the previous section; on the almost-sure event of path continuity it is
the full supremum.
Step 3: a fast subsequence.
Since
\[
\|\phi_n - \phi_m\|_{\mathcal{V}} \leq \|\phi_n - f\|_{\mathcal{V}} + \|f - \phi_m\|_{\mathcal{V}} \to 0
\]
as \(n, m \to \infty\), we may choose indices \(n_1 \lt n_2 \lt \cdots\) such that
\(\|\phi_n - \phi_m\|_{\mathcal{V}}^2 \lt 2^{-3k}\) whenever \(n, m \geq n_k\). Taking
\(\varepsilon = 2^{-k}\) in Step 2,
\[
\mathbb{P}\Bigl[ \sup_{0 \leq t \leq T} |I_{n_{k+1}}(t) - I_{n_k}(t)| \geq 2^{-k} \Bigr]
\leq 2^{2k} \cdot 2^{-3k} = 2^{-k} .
\]
Step 4: Borel-Cantelli.
The bounds \(2^{-k}\) are summable, so by the
lemma, for almost every
\(\omega\) there exists \(k_1(\omega)\) such that, for all \(k \geq k_1(\omega)\),
\[
\sup_{0 \leq t \leq T} \bigl| I_{n_{k+1}}(t, \omega) - I_{n_k}(t, \omega) \bigr| \lt 2^{-k} .
\]
Step 5: uniform convergence.
Work on the almost-sure event where the
conclusion of Step 4 holds and every path \(I_{n_k}(\cdot, \omega)\) is continuous — a
countable union of null sets is discarded. For \(l > k \geq k_1(\omega)\) and every
\(t \in [0, T]\), the triangle inequality along the chain of consecutive differences gives
\[
\bigl| I_{n_l}(t, \omega) - I_{n_k}(t, \omega) \bigr|
\leq \sum_{j = k}^{l - 1} 2^{-j}
\lt 2^{-k + 1} ,
\]
so the sequence \(\{I_{n_k}(t, \omega)\}_k\) is Cauchy in \(\mathbb{R}\), uniformly in
\(t\). By completeness of \(\mathbb{R}\) the limit
\(J_t(\omega) = \lim_{k} I_{n_k}(t, \omega)\) exists for every \(t\), and letting
\(l \to \infty\) in the display,
\(\sup_{0 \leq t \leq T} |J_t(\omega) - I_{n_k}(t, \omega)| \leq 2^{-k+1}\) for
\(k \geq k_1(\omega)\): the convergence is uniform on \([0, T]\). On the discarded null set,
set \(J_t(\omega) = 0\) for all \(t\).
Step 6: continuity of the limit.
Fix such an \(\omega\), a point
\(t \in [0, T]\), and \(\varepsilon > 0\). Choose \(k \geq k_1(\omega)\) with
\(2^{-k+1} \lt \varepsilon / 3\), and then \(\delta > 0\) such that
\(|I_{n_k}(s, \omega) - I_{n_k}(t, \omega)| \lt \varepsilon / 3\) whenever
\(|s - t| \lt \delta\), by continuity of \(I_{n_k}(\cdot, \omega)\). For such \(s\),
\[
|J_s - J_t|
\leq |J_s - I_{n_k}(s)| + |I_{n_k}(s) - I_{n_k}(t)| + |I_{n_k}(t) - J_t|
\lt \varepsilon ,
\]
suppressing \(\omega\). Hence \(t \mapsto J_t(\omega)\) is continuous — the classical fact
that a uniform limit of continuous functions is continuous, proved here inline in its
entirety.
Step 7: \(J\) is a version.
Fix \(t \in (0, T]\); at \(t = 0\) the integral
is \(0\) under the convention adopted with the martingale theorem, and \(J_0 = 0\) by
construction. As in the proof of the properties theorem, \(\phi_{n_k} \mathbf{1}_{[0, t)}\) is
bounded elementary and admissible for \(f\) on \([0, t]\), so
\(I_{n_k}(t) \to \int_0^t f\, dw_s\) in \(L^2(\mathbb{P})\); passing to a further
subsequence converging almost surely,
and recalling that \(I_{n_k}(t) \to J_t\) almost surely by Step 5, the two almost-sure limits
coincide almost surely. Hence \(J_t\) is a representative of the class
\(\int_0^t f\, dw_s\), which is the displayed claim.
The theorem earns an immediate reward. The continuous version is a martingale whose paths satisfy
the hypothesis of Doob's inequality, and feeding the Itô isometry through it yields a bound on
the entire trajectory of the integral in terms of the plain size of the integrand — the estimate
that closes the circle of this page's first four sections.
Theorem: Maximal Bound for the Itô Integral
Let \(f \in \mathcal{V}(0, T)\) and let \(\{J_t\}_{0 \leq t \leq T}\) be a continuous version
of the integral process. Then for every \(\lambda > 0\),
\[
\mathbb{P}\Bigl[ \sup_{0 \leq t \leq T} |J_t| \geq \lambda \Bigr]
\leq \frac{1}{\lambda^2}\, \mathbb{E}\Bigl[ \int_0^T f(s, \omega)^2\, ds \Bigr] .
\]
Proof
The process \(J\) is a martingale with respect to \(\{\mathcal{F}_t\}\). For (M1), each
\(J_t\) is \(\mathcal{F}\)-measurable by construction and is a representative of
\(\int_0^t f\, dw_s\), hence \(\mathcal{F}_t\)-measurable by part (iv) of the
properties theorem
— it is precisely here that the every-representative form of (iv), and behind it the null-set
convention, pays for itself: adaptedness costs nothing to transfer. For (M2) and (M3), each
\(J_t\) is almost surely equal to the corresponding value of the martingale of the
martingale theorem,
integrability transfers, and conditional expectations of almost-surely equal variables agree,
so \(\mathbb{E}[J_s \mid \mathcal{F}_t] = J_t\) almost surely for \(t \leq s\). The paths of
\(J\) are continuous almost surely, so
Doob's inequality
applies with \(p = 2\), and with the
Itô isometry
at time \(T\),
\[
\mathbb{P}\Bigl[ \sup_{0 \leq t \leq T} |J_t| \geq \lambda \Bigr]
\leq \frac{1}{\lambda^2}\, \mathbb{E}\bigl[ J_T^2 \bigr]
= \frac{1}{\lambda^2}\, \mathbb{E}\Bigl[ \int_0^T f(s, \omega)^2\, ds \Bigr] .
\]
From this point on, \(\int_0^t f\, dw_s\), viewed as a process in \(t\), will always denote a
continuous version; since any two continuous versions agree for all \(t\) at once almost surely,
the choice is immaterial and the article the is deserved. Every pathwise statement of
stochastic calculus — beginning with the Itô formula on the next page — is a statement about
this object.
What a Sampler Actually Samples
Every numerical scheme that simulates a stochastic differential equation — from
Euler-Maruyama in computational finance to the samplers that integrate the reverse-time
dynamics of diffusion models —
outputs a trajectory: a single continuous curve per random seed. The object such a scheme
approximates is not the family of \(L^2\) classes constructed on the previous page, which has
no trajectories at all, but the continuous version whose existence this section proved. The
theorem is, in that sense, the license under which path simulation operates: it guarantees
that there is a path-valued object to converge to, and the maximal bound above is the shape
of estimate — uniform over the whole time horizon — by which such convergence is measured.
The integral is now everything the previous page promised: a continuous, square-integrable
martingale, isometric to its integrand, vanishing in mean, adapted to the information that
generated it. What remains is to survey how far the construction reaches — to several driving
Brownian motions, to larger filtrations, to integrands beyond \(\mathcal{V}\) — and to settle an
account left open since the previous page's Section 2: what becomes of the calculus if the
evaluation point abandons causality. That is the Stratonovich question, and it closes the page.
Extensions and the Stratonovich Question
The theory now standing was built under four commitments: one driving Brownian motion, its own
filtration, square-integrable integrands, and the left endpoint. The first three are technical
restrictions, and each can be relaxed; we survey the relaxations with precision about what has to
be re-verified, because the verification burden is where the mathematics lives. The fourth is not
a restriction but a decision, made in Section 2 of the previous page and payable now: the page
closes by settling the account with the midpoint convention it set aside there.
Larger Filtrations and Several Driving Motions
Read back through the construction and mark every place the Brownian filtration entered. It
played exactly two mathematical roles: the integrand was required to be
\(\mathcal{F}_t\)-adapted, and each increment \(w_v - w_u\) was
independent of \(\mathcal{F}_u\)
— the input consumed by the elementary isometry, by the zero-mean and martingale computations of
this page, and by nothing else; the null-set convention was bookkeeping, not mathematics. It
follows that the entire theory survives, verbatim, if \(\{\mathcal{F}_t\}\) is replaced by any
larger filtration \(\{\mathcal{H}_t\}\), with \(\mathcal{F}_t \subseteq \mathcal{H}_t\) and each
\(\mathcal{H}_t\) containing the \(\mathbb{P}\)-null sets as before, such that each increment
\(w_v - w_u\) remains
independent of \(\mathcal{H}_u\): integrands may then depend on more information than the history
of \(w\), so long as the extra information never foresees the noise. We stress that independence
of the increment, not merely the weaker conditional centering
\(\mathbb{E}[w_v - w_u \mid \mathcal{H}_u] = 0\), is what the diagonal of the elementary isometry
consumed — the factorization
\(\mathbb{E}[e_j^2 (\Delta w_j)^2] = \mathbb{E}[e_j^2]\, \Delta t_j\) requires the squared
increment, not only the increment, to decouple from the past.
The example that makes the enlargement indispensable is several Brownian motions at once. Let
\(\mathbf{w}_t = (w_t^{(1)}, \ldots, w_t^{(n)})\) be \(n\)-dimensional standard
Brownian motion,
and let \(\mathcal{F}_t^{(n)}\) be the filtration generated by all components up to time
\(t\) — the
Brownian filtration
with the generating events ranging over every coordinate, null sets adjoined as before. An
integrand may now read the entire vector history, and the single fact to check is that each
component's increment \(w_v^{(j)} - w_u^{(j)}\) is independent of the joint past
\(\mathcal{F}_u^{(n)}\). The verification is the Gaussian block computation of the increment
lemma over again: the covariance of the increment with an earlier value of the same
component vanishes by the formula \(\min(t_i, v) - \min(t_i, u) = 0\), and with an earlier value
of a different component because the covariance matrix \(\min(s, t)\, I_n\) in (W1) has
no cross-component entries; the passage from generating events to the full \(\sigma\)-algebra is
the same monotone-class step
taken on faith there. With that single input re-verified, each component is a legitimate
integrator for \(\mathcal{F}_t^{(n)}\)-adapted integrands, and matrix-valued integrands assemble
the components into one object.
Definition: The Multi-Dimensional Itô Integral
Let \(0 \leq S \lt T\), let \(\mathbf{w}_t = (w_t^{(1)}, \ldots, w_t^{(n)})\) be
\(n\)-dimensional standard Brownian motion, and let \(\{\mathcal{F}_t^{(n)}\}\) be the
filtration generated by all of its components up to time \(t\), null sets adjoined. Let
\(\mathcal{V}^{m \times n}(S, T)\) denote the class of
\(m \times n\) matrix-valued processes
\(\mathbf{v}(t, \omega) = [v_{ij}(t, \omega)]\) whose every entry satisfies conditions
(V1)-(V3) of the
class \(\mathcal{V}(S, T)\)
with \(\mathcal{F}_t^{(n)}\) in place of \(\mathcal{F}_t\). For
\(\mathbf{v} \in \mathcal{V}^{m \times n}(S, T)\), the
multi-dimensional Itô integral
\[
\int_S^T \mathbf{v}(t, \omega)\, d\mathbf{w}_t
\]
is the \(\mathbb{R}^m\)-valued random vector whose \(i\)-th component is
\[
\sum_{j=1}^{n} \int_S^T v_{ij}(t, \omega)\, dw_t^{(j)} ,
\]
each summand being the one-dimensional Itô integral driven by \(w^{(j)}\), constructed with
respect to the filtration \(\{\mathcal{F}_t^{(n)}\}\).
The notation is chosen so that the mnemonic is the statement: \(\mathbf{v}\, d\mathbf{w}\) is a
matrix multiplying a column of differentials, row by row. Integrands like
\(v(t, \omega) = w_t^{(2)}\) driving \(dw_t^{(1)}\) — one noise source modulating exposure to
another, unreachable in the single-component theory because \(w^{(2)}\) is not measurable for the
filtration of \(w^{(1)}\) alone — are now admissible, and every result of this page applies to
each component sum: zero mean, the isometry summand by summand, the martingale property with
respect to \(\{\mathcal{F}_t^{(n)}\}\), and a continuous version. The multi-dimensional Itô
formula, which the coming pages develop, consumes precisely this object.
Weaker Integrability
Condition (V3) can also be relaxed: it suffices to demand
\(\mathbb{P}\bigl[ \int_S^T f(t, \omega)^2\, dt \lt \infty \bigr] = 1\), with no expectation at
all. For such integrands one can still produce approximating step processes, but the
approximation and the resulting integral converge only in the sense of
convergence in probability,
and the price is steep: without (V3) the isometry has no finite right-hand side, the zero-mean
property may fail, and the integral need not be a martingale — it retains only a localized
remnant of the property, the notion of a local martingale, which we leave undefined. We
record this extension at statement level and take none of it on: the class \(\mathcal{V}\) and
its martingale calculus suffice for everything on our horizon, and when the weaker class is
eventually wanted, the honest construction will be built, not borrowed.
The Stratonovich Question
One account remains open. Section 2 of the previous page discovered that the would-be integral
of \(w\) against itself depends on the evaluation point, the discrepancy between right and left
endpoints being exactly the quadratic variation; it committed to the left endpoint and named the
midpoint alternative — the Stratonovich integral, written
\(\int f \circ dw_t\) — with a promise that it would reappear. It reappears now, because for the
test integrand everything is computable. The
first computation
gave \(\int_0^T w_s\, dw_s = \tfrac{1}{2} w_T^2 - \tfrac{1}{2} T\); the right-endpoint sums
exceed the left by the quadratic sums converging to \(T\); and the average of the two therefore
converges to the classical value \(\tfrac{1}{2} w_T^2\). The midpoint sums can be shown to share
this limit — a small estimate we record at statement level — so
\(\int_0^T w_s \circ dw_s = \tfrac{1}{2} w_T^2\): the Stratonovich integral of \(w\) obeys the
classical chain rule on the nose, and the Itô-Stratonovich discrepancy for this integrand is
exactly half the quadratic variation, \(\tfrac{1}{2} T\), rigorously accounted for by the
path properties theorem.
This is no isolated coincidence, and the deeper reason lies in the modeling question that opened
the previous page: which integral makes the white-noise equation an honest idealization?
Approximate the rough path by smooth ones — say, the polygonal interpolations of \(w\) along
ever finer dyadic partitions, which converge to \(w\) uniformly on bounded intervals almost
surely — and for each \(\omega\) solve the ordinary differential equation obtained by using the
derivative of the interpolation as the noise. The Wong-Zakai approximation theorem, which we
state without proof — it rests on a finer analysis of pathwise approximation than this page
undertakes — says the solutions converge, and the limit solves the equation in the
Stratonovich sense. Smooth physics,
pushed to the white-noise limit, lands on the midpoint calculus: classical mechanisms have no
arrow-of-time asymmetry inside an infinitesimal interval, and the midpoint is where that symmetry
lives.
Yet nothing is lost by refusing the midpoint, because the two calculi are intertranslatable. At
statement level — the displays below are shorthand for integral equations, and the notion of
solving them is itself the business of a coming page — the Stratonovich equation
\[
X_t = X_0 + \int_0^t b(s, X_s)\, ds + \int_0^t \sigma(s, X_s) \circ dw_s
\]
describes the same process as the Itô equation with a corrected drift,
\[
X_t = X_0 + \int_0^t \Bigl( b(s, X_s)
+ \tfrac{1}{2}\, \sigma'(s, X_s)\, \sigma(s, X_s) \Bigr) ds
+ \int_0^t \sigma(s, X_s)\, dw_s ,
\]
where \(\sigma'\) denotes the derivative of \(\sigma(t, x)\) in \(x\); the correction
\(\tfrac{1}{2} \sigma' \sigma\) is the equation-level face of the half-quadratic-variation gap
computed above, and its derivation is an exercise in the Itô formula, deferred to the page that
proves it. When \(\sigma\) does not depend on \(x\) — additive noise — the correction vanishes
and the two interpretations coincide.
The choice, then, is a genuine trade. The Stratonovich calculus keeps the classical chain rule
and transforms cleanly under changes of variables, which makes it the natural language where
geometry dominates — stochastic calculus on manifolds is built on it — and Wong-Zakai makes it
the limit of smooth physical models. The Itô calculus pays a correction term in its chain rule
and receives, in exchange, everything this page proved: the integral never reads the future, has
mean zero, is an isometry, is a martingale, obeys Doob's inequality and the maximal bound, and
admits a continuous version. For estimation, prediction, and control — settings where the
filtration is the point — the martingale structure is not a convenience but the subject matter,
and it is unavailable on the Stratonovich side. Since the drift correction translates freely
between the two, we commit to Itô and lose nothing: whenever a model arrives in Stratonovich
form, the corrected drift converts it.
The commitment fixes the direction of the track. The previous page ended by reading the pattern
\(d(w_t^2) = 2 w_t\, dw_t + dt\) off its one computation and calling it the seed of a calculus;
this page has grown the integral into a continuous martingale on which that calculus can act path
by path. What remains is the calculus itself: the Itô formula, the chain rule whose correction
term is the quadratic variation this page has been paying and collecting throughout, and the
gateway to differential equations driven by noise. That is the next page's work.