The Toolkit
The previous page stated the existence and uniqueness theorem and proved its uniqueness half. Under the linear growth condition (H1) and the Lipschitz condition (H2), two solutions with the same initial value are indistinguishable. This page proves the other half. A solution must now be manufactured from the coefficients, the noise, and the initial value alone, and the whole of it is a construction.
The setting carries over unchanged. Throughout, \(w_t\) is the standard one-dimensional Brownian motion, \(Z\) is a square-integrable initial value independent of the entire motion, \(\{\mathcal{H}_t\}\) is the enlarged filtration built from \(Z\), the motion, and the null sets, and \(\mathcal{V}(0, T)\) is the integrand class over this filtration, with every quoted result taken in its enlarged version, as the previous page declared once and for all. The coefficients \(b\) and \(\sigma\) satisfy (H1) and (H2) with the constants \(C\) and \(D\) fixed there, and the target is a process satisfying clauses (S1) through (S4) of the solution definition. The horizon \(T \gt 0\) stays fixed, and \(\lambda\) denotes Lebesgue measure on \([0, T]\), so that \(L^2(\lambda \otimes \mathbb{P})\) is the space in which integrands and processes are measured over the whole strip \([0, T] \times \Omega\).
The construction runs in two stages, and each stage runs on inequalities. The first stage iterates the equation and shows that successive iterates contract in mean square at factorial speed, by the same isometry-plus-Gronwall mechanism that drove the uniqueness proof. The second stage upgrades this mean-square convergence to convergence of the paths themselves, uniformly in time, and here a new kind of statement is needed. A mean square is an average over \(\Omega\), while a statement about paths concerns the probability of an event. The bridge between the two is one inequality, used everywhere in probability, that this track has never owned. We make it ours first.
Tail Bounds from Moments
The observation is that a nonnegative random variable cannot be large with large probability without its expectation noticing. Cutting the variable down at a threshold can only shrink it, and what remains still dominates the threshold times the indicator of the exceedance.
(i)
(Markov) Let \(Y\) be a nonnegative random variable and \(a \gt 0\). Then
\[ \mathbb{P}\bigl[ Y \geq a \bigr] \leq \frac{1}{a}\, \mathbb{E}[Y] . \]
(ii)
(Chebyshev) Let \(X\) be any random variable and \(a \gt 0\). Then
\[ \mathbb{P}\bigl[ |X| \geq a \bigr] \leq \frac{1}{a^2}\, \mathbb{E}\bigl[ X^2 \bigr] . \]
(i).
The pointwise inequality
\[ a\, \mathbf{1}_{\{ Y \geq a \}}(\omega) \leq Y(\omega) \]
holds for every \(\omega\). On the event \(\{Y \geq a\}\) it reads \(a \leq Y(\omega)\), which is the membership condition itself, and off the event it reads \(0 \leq Y(\omega)\), which is nonnegativity. The event is measurable because \(Y\) is, expectation is monotone, and the expectation of \(a\) times an indicator is \(a\) times the probability of its event, so \(a\, \mathbb{P}[Y \geq a] \leq \mathbb{E}[Y]\). Dividing by \(a \gt 0\) finishes. No integrability is assumed. When \(\mathbb{E}[Y]\) is infinite the bound holds and says nothing.
(ii).
Squaring is strictly increasing on \([0, \infty)\), so the events
\(\{ |X| \geq a \}\) and \(\{ X^2 \geq a^2 \}\) are equal. Applying (i) to
the nonnegative variable \(X^2\) at the threshold \(a^2\) gives
\(\mathbb{P}[X^2 \geq a^2] \leq a^{-2}\, \mathbb{E}[X^2]\).
Two remarks position the theorem for use. Since \(\{Y \gt a\}\) is contained in \(\{Y \geq a\}\), the same bounds hold for strict exceedances, and it is the strict form that the arguments ahead consume. And applying (ii) to \(X - \mathbb{E}[X]\), when the mean exists, bounds the probability of a deviation from the mean by the variance over the threshold squared, which is the statement the name Chebyshev most often attaches to. This page needs the uncentered form.
What Is Already in Hand
The rest of the kit exists, and it is worth laying out on the bench in the order of use. The mean-square stage runs on three estimates and one completeness theorem. The Cauchy-Schwarz case of Hölder's inequality controls the drift, the Itô isometry converts the mean square of the noise term into a time integral, and Gronwall's inequality from the previous page turns the resulting self-referential estimates into explicit bounds. The division of labor is exactly that of the uniqueness proof, and the membership checks that license the isometry will be repeated, not inherited, since the iterates are new processes. When the mean-square estimates are in place, the iterates form a Cauchy sequence in \(L^2(\lambda \otimes \mathbb{P})\), and the Riesz-Fischer theorem supplies a limit in that space, since the product measure is a measure like any other.
The pathwise stage runs on three more tools. The Chebyshev inequality just proved converts the factorial-speed mean-square estimates into summable bounds on tail probabilities of the drift differences. The maximal bound for the Itô integral does the same for the noise differences, and does something Chebyshev alone cannot. For an integrand in \(\mathcal{V}(0, T)\) and a continuous version of its integral process, it bounds the tail of the running supremum over the whole horizon by the mean energy of the integrand over the threshold squared, which is exactly the uniform-in-time control that path convergence requires. Finally, the Borel-Cantelli lemma converts summable tail probabilities into an almost sure statement, that only finitely many of the exceedances ever happen. Chained together, these three turn factorial decay of mean squares into uniform convergence of almost every path.
The kit is complete. The scheme that consumes it comes next.