Why Conditional Expectation?
Elementary probability offers two distinct constructions that both go by the name
conditional expectation. Given an event \(A\) with
\(\mathbb{P}(A) \gt 0\), one writes
\[
\mathbb{E}[X \mid A] = \frac{1}{\mathbb{P}(A)} \int_A X \, d\mathbb{P},
\]
which is a single number, the average of \(X\) restricted to the event \(A\). Given a random
variable \(Y\) such that \((X, Y)\) has a joint density \(p(x, y)\), one writes
\[
\mathbb{E}[X \mid Y = y] = \int x \, p(x \mid y) \, dx,
\]
which is a function of \(y\).
These two formulas address visibly different situations. The first averages over a
positive-probability event, and the second averages along a measure-zero fibre using a conditional
density. Neither formula reduces to the other, and the second one is not even well-defined when
\(p(x \mid y)\) fails to exist as a function (for example, when \(Y\) is mixed discrete-continuous
or supported on a fractal).
A unifying definition exists, and its existence is the headline payoff of the
Radon-Nikodym theorem.
We read a sub-\(\sigma\)-algebra \(\mathcal{G} \subseteq \mathcal{F}\) as "the information
available to an observer". Given an integrable random variable \(X\) and such a \(\mathcal{G}\),
there is a \(\mathcal{G}\)-measurable random variable, unique up to almost-sure equality and
denoted \(\mathbb{E}[X \mid \mathcal{G}]\), that simultaneously specialises to both classical
formulas and continues to make sense in every intermediate situation. The unifying object is a
function on \(\Omega\) rather than a number. For square-integrable \(X\), its values give the
best forecast of \(X\), in the mean-square sense, that an observer with information
\(\mathcal{G}\) can make.
The construction is short, because all the analytic machinery was set up in
Signed Measures & Radon-Nikodym Theorem,
following the preview given in
Limit Theorems & Product Measures.
This page collects the payoff.
The conditional expectation also underlies, in idealised form, several core constructions of
machine learning, which the closing section takes up in detail. In
expectation-maximisation, the E-step is exactly a conditional expectation. The Bellman equation of
reinforcement learning rests on the tower property. Under the usual assumption that a new
observation is conditionally independent of the observed data \(D\) given the parameter \(\theta\),
the Bayesian posterior predictive
\(p(x_{\text{new}} \mid D) = \mathbb{E}[p(x_{\text{new}} \mid \theta) \mid D]\) is a conditional
expectation given \(D\), and the
ELBO of variational inference replaces an
intractable conditional measure by a tractable surrogate.
The rigorous object built here seldom appears directly in code, where Monte Carlo,
single-trajectory stochastic approximation, or variational surrogates stand in for it. It is
nonetheless the object those approximations approximate.
Before turning to the construction, we fix the notation for the rest of the page. The underlying
probability space is \((\Omega, \mathcal{F}, \mathbb{P})\), and \(\mathcal{G}\) always denotes a
sub-\(\sigma\)-algebra of \(\mathcal{F}\), a smaller collection of measurable sets representing
partial information. The restriction of \(\mathbb{P}\) to \(\mathcal{G}\) is written
\(\mathbb{P}|_{\mathcal{G}}\), the same probability measure regarded as defined only on the smaller
\(\sigma\)-algebra. The object we are going to construct is denoted
\(\mathbb{E}[X \mid \mathcal{G}]\), and when \(\mathcal{G} = \sigma(Y)\) we abbreviate it as
\(\mathbb{E}[X \mid Y]\).
Throughout, "\(\mathbb{P}\)-a.s." means almost surely with respect to \(\mathbb{P}\). In the
probabilistic context we use this qualifier in place of "\(\mathbb{P}\)-a.e.", consistent with
earlier pages of this section.
Definition via Radon-Nikodym
The construction proceeds in three steps. We first associate to each integrable random variable
\(X\) and each sub-\(\sigma\)-algebra \(\mathcal{G}\) a finite signed measure \(\nu_X\) on
\(\mathcal{G}\). We then verify that \(\nu_X\) is absolutely continuous with respect to
\(\mathbb{P}|_{\mathcal{G}}\), so that the Radon-Nikodym theorem applies. The resulting
Radon-Nikodym derivative is, by definition, the conditional expectation. The discrete case is
recovered immediately, and the continuous case is identified as requiring one further layer of
machinery, the framework of regular conditional distributions and the disintegration theorem.
The Signed Measure Associated with \(X\)
Definition: The Signed Measure \(\nu_X\)
Let \(X \in L^1(\Omega, \mathcal{F}, \mathbb{P})\) and let
\(\mathcal{G} \subseteq \mathcal{F}\) be a sub-\(\sigma\)-algebra. Define
\[
\nu_X : \mathcal{G} \to \mathbb{R}, \quad
\nu_X(A) = \int_A X \, d\mathbb{P}, \quad A \in \mathcal{G}.
\]
Three properties of \(\nu_X\) must be checked before the Radon-Nikodym theorem can be invoked:
that \(\nu_X\) is a signed measure on \((\Omega, \mathcal{G})\), that it is finite, and
that it is absolutely continuous with respect to \(\mathbb{P}|_{\mathcal{G}}\).
Verification (signed measure).
We check the conditions of
signed measure.
Clearly \(\nu_X(\emptyset) = \int_\emptyset X \, d\mathbb{P} = 0\). For countable additivity,
let \((A_n)_{n \geq 1}\) be a sequence of pairwise disjoint sets in \(\mathcal{G}\) and write
\(A = \bigsqcup_n A_n\). The sequence \(S_N = \sum_{n=1}^N X \mathbf{1}_{A_n}\) converges
\(\mathbb{P}\)-a.s. to \(X \mathbf{1}_A\), and is dominated in absolute value by
\(|X| \in L^1(\mathbb{P})\). The
dominated convergence theorem
gives
\[
\begin{align*}
\nu_X(A)
&= \int_A X \, d\mathbb{P} \\\\
&= \int_\Omega X \mathbf{1}_A \, d\mathbb{P} \\\\
&= \lim_{N \to \infty} \sum_{n=1}^N \int_{A_n} X \, d\mathbb{P} \\\\
&= \sum_{n=1}^\infty \nu_X(A_n),
\end{align*}
\]
and the series converges absolutely because
\[
\begin{align*}
\sum_n |\nu_X(A_n)|
&\leq \sum_n \int_{A_n} |X| \, d\mathbb{P} \\\\
&= \int_A |X| \, d\mathbb{P} \\\\
&\leq \mathbb{E}[|X|] \\\\
&\lt \infty.
\end{align*}
\]
Since \(X \in L^1\), \(\nu_X\) takes values in \(\mathbb{R}\) (never \(\pm \infty\)), so the
sign-restriction condition is trivially satisfied.
Verification (finiteness).
The Jordan decomposition
gives \(\nu_X = \nu_X^+ - \nu_X^-\), where both parts are non-negative measures on
\(\mathcal{G}\). If \(\Omega = P \sqcup N\) is a
Hahn decomposition
for \(\nu_X\), with \(P, N \in \mathcal{G}\), the proof of the Jordan decomposition constructs
the two parts as \(\nu_X^+(A) = \nu_X(A \cap P)\) and \(\nu_X^-(A) = -\nu_X(A \cap N)\). The
total variation
\(|\nu_X| = \nu_X^+ + \nu_X^-\) therefore satisfies
\[
\begin{align*}
|\nu_X|(\Omega)
&= \nu_X(P) - \nu_X(N) \\\\
&= \int_P X \, d\mathbb{P} - \int_N X \, d\mathbb{P} \\\\
&\leq \int_P |X| \, d\mathbb{P} + \int_N |X| \, d\mathbb{P} \\\\
&= \mathbb{E}[|X|] \\\\
&\lt \infty.
\end{align*}
\]
Hence \(\nu_X\) is a finite signed measure. The inequality can be strict, because \(P\) and
\(N\) belong to \(\mathcal{G}\) while \(X\) is only \(\mathcal{F}\)-measurable and may change
sign inside them.
Verification (absolute continuity).
Let \(A \in \mathcal{G}\) with \(\mathbb{P}|_{\mathcal{G}}(A) = 0\), that is,
\(\mathbb{P}(A) = 0\) (the restricted measure agrees with \(\mathbb{P}\) on
\(\mathcal{G}\)-sets by definition). Then \(X \mathbf{1}_A = 0\) \(\mathbb{P}\)-a.s., whence
\(\nu_X(A) = \int_A X \, d\mathbb{P} = 0\). By the definition of
absolute continuity,
\(\nu_X \ll \mathbb{P}|_{\mathcal{G}}\).
The ingredients for the Radon-Nikodym theorem are now in place. The measure
\(\mathbb{P}|_{\mathcal{G}}\) is finite (hence \(\sigma\)-finite) and non-negative on
\((\Omega, \mathcal{G})\), and \(\nu_X\) is a finite signed measure absolutely continuous with
respect to it. The theorem is stated and proved for non-negative \(\nu\) only, and the proof below
applies it to the two Jordan parts of \(\nu_X\). The \(\mathcal{G}\)-measurable density so obtained
represents \(\nu_X\) as an integral against \(\mathbb{P}|_{\mathcal{G}}\), and we take it as the
definition of the conditional expectation.
The Conditional Expectation
Theorem & Definition: Conditional Expectation
Let \(X \in L^1(\Omega, \mathcal{F}, \mathbb{P})\) and \(\mathcal{G} \subseteq \mathcal{F}\) a
sub-\(\sigma\)-algebra. There exists a \(\mathcal{G}\)-measurable function
\(Y \in L^1(\Omega, \mathcal{G}, \mathbb{P}|_{\mathcal{G}})\), unique up to \(\mathbb{P}\)-a.s.
equality, satisfying the averaging identity
\[
\int_A Y \, d\mathbb{P} = \int_A X \, d\mathbb{P} \quad
\text{for every } A \in \mathcal{G}. \tag{$\ast$}
\]
Any such \(Y\) is called a version of the conditional expectation of \(X\) given
\(\mathcal{G}\), and we write
\[
\mathbb{E}[X \mid \mathcal{G}] = Y,
\]
or, equivalently,
\[
\mathbb{E}[X \mid \mathcal{G}] = \frac{d\nu_X}{d \mathbb{P}|_{\mathcal{G}}}.
\]
Proof.
By the verifications above, \(\nu_X\) is a finite signed measure on \((\Omega, \mathcal{G})\)
with \(\nu_X \ll \mathbb{P}|_{\mathcal{G}}\). Both Jordan parts inherit the absolute
continuity. With \(P, N\) the Hahn sets used above, if \(\mathbb{P}(A) = 0\) then
\(A \cap P\) and \(A \cap N\) are \(\mathbb{P}\)-null sets of \(\mathcal{G}\), so
\(\nu_X^+(A) = \nu_X(A \cap P) = 0\) and \(\nu_X^-(A) = -\nu_X(A \cap N) = 0\). Both parts
are finite, hence \(\sigma\)-finite, by the finiteness verification. Apply the
Radon-Nikodym theorem
separately to the Jordan parts \(\nu_X^+, \nu_X^-\) of \(\nu_X\). The resulting non-negative
densities \(f^+, f^-\) belong to \(L^1(\mathbb{P}|_{\mathcal{G}})\) because
\(\int f^\pm \, d\mathbb{P}|_{\mathcal{G}} = \nu_X^\pm(\Omega) \lt \infty\). Being
integrable, \(f^+\) and \(f^-\) are finite \(\mathbb{P}\)-a.s., and we redefine both as \(0\)
on the \(\mathbb{P}\)-null set in \(\mathcal{G}\) where either is infinite. This change
affects none of the integrals. Set \(Y = f^+ - f^-\). Then \(Y\) is
\(\mathcal{G}\)-measurable, integrable, and
\[
\begin{align*}
\int_A Y \, d\mathbb{P}
&= \int_A (f^+ - f^-) \, d\mathbb{P}|_{\mathcal{G}} \\\\
&= \nu_X^+(A) - \nu_X^-(A) \\\\
&= \nu_X(A) \\\\
&= \int_A X \, d\mathbb{P}
\end{align*}
\]
for every \(A \in \mathcal{G}\), establishing (\(\ast\)). The first equality uses the fact,
which we take as given, that a \(\mathcal{G}\)-measurable function has the same integral
against \(\mathbb{P}\) and against \(\mathbb{P}|_{\mathcal{G}}\). The two measures agree on
\(\mathcal{G}\), hence on \(\mathcal{G}\)-measurable simple functions, and the integral is
built from these.
Uniqueness up to \(\mathbb{P}\)-a.s. equality can be checked directly. If \(Y'\) is another
version, then \(\int_A (Y - Y') \, d\mathbb{P} = 0\) for every \(A \in \mathcal{G}\). Taking
\(A = \{Y - Y' \gt 0\} \in \mathcal{G}\) gives a strictly positive integrand on \(A\) with zero
integral, which forces \(\mathbb{P}(A) = 0\). The set \(\{Y - Y' \lt 0\}\) is handled
symmetrically, and hence \(Y = Y'\) \(\mathbb{P}\)-a.s.
Two remarks on this definition are essential, and both will be invoked repeatedly.
The averaging identity is the working characterisation. Although the construction
goes through Radon-Nikodym, most subsequent proofs on this page use the averaging identity
(\(\ast\)) directly, and the rest build on properties proved that way. The pattern is the same each
time. To verify that a candidate function \(Y\) is a version of \(\mathbb{E}[X \mid \mathcal{G}]\),
one shows that \(Y\) is \(\mathcal{G}\)-measurable and that
\(\int_A Y \, d\mathbb{P} = \int_A X \, d\mathbb{P}\) for all \(A \in \mathcal{G}\). The
a.s.-uniqueness clause then identifies \(Y\) as the conditional expectation. Radon-Nikodym is the
existence engine, and the averaging identity is the daily tool.
The values of \(\mathbb{E}[X \mid \mathcal{G}](\omega)\) are defined only up to a
\(\mathbb{P}\)-null set. Two versions of \(\mathbb{E}[X \mid \mathcal{G}]\) can disagree
on any set \(N \in \mathcal{G}\) with \(\mathbb{P}(N) = 0\), and they will both be valid
representatives. A statement such as "\(\mathbb{E}[X \mid \mathcal{G}](\omega_0) = c\)" for a
particular \(\omega_0 \in \Omega\) is therefore not meaningful in isolation whenever \(\omega_0\)
lies in a \(\mathbb{P}\)-null set of \(\mathcal{G}\), as every point does when \(\mathcal{G}\) is
generated by a continuous random variable. Such a statement acquires meaning only after a specific
version has been fixed.
This subtlety is the seed of regular conditional distributions. Their central object is a "regular
version", a coherent choice of representative that behaves well as a function of \(\omega\) and
yields an honest probability measure on \(\mathcal{F}\) for each \(\omega\).
Recovery of the Discrete Case
Let \(\{A_i\}_{i \geq 1}\) be a countable measurable partition of \(\Omega\) and let
\(\mathcal{G} = \sigma(\{A_i\}_{i \geq 1})\) be the sub-\(\sigma\)-algebra it generates.
The elements of \(\mathcal{G}\) are precisely the countable unions of the partition
blocks. A \(\mathcal{G}\)-measurable function is constant on each \(A_i\), so any
candidate version of \(\mathbb{E}[X \mid \mathcal{G}]\) is determined by its constant
value on each block.
Proposition: Discrete Case
With \(\mathcal{G} = \sigma(\{A_i\}_{i \geq 1})\) for a countable measurable partition
\(\{A_i\}\), and for any \(X \in L^1(\mathbb{P})\), one has
\[
\begin{align*}
\mathbb{E}[X \mid \mathcal{G}](\omega)
&= \frac{1}{\mathbb{P}(A_i)} \int_{A_i} X \, d\mathbb{P} \\\\
&= \mathbb{E}[X \mid A_i] \quad \text{for } \omega \in A_i,
\end{align*}
\]
on every block \(A_i\) with \(\mathbb{P}(A_i) \gt 0\). On blocks with \(\mathbb{P}(A_i) = 0\),
the value may be any constant (consistent with a.s.-uniqueness).
Proof.
Define \(Y(\omega) = \mathbb{E}[X \mid A_i]\) for \(\omega \in A_i\) when
\(\mathbb{P}(A_i) \gt 0\), and \(Y(\omega) = 0\) otherwise. Then \(Y\) is constant on each
\(A_i\) and is therefore \(\mathcal{G}\)-measurable. It is integrable, because
\[
\begin{align*}
\sum_{i : \, \mathbb{P}(A_i) \gt 0} |\mathbb{E}[X \mid A_i]| \, \mathbb{P}(A_i)
&\leq \sum_i \int_{A_i} |X| \, d\mathbb{P} \\\\
&\leq \mathbb{E}[|X|].
\end{align*}
\]
To verify the averaging identity, let \(A \in \mathcal{G}\). Then
\(A = \bigsqcup_{i \in I} A_i\) for some countable index set \(I\), and
\[
\begin{align*}
\int_A Y \, d\mathbb{P}
&= \sum_{i \in I, \, \mathbb{P}(A_i) \gt 0} \mathbb{E}[X \mid A_i] \cdot \mathbb{P}(A_i) \\\\
&= \sum_{i \in I, \, \mathbb{P}(A_i) \gt 0} \int_{A_i} X \, d\mathbb{P} \\\\
&= \int_A X \, d\mathbb{P},
\end{align*}
\]
where the first and last equalities use the countable additivity verified above, applied to the
integrable functions \(Y\) and \(X\). The last one also discards blocks of probability zero,
which contribute nothing to the integral of \(X\) either. By the a.s.-uniqueness clause of the
conditional expectation, \(Y\) is a version of \(\mathbb{E}[X \mid \mathcal{G}]\).
The proposition recovers the elementary conditional expectation given an event, the first formula
of the introduction. When \(\mathcal{G}\) is generated by a countable partition, the abstract
definition reduces, block by block, to that formula. The novelty of the abstract definition lies in
the cases that the elementary formula does not cover.
The Continuous Case Requires One More Layer
Let \(Y\) be a real-valued random variable and consider \(\mathcal{G} = \sigma(Y)\), the
sub-\(\sigma\)-algebra generated by \(Y\). The conditional expectation
\(\mathbb{E}[X \mid \sigma(Y)]\), abbreviated \(\mathbb{E}[X \mid Y]\), is a
\(\sigma(Y)\)-measurable random variable on \(\Omega\). The Doob-Dynkin lemma, which we use
without proof, states that every \(\sigma(Y)\)-measurable real-valued random variable has the form
\(g(Y)\) for some Borel-measurable \(g : \mathbb{R} \to \mathbb{R}\). Applied to
\(\mathbb{E}[X \mid Y]\), it yields such a \(g\), determined up to null sets of the
distribution \(P_Y\)
on \(\mathbb{R}\). The function \(g\) is what one would like to call
"\(y \mapsto \mathbb{E}[X \mid Y = y]\)", a deterministic forecast of \(X\) for each observed
value of \(Y\).
The notation \(\mathbb{E}[X \mid Y = y]\) thus has a meaning, with one caveat. As a function, \(g\)
is determined only up to \(P_Y\)-null sets. For continuous \(Y\), every singleton \(\{y_0\}\) has
\(P_Y\)-measure zero, so the value \(g(y_0)\) at any specific point is not determined by the
abstract definition. Two versions of \(g\) can disagree on a \(P_Y\)-null set and both remain
valid, so pointwise statements are again meaningful only up to such null sets.
When \((X, Y)\) has a joint density \(p(x, y)\), the second formula of the introduction supplies a
version of \(g\). Write \(p_Y(y) = \int p(x, y) \, dx\) for the marginal density and
\(p(x \mid y) = p(x, y) / p_Y(y)\) where \(p_Y(y) \gt 0\). Set
\(g(y) = \int x \, p(x \mid y) \, dx\) where \(p_Y(y) \gt 0\) and the integral converges
absolutely, and \(g(y) = 0\) elsewhere. Every set of \(\sigma(Y)\) has the form \(\{Y \in B\}\)
with \(B\) Borel, and
\[
\begin{align*}
\int_{\{Y \in B\}} g(Y) \, d\mathbb{P}
&= \int_B g(y) \, p_Y(y) \, dy \\\\
&= \int_B \int x \, p(x, y) \, dx \, dy \\\\
&= \int_{\{Y \in B\}} X \, d\mathbb{P}.
\end{align*}
\]
The second equality is the definition of \(g\), together with the fact that \(p(x, y) = 0\) for
almost every \(x\) wherever \(p_Y(y) = 0\). The first and third express expectations through the
densities of \(Y\) and of \((X, Y)\), a standard step that we do not verify here. The third also
writes the double integral as an iterated one by
Fubini's theorem,
legitimate because \(X\) is integrable. The same computation with \(|x|\) in place of \(x\) bounds
\(\mathbb{E}[|g(Y)|]\) by \(\mathbb{E}[|X|]\). Hence \(g(Y)\) satisfies the averaging identity on
\(\sigma(Y)\) and is a version of \(\mathbb{E}[X \mid Y]\).
Most ML uses of the notation call for more. Bayesian inference over continuous parameters and the
Bellman equation on a continuous state space work with a conditional distribution given
\(Y = y\), not with a single conditional expectation. They require the versions of
\(\mathbb{P}(B \mid Y = y) = \mathbb{E}[\mathbf{1}_B \mid Y = y]\) to be chosen simultaneously for
all \(B \in \mathcal{F}\), so that \(B \mapsto \mathbb{P}(B \mid Y = y)\) is an honest probability
measure on \((\Omega, \mathcal{F})\) for every \(y\). Given that choice,
\(\mathbb{E}[X \mid Y = y]\) can be computed as an integral
\(\int X \, d\mathbb{P}(\cdot \mid Y = y)\) in the elementary sense.
Such a coherent choice is called a regular conditional distribution. Existence is
not automatic, since it requires a regularity hypothesis on the space carrying the conditional
measures (typically that \((\Omega, \mathcal{F})\) is a standard Borel space, the framework of the
disintegration theorem). The construction of regular conditional distributions is the central topic
of a separate strand of measure-theoretic probability, and we do not develop it on this page.
Nothing proved so far depends on this deferral. The abstract \(\sigma(Y)\)-measurable function
\(\mathbb{E}[X \mid Y]\) exists and is unique a.s. by the construction above. Only its
pointwise-coherent reading as a function of \(y\) requires additional machinery, together with
the conditional measures \(\mathbb{P}(\cdot \mid Y = y)\) that allow integrals over the fibre to
be computed directly.
Properties of Conditional Expectation
The averaging identity (\(\ast\)), together with the a.s.-uniqueness clause, remains the working
tool for the algebraic properties, the projection theorem, and the tower property, used in the
pattern described after the definition. The inequalities are then derived from the algebraic
properties. We collect the algebraic properties first, then inequalities, then the geometric
\(L^2\) characterisation, and finally the tower property, the structural identity behind martingale
theory and dynamic programming.
We begin with a note on scope. Throughout this page, \(X \in L^1(\mathbb{P})\) is integrable, and
every instance of \(\mathbb{E}[X \mid \mathcal{G}]\) is consequently a finite-valued random
variable. Some standard treatments first define \(\mathbb{E}[X \mid \mathcal{G}]\) for non-negative
\(X\) (allowing the value \(+\infty\) via the \(\sigma\)-finite Radon-Nikodym theorem) and then
extend to \(L^1\) via the decomposition \(X = X^+ - X^-\). On \(L^1\), the resulting object
coincides with ours, and the \(L^1\)-first restriction adopted here keeps every quantity on the
page finite by construction.
Two intermediate steps handle possibly infinite quantities before this finiteness is secured. The
existence proof receives \([0, \infty]\)-valued densities from the Radon-Nikodym theorem and makes
them finite by a change on a null set. In Step 3 of the take-out rule, the monotone convergence
theorem delivers an identity in \([0, \infty]\), and the hypothesis \(ZX \in L^1(\mathbb{P})\)
makes both sides finite.
Algebraic Properties
Theorem: Linearity
Let \(X, Y \in L^1(\mathbb{P})\) and \(a, b \in \mathbb{R}\). Then
\[
\mathbb{E}[aX + bY \mid \mathcal{G}] = a\, \mathbb{E}[X \mid \mathcal{G}]
+ b\, \mathbb{E}[Y \mid \mathcal{G}] \quad \mathbb{P}\text{-a.s.}
\]
Proof.
The function \(Z = a\, \mathbb{E}[X \mid \mathcal{G}] + b\, \mathbb{E}[Y \mid \mathcal{G}]\) is
\(\mathcal{G}\)-measurable (linear combination of \(\mathcal{G}\)-measurable functions) and
integrable. For \(A \in \mathcal{G}\),
\[
\begin{align*}
\int_A Z \, d\mathbb{P}
&= a \int_A \mathbb{E}[X \mid \mathcal{G}] \, d\mathbb{P} + b \int_A \mathbb{E}[Y \mid \mathcal{G}] \, d\mathbb{P} \\\\
&= a \int_A X \, d\mathbb{P} + b \int_A Y \, d\mathbb{P} \\\\
&= \int_A (aX + bY) \, d\mathbb{P},
\end{align*}
\]
using linearity of the Lebesgue integral and the averaging identity for each summand. The
a.s.-uniqueness clause identifies \(Z\) as a version of
\(\mathbb{E}[aX + bY \mid \mathcal{G}]\).
Theorem: Monotonicity
If \(X, Y \in L^1(\mathbb{P})\) and \(X \leq Y\) \(\mathbb{P}\)-a.s., then
\[
\mathbb{E}[X \mid \mathcal{G}] \leq \mathbb{E}[Y \mid \mathcal{G}]
\quad \mathbb{P}\text{-a.s.}
\]
Proof.
Set \(D = \mathbb{E}[X \mid \mathcal{G}] - \mathbb{E}[Y \mid \mathcal{G}]\), a
\(\mathcal{G}\)-measurable function. We want to show \(D \leq 0\) \(\mathbb{P}\)-a.s. Let
\(A = \{D \gt 0\} \in \mathcal{G}\). By linearity and the averaging identity,
\[
\begin{align*}
\int_A D \, d\mathbb{P}
&= \int_A \mathbb{E}[X \mid \mathcal{G}] \, d\mathbb{P} - \int_A \mathbb{E}[Y \mid \mathcal{G}] \, d\mathbb{P} \\\\
&= \int_A X \, d\mathbb{P} - \int_A Y \, d\mathbb{P} \\\\
&= \int_A (X - Y) \, d\mathbb{P} \\\\
&\leq 0,
\end{align*}
\]
since \(X - Y \leq 0\) \(\mathbb{P}\)-a.s. But \(D \gt 0\) on \(A\), so
\(\int_A D \, d\mathbb{P} \geq 0\), with strict inequality unless \(\mathbb{P}(A) = 0\). Hence
\(\mathbb{P}(A) = 0\), that is, \(D \leq 0\) \(\mathbb{P}\)-a.s.
Theorem: Take-out (Pull-out) of \(\mathcal{G}\)-Measurable Factors
Let \(X \in L^1(\mathbb{P})\) and let \(Z\) be a \(\mathcal{G}\)-measurable random variable
such that \(ZX \in L^1(\mathbb{P})\). Then
\[
\mathbb{E}[ZX \mid \mathcal{G}] = Z \cdot \mathbb{E}[X \mid \mathcal{G}]
\quad \mathbb{P}\text{-a.s.}
\]
Proof (standard three-step extension).
We verify the averaging identity for \(Z \cdot \mathbb{E}[X \mid \mathcal{G}]\) in
three stages: indicator, simple, then general \(\mathcal{G}\)-measurable.
Step 1 (indicator). Let \(Z = \mathbf{1}_B\) for \(B \in \mathcal{G}\). For any
\(A \in \mathcal{G}\), \(A \cap B \in \mathcal{G}\), so
\[
\begin{align*}
\int_A \mathbf{1}_B \cdot \mathbb{E}[X \mid \mathcal{G}] \, d\mathbb{P}
&= \int_{A \cap B} \mathbb{E}[X \mid \mathcal{G}] \, d\mathbb{P} \\\\
&= \int_{A \cap B} X \, d\mathbb{P} \\\\
&= \int_A \mathbf{1}_B X \, d\mathbb{P},
\end{align*}
\]
and \(\mathbf{1}_B \cdot \mathbb{E}[X \mid \mathcal{G}]\) is \(\mathcal{G}\)-measurable as a
product of \(\mathcal{G}\)-measurable functions. The averaging identity holds.
Step 2 (simple non-negative). By linearity (already proved), the identity
extends to non-negative simple \(Z = \sum_{k=1}^n c_k \mathbf{1}_{B_k}\) with \(B_k \in \mathcal{G}\)
and \(c_k \geq 0\).
Step 3 (general). First take \(X \geq 0\). Let \(Z \geq 0\) be
\(\mathcal{G}\)-measurable, and choose non-negative simple \(\mathcal{G}\)-measurable functions
\(Z_n \uparrow Z\) (the standard simple-function approximation, whose existence we assume).
Then \(Z_n X \uparrow ZX\) \(\mathbb{P}\)-a.s., and
\(Z_n \cdot \mathbb{E}[X \mid \mathcal{G}] \uparrow Z \cdot \mathbb{E}[X \mid \mathcal{G}]\)
\(\mathbb{P}\)-a.s. (using \(\mathbb{E}[X \mid \mathcal{G}] \geq 0\) by monotonicity, since
\(X \geq 0\)). The
monotone convergence theorem
applied to both sides of the Step-2 identity, integrated over an arbitrary
\(A \in \mathcal{G}\), gives
\[
\int_A Z \cdot \mathbb{E}[X \mid \mathcal{G}] \, d\mathbb{P}
= \int_A ZX \, d\mathbb{P}.
\]
Taking \(A = \Omega\) shows that \(Z \cdot \mathbb{E}[X \mid \mathcal{G}]\) is integrable
whenever \(ZX\) is, and the identity then makes it a version of
\(\mathbb{E}[ZX \mid \mathcal{G}]\).
For general \(X \in L^1\), decompose \(X = X^+ - X^-\) and \(Z = Z^+ - Z^-\) and apply the
non-negative case to each of the four products. The integrability hypothesis \(ZX \in L^1\)
ensures that each piece is integrable, since \(|Z^\pm X^\pm| \leq |ZX|\). Linearity (already
proved for conditional expectation) reassembles the four pieces. Therefore
\(Z \cdot \mathbb{E}[X \mid \mathcal{G}]\) is a version of \(\mathbb{E}[ZX \mid \mathcal{G}]\).
Theorem: Independence Collapse
If \(X \in L^1(\mathbb{P})\) and \(\sigma(X)\) is independent of \(\mathcal{G}\), then
\[
\mathbb{E}[X \mid \mathcal{G}] = \mathbb{E}[X] \quad \mathbb{P}\text{-a.s.}
\]
Proof.
The constant function \(\mathbb{E}[X]\) is \(\mathcal{G}\)-measurable. For
\(A \in \mathcal{G}\), the random variables \(X\) and \(\mathbf{1}_A\) are independent, because
every event \(\{X \in B_1\}\) lies in \(\sigma(X)\) and every event
\(\{\mathbf{1}_A \in B_2\}\) lies in \(\mathcal{G}\). By the
product formula for expectations,
\(\mathbb{E}[X \mathbf{1}_A] = \mathbb{E}[X] \mathbb{E}[\mathbf{1}_A]
= \mathbb{E}[X] \cdot \mathbb{P}(A)\), hence
\[
\begin{align*}
\int_A \mathbb{E}[X] \, d\mathbb{P}
&= \mathbb{E}[X] \cdot \mathbb{P}(A) \\\\
&= \mathbb{E}[X \mathbf{1}_A] \\\\
&= \int_A X \, d\mathbb{P}.
\end{align*}
\]
The a.s.-uniqueness clause identifies the constant \(\mathbb{E}[X]\) as a version of
\(\mathbb{E}[X \mid \mathcal{G}]\).
Inequalities
Theorem: Jensen's Inequality for Conditional Expectation
Let \(\varphi : \mathbb{R} \to \mathbb{R}\) be convex, and let \(X \in L^1(\mathbb{P})\) with
\(\varphi(X) \in L^1(\mathbb{P})\). Then
\[
\varphi\big(\mathbb{E}[X \mid \mathcal{G}]\big) \leq
\mathbb{E}[\varphi(X) \mid \mathcal{G}]
\quad \mathbb{P}\text{-a.s.}
\]
Proof (supporting-line argument).
For a convex function \(\varphi : \mathbb{R} \to \mathbb{R}\), every point
\(x_0 \in \mathbb{R}\) admits a supporting affine function. That is, there exist
\(a, b \in \mathbb{R}\) (depending on \(x_0\)) with \(\varphi(x_0) = a x_0 + b\) and
\(\varphi(x) \geq a x + b\) for all \(x \in \mathbb{R}\). Moreover, since \(\varphi\) is convex
on all of \(\mathbb{R}\), it is the pointwise supremum of a countable family of affine
functions. Explicitly, there exist sequences \((a_n), (b_n) \subset \mathbb{R}\) with
\[
\varphi(x) = \sup_{n \in \mathbb{N}} (a_n x + b_n) \quad \text{for all } x \in \mathbb{R}.
\]
One construction takes an affine function supporting \(\varphi\) at each rational \(x_0\). Any
slope between the left and right derivatives of \(\varphi\) at \(x_0\) gives one. We rely
without proof on two standard facts about a convex function on \(\mathbb{R}\). Its one-sided
derivatives exist and are bounded on bounded intervals, and \(\varphi\) is continuous. Every
supporting function lies below \(\varphi\), so the supremum is at most \(\varphi\). Conversely,
fix \(x\) and let rationals \(q_k \to x\). The supporting function at \(q_k\), evaluated at
\(x\), equals \(\varphi(q_k) + a_k (x - q_k)\), where \(a_k\) is the slope chosen at \(q_k\).
The slopes \(a_k\) are bounded, so this value tends to \(\varphi(x)\) by continuity.
For each \(n\), apply linearity and monotonicity of conditional expectation to the affine
inequality \(\varphi(X) \geq a_n X + b_n\):
\[
\begin{align*}
\mathbb{E}[\varphi(X) \mid \mathcal{G}]
&\geq \mathbb{E}[a_n X + b_n \mid \mathcal{G}] \\\\
&= a_n\, \mathbb{E}[X \mid \mathcal{G}] + b_n \quad \mathbb{P}\text{-a.s.}
\end{align*}
\]
Here the constant \(b_n\) is its own conditional expectation, being \(\mathcal{G}\)-measurable
and satisfying (\(\ast\)) trivially. The exceptional null set may depend on \(n\), but the
union over the countable index set is still null. Outside this single null set,
\[
\begin{align*}
\mathbb{E}[\varphi(X) \mid \mathcal{G}]
&\geq \sup_n \big( a_n \mathbb{E}[X \mid \mathcal{G}] + b_n \big) \\\\
&= \varphi\big(\mathbb{E}[X \mid \mathcal{G}]\big),
\end{align*}
\]
which is the asserted inequality.
Theorem: \(L^p\) Contraction
For \(1 \leq p \lt \infty\) and \(X \in L^p(\Omega, \mathcal{F}, \mathbb{P})\),
\[
\big\| \mathbb{E}[X \mid \mathcal{G}] \big\|_p \leq \|X\|_p.
\]
In particular, conditional expectation is a contraction on \(L^p(\mathbb{P})\).
Proof.
On a probability space \(L^p(\mathbb{P}) \subseteq L^1(\mathbb{P})\), as shown in
our measure-theoretic treatment of expectation,
so \(\mathbb{E}[X \mid \mathcal{G}]\) is defined. The function \(\varphi(t) = |t|^p\) is convex
on \(\mathbb{R}\) for \(p \geq 1\), and \(\varphi(X) = |X|^p \in L^1(\mathbb{P})\) by the
assumption \(X \in L^p(\mathbb{P})\). Apply Jensen's inequality for conditional expectation:
\[
\big| \mathbb{E}[X \mid \mathcal{G}] \big|^p \leq
\mathbb{E}[|X|^p \mid \mathcal{G}] \quad \mathbb{P}\text{-a.s.}
\]
Take expectations of both sides. On the right, the
tower-with-trivial-\(\sigma\)-algebra identity
\(\mathbb{E}[\mathbb{E}[Z \mid \mathcal{G}]] = \mathbb{E}[Z]\) (the averaging identity applied
to \(A = \Omega \in \mathcal{G}\)) gives
\(\mathbb{E}[\mathbb{E}[|X|^p \mid \mathcal{G}]] = \mathbb{E}[|X|^p]\). Hence
\[
\begin{align*}
\big\| \mathbb{E}[X \mid \mathcal{G}] \big\|_p^p
&= \mathbb{E}\big[ \big| \mathbb{E}[X \mid \mathcal{G}] \big|^p \big] \\\\
&\leq \mathbb{E}[|X|^p] \\\\
&= \|X\|_p^p,
\end{align*}
\]
and taking \(p\)-th roots gives the claim.
The \(L^2\) Projection Characterisation
At \(p = 2\) the contraction acquires geometric content. The space
\(L^2(\Omega, \mathcal{G}, \mathbb{P}|_{\mathcal{G}})\) of square-integrable
\(\mathcal{G}\)-measurable functions sits inside \(L^2(\Omega, \mathcal{F}, \mathbb{P})\) as a
closed linear subspace, namely the classes that have a \(\mathcal{G}\)-measurable representative.
It is closed because an \(L^2\)-convergent sequence of \(\mathcal{G}\)-measurable functions has a
subsequence converging a.s.
The set where this subsequence converges belongs to \(\mathcal{G}\) and has probability one, so the
pointwise limit on that set, extended by \(0\) elsewhere, is a \(\mathcal{G}\)-measurable
representative of the \(L^2\) limit. Conditional expectation, restricted to \(L^2\), is precisely
the orthogonal projection onto this subspace.
Theorem: Conditional Expectation as \(L^2\) Projection
Let \(X \in L^2(\Omega, \mathcal{F}, \mathbb{P})\). Then \(\mathbb{E}[X \mid \mathcal{G}]\) is
the orthogonal projection of \(X\) onto the closed subspace
\(L^2(\Omega, \mathcal{G}, \mathbb{P}|_{\mathcal{G}}) \subseteq L^2(\Omega, \mathcal{F}, \mathbb{P})\).
Equivalently, \(\mathbb{E}[X \mid \mathcal{G}]\) is, up to \(\mathbb{P}\)-a.s. equality, the
unique \(\mathcal{G}\)-measurable square-integrable function minimising
\[
\mathbb{E}\big[ (X - Y)^2 \big] \quad \text{over } Y \in L^2(\Omega, \mathcal{G}, \mathbb{P}|_{\mathcal{G}}).
\]
Proof.
The Hilbert projection theorem
applied to the closed subspace \(M = L^2(\Omega, \mathcal{G}, \mathbb{P}|_{\mathcal{G}})\) of
the Hilbert space \(L^2(\Omega, \mathcal{F}, \mathbb{P})\), complete by
Riesz-Fischer,
produces a unique element \(P_M(X) \in M\) such that \(X - P_M(X) \perp M\). The same element
is the unique minimiser of \(\|X - Y\|_{L^2}\) over \(Y \in M\). Orthogonality means
\(\langle X - P_M(X), Z \rangle_{L^2} = 0\) for every \(Z \in M\), that is,
\[
\mathbb{E}\big[ (X - P_M(X)) \cdot Z \big] = 0 \quad \text{for all }
Z \in L^2(\Omega, \mathcal{G}, \mathbb{P}|_{\mathcal{G}}). \tag{$\dagger$}
\]
Specialising \(Z = \mathbf{1}_A\) for \(A \in \mathcal{G}\) (which is bounded, hence in
\(L^2\), and \(\mathcal{G}\)-measurable), (\(\dagger\)) reduces to
\[
\int_A (X - P_M(X)) \, d\mathbb{P} = 0,
\]
that is,
\[
\int_A P_M(X) \, d\mathbb{P} = \int_A X \, d\mathbb{P}.
\]
Thus \(P_M(X)\) is \(\mathcal{G}\)-measurable, integrable (since \(L^2 \subseteq L^1\) on a
finite measure space), and satisfies the averaging identity on every \(A \in \mathcal{G}\), so
the a.s.-uniqueness clause makes \(P_M(X)\) a version of \(\mathbb{E}[X \mid \mathcal{G}]\).
The \(L^2\) projection identification is the geometric face of conditional expectation. It explains
in one stroke why \(\mathbb{E}[X \mid \mathcal{G}]\) is the minimum-mean-square forecast of \(X\)
based on the information \(\mathcal{G}\), because orthogonal projection minimises distance and
squared \(L^2\)-distance is mean-square error.
The "best linear predictor" of classical statistics (linear regression, the Wiener filter, the
Kalman update) is the orthogonal projection onto a smaller subspace, the affine functions of the
observations. It agrees with the conditional expectation when the variables are jointly Gaussian, a
fact we do not prove here. In general it is the best approximation of the conditional expectation
within that subspace, because projecting onto a subspace of \(L^2(\mathcal{G})\) factors through
the projection onto \(L^2(\mathcal{G})\). The same geometric picture also makes the next property,
the tower property, visually obvious, as we note after its proof.
The Tower Property
Theorem: Tower Property
Let \(\mathcal{H} \subseteq \mathcal{G} \subseteq \mathcal{F}\) be sub-\(\sigma\)-algebras and
\(X \in L^1(\mathbb{P})\). Then
\[
\mathbb{E}\big[\, \mathbb{E}[X \mid \mathcal{G}] \,\big|\, \mathcal{H} \,\big]
= \mathbb{E}[X \mid \mathcal{H}] \quad \mathbb{P}\text{-a.s.}
\]
In particular, \(\mathbb{E}\big[\mathbb{E}[X \mid \mathcal{G}]\big] = \mathbb{E}[X]\) (taking
\(\mathcal{H} = \{\emptyset, \Omega\}\)).
Proof.
Write \(W = \mathbb{E}[X \mid \mathcal{G}]\) and let \(V = \mathbb{E}[W \mid \mathcal{H}]\),
which is \(\mathcal{H}\)-measurable and integrable by construction. We verify that \(V\)
satisfies the averaging identity for \(\mathbb{E}[X \mid \mathcal{H}]\). For every
\(A \in \mathcal{H}\),
\[
\int_A V \, d\mathbb{P}
\stackrel{(\mathrm{i})}{=} \int_A W \, d\mathbb{P}
\stackrel{(\mathrm{ii})}{=} \int_A X \, d\mathbb{P},
\]
where (i) is the averaging identity for \(V = \mathbb{E}[W \mid \mathcal{H}]\) on
\(A \in \mathcal{H}\), and (ii) is the averaging identity for
\(W = \mathbb{E}[X \mid \mathcal{G}]\) on \(A\), valid because
\(A \in \mathcal{H} \subseteq \mathcal{G}\). The a.s.-uniqueness clause identifies \(V\) as a
version of \(\mathbb{E}[X \mid \mathcal{H}]\).
The tower property is the structural identity that drives iterated conditioning. Read in the
projection picture, for \(X \in L^2\), it says that projecting onto \(L^2(\mathcal{G})\) and then
onto the smaller subspace \(L^2(\mathcal{H})\) produces the same vector as projecting directly onto
\(L^2(\mathcal{H})\). Read in dynamic-programming terms, it says that the value of \(X\) under
coarse information \(\mathcal{H}\) can be computed by first computing the value under finer
information \(\mathcal{G}\) and then averaging that over \(\mathcal{H}\). This is the idealised
recursion structure of the Bellman equation, taken up in the next section.
Conditional Expectation in Practice
The first three ML scenarios below correspond, in idealised form, to a conditional expectation,
and the fourth, variational inference, is what replaces it when the conditional measure is
inaccessible. Where a conditional expectation appears, we identify the sub-\(\sigma\)-algebra
and read off which property of the previous section, if any, is being invoked. In code, the
object so identified is typically reached only through a density assumption combined with a
sampling- or surrogate-based approximation.
Expectation-Maximisation
The expectation-maximisation (EM) algorithm fits a parametric model \(p(x, z \mid \theta)\) with
observed data \(X\) and latent variable \(Z\) by alternating between two steps. Given a current
parameter estimate \(\theta^{(t)}\), the E-step computes
\[
Q(\theta \mid \theta^{(t)}) = \mathbb{E}\big[ \log p(X, Z \mid \theta) \,\big|\, X, \theta^{(t)} \big],
\]
where the conditional expectation is taken with respect to the conditional distribution of \(Z\)
given the observed \(X\) under parameter \(\theta^{(t)}\). The parameter \(\theta^{(t)}\) selects
the probability measure and is not a random variable. Only \(X\) is conditioned on, so
\(\mathcal{G} = \sigma(X)\). The M-step sets
\(\theta^{(t+1)} = \arg\max_\theta Q(\theta \mid \theta^{(t)})\).
The E-step is, in its idealised (density-based) form, a conditional expectation. When the
conditional density \(p(z \mid x, \theta^{(t)})\) exists, the expectation reduces to the explicit
integral \(\int \log p(x, z \mid \theta) \, p(z \mid x, \theta^{(t)}) \, dz\) that is implemented
in code, as a discrete sum when \(z\) takes finitely many values and as a Monte Carlo estimate
otherwise. The M-step is a finite-dimensional optimisation.
The monotonic improvement of the marginal log-likelihood \(\ell(\theta) = \log p(X \mid \theta)\)
under EM iterations follows from Jensen's inequality. Write
\(p(z \mid x, \theta) = p(x, z \mid \theta) / p(x \mid \theta)\), so that
\(\log p(x \mid \theta) = \log p(x, z \mid \theta) - \log p(z \mid x, \theta)\) for every \(z\)
with \(p(x, z \mid \theta) \gt 0\), provided \(p(x \mid \theta) \gt 0\). By the assumption, stated
below, that the two conditional densities of \(z\) at \(\theta\) and at \(\theta^{(t)}\) are
positive on the same set, this covers \(p(\cdot \mid x, \theta^{(t)})\)-almost every \(z\). We take
conditional expectations of both sides given \(X\) under \(p(\cdot \mid X, \theta^{(t)})\). Since
the left side does not depend on \(z\), it is unchanged, and we obtain
\[
\log p(X \mid \theta)
= \underbrace{\mathbb{E}\big[ \log p(X, Z \mid \theta) \,\big|\, X, \theta^{(t)} \big]}_{Q(\theta \mid \theta^{(t)})}
- \underbrace{\mathbb{E}\big[ \log p(Z \mid X, \theta) \,\big|\, X, \theta^{(t)} \big]}_{H(\theta \mid \theta^{(t)})}.
\]
Subtracting the same identity at \(\theta = \theta^{(t)}\) gives
\[
\ell(\theta) - \ell(\theta^{(t)})
= \big[ Q(\theta \mid \theta^{(t)}) - Q(\theta^{(t)} \mid \theta^{(t)}) \big]
+ \big[ H(\theta^{(t)} \mid \theta^{(t)}) - H(\theta \mid \theta^{(t)}) \big].
\]
The first bracket is non-negative for \(\theta = \theta^{(t+1)}\) by definition of the M-step. The
second bracket is non-negative by Jensen's inequality for \(\varphi(u) = -\log u\). This
\(\varphi\) is convex on \((0, \infty)\) only, so the conditional theorem of the previous section,
stated for \(\varphi\) on all of \(\mathbb{R}\), does not apply verbatim. For a fixed observed
value \(x\), however, the conditional expectation is an integral against the density
\(p(z \mid x, \theta^{(t)})\), and the
elementary Jensen inequality
on the interval \((0, \infty)\) suffices. We assume that \(p(z \mid x, \theta)\) and
\(p(z \mid x, \theta^{(t)})\) are positive on the same set of \(z\), so that the ratio below takes
values in \((0, \infty)\), and that the logarithms involved are integrable. With
\(H(\theta^{(t)} \mid \theta^{(t)}) - H(\theta \mid \theta^{(t)})
= \mathbb{E}\big[ -\log( p(Z \mid X, \theta) / p(Z \mid X, \theta^{(t)}) ) \,\big|\, X, \theta^{(t)} \big]\),
the inequality gives
\[
\begin{align*}
\mathbb{E}\big[ -\log( p(Z \mid X, \theta) / p(Z \mid X, \theta^{(t)}) ) \,\big|\, X, \theta^{(t)} \big]
&\geq -\log \mathbb{E}\big[ p(Z \mid X, \theta) / p(Z \mid X, \theta^{(t)}) \,\big|\, X, \theta^{(t)} \big] \\\\
&= -\log 1 \\\\
&= 0,
\end{align*}
\]
where the inner expectation evaluates to \(1\) by the explicit calculation
\[
\begin{align*}
\mathbb{E}\big[ p(Z \mid X, \theta) / p(Z \mid X, \theta^{(t)}) \,\big|\, X, \theta^{(t)} \big]
&= \int \frac{p(z \mid x, \theta)}{p(z \mid x, \theta^{(t)})} \, p(z \mid x, \theta^{(t)}) \, dz \\\\
&= \int p(z \mid x, \theta) \, dz \\\\
&= 1,
\end{align*}
\]
in which the conditioning density cancels and the remaining integrand is a probability density that
integrates to \(1\). This is the same algebraic structure as the importance-sampling identity
\(\mathbb{E}_q[f(Z) \, p(Z)/q(Z)] = \mathbb{E}_p[f(Z)]\), valid when \(q(z) \gt 0\) wherever
\(f(z)\, p(z) \neq 0\) and \(f\) is integrable under \(p\).
Both brackets in the earlier decomposition are non-negative, so
\(\ell(\theta^{(t+1)}) \geq \ell(\theta^{(t)})\). EM never decreases the marginal log-likelihood.
Reinforcement Learning: Value Functions and the Bellman Equation
In a Markov decision process with policy \(\pi\), the state-value function
\[
V_\pi(s) = \mathbb{E}\Big[\, \sum_{t=0}^\infty \gamma^t R_{t+1} \,\Big|\, S_0 = s \,\Big]
\]
is a conditional expectation of the discounted return given the initial state. The
Bellman expectation equation
\[
V_\pi(s) = \mathbb{E}\big[ R_1 + \gamma V_\pi(S_1) \,\big|\, S_0 = s \big]
\]
rests on the tower property. Assume bounded rewards and \(\gamma \lt 1\), so that the discounted
return \(X = \sum_t \gamma^t R_{t+1}\) is integrable, and a countable state space with
\(\mathbb{P}(S_0 = s) \gt 0\) for every state \(s\), so that conditioning on \(S_0 = s\) is the
elementary formula recovered in the discrete case. The \(\sigma\)-algebra structure is
\(\sigma(S_0) \subseteq \sigma(S_0, S_1)\), and the tower property gives
\[
\mathbb{E}[X \mid \sigma(S_0)] = \mathbb{E}\big[ \mathbb{E}[X \mid \sigma(S_0, S_1)] \,\big|\, \sigma(S_0) \big].
\]
The inner conditional expectation equals
\(\mathbb{E}[R_1 \mid \sigma(S_0, S_1)] + \gamma V_\pi(S_1)\). This step uses linearity together
with the Markov property and time-homogeneity of the process. The conditional expectation, given
\((S_0, S_1)\), of the discounted return from time \(1\) onward depends on \(S_1\) alone, through
the same function \(V_\pi\). Both properties belong to the MDP model and do not follow from the
tower property. A second use of the tower property,
\[
\mathbb{E}\big[ \mathbb{E}[R_1 \mid \sigma(S_0, S_1)] \,\big|\, \sigma(S_0) \big] = \mathbb{E}[R_1 \mid \sigma(S_0)],
\]
then yields the Bellman equation.
The full development of value functions, Bellman optimality, and the algorithmic apparatus (value
iteration, policy iteration, Q-learning) is on the
reinforcement learning page.
Bayesian Posterior Predictive
Given a Bayesian model with parameter \(\theta\), prior \(\pi(\theta)\), and observed data \(D\),
the posterior predictive distribution for a new observation \(X_{\text{new}}\) is
\[
p(x_{\text{new}} \mid D) = \mathbb{E}\big[ p(x_{\text{new}} \mid \theta) \,\big|\, D \big],
\]
a conditional expectation given the data. In the Bayesian model, \(\theta\) is a random variable on
the same probability space as \(D\) and \(X_{\text{new}}\), and \(\mathcal{G} = \sigma(D)\). The
identity is, in idealised form, the tower property for \(\sigma(D) \subseteq \sigma(\theta, D)\),
combined with the modelling assumption that \(X_{\text{new}}\) is conditionally independent of
\(D\) given \(\theta\). That assumption replaces \(p(x_{\text{new}} \mid \theta, D)\) by
\(p(x_{\text{new}} \mid \theta)\), and the outer conditioning on \(D\) "marginalises over
uncertainty in \(\theta\)".
When \(\theta\) is a continuous parameter, the pointwise reading \(p(x_{\text{new}} \mid D)\)
requires the regular-conditional-distribution machinery, which we do not develop on this page. In
practice this expectation is approximated by Markov chain Monte Carlo or by variational surrogates.
The take-out property licenses pulling deterministic functions of the data outside the conditional
expectation, and the tower property licenses hierarchical decompositions (for example, predicting
via an intermediate latent layer).
Variational Inference and the ELBO
Variational inference approximates an
intractable posterior \(p(z \mid x)\) by a tractable surrogate \(q(z \mid x)\). The
evidence lower bound
\(\mathrm{ELBO}(q) = \mathbb{E}_{q(z \mid x)}[\log p(x, z) - \log q(z \mid x)]\) is an expectation
against \(q\), not against \(p(\cdot \mid x)\). The
exact identity
\(\log p(x) = \mathrm{ELBO}(q) + D_{\mathrm{KL}}(q \,\|\, p(\cdot \mid x))\), valid when
\(q \ll p(\cdot \mid x)\) and the logarithms are \(q\)-integrable, certifies that maximising the
ELBO is equivalent to minimising the KL to the posterior. From the conditional-expectation
perspective, this is a manipulation that replaces the inaccessible conditional measure with a
tractable one. The surrogate gives up the averaging identity (\(\ast\)), which only the true
conditional measure satisfies, and the KL term records the discrepancy. The full development is on
the variational inference page.
Rigorous Foundation, Approximate Practice
A broader observation is worth recording, even at the cost of leaving rigorously verified
ground. The applications surveyed above all share a common pattern. Machine learning
implements a finite-sample, density-based approximation of an object whose rigorous existence
is licensed by the conditional expectation construction of this page, but the rigorous
construction itself is rarely instantiated in code. Monte Carlo replaces the integral, a
single sampled trajectory replaces the Bellman expectation, and a variational surrogate
replaces the intractable posterior. These approximations have carried machine learning through
a successful empirical era.
Contemporary large language models show several unstable behaviours, among them the
inconsistency of long chains of probabilistic reasoning, brittleness under distribution
shift, and difficulty in calibrating uncertainty. It is at least worth asking whether some of
these behaviours are connected to the absence of a rigorous measure-theoretic substrate
underneath the approximations.
The connection is hypothesised, not proved. The question of how rigorous mathematical structure
should enter the foundations of AI systems remains an active research direction, with measure
theory, topology, and differential geometry among the candidate frameworks. This curriculum is
built on the working assumption that the rigorous mathematical layer will become increasingly
relevant as the field matures. A reader who has internalised the construction on this page is
then better positioned to follow that line of development as it unfolds.