Conditional Expectation

Why Conditional Expectation? Definition via Radon-Nikodym Properties of Conditional Expectation Conditional Expectation in Practice

Why Conditional Expectation?

Elementary probability offers two distinct constructions that both go by the name conditional expectation. Given an event \(A\) with \(\mathbb{P}(A) \gt 0\), one writes \[ \mathbb{E}[X \mid A] = \frac{1}{\mathbb{P}(A)} \int_A X \, d\mathbb{P}, \] which is a single number, the average of \(X\) restricted to the event \(A\). Given a random variable \(Y\) such that \((X, Y)\) has a joint density \(p(x, y)\), one writes \[ \mathbb{E}[X \mid Y = y] = \int x \, p(x \mid y) \, dx, \] which is a function of \(y\).

These two formulas address visibly different situations. The first averages over a positive-probability event, and the second averages along a measure-zero fibre using a conditional density. Neither formula reduces to the other, and the second one is not even well-defined when \(p(x \mid y)\) fails to exist as a function (for example, when \(Y\) is mixed discrete-continuous or supported on a fractal).

A unifying definition exists, and its existence is the headline payoff of the Radon-Nikodym theorem. We read a sub-\(\sigma\)-algebra \(\mathcal{G} \subseteq \mathcal{F}\) as "the information available to an observer". Given an integrable random variable \(X\) and such a \(\mathcal{G}\), there is a \(\mathcal{G}\)-measurable random variable, unique up to almost-sure equality and denoted \(\mathbb{E}[X \mid \mathcal{G}]\), that simultaneously specialises to both classical formulas and continues to make sense in every intermediate situation. The unifying object is a function on \(\Omega\) rather than a number. For square-integrable \(X\), its values give the best forecast of \(X\), in the mean-square sense, that an observer with information \(\mathcal{G}\) can make.

The construction is short, because all the analytic machinery was set up in Signed Measures & Radon-Nikodym Theorem, following the preview given in Limit Theorems & Product Measures. This page collects the payoff.

The conditional expectation also underlies, in idealised form, several core constructions of machine learning, which the closing section takes up in detail. In expectation-maximisation, the E-step is exactly a conditional expectation. The Bellman equation of reinforcement learning rests on the tower property. Under the usual assumption that a new observation is conditionally independent of the observed data \(D\) given the parameter \(\theta\), the Bayesian posterior predictive \(p(x_{\text{new}} \mid D) = \mathbb{E}[p(x_{\text{new}} \mid \theta) \mid D]\) is a conditional expectation given \(D\), and the ELBO of variational inference replaces an intractable conditional measure by a tractable surrogate.

The rigorous object built here seldom appears directly in code, where Monte Carlo, single-trajectory stochastic approximation, or variational surrogates stand in for it. It is nonetheless the object those approximations approximate.

Before turning to the construction, we fix the notation for the rest of the page. The underlying probability space is \((\Omega, \mathcal{F}, \mathbb{P})\), and \(\mathcal{G}\) always denotes a sub-\(\sigma\)-algebra of \(\mathcal{F}\), a smaller collection of measurable sets representing partial information. The restriction of \(\mathbb{P}\) to \(\mathcal{G}\) is written \(\mathbb{P}|_{\mathcal{G}}\), the same probability measure regarded as defined only on the smaller \(\sigma\)-algebra. The object we are going to construct is denoted \(\mathbb{E}[X \mid \mathcal{G}]\), and when \(\mathcal{G} = \sigma(Y)\) we abbreviate it as \(\mathbb{E}[X \mid Y]\).

Throughout, "\(\mathbb{P}\)-a.s." means almost surely with respect to \(\mathbb{P}\). In the probabilistic context we use this qualifier in place of "\(\mathbb{P}\)-a.e.", consistent with earlier pages of this section.

Definition via Radon-Nikodym

The construction proceeds in three steps. We first associate to each integrable random variable \(X\) and each sub-\(\sigma\)-algebra \(\mathcal{G}\) a finite signed measure \(\nu_X\) on \(\mathcal{G}\). We then verify that \(\nu_X\) is absolutely continuous with respect to \(\mathbb{P}|_{\mathcal{G}}\), so that the Radon-Nikodym theorem applies. The resulting Radon-Nikodym derivative is, by definition, the conditional expectation. The discrete case is recovered immediately, and the continuous case is identified as requiring one further layer of machinery, the framework of regular conditional distributions and the disintegration theorem.

The Signed Measure Associated with \(X\)

Definition: The Signed Measure \(\nu_X\)

Let \(X \in L^1(\Omega, \mathcal{F}, \mathbb{P})\) and let \(\mathcal{G} \subseteq \mathcal{F}\) be a sub-\(\sigma\)-algebra. Define \[ \nu_X : \mathcal{G} \to \mathbb{R}, \quad \nu_X(A) = \int_A X \, d\mathbb{P}, \quad A \in \mathcal{G}. \]

Three properties of \(\nu_X\) must be checked before the Radon-Nikodym theorem can be invoked: that \(\nu_X\) is a signed measure on \((\Omega, \mathcal{G})\), that it is finite, and that it is absolutely continuous with respect to \(\mathbb{P}|_{\mathcal{G}}\).

Verification (signed measure).

We check the conditions of signed measure. Clearly \(\nu_X(\emptyset) = \int_\emptyset X \, d\mathbb{P} = 0\). For countable additivity, let \((A_n)_{n \geq 1}\) be a sequence of pairwise disjoint sets in \(\mathcal{G}\) and write \(A = \bigsqcup_n A_n\). The sequence \(S_N = \sum_{n=1}^N X \mathbf{1}_{A_n}\) converges \(\mathbb{P}\)-a.s. to \(X \mathbf{1}_A\), and is dominated in absolute value by \(|X| \in L^1(\mathbb{P})\). The dominated convergence theorem gives \[ \begin{align*} \nu_X(A) &= \int_A X \, d\mathbb{P} \\\\ &= \int_\Omega X \mathbf{1}_A \, d\mathbb{P} \\\\ &= \lim_{N \to \infty} \sum_{n=1}^N \int_{A_n} X \, d\mathbb{P} \\\\ &= \sum_{n=1}^\infty \nu_X(A_n), \end{align*} \] and the series converges absolutely because \[ \begin{align*} \sum_n |\nu_X(A_n)| &\leq \sum_n \int_{A_n} |X| \, d\mathbb{P} \\\\ &= \int_A |X| \, d\mathbb{P} \\\\ &\leq \mathbb{E}[|X|] \\\\ &\lt \infty. \end{align*} \]

Since \(X \in L^1\), \(\nu_X\) takes values in \(\mathbb{R}\) (never \(\pm \infty\)), so the sign-restriction condition is trivially satisfied.

Verification (finiteness).

The Jordan decomposition gives \(\nu_X = \nu_X^+ - \nu_X^-\), where both parts are non-negative measures on \(\mathcal{G}\). If \(\Omega = P \sqcup N\) is a Hahn decomposition for \(\nu_X\), with \(P, N \in \mathcal{G}\), the proof of the Jordan decomposition constructs the two parts as \(\nu_X^+(A) = \nu_X(A \cap P)\) and \(\nu_X^-(A) = -\nu_X(A \cap N)\). The total variation \(|\nu_X| = \nu_X^+ + \nu_X^-\) therefore satisfies \[ \begin{align*} |\nu_X|(\Omega) &= \nu_X(P) - \nu_X(N) \\\\ &= \int_P X \, d\mathbb{P} - \int_N X \, d\mathbb{P} \\\\ &\leq \int_P |X| \, d\mathbb{P} + \int_N |X| \, d\mathbb{P} \\\\ &= \mathbb{E}[|X|] \\\\ &\lt \infty. \end{align*} \] Hence \(\nu_X\) is a finite signed measure. The inequality can be strict, because \(P\) and \(N\) belong to \(\mathcal{G}\) while \(X\) is only \(\mathcal{F}\)-measurable and may change sign inside them.

Verification (absolute continuity).

Let \(A \in \mathcal{G}\) with \(\mathbb{P}|_{\mathcal{G}}(A) = 0\), that is, \(\mathbb{P}(A) = 0\) (the restricted measure agrees with \(\mathbb{P}\) on \(\mathcal{G}\)-sets by definition). Then \(X \mathbf{1}_A = 0\) \(\mathbb{P}\)-a.s., whence \(\nu_X(A) = \int_A X \, d\mathbb{P} = 0\). By the definition of absolute continuity, \(\nu_X \ll \mathbb{P}|_{\mathcal{G}}\).

The ingredients for the Radon-Nikodym theorem are now in place. The measure \(\mathbb{P}|_{\mathcal{G}}\) is finite (hence \(\sigma\)-finite) and non-negative on \((\Omega, \mathcal{G})\), and \(\nu_X\) is a finite signed measure absolutely continuous with respect to it. The theorem is stated and proved for non-negative \(\nu\) only, and the proof below applies it to the two Jordan parts of \(\nu_X\). The \(\mathcal{G}\)-measurable density so obtained represents \(\nu_X\) as an integral against \(\mathbb{P}|_{\mathcal{G}}\), and we take it as the definition of the conditional expectation.

The Conditional Expectation

Theorem & Definition: Conditional Expectation

Let \(X \in L^1(\Omega, \mathcal{F}, \mathbb{P})\) and \(\mathcal{G} \subseteq \mathcal{F}\) a sub-\(\sigma\)-algebra. There exists a \(\mathcal{G}\)-measurable function \(Y \in L^1(\Omega, \mathcal{G}, \mathbb{P}|_{\mathcal{G}})\), unique up to \(\mathbb{P}\)-a.s. equality, satisfying the averaging identity \[ \int_A Y \, d\mathbb{P} = \int_A X \, d\mathbb{P} \quad \text{for every } A \in \mathcal{G}. \tag{$\ast$} \] Any such \(Y\) is called a version of the conditional expectation of \(X\) given \(\mathcal{G}\), and we write \[ \mathbb{E}[X \mid \mathcal{G}] = Y, \] or, equivalently, \[ \mathbb{E}[X \mid \mathcal{G}] = \frac{d\nu_X}{d \mathbb{P}|_{\mathcal{G}}}. \]

Proof.

By the verifications above, \(\nu_X\) is a finite signed measure on \((\Omega, \mathcal{G})\) with \(\nu_X \ll \mathbb{P}|_{\mathcal{G}}\). Both Jordan parts inherit the absolute continuity. With \(P, N\) the Hahn sets used above, if \(\mathbb{P}(A) = 0\) then \(A \cap P\) and \(A \cap N\) are \(\mathbb{P}\)-null sets of \(\mathcal{G}\), so \(\nu_X^+(A) = \nu_X(A \cap P) = 0\) and \(\nu_X^-(A) = -\nu_X(A \cap N) = 0\). Both parts are finite, hence \(\sigma\)-finite, by the finiteness verification. Apply the Radon-Nikodym theorem separately to the Jordan parts \(\nu_X^+, \nu_X^-\) of \(\nu_X\). The resulting non-negative densities \(f^+, f^-\) belong to \(L^1(\mathbb{P}|_{\mathcal{G}})\) because \(\int f^\pm \, d\mathbb{P}|_{\mathcal{G}} = \nu_X^\pm(\Omega) \lt \infty\). Being integrable, \(f^+\) and \(f^-\) are finite \(\mathbb{P}\)-a.s., and we redefine both as \(0\) on the \(\mathbb{P}\)-null set in \(\mathcal{G}\) where either is infinite. This change affects none of the integrals. Set \(Y = f^+ - f^-\). Then \(Y\) is \(\mathcal{G}\)-measurable, integrable, and \[ \begin{align*} \int_A Y \, d\mathbb{P} &= \int_A (f^+ - f^-) \, d\mathbb{P}|_{\mathcal{G}} \\\\ &= \nu_X^+(A) - \nu_X^-(A) \\\\ &= \nu_X(A) \\\\ &= \int_A X \, d\mathbb{P} \end{align*} \] for every \(A \in \mathcal{G}\), establishing (\(\ast\)). The first equality uses the fact, which we take as given, that a \(\mathcal{G}\)-measurable function has the same integral against \(\mathbb{P}\) and against \(\mathbb{P}|_{\mathcal{G}}\). The two measures agree on \(\mathcal{G}\), hence on \(\mathcal{G}\)-measurable simple functions, and the integral is built from these.

Uniqueness up to \(\mathbb{P}\)-a.s. equality can be checked directly. If \(Y'\) is another version, then \(\int_A (Y - Y') \, d\mathbb{P} = 0\) for every \(A \in \mathcal{G}\). Taking \(A = \{Y - Y' \gt 0\} \in \mathcal{G}\) gives a strictly positive integrand on \(A\) with zero integral, which forces \(\mathbb{P}(A) = 0\). The set \(\{Y - Y' \lt 0\}\) is handled symmetrically, and hence \(Y = Y'\) \(\mathbb{P}\)-a.s.

Two remarks on this definition are essential, and both will be invoked repeatedly.

The averaging identity is the working characterisation. Although the construction goes through Radon-Nikodym, most subsequent proofs on this page use the averaging identity (\(\ast\)) directly, and the rest build on properties proved that way. The pattern is the same each time. To verify that a candidate function \(Y\) is a version of \(\mathbb{E}[X \mid \mathcal{G}]\), one shows that \(Y\) is \(\mathcal{G}\)-measurable and that \(\int_A Y \, d\mathbb{P} = \int_A X \, d\mathbb{P}\) for all \(A \in \mathcal{G}\). The a.s.-uniqueness clause then identifies \(Y\) as the conditional expectation. Radon-Nikodym is the existence engine, and the averaging identity is the daily tool.

The values of \(\mathbb{E}[X \mid \mathcal{G}](\omega)\) are defined only up to a \(\mathbb{P}\)-null set. Two versions of \(\mathbb{E}[X \mid \mathcal{G}]\) can disagree on any set \(N \in \mathcal{G}\) with \(\mathbb{P}(N) = 0\), and they will both be valid representatives. A statement such as "\(\mathbb{E}[X \mid \mathcal{G}](\omega_0) = c\)" for a particular \(\omega_0 \in \Omega\) is therefore not meaningful in isolation whenever \(\omega_0\) lies in a \(\mathbb{P}\)-null set of \(\mathcal{G}\), as every point does when \(\mathcal{G}\) is generated by a continuous random variable. Such a statement acquires meaning only after a specific version has been fixed.

This subtlety is the seed of regular conditional distributions. Their central object is a "regular version", a coherent choice of representative that behaves well as a function of \(\omega\) and yields an honest probability measure on \(\mathcal{F}\) for each \(\omega\).

Recovery of the Discrete Case

Let \(\{A_i\}_{i \geq 1}\) be a countable measurable partition of \(\Omega\) and let \(\mathcal{G} = \sigma(\{A_i\}_{i \geq 1})\) be the sub-\(\sigma\)-algebra it generates. The elements of \(\mathcal{G}\) are precisely the countable unions of the partition blocks. A \(\mathcal{G}\)-measurable function is constant on each \(A_i\), so any candidate version of \(\mathbb{E}[X \mid \mathcal{G}]\) is determined by its constant value on each block.

Proposition: Discrete Case

With \(\mathcal{G} = \sigma(\{A_i\}_{i \geq 1})\) for a countable measurable partition \(\{A_i\}\), and for any \(X \in L^1(\mathbb{P})\), one has \[ \begin{align*} \mathbb{E}[X \mid \mathcal{G}](\omega) &= \frac{1}{\mathbb{P}(A_i)} \int_{A_i} X \, d\mathbb{P} \\\\ &= \mathbb{E}[X \mid A_i] \quad \text{for } \omega \in A_i, \end{align*} \] on every block \(A_i\) with \(\mathbb{P}(A_i) \gt 0\). On blocks with \(\mathbb{P}(A_i) = 0\), the value may be any constant (consistent with a.s.-uniqueness).

Proof.

Define \(Y(\omega) = \mathbb{E}[X \mid A_i]\) for \(\omega \in A_i\) when \(\mathbb{P}(A_i) \gt 0\), and \(Y(\omega) = 0\) otherwise. Then \(Y\) is constant on each \(A_i\) and is therefore \(\mathcal{G}\)-measurable. It is integrable, because \[ \begin{align*} \sum_{i : \, \mathbb{P}(A_i) \gt 0} |\mathbb{E}[X \mid A_i]| \, \mathbb{P}(A_i) &\leq \sum_i \int_{A_i} |X| \, d\mathbb{P} \\\\ &\leq \mathbb{E}[|X|]. \end{align*} \] To verify the averaging identity, let \(A \in \mathcal{G}\). Then \(A = \bigsqcup_{i \in I} A_i\) for some countable index set \(I\), and \[ \begin{align*} \int_A Y \, d\mathbb{P} &= \sum_{i \in I, \, \mathbb{P}(A_i) \gt 0} \mathbb{E}[X \mid A_i] \cdot \mathbb{P}(A_i) \\\\ &= \sum_{i \in I, \, \mathbb{P}(A_i) \gt 0} \int_{A_i} X \, d\mathbb{P} \\\\ &= \int_A X \, d\mathbb{P}, \end{align*} \] where the first and last equalities use the countable additivity verified above, applied to the integrable functions \(Y\) and \(X\). The last one also discards blocks of probability zero, which contribute nothing to the integral of \(X\) either. By the a.s.-uniqueness clause of the conditional expectation, \(Y\) is a version of \(\mathbb{E}[X \mid \mathcal{G}]\).

The proposition recovers the elementary conditional expectation given an event, the first formula of the introduction. When \(\mathcal{G}\) is generated by a countable partition, the abstract definition reduces, block by block, to that formula. The novelty of the abstract definition lies in the cases that the elementary formula does not cover.

The Continuous Case Requires One More Layer

Let \(Y\) be a real-valued random variable and consider \(\mathcal{G} = \sigma(Y)\), the sub-\(\sigma\)-algebra generated by \(Y\). The conditional expectation \(\mathbb{E}[X \mid \sigma(Y)]\), abbreviated \(\mathbb{E}[X \mid Y]\), is a \(\sigma(Y)\)-measurable random variable on \(\Omega\). The Doob-Dynkin lemma, which we use without proof, states that every \(\sigma(Y)\)-measurable real-valued random variable has the form \(g(Y)\) for some Borel-measurable \(g : \mathbb{R} \to \mathbb{R}\). Applied to \(\mathbb{E}[X \mid Y]\), it yields such a \(g\), determined up to null sets of the distribution \(P_Y\) on \(\mathbb{R}\). The function \(g\) is what one would like to call "\(y \mapsto \mathbb{E}[X \mid Y = y]\)", a deterministic forecast of \(X\) for each observed value of \(Y\).

The notation \(\mathbb{E}[X \mid Y = y]\) thus has a meaning, with one caveat. As a function, \(g\) is determined only up to \(P_Y\)-null sets. For continuous \(Y\), every singleton \(\{y_0\}\) has \(P_Y\)-measure zero, so the value \(g(y_0)\) at any specific point is not determined by the abstract definition. Two versions of \(g\) can disagree on a \(P_Y\)-null set and both remain valid, so pointwise statements are again meaningful only up to such null sets.

When \((X, Y)\) has a joint density \(p(x, y)\), the second formula of the introduction supplies a version of \(g\). Write \(p_Y(y) = \int p(x, y) \, dx\) for the marginal density and \(p(x \mid y) = p(x, y) / p_Y(y)\) where \(p_Y(y) \gt 0\). Set \(g(y) = \int x \, p(x \mid y) \, dx\) where \(p_Y(y) \gt 0\) and the integral converges absolutely, and \(g(y) = 0\) elsewhere. Every set of \(\sigma(Y)\) has the form \(\{Y \in B\}\) with \(B\) Borel, and \[ \begin{align*} \int_{\{Y \in B\}} g(Y) \, d\mathbb{P} &= \int_B g(y) \, p_Y(y) \, dy \\\\ &= \int_B \int x \, p(x, y) \, dx \, dy \\\\ &= \int_{\{Y \in B\}} X \, d\mathbb{P}. \end{align*} \] The second equality is the definition of \(g\), together with the fact that \(p(x, y) = 0\) for almost every \(x\) wherever \(p_Y(y) = 0\). The first and third express expectations through the densities of \(Y\) and of \((X, Y)\), a standard step that we do not verify here. The third also writes the double integral as an iterated one by Fubini's theorem, legitimate because \(X\) is integrable. The same computation with \(|x|\) in place of \(x\) bounds \(\mathbb{E}[|g(Y)|]\) by \(\mathbb{E}[|X|]\). Hence \(g(Y)\) satisfies the averaging identity on \(\sigma(Y)\) and is a version of \(\mathbb{E}[X \mid Y]\).

Most ML uses of the notation call for more. Bayesian inference over continuous parameters and the Bellman equation on a continuous state space work with a conditional distribution given \(Y = y\), not with a single conditional expectation. They require the versions of \(\mathbb{P}(B \mid Y = y) = \mathbb{E}[\mathbf{1}_B \mid Y = y]\) to be chosen simultaneously for all \(B \in \mathcal{F}\), so that \(B \mapsto \mathbb{P}(B \mid Y = y)\) is an honest probability measure on \((\Omega, \mathcal{F})\) for every \(y\). Given that choice, \(\mathbb{E}[X \mid Y = y]\) can be computed as an integral \(\int X \, d\mathbb{P}(\cdot \mid Y = y)\) in the elementary sense.

Such a coherent choice is called a regular conditional distribution. Existence is not automatic, since it requires a regularity hypothesis on the space carrying the conditional measures (typically that \((\Omega, \mathcal{F})\) is a standard Borel space, the framework of the disintegration theorem). The construction of regular conditional distributions is the central topic of a separate strand of measure-theoretic probability, and we do not develop it on this page.

Nothing proved so far depends on this deferral. The abstract \(\sigma(Y)\)-measurable function \(\mathbb{E}[X \mid Y]\) exists and is unique a.s. by the construction above. Only its pointwise-coherent reading as a function of \(y\) requires additional machinery, together with the conditional measures \(\mathbb{P}(\cdot \mid Y = y)\) that allow integrals over the fibre to be computed directly.

Properties of Conditional Expectation

The averaging identity (\(\ast\)), together with the a.s.-uniqueness clause, remains the working tool for the algebraic properties, the projection theorem, and the tower property, used in the pattern described after the definition. The inequalities are then derived from the algebraic properties. We collect the algebraic properties first, then inequalities, then the geometric \(L^2\) characterisation, and finally the tower property, the structural identity behind martingale theory and dynamic programming.

We begin with a note on scope. Throughout this page, \(X \in L^1(\mathbb{P})\) is integrable, and every instance of \(\mathbb{E}[X \mid \mathcal{G}]\) is consequently a finite-valued random variable. Some standard treatments first define \(\mathbb{E}[X \mid \mathcal{G}]\) for non-negative \(X\) (allowing the value \(+\infty\) via the \(\sigma\)-finite Radon-Nikodym theorem) and then extend to \(L^1\) via the decomposition \(X = X^+ - X^-\). On \(L^1\), the resulting object coincides with ours, and the \(L^1\)-first restriction adopted here keeps every quantity on the page finite by construction.

Two intermediate steps handle possibly infinite quantities before this finiteness is secured. The existence proof receives \([0, \infty]\)-valued densities from the Radon-Nikodym theorem and makes them finite by a change on a null set. In Step 3 of the take-out rule, the monotone convergence theorem delivers an identity in \([0, \infty]\), and the hypothesis \(ZX \in L^1(\mathbb{P})\) makes both sides finite.

Algebraic Properties

Theorem: Linearity

Let \(X, Y \in L^1(\mathbb{P})\) and \(a, b \in \mathbb{R}\). Then \[ \mathbb{E}[aX + bY \mid \mathcal{G}] = a\, \mathbb{E}[X \mid \mathcal{G}] + b\, \mathbb{E}[Y \mid \mathcal{G}] \quad \mathbb{P}\text{-a.s.} \]

Proof.

The function \(Z = a\, \mathbb{E}[X \mid \mathcal{G}] + b\, \mathbb{E}[Y \mid \mathcal{G}]\) is \(\mathcal{G}\)-measurable (linear combination of \(\mathcal{G}\)-measurable functions) and integrable. For \(A \in \mathcal{G}\), \[ \begin{align*} \int_A Z \, d\mathbb{P} &= a \int_A \mathbb{E}[X \mid \mathcal{G}] \, d\mathbb{P} + b \int_A \mathbb{E}[Y \mid \mathcal{G}] \, d\mathbb{P} \\\\ &= a \int_A X \, d\mathbb{P} + b \int_A Y \, d\mathbb{P} \\\\ &= \int_A (aX + bY) \, d\mathbb{P}, \end{align*} \] using linearity of the Lebesgue integral and the averaging identity for each summand. The a.s.-uniqueness clause identifies \(Z\) as a version of \(\mathbb{E}[aX + bY \mid \mathcal{G}]\).

Theorem: Monotonicity

If \(X, Y \in L^1(\mathbb{P})\) and \(X \leq Y\) \(\mathbb{P}\)-a.s., then \[ \mathbb{E}[X \mid \mathcal{G}] \leq \mathbb{E}[Y \mid \mathcal{G}] \quad \mathbb{P}\text{-a.s.} \]

Proof.

Set \(D = \mathbb{E}[X \mid \mathcal{G}] - \mathbb{E}[Y \mid \mathcal{G}]\), a \(\mathcal{G}\)-measurable function. We want to show \(D \leq 0\) \(\mathbb{P}\)-a.s. Let \(A = \{D \gt 0\} \in \mathcal{G}\). By linearity and the averaging identity, \[ \begin{align*} \int_A D \, d\mathbb{P} &= \int_A \mathbb{E}[X \mid \mathcal{G}] \, d\mathbb{P} - \int_A \mathbb{E}[Y \mid \mathcal{G}] \, d\mathbb{P} \\\\ &= \int_A X \, d\mathbb{P} - \int_A Y \, d\mathbb{P} \\\\ &= \int_A (X - Y) \, d\mathbb{P} \\\\ &\leq 0, \end{align*} \] since \(X - Y \leq 0\) \(\mathbb{P}\)-a.s. But \(D \gt 0\) on \(A\), so \(\int_A D \, d\mathbb{P} \geq 0\), with strict inequality unless \(\mathbb{P}(A) = 0\). Hence \(\mathbb{P}(A) = 0\), that is, \(D \leq 0\) \(\mathbb{P}\)-a.s.

Theorem: Take-out (Pull-out) of \(\mathcal{G}\)-Measurable Factors

Let \(X \in L^1(\mathbb{P})\) and let \(Z\) be a \(\mathcal{G}\)-measurable random variable such that \(ZX \in L^1(\mathbb{P})\). Then \[ \mathbb{E}[ZX \mid \mathcal{G}] = Z \cdot \mathbb{E}[X \mid \mathcal{G}] \quad \mathbb{P}\text{-a.s.} \]

Proof (standard three-step extension).

We verify the averaging identity for \(Z \cdot \mathbb{E}[X \mid \mathcal{G}]\) in three stages: indicator, simple, then general \(\mathcal{G}\)-measurable.

Step 1 (indicator). Let \(Z = \mathbf{1}_B\) for \(B \in \mathcal{G}\). For any \(A \in \mathcal{G}\), \(A \cap B \in \mathcal{G}\), so \[ \begin{align*} \int_A \mathbf{1}_B \cdot \mathbb{E}[X \mid \mathcal{G}] \, d\mathbb{P} &= \int_{A \cap B} \mathbb{E}[X \mid \mathcal{G}] \, d\mathbb{P} \\\\ &= \int_{A \cap B} X \, d\mathbb{P} \\\\ &= \int_A \mathbf{1}_B X \, d\mathbb{P}, \end{align*} \] and \(\mathbf{1}_B \cdot \mathbb{E}[X \mid \mathcal{G}]\) is \(\mathcal{G}\)-measurable as a product of \(\mathcal{G}\)-measurable functions. The averaging identity holds.

Step 2 (simple non-negative). By linearity (already proved), the identity extends to non-negative simple \(Z = \sum_{k=1}^n c_k \mathbf{1}_{B_k}\) with \(B_k \in \mathcal{G}\) and \(c_k \geq 0\).

Step 3 (general). First take \(X \geq 0\). Let \(Z \geq 0\) be \(\mathcal{G}\)-measurable, and choose non-negative simple \(\mathcal{G}\)-measurable functions \(Z_n \uparrow Z\) (the standard simple-function approximation, whose existence we assume). Then \(Z_n X \uparrow ZX\) \(\mathbb{P}\)-a.s., and \(Z_n \cdot \mathbb{E}[X \mid \mathcal{G}] \uparrow Z \cdot \mathbb{E}[X \mid \mathcal{G}]\) \(\mathbb{P}\)-a.s. (using \(\mathbb{E}[X \mid \mathcal{G}] \geq 0\) by monotonicity, since \(X \geq 0\)). The monotone convergence theorem applied to both sides of the Step-2 identity, integrated over an arbitrary \(A \in \mathcal{G}\), gives \[ \int_A Z \cdot \mathbb{E}[X \mid \mathcal{G}] \, d\mathbb{P} = \int_A ZX \, d\mathbb{P}. \] Taking \(A = \Omega\) shows that \(Z \cdot \mathbb{E}[X \mid \mathcal{G}]\) is integrable whenever \(ZX\) is, and the identity then makes it a version of \(\mathbb{E}[ZX \mid \mathcal{G}]\).

For general \(X \in L^1\), decompose \(X = X^+ - X^-\) and \(Z = Z^+ - Z^-\) and apply the non-negative case to each of the four products. The integrability hypothesis \(ZX \in L^1\) ensures that each piece is integrable, since \(|Z^\pm X^\pm| \leq |ZX|\). Linearity (already proved for conditional expectation) reassembles the four pieces. Therefore \(Z \cdot \mathbb{E}[X \mid \mathcal{G}]\) is a version of \(\mathbb{E}[ZX \mid \mathcal{G}]\).

Theorem: Independence Collapse

If \(X \in L^1(\mathbb{P})\) and \(\sigma(X)\) is independent of \(\mathcal{G}\), then \[ \mathbb{E}[X \mid \mathcal{G}] = \mathbb{E}[X] \quad \mathbb{P}\text{-a.s.} \]

Proof.

The constant function \(\mathbb{E}[X]\) is \(\mathcal{G}\)-measurable. For \(A \in \mathcal{G}\), the random variables \(X\) and \(\mathbf{1}_A\) are independent, because every event \(\{X \in B_1\}\) lies in \(\sigma(X)\) and every event \(\{\mathbf{1}_A \in B_2\}\) lies in \(\mathcal{G}\). By the product formula for expectations, \(\mathbb{E}[X \mathbf{1}_A] = \mathbb{E}[X] \mathbb{E}[\mathbf{1}_A] = \mathbb{E}[X] \cdot \mathbb{P}(A)\), hence \[ \begin{align*} \int_A \mathbb{E}[X] \, d\mathbb{P} &= \mathbb{E}[X] \cdot \mathbb{P}(A) \\\\ &= \mathbb{E}[X \mathbf{1}_A] \\\\ &= \int_A X \, d\mathbb{P}. \end{align*} \] The a.s.-uniqueness clause identifies the constant \(\mathbb{E}[X]\) as a version of \(\mathbb{E}[X \mid \mathcal{G}]\).

Inequalities

Theorem: Jensen's Inequality for Conditional Expectation

Let \(\varphi : \mathbb{R} \to \mathbb{R}\) be convex, and let \(X \in L^1(\mathbb{P})\) with \(\varphi(X) \in L^1(\mathbb{P})\). Then \[ \varphi\big(\mathbb{E}[X \mid \mathcal{G}]\big) \leq \mathbb{E}[\varphi(X) \mid \mathcal{G}] \quad \mathbb{P}\text{-a.s.} \]

Proof (supporting-line argument).

For a convex function \(\varphi : \mathbb{R} \to \mathbb{R}\), every point \(x_0 \in \mathbb{R}\) admits a supporting affine function. That is, there exist \(a, b \in \mathbb{R}\) (depending on \(x_0\)) with \(\varphi(x_0) = a x_0 + b\) and \(\varphi(x) \geq a x + b\) for all \(x \in \mathbb{R}\). Moreover, since \(\varphi\) is convex on all of \(\mathbb{R}\), it is the pointwise supremum of a countable family of affine functions. Explicitly, there exist sequences \((a_n), (b_n) \subset \mathbb{R}\) with \[ \varphi(x) = \sup_{n \in \mathbb{N}} (a_n x + b_n) \quad \text{for all } x \in \mathbb{R}. \]

One construction takes an affine function supporting \(\varphi\) at each rational \(x_0\). Any slope between the left and right derivatives of \(\varphi\) at \(x_0\) gives one. We rely without proof on two standard facts about a convex function on \(\mathbb{R}\). Its one-sided derivatives exist and are bounded on bounded intervals, and \(\varphi\) is continuous. Every supporting function lies below \(\varphi\), so the supremum is at most \(\varphi\). Conversely, fix \(x\) and let rationals \(q_k \to x\). The supporting function at \(q_k\), evaluated at \(x\), equals \(\varphi(q_k) + a_k (x - q_k)\), where \(a_k\) is the slope chosen at \(q_k\). The slopes \(a_k\) are bounded, so this value tends to \(\varphi(x)\) by continuity.

For each \(n\), apply linearity and monotonicity of conditional expectation to the affine inequality \(\varphi(X) \geq a_n X + b_n\): \[ \begin{align*} \mathbb{E}[\varphi(X) \mid \mathcal{G}] &\geq \mathbb{E}[a_n X + b_n \mid \mathcal{G}] \\\\ &= a_n\, \mathbb{E}[X \mid \mathcal{G}] + b_n \quad \mathbb{P}\text{-a.s.} \end{align*} \] Here the constant \(b_n\) is its own conditional expectation, being \(\mathcal{G}\)-measurable and satisfying (\(\ast\)) trivially. The exceptional null set may depend on \(n\), but the union over the countable index set is still null. Outside this single null set, \[ \begin{align*} \mathbb{E}[\varphi(X) \mid \mathcal{G}] &\geq \sup_n \big( a_n \mathbb{E}[X \mid \mathcal{G}] + b_n \big) \\\\ &= \varphi\big(\mathbb{E}[X \mid \mathcal{G}]\big), \end{align*} \] which is the asserted inequality.

Theorem: \(L^p\) Contraction

For \(1 \leq p \lt \infty\) and \(X \in L^p(\Omega, \mathcal{F}, \mathbb{P})\), \[ \big\| \mathbb{E}[X \mid \mathcal{G}] \big\|_p \leq \|X\|_p. \] In particular, conditional expectation is a contraction on \(L^p(\mathbb{P})\).

Proof.

On a probability space \(L^p(\mathbb{P}) \subseteq L^1(\mathbb{P})\), as shown in our measure-theoretic treatment of expectation, so \(\mathbb{E}[X \mid \mathcal{G}]\) is defined. The function \(\varphi(t) = |t|^p\) is convex on \(\mathbb{R}\) for \(p \geq 1\), and \(\varphi(X) = |X|^p \in L^1(\mathbb{P})\) by the assumption \(X \in L^p(\mathbb{P})\). Apply Jensen's inequality for conditional expectation: \[ \big| \mathbb{E}[X \mid \mathcal{G}] \big|^p \leq \mathbb{E}[|X|^p \mid \mathcal{G}] \quad \mathbb{P}\text{-a.s.} \] Take expectations of both sides. On the right, the tower-with-trivial-\(\sigma\)-algebra identity \(\mathbb{E}[\mathbb{E}[Z \mid \mathcal{G}]] = \mathbb{E}[Z]\) (the averaging identity applied to \(A = \Omega \in \mathcal{G}\)) gives \(\mathbb{E}[\mathbb{E}[|X|^p \mid \mathcal{G}]] = \mathbb{E}[|X|^p]\). Hence \[ \begin{align*} \big\| \mathbb{E}[X \mid \mathcal{G}] \big\|_p^p &= \mathbb{E}\big[ \big| \mathbb{E}[X \mid \mathcal{G}] \big|^p \big] \\\\ &\leq \mathbb{E}[|X|^p] \\\\ &= \|X\|_p^p, \end{align*} \] and taking \(p\)-th roots gives the claim.

The \(L^2\) Projection Characterisation

At \(p = 2\) the contraction acquires geometric content. The space \(L^2(\Omega, \mathcal{G}, \mathbb{P}|_{\mathcal{G}})\) of square-integrable \(\mathcal{G}\)-measurable functions sits inside \(L^2(\Omega, \mathcal{F}, \mathbb{P})\) as a closed linear subspace, namely the classes that have a \(\mathcal{G}\)-measurable representative. It is closed because an \(L^2\)-convergent sequence of \(\mathcal{G}\)-measurable functions has a subsequence converging a.s. The set where this subsequence converges belongs to \(\mathcal{G}\) and has probability one, so the pointwise limit on that set, extended by \(0\) elsewhere, is a \(\mathcal{G}\)-measurable representative of the \(L^2\) limit. Conditional expectation, restricted to \(L^2\), is precisely the orthogonal projection onto this subspace.

Theorem: Conditional Expectation as \(L^2\) Projection

Let \(X \in L^2(\Omega, \mathcal{F}, \mathbb{P})\). Then \(\mathbb{E}[X \mid \mathcal{G}]\) is the orthogonal projection of \(X\) onto the closed subspace \(L^2(\Omega, \mathcal{G}, \mathbb{P}|_{\mathcal{G}}) \subseteq L^2(\Omega, \mathcal{F}, \mathbb{P})\). Equivalently, \(\mathbb{E}[X \mid \mathcal{G}]\) is, up to \(\mathbb{P}\)-a.s. equality, the unique \(\mathcal{G}\)-measurable square-integrable function minimising \[ \mathbb{E}\big[ (X - Y)^2 \big] \quad \text{over } Y \in L^2(\Omega, \mathcal{G}, \mathbb{P}|_{\mathcal{G}}). \]

Proof.

The Hilbert projection theorem applied to the closed subspace \(M = L^2(\Omega, \mathcal{G}, \mathbb{P}|_{\mathcal{G}})\) of the Hilbert space \(L^2(\Omega, \mathcal{F}, \mathbb{P})\), complete by Riesz-Fischer, produces a unique element \(P_M(X) \in M\) such that \(X - P_M(X) \perp M\). The same element is the unique minimiser of \(\|X - Y\|_{L^2}\) over \(Y \in M\). Orthogonality means \(\langle X - P_M(X), Z \rangle_{L^2} = 0\) for every \(Z \in M\), that is, \[ \mathbb{E}\big[ (X - P_M(X)) \cdot Z \big] = 0 \quad \text{for all } Z \in L^2(\Omega, \mathcal{G}, \mathbb{P}|_{\mathcal{G}}). \tag{$\dagger$} \]

Specialising \(Z = \mathbf{1}_A\) for \(A \in \mathcal{G}\) (which is bounded, hence in \(L^2\), and \(\mathcal{G}\)-measurable), (\(\dagger\)) reduces to \[ \int_A (X - P_M(X)) \, d\mathbb{P} = 0, \] that is, \[ \int_A P_M(X) \, d\mathbb{P} = \int_A X \, d\mathbb{P}. \] Thus \(P_M(X)\) is \(\mathcal{G}\)-measurable, integrable (since \(L^2 \subseteq L^1\) on a finite measure space), and satisfies the averaging identity on every \(A \in \mathcal{G}\), so the a.s.-uniqueness clause makes \(P_M(X)\) a version of \(\mathbb{E}[X \mid \mathcal{G}]\).

The \(L^2\) projection identification is the geometric face of conditional expectation. It explains in one stroke why \(\mathbb{E}[X \mid \mathcal{G}]\) is the minimum-mean-square forecast of \(X\) based on the information \(\mathcal{G}\), because orthogonal projection minimises distance and squared \(L^2\)-distance is mean-square error.

The "best linear predictor" of classical statistics (linear regression, the Wiener filter, the Kalman update) is the orthogonal projection onto a smaller subspace, the affine functions of the observations. It agrees with the conditional expectation when the variables are jointly Gaussian, a fact we do not prove here. In general it is the best approximation of the conditional expectation within that subspace, because projecting onto a subspace of \(L^2(\mathcal{G})\) factors through the projection onto \(L^2(\mathcal{G})\). The same geometric picture also makes the next property, the tower property, visually obvious, as we note after its proof.

The Tower Property

Theorem: Tower Property

Let \(\mathcal{H} \subseteq \mathcal{G} \subseteq \mathcal{F}\) be sub-\(\sigma\)-algebras and \(X \in L^1(\mathbb{P})\). Then \[ \mathbb{E}\big[\, \mathbb{E}[X \mid \mathcal{G}] \,\big|\, \mathcal{H} \,\big] = \mathbb{E}[X \mid \mathcal{H}] \quad \mathbb{P}\text{-a.s.} \] In particular, \(\mathbb{E}\big[\mathbb{E}[X \mid \mathcal{G}]\big] = \mathbb{E}[X]\) (taking \(\mathcal{H} = \{\emptyset, \Omega\}\)).

Proof.

Write \(W = \mathbb{E}[X \mid \mathcal{G}]\) and let \(V = \mathbb{E}[W \mid \mathcal{H}]\), which is \(\mathcal{H}\)-measurable and integrable by construction. We verify that \(V\) satisfies the averaging identity for \(\mathbb{E}[X \mid \mathcal{H}]\). For every \(A \in \mathcal{H}\), \[ \int_A V \, d\mathbb{P} \stackrel{(\mathrm{i})}{=} \int_A W \, d\mathbb{P} \stackrel{(\mathrm{ii})}{=} \int_A X \, d\mathbb{P}, \] where (i) is the averaging identity for \(V = \mathbb{E}[W \mid \mathcal{H}]\) on \(A \in \mathcal{H}\), and (ii) is the averaging identity for \(W = \mathbb{E}[X \mid \mathcal{G}]\) on \(A\), valid because \(A \in \mathcal{H} \subseteq \mathcal{G}\). The a.s.-uniqueness clause identifies \(V\) as a version of \(\mathbb{E}[X \mid \mathcal{H}]\).

The tower property is the structural identity that drives iterated conditioning. Read in the projection picture, for \(X \in L^2\), it says that projecting onto \(L^2(\mathcal{G})\) and then onto the smaller subspace \(L^2(\mathcal{H})\) produces the same vector as projecting directly onto \(L^2(\mathcal{H})\). Read in dynamic-programming terms, it says that the value of \(X\) under coarse information \(\mathcal{H}\) can be computed by first computing the value under finer information \(\mathcal{G}\) and then averaging that over \(\mathcal{H}\). This is the idealised recursion structure of the Bellman equation, taken up in the next section.

Conditional Expectation in Practice

The first three ML scenarios below correspond, in idealised form, to a conditional expectation, and the fourth, variational inference, is what replaces it when the conditional measure is inaccessible. Where a conditional expectation appears, we identify the sub-\(\sigma\)-algebra and read off which property of the previous section, if any, is being invoked. In code, the object so identified is typically reached only through a density assumption combined with a sampling- or surrogate-based approximation.

Expectation-Maximisation

The expectation-maximisation (EM) algorithm fits a parametric model \(p(x, z \mid \theta)\) with observed data \(X\) and latent variable \(Z\) by alternating between two steps. Given a current parameter estimate \(\theta^{(t)}\), the E-step computes \[ Q(\theta \mid \theta^{(t)}) = \mathbb{E}\big[ \log p(X, Z \mid \theta) \,\big|\, X, \theta^{(t)} \big], \] where the conditional expectation is taken with respect to the conditional distribution of \(Z\) given the observed \(X\) under parameter \(\theta^{(t)}\). The parameter \(\theta^{(t)}\) selects the probability measure and is not a random variable. Only \(X\) is conditioned on, so \(\mathcal{G} = \sigma(X)\). The M-step sets \(\theta^{(t+1)} = \arg\max_\theta Q(\theta \mid \theta^{(t)})\).

The E-step is, in its idealised (density-based) form, a conditional expectation. When the conditional density \(p(z \mid x, \theta^{(t)})\) exists, the expectation reduces to the explicit integral \(\int \log p(x, z \mid \theta) \, p(z \mid x, \theta^{(t)}) \, dz\) that is implemented in code, as a discrete sum when \(z\) takes finitely many values and as a Monte Carlo estimate otherwise. The M-step is a finite-dimensional optimisation.

The monotonic improvement of the marginal log-likelihood \(\ell(\theta) = \log p(X \mid \theta)\) under EM iterations follows from Jensen's inequality. Write \(p(z \mid x, \theta) = p(x, z \mid \theta) / p(x \mid \theta)\), so that \(\log p(x \mid \theta) = \log p(x, z \mid \theta) - \log p(z \mid x, \theta)\) for every \(z\) with \(p(x, z \mid \theta) \gt 0\), provided \(p(x \mid \theta) \gt 0\). By the assumption, stated below, that the two conditional densities of \(z\) at \(\theta\) and at \(\theta^{(t)}\) are positive on the same set, this covers \(p(\cdot \mid x, \theta^{(t)})\)-almost every \(z\). We take conditional expectations of both sides given \(X\) under \(p(\cdot \mid X, \theta^{(t)})\). Since the left side does not depend on \(z\), it is unchanged, and we obtain \[ \log p(X \mid \theta) = \underbrace{\mathbb{E}\big[ \log p(X, Z \mid \theta) \,\big|\, X, \theta^{(t)} \big]}_{Q(\theta \mid \theta^{(t)})} - \underbrace{\mathbb{E}\big[ \log p(Z \mid X, \theta) \,\big|\, X, \theta^{(t)} \big]}_{H(\theta \mid \theta^{(t)})}. \] Subtracting the same identity at \(\theta = \theta^{(t)}\) gives \[ \ell(\theta) - \ell(\theta^{(t)}) = \big[ Q(\theta \mid \theta^{(t)}) - Q(\theta^{(t)} \mid \theta^{(t)}) \big] + \big[ H(\theta^{(t)} \mid \theta^{(t)}) - H(\theta \mid \theta^{(t)}) \big]. \]

The first bracket is non-negative for \(\theta = \theta^{(t+1)}\) by definition of the M-step. The second bracket is non-negative by Jensen's inequality for \(\varphi(u) = -\log u\). This \(\varphi\) is convex on \((0, \infty)\) only, so the conditional theorem of the previous section, stated for \(\varphi\) on all of \(\mathbb{R}\), does not apply verbatim. For a fixed observed value \(x\), however, the conditional expectation is an integral against the density \(p(z \mid x, \theta^{(t)})\), and the elementary Jensen inequality on the interval \((0, \infty)\) suffices. We assume that \(p(z \mid x, \theta)\) and \(p(z \mid x, \theta^{(t)})\) are positive on the same set of \(z\), so that the ratio below takes values in \((0, \infty)\), and that the logarithms involved are integrable. With \(H(\theta^{(t)} \mid \theta^{(t)}) - H(\theta \mid \theta^{(t)}) = \mathbb{E}\big[ -\log( p(Z \mid X, \theta) / p(Z \mid X, \theta^{(t)}) ) \,\big|\, X, \theta^{(t)} \big]\), the inequality gives \[ \begin{align*} \mathbb{E}\big[ -\log( p(Z \mid X, \theta) / p(Z \mid X, \theta^{(t)}) ) \,\big|\, X, \theta^{(t)} \big] &\geq -\log \mathbb{E}\big[ p(Z \mid X, \theta) / p(Z \mid X, \theta^{(t)}) \,\big|\, X, \theta^{(t)} \big] \\\\ &= -\log 1 \\\\ &= 0, \end{align*} \] where the inner expectation evaluates to \(1\) by the explicit calculation \[ \begin{align*} \mathbb{E}\big[ p(Z \mid X, \theta) / p(Z \mid X, \theta^{(t)}) \,\big|\, X, \theta^{(t)} \big] &= \int \frac{p(z \mid x, \theta)}{p(z \mid x, \theta^{(t)})} \, p(z \mid x, \theta^{(t)}) \, dz \\\\ &= \int p(z \mid x, \theta) \, dz \\\\ &= 1, \end{align*} \] in which the conditioning density cancels and the remaining integrand is a probability density that integrates to \(1\). This is the same algebraic structure as the importance-sampling identity \(\mathbb{E}_q[f(Z) \, p(Z)/q(Z)] = \mathbb{E}_p[f(Z)]\), valid when \(q(z) \gt 0\) wherever \(f(z)\, p(z) \neq 0\) and \(f\) is integrable under \(p\).

Both brackets in the earlier decomposition are non-negative, so \(\ell(\theta^{(t+1)}) \geq \ell(\theta^{(t)})\). EM never decreases the marginal log-likelihood.

Reinforcement Learning: Value Functions and the Bellman Equation

In a Markov decision process with policy \(\pi\), the state-value function \[ V_\pi(s) = \mathbb{E}\Big[\, \sum_{t=0}^\infty \gamma^t R_{t+1} \,\Big|\, S_0 = s \,\Big] \] is a conditional expectation of the discounted return given the initial state. The Bellman expectation equation \[ V_\pi(s) = \mathbb{E}\big[ R_1 + \gamma V_\pi(S_1) \,\big|\, S_0 = s \big] \] rests on the tower property. Assume bounded rewards and \(\gamma \lt 1\), so that the discounted return \(X = \sum_t \gamma^t R_{t+1}\) is integrable, and a countable state space with \(\mathbb{P}(S_0 = s) \gt 0\) for every state \(s\), so that conditioning on \(S_0 = s\) is the elementary formula recovered in the discrete case. The \(\sigma\)-algebra structure is \(\sigma(S_0) \subseteq \sigma(S_0, S_1)\), and the tower property gives \[ \mathbb{E}[X \mid \sigma(S_0)] = \mathbb{E}\big[ \mathbb{E}[X \mid \sigma(S_0, S_1)] \,\big|\, \sigma(S_0) \big]. \] The inner conditional expectation equals \(\mathbb{E}[R_1 \mid \sigma(S_0, S_1)] + \gamma V_\pi(S_1)\). This step uses linearity together with the Markov property and time-homogeneity of the process. The conditional expectation, given \((S_0, S_1)\), of the discounted return from time \(1\) onward depends on \(S_1\) alone, through the same function \(V_\pi\). Both properties belong to the MDP model and do not follow from the tower property. A second use of the tower property, \[ \mathbb{E}\big[ \mathbb{E}[R_1 \mid \sigma(S_0, S_1)] \,\big|\, \sigma(S_0) \big] = \mathbb{E}[R_1 \mid \sigma(S_0)], \] then yields the Bellman equation.

The full development of value functions, Bellman optimality, and the algorithmic apparatus (value iteration, policy iteration, Q-learning) is on the reinforcement learning page.

Bayesian Posterior Predictive

Given a Bayesian model with parameter \(\theta\), prior \(\pi(\theta)\), and observed data \(D\), the posterior predictive distribution for a new observation \(X_{\text{new}}\) is \[ p(x_{\text{new}} \mid D) = \mathbb{E}\big[ p(x_{\text{new}} \mid \theta) \,\big|\, D \big], \] a conditional expectation given the data. In the Bayesian model, \(\theta\) is a random variable on the same probability space as \(D\) and \(X_{\text{new}}\), and \(\mathcal{G} = \sigma(D)\). The identity is, in idealised form, the tower property for \(\sigma(D) \subseteq \sigma(\theta, D)\), combined with the modelling assumption that \(X_{\text{new}}\) is conditionally independent of \(D\) given \(\theta\). That assumption replaces \(p(x_{\text{new}} \mid \theta, D)\) by \(p(x_{\text{new}} \mid \theta)\), and the outer conditioning on \(D\) "marginalises over uncertainty in \(\theta\)".

When \(\theta\) is a continuous parameter, the pointwise reading \(p(x_{\text{new}} \mid D)\) requires the regular-conditional-distribution machinery, which we do not develop on this page. In practice this expectation is approximated by Markov chain Monte Carlo or by variational surrogates. The take-out property licenses pulling deterministic functions of the data outside the conditional expectation, and the tower property licenses hierarchical decompositions (for example, predicting via an intermediate latent layer).

Variational Inference and the ELBO

Variational inference approximates an intractable posterior \(p(z \mid x)\) by a tractable surrogate \(q(z \mid x)\). The evidence lower bound \(\mathrm{ELBO}(q) = \mathbb{E}_{q(z \mid x)}[\log p(x, z) - \log q(z \mid x)]\) is an expectation against \(q\), not against \(p(\cdot \mid x)\). The exact identity \(\log p(x) = \mathrm{ELBO}(q) + D_{\mathrm{KL}}(q \,\|\, p(\cdot \mid x))\), valid when \(q \ll p(\cdot \mid x)\) and the logarithms are \(q\)-integrable, certifies that maximising the ELBO is equivalent to minimising the KL to the posterior. From the conditional-expectation perspective, this is a manipulation that replaces the inaccessible conditional measure with a tractable one. The surrogate gives up the averaging identity (\(\ast\)), which only the true conditional measure satisfies, and the KL term records the discrepancy. The full development is on the variational inference page.

Rigorous Foundation, Approximate Practice

A broader observation is worth recording, even at the cost of leaving rigorously verified ground. The applications surveyed above all share a common pattern. Machine learning implements a finite-sample, density-based approximation of an object whose rigorous existence is licensed by the conditional expectation construction of this page, but the rigorous construction itself is rarely instantiated in code. Monte Carlo replaces the integral, a single sampled trajectory replaces the Bellman expectation, and a variational surrogate replaces the intractable posterior. These approximations have carried machine learning through a successful empirical era.

Contemporary large language models show several unstable behaviours, among them the inconsistency of long chains of probabilistic reasoning, brittleness under distribution shift, and difficulty in calibrating uncertainty. It is at least worth asking whether some of these behaviours are connected to the absence of a rigorous measure-theoretic substrate underneath the approximations.

The connection is hypothesised, not proved. The question of how rigorous mathematical structure should enter the foundations of AI systems remains an active research direction, with measure theory, topology, and differential geometry among the candidate frameworks. This curriculum is built on the working assumption that the rigorous mathematical layer will become increasingly relevant as the field matures. A reader who has internalised the construction on this page is then better positioned to follow that line of development as it unfolds.