Random Variables
In our development of Basic Probability Ideas, we worked with
probability in terms of events, which are subsets of a sample space. While this framework is logically
complete, it is insufficient for the quantitative demands of statistics and machine learning. We need
to associate numerical values with outcomes so that we can compute averages, measure spread,
and apply the tools of calculus. A random variable is precisely this bridge, a
numerical function on the sample space.
Definition: Random Variable
A random variable is a function \(X: S \to \mathbb{R}\) that assigns a numerical
value to each outcome in the sample space \(S\). We denote random variables by capital letters
(\(X, Y, Z\)) and specific values they take by the corresponding lowercase letters (\(x, y, z\)).
A Note on Measurability
Strictly speaking, the function \(X\) must satisfy a technical measurability
condition so that \(\{X \in B\}\) is a valid event for every Borel set \(B \subseteq \mathbb{R}\).
This condition is what makes probabilities like \(P(X \leq x)\) and \(P(a \leq X \leq b)\)
well-defined. For the discrete and continuous random variables studied here, this condition holds
for every distribution we consider, so we take it for granted. The full formulation is developed in
Measure-Theoretic Probability, where a
random variable is recast as a measurable function from a probability space to the real line
equipped with its Borel \(\sigma\)-algebra.
We study two fundamental types of random variables, discrete and continuous, and the distinction
determines which mathematical tool, summation or integration, is used to analyze them. The two types
do not exhaust all random variables: a variable that equals \(0\) with probability \(\tfrac{1}{2}\)
and is otherwise uniform on \([0, 1]\) is neither.
Discrete Random Variables
A random variable is discrete if the set of values it can take is finite or countably
infinite (for example, \(\{0, 1, 2, \ldots\}\)). The probability structure of a discrete random
variable is completely characterized by its probability mass function.
Definition: Probability Mass Function (p.m.f.)
The probability mass function of a discrete random variable \(X\) is the function
\[
f(x) = P(X = x)
\]
satisfying:
- \(f(x) \geq 0\) for all \(x\).
- \(\sum_{x} f(x) = 1\), where the sum is over all possible values of \(X\).
To answer questions of the form "what is the probability that \(X\) is at most \(x\)?",
we accumulate the mass function into a running total. This cumulative perspective is
especially useful for computing probabilities over intervals.
Definition: Cumulative Distribution Function (Discrete Case)
The cumulative distribution function (c.d.f.) of a discrete random variable \(X\) is
\[
F(x) = P(X \leq x) = \sum_{k \leq x} f(k).
\]
For an integer-valued \(X\) and integers \(a \leq b\), this gives
\(P(a \leq X \leq b) = F(b) - F(a-1)\).
More generally, \(P(a \leq X \leq b) = F(b) - P(X \lt a)\), where \(P(X \lt a)\) excludes the
mass at \(x = a\) itself.
Continuous Random Variables
A random variable is continuous if its probabilities are given by a density
function: instead of assigning mass to individual points, we obtain the probability that \(X\)
falls in an interval by integrating the density over that interval. Taking the degenerate
interval \([x, x]\) forces \(P(X = x) = 0\) for every value \(x\). With the countable additivity that
infinite sample spaces require, the set of values \(X\) can take is therefore uncountably infinite,
since countably many points of probability zero cannot carry total probability one.
Definition: Probability Density Function (p.d.f.)
The probability density function of a continuous random variable \(X\) is a
function \(f(x)\) satisfying:
- \(f(x) \geq 0\) for all \(x\).
- \(\int_{-\infty}^{\infty} f(x)\,dx = 1\).
- \(P(a \leq X \leq b) = \int_{a}^{b} f(x)\,dx\) for any \(a \leq b\).
Note that \(f(x)\) itself is not a probability. It is a density. In particular, \(f(x)\)
can exceed 1 (for example, the uniform distribution on \([0, 0.5]\) has density \(f(x) = 2\)).
Definition: Cumulative Distribution Function (Continuous Case)
The cumulative distribution function of a continuous random variable \(X\) is
\[
F(x) = P(X \leq x) = \int_{-\infty}^{x} f(u)\,du.
\]
By the Fundamental Theorem of Calculus, the density is recovered as:
\[
f(x) = \frac{dF(x)}{dx}
\]
at every point of continuity of \(f\). Furthermore, \(P(a \leq X \leq b) = F(b) - F(a)\).
Note that for a continuous random variable, \(P(X = a) = \int_a^a f(x)\,dx = 0\), so
\(P(a \leq X \leq b) = P(a \lt X \lt b)\). The distinction between strict and non-strict inequalities
matters only at values that carry positive probability, and a continuous random variable has none.
With the language of random variables and their distributions established, we can now ask
the two most fundamental questions about any distribution: where is its center, and
how spread out is it? These are captured by the expected value and
variance, respectively.
Expected Value
The expected value (or mean) of a random variable provides a
single number that summarizes the "center" of its distribution. It is a weighted average of all
possible values, with the weights supplied by the mass function in the discrete case and by the
density in the continuous case. This concept is indispensable in machine learning. Risks are
expected losses. Under squared loss, the best prediction of a target with finite variance is a
conditional expectation, and training algorithms typically minimize an empirical estimate of the
risk, often with a regularization term added.
Definition: Expected Value
The expected value of a random variable \(X\) is defined as:
\[
\mathbb{E}[X] = \mu =
\begin{cases}
\displaystyle\sum_{x} x\, f(x) & \text{if } X \text{ is discrete} \\\\
\displaystyle\int_{-\infty}^{\infty} x\, f(x)\,dx & \text{if } X \text{ is continuous}
\end{cases}
\]
provided the sum or integral converges absolutely.
The expected value can be interpreted as the center of gravity of the distribution. If
the distribution is symmetric about some point \(c\) and the mean exists, then \(\mathbb{E}[X] = c\).
In that case \(c\) is also a median of the distribution.
Computing \(\mathbb{E}[g(X)]\) from the definition appears to require the mass function or density
of the new random variable \(g(X)\). It does not. The distribution of \(X\) already carries
everything needed, and the values of \(g\) can be averaged directly against it.
Theorem: Expectation of a Function of a Random Variable
Let \(X\) be a random variable with mass function or density \(f\), and let
\(g : \mathbb{R} \to \mathbb{R}\) be a function for which \(g(X)\) is again a random variable. Then
\[
\mathbb{E}[g(X)] =
\begin{cases}
\displaystyle\sum_{x} g(x)\, f(x) & \text{if } X \text{ is discrete} \\\\
\displaystyle\int_{-\infty}^{\infty} g(x)\, f(x)\,dx & \text{if } X \text{ is continuous}
\end{cases}
\]
provided the sum or integral above converges absolutely. No description of the distribution of
\(g(X)\) is required.
Proof (discrete case):
Write \(Y = g(X)\), which takes at most as many values as \(X\) and is therefore discrete, and let
\(\mathcal{Y}\) be its set of values. For each \(y \in \mathcal{Y}\), the event \(\{Y = y\}\) is
the disjoint union of the events \(\{X = x\}\) over those \(x\) with \(g(x) = y\), so
\(P(Y = y) = \sum_{x : g(x) = y} f(x)\). Hence
\[
\begin{align*}
\mathbb{E}[Y] &= \sum_{y \in \mathcal{Y}} y\, P(Y = y) \\\\
&= \sum_{y \in \mathcal{Y}} \sum_{x : g(x) = y} y\, f(x) \\\\
&= \sum_{y \in \mathcal{Y}} \sum_{x : g(x) = y} g(x)\, f(x) \\\\
&= \sum_{x} g(x)\, f(x).
\end{align*}
\]
The third equality substitutes \(g(x)\) for \(y\), which is an identity on the index set of the
inner sum. The last one uses that the sets \(\{x : g(x) = y\}\), as \(y\) ranges over
\(\mathcal{Y}\), partition the values of \(X\). The same regrouping applied to \(|g|\) has
non-negative terms throughout, so it is valid with no hypothesis at all and gives
\(\sum_{y} |y|\, P(Y = y) = \sum_{x} |g(x)|\, f(x)\). Therefore \(\mathbb{E}[Y]\) exists exactly
when the sum in the statement converges absolutely, and under that hypothesis the rearrangement of
the signed terms above is legitimate.
The continuous case is a different kind of statement. The set \(\{x : g(x) \leq t\}\) need not
be an interval, so the substitution rules of one-variable calculus do not reach it, and the
density of \(g(X)\) need not exist even when that of \(X\) does. We take the identity as given
here and prove it once expectation is constructed as an integral with respect to the
distribution of \(X\).
One of the most useful properties of expectation is its linearity, which holds
regardless of whether the random variables involved are independent.
Theorem: Linearity of Expectation
Let \(X\) and \(Y\) be random variables on the same sample space, and let
\(a, b, c \in \mathbb{R}\) be constants. Then
\[
\mathbb{E}[aX + bY + c] = a\,\mathbb{E}[X] + b\,\mathbb{E}[Y] + c,
\]
whenever the expectations on the right are well-defined. In particular,
\(\mathbb{E}[X + Y] = \mathbb{E}[X] + \mathbb{E}[Y]\) and
\(\mathbb{E}[aX + c] = a\,\mathbb{E}[X] + c\).
Remark. The additivity \(\mathbb{E}[X + Y] = \mathbb{E}[X] + \mathbb{E}[Y]\) holds
without any assumption of independence between \(X\) and \(Y\). That fact will be central
in the analyses of estimators, gradient noise, and bias-variance decompositions.
Proof (scalar case \(\mathbb{E}[aX + c] = a\,\mathbb{E}[X] + c\)):
For continuous \(X\) with density \(f\),
\[
\begin{align*}
\mathbb{E}[aX + c] &= \int_{-\infty}^{\infty} (ax + c)\, f(x)\,dx \\\\
&= a\int_{-\infty}^{\infty} x\, f(x)\,dx + c\int_{-\infty}^{\infty} f(x)\,dx \\\\
&= a\,\mathbb{E}[X] + c,
\end{align*}
\]
using \(\int f(x)\,dx = 1\) in the last step. The first equality is the
Expectation of a Function of a Random Variable
applied to \(g(x) = ax + c\), whose integral against \(f\) converges absolutely because \(|ax + c| \leq |a|\,|x| + |c|\) and
\(\mathbb{E}[X]\) exists. It integrates \(ax + c\) against the density of \(X\) itself, so it needs no density for
\(aX + c\), which has none when \(a = 0\). The continuous half of that theorem is taken as given on this page.
For discrete \(X\) with p.m.f. \(f\), the same argument with sums gives
\[
\begin{align*}
\mathbb{E}[aX + c] &= \sum_{x}(ax + c) f(x) \\\\
&= a\sum_x x\,f(x) + c\sum_x f(x) \\\\
&= a\,\mathbb{E}[X] + c.
\end{align*}
\]
Here the first equality regroups the sum. When \(a \neq 0\) the map \(x \mapsto ax + c\) is
injective, so the values of \(aX + c\) correspond one to one with those of \(X\) and
\(\sum_z z\,P(aX + c = z)\) is the displayed sum term by term. When \(a = 0\) the identity reads
\(\mathbb{E}[c] = c\).
The bivariate identity \(\mathbb{E}[X + Y] = \mathbb{E}[X] + \mathbb{E}[Y]\) requires the
joint distribution of \((X, Y)\), which we postpone until joint and multivariate distributions
are introduced. The argument runs in parallel, with the single sum or integral replaced by a
double sum or integral over the joint distribution. Crucially, it does not require \(X\) and
\(Y\) to be independent.
Knowing the center of a distribution is valuable, but it tells us nothing about how concentrated
or dispersed the values are around that center. Two distributions can share the same mean yet
differ dramatically in spread. To quantify this spread, we introduce the variance.
Variance
The variance measures the expected squared deviation of a random variable from its
mean. A small variance indicates that the values of \(X\) tend to cluster tightly around \(\mu\),
while a large variance indicates wide dispersion. In machine learning, variance appears
everywhere: in the bias-variance tradeoff, in gradient noise during stochastic optimization,
and in the uncertainty quantification of Bayesian predictions.
Definition: Variance and Standard Deviation
The variance of a random variable \(X\) with mean \(\mu = \mathbb{E}[X]\) is:
\[
\operatorname{Var}(X) = \sigma^2 = \mathbb{E}\bigl[(X - \mu)^2\bigr] \geq 0.
\]
The standard deviation is \(\sigma = \sqrt{\operatorname{Var}(X)}\), which has the same
units as \(X\).
The definition involves the unknown quantity \(\mathbb{E}[X]\) inside the expectation. Expanding the square
yields a computationally convenient alternative.
Proposition: Computational Identity for Variance
Let \(X\) be a random variable with mass function or density \(f\), and suppose that
\(\sum_{x} x^2 f(x)\) (discrete case) or \(\int_{-\infty}^{\infty} x^2 f(x)\,dx\) (continuous case) is
finite. Then the mean \(\mu = \mathbb{E}[X]\) exists, \(\operatorname{Var}(X)\) is finite, and
\[
\operatorname{Var}(X) = \mathbb{E}[X^2] - \mu^2.
\]
Proof:
We give the discrete case. Since \(|x| \leq 1 + x^2\) for every real \(x\),
\(\sum_{x} |x|\, f(x) \leq \sum_{x} f(x) + \sum_{x} x^2 f(x) \lt \infty\), so \(\mu\) exists. The
Expectation of a Function of a Random Variable
applied to \(g(x) = x^2\) identifies the finite sum \(\sum_{x} x^2 f(x)\) with \(\mathbb{E}[X^2]\).
Applied to \(g(x) = (x - \mu)^2\), whose sum against \(f\) converges because
\((x - \mu)^2 \leq x^2 + 2|\mu|\,|x| + \mu^2\), the same theorem gives
\[
\begin{align*}
\operatorname{Var}(X) &= \mathbb{E}\bigl[(X - \mu)^2\bigr] \\\\
&= \sum_{x} (x - \mu)^2 f(x) \\\\
&= \sum_{x} x^2 f(x) - 2\mu \sum_{x} x\, f(x) \\\\
&\quad + \mu^2 \sum_{x} f(x) \\\\
&= \mathbb{E}[X^2] - 2\mu^2 + \mu^2 \\\\
&= \mathbb{E}[X^2] - \mu^2.
\end{align*}
\]
The sum splits into three because each of the three series converges absolutely. The continuous case runs
verbatim with integrals in place of sums, and rests on the continuous half of that theorem, which the
page takes as given.
That is, the variance equals the mean of the square minus the square of the mean. Note
that \(\operatorname{Var}(c) = 0\) for any constant \(c\), since a constant has no spread.
The following result describes how variance transforms under linear operations. Variance is
affected by scaling but not by translation, in contrast to expectation, which shifts along with
the additive constant.
Proof:
By the scalar case of
linearity of expectation,
\(\mathbb{E}[aX + b] = a\mu + b\), so the deviation of \(aX + b\) from its mean is the random variable
\((aX + b) - (a\mu + b) = a(X - \mu)\). Put \(Z = (X - \mu)^2\), whose expectation is
\(\operatorname{Var}(X)\). The scalar case of linearity, applied now to \(Z\), gives
\[
\begin{align*}
\operatorname{Var}(aX + b) &= \mathbb{E}\bigl[a^2 (X - \mu)^2\bigr] \\\\
&= \mathbb{E}\bigl[a^2 Z\bigr] \\\\
&= a^2\,\mathbb{E}[Z] \\\\
&= a^2\,\operatorname{Var}(X).
\end{align*}
\]
Insight: Random Variables in Machine Learning
The framework of random variables, expectations, and variances is the language in which virtually
all of machine learning is written. A model's risk is an expectation, the
expected loss \(\mathbb{E}[\ell(Y, \hat{Y})]\) over the data distribution, and training typically
minimizes an empirical estimate of it. The
bias-variance decomposition shows that a model's expected squared prediction error
decomposes as \(\text{Bias}^2 + \text{Variance} + \text{Irreducible Noise}\), directly using the
concepts defined here. In the pages ahead, we will study specific families of distributions that
serve as building blocks for probabilistic models throughout statistics and machine learning. They
include the Gamma and Beta and
Gaussian families, among others.