Random Variables

Random Variables Expected Value Variance

Random Variables

In our development of Basic Probability Ideas, we worked with probability in terms of events, which are subsets of a sample space. While this framework is logically complete, it is insufficient for the quantitative demands of statistics and machine learning. We need to associate numerical values with outcomes so that we can compute averages, measure spread, and apply the tools of calculus. A random variable is precisely this bridge, a numerical function on the sample space.

Definition: Random Variable

A random variable is a function \(X: S \to \mathbb{R}\) that assigns a numerical value to each outcome in the sample space \(S\). We denote random variables by capital letters (\(X, Y, Z\)) and specific values they take by the corresponding lowercase letters (\(x, y, z\)).

A Note on Measurability

Strictly speaking, the function \(X\) must satisfy a technical measurability condition so that \(\{X \in B\}\) is a valid event for every Borel set \(B \subseteq \mathbb{R}\). This condition is what makes probabilities like \(P(X \leq x)\) and \(P(a \leq X \leq b)\) well-defined. For the discrete and continuous random variables studied here, this condition holds for every distribution we consider, so we take it for granted. The full formulation is developed in Measure-Theoretic Probability, where a random variable is recast as a measurable function from a probability space to the real line equipped with its Borel \(\sigma\)-algebra.

We study two fundamental types of random variables, discrete and continuous, and the distinction determines which mathematical tool, summation or integration, is used to analyze them. The two types do not exhaust all random variables: a variable that equals \(0\) with probability \(\tfrac{1}{2}\) and is otherwise uniform on \([0, 1]\) is neither.

Discrete Random Variables

A random variable is discrete if the set of values it can take is finite or countably infinite (for example, \(\{0, 1, 2, \ldots\}\)). The probability structure of a discrete random variable is completely characterized by its probability mass function.

Definition: Probability Mass Function (p.m.f.)

The probability mass function of a discrete random variable \(X\) is the function \[ f(x) = P(X = x) \] satisfying:

  1. \(f(x) \geq 0\) for all \(x\).
  2. \(\sum_{x} f(x) = 1\), where the sum is over all possible values of \(X\).

To answer questions of the form "what is the probability that \(X\) is at most \(x\)?", we accumulate the mass function into a running total. This cumulative perspective is especially useful for computing probabilities over intervals.

Definition: Cumulative Distribution Function (Discrete Case)

The cumulative distribution function (c.d.f.) of a discrete random variable \(X\) is \[ F(x) = P(X \leq x) = \sum_{k \leq x} f(k). \] For an integer-valued \(X\) and integers \(a \leq b\), this gives \(P(a \leq X \leq b) = F(b) - F(a-1)\). More generally, \(P(a \leq X \leq b) = F(b) - P(X \lt a)\), where \(P(X \lt a)\) excludes the mass at \(x = a\) itself.

Continuous Random Variables

A random variable is continuous if its probabilities are given by a density function: instead of assigning mass to individual points, we obtain the probability that \(X\) falls in an interval by integrating the density over that interval. Taking the degenerate interval \([x, x]\) forces \(P(X = x) = 0\) for every value \(x\). With the countable additivity that infinite sample spaces require, the set of values \(X\) can take is therefore uncountably infinite, since countably many points of probability zero cannot carry total probability one.

Definition: Probability Density Function (p.d.f.)

The probability density function of a continuous random variable \(X\) is a function \(f(x)\) satisfying:

  1. \(f(x) \geq 0\) for all \(x\).
  2. \(\int_{-\infty}^{\infty} f(x)\,dx = 1\).
  3. \(P(a \leq X \leq b) = \int_{a}^{b} f(x)\,dx\) for any \(a \leq b\).

Note that \(f(x)\) itself is not a probability. It is a density. In particular, \(f(x)\) can exceed 1 (for example, the uniform distribution on \([0, 0.5]\) has density \(f(x) = 2\)).

Definition: Cumulative Distribution Function (Continuous Case)

The cumulative distribution function of a continuous random variable \(X\) is \[ F(x) = P(X \leq x) = \int_{-\infty}^{x} f(u)\,du. \]

By the Fundamental Theorem of Calculus, the density is recovered as: \[ f(x) = \frac{dF(x)}{dx} \] at every point of continuity of \(f\). Furthermore, \(P(a \leq X \leq b) = F(b) - F(a)\).

Note that for a continuous random variable, \(P(X = a) = \int_a^a f(x)\,dx = 0\), so \(P(a \leq X \leq b) = P(a \lt X \lt b)\). The distinction between strict and non-strict inequalities matters only at values that carry positive probability, and a continuous random variable has none.

With the language of random variables and their distributions established, we can now ask the two most fundamental questions about any distribution: where is its center, and how spread out is it? These are captured by the expected value and variance, respectively.

Expected Value

The expected value (or mean) of a random variable provides a single number that summarizes the "center" of its distribution. It is a weighted average of all possible values, with the weights supplied by the mass function in the discrete case and by the density in the continuous case. This concept is indispensable in machine learning. Risks are expected losses. Under squared loss, the best prediction of a target with finite variance is a conditional expectation, and training algorithms typically minimize an empirical estimate of the risk, often with a regularization term added.

Definition: Expected Value

The expected value of a random variable \(X\) is defined as: \[ \mathbb{E}[X] = \mu = \begin{cases} \displaystyle\sum_{x} x\, f(x) & \text{if } X \text{ is discrete} \\\\ \displaystyle\int_{-\infty}^{\infty} x\, f(x)\,dx & \text{if } X \text{ is continuous} \end{cases} \] provided the sum or integral converges absolutely.

The expected value can be interpreted as the center of gravity of the distribution. If the distribution is symmetric about some point \(c\) and the mean exists, then \(\mathbb{E}[X] = c\). In that case \(c\) is also a median of the distribution.

Computing \(\mathbb{E}[g(X)]\) from the definition appears to require the mass function or density of the new random variable \(g(X)\). It does not. The distribution of \(X\) already carries everything needed, and the values of \(g\) can be averaged directly against it.

Theorem: Expectation of a Function of a Random Variable

Let \(X\) be a random variable with mass function or density \(f\), and let \(g : \mathbb{R} \to \mathbb{R}\) be a function for which \(g(X)\) is again a random variable. Then \[ \mathbb{E}[g(X)] = \begin{cases} \displaystyle\sum_{x} g(x)\, f(x) & \text{if } X \text{ is discrete} \\\\ \displaystyle\int_{-\infty}^{\infty} g(x)\, f(x)\,dx & \text{if } X \text{ is continuous} \end{cases} \] provided the sum or integral above converges absolutely. No description of the distribution of \(g(X)\) is required.

Proof (discrete case):

Write \(Y = g(X)\), which takes at most as many values as \(X\) and is therefore discrete, and let \(\mathcal{Y}\) be its set of values. For each \(y \in \mathcal{Y}\), the event \(\{Y = y\}\) is the disjoint union of the events \(\{X = x\}\) over those \(x\) with \(g(x) = y\), so \(P(Y = y) = \sum_{x : g(x) = y} f(x)\). Hence \[ \begin{align*} \mathbb{E}[Y] &= \sum_{y \in \mathcal{Y}} y\, P(Y = y) \\\\ &= \sum_{y \in \mathcal{Y}} \sum_{x : g(x) = y} y\, f(x) \\\\ &= \sum_{y \in \mathcal{Y}} \sum_{x : g(x) = y} g(x)\, f(x) \\\\ &= \sum_{x} g(x)\, f(x). \end{align*} \] The third equality substitutes \(g(x)\) for \(y\), which is an identity on the index set of the inner sum. The last one uses that the sets \(\{x : g(x) = y\}\), as \(y\) ranges over \(\mathcal{Y}\), partition the values of \(X\). The same regrouping applied to \(|g|\) has non-negative terms throughout, so it is valid with no hypothesis at all and gives \(\sum_{y} |y|\, P(Y = y) = \sum_{x} |g(x)|\, f(x)\). Therefore \(\mathbb{E}[Y]\) exists exactly when the sum in the statement converges absolutely, and under that hypothesis the rearrangement of the signed terms above is legitimate.

The continuous case is a different kind of statement. The set \(\{x : g(x) \leq t\}\) need not be an interval, so the substitution rules of one-variable calculus do not reach it, and the density of \(g(X)\) need not exist even when that of \(X\) does. We take the identity as given here and prove it once expectation is constructed as an integral with respect to the distribution of \(X\).

One of the most useful properties of expectation is its linearity, which holds regardless of whether the random variables involved are independent.

Theorem: Linearity of Expectation

Let \(X\) and \(Y\) be random variables on the same sample space, and let \(a, b, c \in \mathbb{R}\) be constants. Then \[ \mathbb{E}[aX + bY + c] = a\,\mathbb{E}[X] + b\,\mathbb{E}[Y] + c, \] whenever the expectations on the right are well-defined. In particular, \(\mathbb{E}[X + Y] = \mathbb{E}[X] + \mathbb{E}[Y]\) and \(\mathbb{E}[aX + c] = a\,\mathbb{E}[X] + c\).

Remark. The additivity \(\mathbb{E}[X + Y] = \mathbb{E}[X] + \mathbb{E}[Y]\) holds without any assumption of independence between \(X\) and \(Y\). That fact will be central in the analyses of estimators, gradient noise, and bias-variance decompositions.

Proof (scalar case \(\mathbb{E}[aX + c] = a\,\mathbb{E}[X] + c\)):

For continuous \(X\) with density \(f\), \[ \begin{align*} \mathbb{E}[aX + c] &= \int_{-\infty}^{\infty} (ax + c)\, f(x)\,dx \\\\ &= a\int_{-\infty}^{\infty} x\, f(x)\,dx + c\int_{-\infty}^{\infty} f(x)\,dx \\\\ &= a\,\mathbb{E}[X] + c, \end{align*} \] using \(\int f(x)\,dx = 1\) in the last step. The first equality is the Expectation of a Function of a Random Variable applied to \(g(x) = ax + c\), whose integral against \(f\) converges absolutely because \(|ax + c| \leq |a|\,|x| + |c|\) and \(\mathbb{E}[X]\) exists. It integrates \(ax + c\) against the density of \(X\) itself, so it needs no density for \(aX + c\), which has none when \(a = 0\). The continuous half of that theorem is taken as given on this page.

For discrete \(X\) with p.m.f. \(f\), the same argument with sums gives \[ \begin{align*} \mathbb{E}[aX + c] &= \sum_{x}(ax + c) f(x) \\\\ &= a\sum_x x\,f(x) + c\sum_x f(x) \\\\ &= a\,\mathbb{E}[X] + c. \end{align*} \] Here the first equality regroups the sum. When \(a \neq 0\) the map \(x \mapsto ax + c\) is injective, so the values of \(aX + c\) correspond one to one with those of \(X\) and \(\sum_z z\,P(aX + c = z)\) is the displayed sum term by term. When \(a = 0\) the identity reads \(\mathbb{E}[c] = c\).

The bivariate identity \(\mathbb{E}[X + Y] = \mathbb{E}[X] + \mathbb{E}[Y]\) requires the joint distribution of \((X, Y)\), which we postpone until joint and multivariate distributions are introduced. The argument runs in parallel, with the single sum or integral replaced by a double sum or integral over the joint distribution. Crucially, it does not require \(X\) and \(Y\) to be independent.

Knowing the center of a distribution is valuable, but it tells us nothing about how concentrated or dispersed the values are around that center. Two distributions can share the same mean yet differ dramatically in spread. To quantify this spread, we introduce the variance.

Variance

The variance measures the expected squared deviation of a random variable from its mean. A small variance indicates that the values of \(X\) tend to cluster tightly around \(\mu\), while a large variance indicates wide dispersion. In machine learning, variance appears everywhere: in the bias-variance tradeoff, in gradient noise during stochastic optimization, and in the uncertainty quantification of Bayesian predictions.

Definition: Variance and Standard Deviation

The variance of a random variable \(X\) with mean \(\mu = \mathbb{E}[X]\) is: \[ \operatorname{Var}(X) = \sigma^2 = \mathbb{E}\bigl[(X - \mu)^2\bigr] \geq 0. \] The standard deviation is \(\sigma = \sqrt{\operatorname{Var}(X)}\), which has the same units as \(X\).

The definition involves the unknown quantity \(\mathbb{E}[X]\) inside the expectation. Expanding the square yields a computationally convenient alternative.

Proposition: Computational Identity for Variance

Let \(X\) be a random variable with mass function or density \(f\), and suppose that \(\sum_{x} x^2 f(x)\) (discrete case) or \(\int_{-\infty}^{\infty} x^2 f(x)\,dx\) (continuous case) is finite. Then the mean \(\mu = \mathbb{E}[X]\) exists, \(\operatorname{Var}(X)\) is finite, and \[ \operatorname{Var}(X) = \mathbb{E}[X^2] - \mu^2. \]

Proof:

We give the discrete case. Since \(|x| \leq 1 + x^2\) for every real \(x\), \(\sum_{x} |x|\, f(x) \leq \sum_{x} f(x) + \sum_{x} x^2 f(x) \lt \infty\), so \(\mu\) exists. The Expectation of a Function of a Random Variable applied to \(g(x) = x^2\) identifies the finite sum \(\sum_{x} x^2 f(x)\) with \(\mathbb{E}[X^2]\). Applied to \(g(x) = (x - \mu)^2\), whose sum against \(f\) converges because \((x - \mu)^2 \leq x^2 + 2|\mu|\,|x| + \mu^2\), the same theorem gives \[ \begin{align*} \operatorname{Var}(X) &= \mathbb{E}\bigl[(X - \mu)^2\bigr] \\\\ &= \sum_{x} (x - \mu)^2 f(x) \\\\ &= \sum_{x} x^2 f(x) - 2\mu \sum_{x} x\, f(x) \\\\ &\quad + \mu^2 \sum_{x} f(x) \\\\ &= \mathbb{E}[X^2] - 2\mu^2 + \mu^2 \\\\ &= \mathbb{E}[X^2] - \mu^2. \end{align*} \] The sum splits into three because each of the three series converges absolutely. The continuous case runs verbatim with integrals in place of sums, and rests on the continuous half of that theorem, which the page takes as given.

That is, the variance equals the mean of the square minus the square of the mean. Note that \(\operatorname{Var}(c) = 0\) for any constant \(c\), since a constant has no spread.

The following result describes how variance transforms under linear operations. Variance is affected by scaling but not by translation, in contrast to expectation, which shifts along with the additive constant.

Theorem: Variance of a Linear Transformation

Let \(X\) be a random variable whose mean \(\mu = \mathbb{E}[X]\) exists and whose variance \(\operatorname{Var}(X)\) is finite. For any constants \(a, b \in \mathbb{R}\), \[ \operatorname{Var}(aX + b) = a^2\,\operatorname{Var}(X). \] The additive constant \(b\) shifts the distribution without changing its spread.

Proof:

By the scalar case of linearity of expectation, \(\mathbb{E}[aX + b] = a\mu + b\), so the deviation of \(aX + b\) from its mean is the random variable \((aX + b) - (a\mu + b) = a(X - \mu)\). Put \(Z = (X - \mu)^2\), whose expectation is \(\operatorname{Var}(X)\). The scalar case of linearity, applied now to \(Z\), gives \[ \begin{align*} \operatorname{Var}(aX + b) &= \mathbb{E}\bigl[a^2 (X - \mu)^2\bigr] \\\\ &= \mathbb{E}\bigl[a^2 Z\bigr] \\\\ &= a^2\,\mathbb{E}[Z] \\\\ &= a^2\,\operatorname{Var}(X). \end{align*} \]

Insight: Random Variables in Machine Learning

The framework of random variables, expectations, and variances is the language in which virtually all of machine learning is written. A model's risk is an expectation, the expected loss \(\mathbb{E}[\ell(Y, \hat{Y})]\) over the data distribution, and training typically minimizes an empirical estimate of it. The bias-variance decomposition shows that a model's expected squared prediction error decomposes as \(\text{Bias}^2 + \text{Variance} + \text{Irreducible Noise}\), directly using the concepts defined here. In the pages ahead, we will study specific families of distributions that serve as building blocks for probabilistic models throughout statistics and machine learning. They include the Gamma and Beta and Gaussian families, among others.