Student's \(t\)-Distribution
On the normal distribution page, we saw that the chi-squared distribution governs
sums of squared standard normals, and that the sample variance \(s^2\) satisfies \((n-1)s^2/\sigma^2 \sim \chi^2_{n-1}\).
A natural question follows: what happens to our inference when we replace the unknown population standard deviation
\(\sigma\) with its estimate \(s\)? The answer is the Student's \(t\)-distribution, which arises as the
ratio of a standard normal to the square root of an independent chi-squared variable divided by its degrees of freedom.
Independence of random variables is used on this page before it has been defined. We take for granted that independent
random variables have a joint density that factors into the product of their individual densities, and that the same
product structure carries over to the change of variables below. The machinery behind these facts is developed later.
Definition: Standard Student's \(t\)-Distribution (ratio form)
Let \(Z \sim \mathcal{N}(0, 1)\) and \(Q \sim \chi^2_\nu\) with \(Z\) and \(Q\) independent. The random variable
\[
T = \frac{Z}{\sqrt{Q/\nu}}
\]
is said to follow the standard Student's \(t\)-distribution with \(\nu \gt 0\)
degrees of freedom, written \(T \sim t_\nu\).
Proposition: Standard \(t\) p.d.f.
The p.d.f. of \(T \sim t_\nu\) is
\[
f_T(t) = \frac{\Gamma\!\bigl((\nu+1)/2\bigr)}{\sqrt{\nu\pi}\,\Gamma(\nu/2)}\,
\left(1 + \frac{t^2}{\nu}\right)^{-(\nu+1)/2}, \quad t \in \mathbb{R}.
\]
Sketch. The density of \(T\) is obtained by writing the joint density of \((Z, Q)\) (which factors as
\(\varphi(z)\,f_{\chi^2_\nu}(q)\) by independence), changing variables to \((T, Q)\) with \(Z = T\sqrt{Q/\nu}\), and
integrating out \(Q\). The integrand reduces to a gamma integral that produces the factor \(\Gamma((\nu+1)/2)\), and the
remaining constants combine into the prefactor above. The factorization is the fact taken for granted at the start of
this page. We also take for granted the change-of-variables rule for joint densities, under which the joint density
picks up the absolute value of the Jacobian determinant of the inverse map. Applying it and evaluating the gamma
integral are routine computations that we do not carry out here.
Definition: Student's \(t\)-Distribution (location-scale form)
For location \(\mu \in \mathbb{R}\) and scale \(\sigma \gt 0\) (not the standard deviation), a random
variable \(Y\) follows the Student's \(t\)-distribution with parameters
\((\mu, \sigma, \nu)\), written \(Y \sim t_\nu(\mu, \sigma^2)\), if \(Y = \mu + \sigma T\) for some
\(T \sim t_\nu\). The corresponding p.d.f. is
\[
\begin{align*}
f(y \mid \mu, \sigma, \nu)
&= \frac{1}{\sigma}\,f_T\!\!\left(\frac{y - \mu}{\sigma}\right) \\\\
&= \frac{\Gamma\!\bigl((\nu+1)/2\bigr)}{\sigma\sqrt{\nu\pi}\,\Gamma(\nu/2)}
\left[1 + \frac{1}{\nu}\!\left(\frac{y - \mu}{\sigma}\right)^{\!2}\right]^{-(\nu+1)/2}.
\end{align*}
\]
The first moments of \(Y \sim t_\nu(\mu, \sigma^2)\) are
\[
\begin{align*}
\operatorname{mode}(Y) &= \mu \quad (\text{any } \nu \gt 0), \\\\
\mathbb{E}[Y] &= \mu \quad (\text{if } \nu \gt 1), \\\\
\operatorname{Var}(Y) &= \frac{\nu\,\sigma^2}{\nu - 2} \quad (\text{if } \nu \gt 2).
\end{align*}
\]
The mean fails to exist for \(\nu \leq 1\), because the defining integral diverges (the Cauchy case \(\nu = 1\) below
works this out in full). When \(\nu \leq 2\) the variance fails as well. The density decays only like
\(|y|^{-(\nu+1)}\) at infinity, so the variance integrand \(y^2 f(y)\) decays like \(|y|^{1-\nu}\), which is
integrable only when \(\nu \gt 2\). The mode, by contrast, is always well defined, since the density is symmetric
about \(\mu\) and unimodal there for every \(\nu \gt 0\).
The key feature of the \(t\)-distribution is its heavy tails. Compared to a normal distribution with
the same location and scale, it assigns substantially more probability mass to extreme values, and estimates that use it
as the noise model are correspondingly more robust to outliers.
The parameter \(\nu\) controls the tail heaviness. As \(\nu \to \infty\), the \(t\)-distribution converges to the
normal distribution \(\mathcal{N}(\mu, \sigma^2)\). In practice, for \(\nu \gg 5\), the \(t\)-distribution is nearly
indistinguishable from the normal and loses its robustness advantage.
Cauchy Distribution
An important special case of the \(t\)-distribution arises when the degrees of freedom parameter equals one.
Definition: Cauchy Distribution
When \(\nu = 1\), the Student's \(t\)-distribution reduces to the Cauchy distribution with location
\(\mu\) and scale \(\gamma \gt 0\) that has p.d.f.:
\[
f(x \mid \mu, \gamma) = \frac{1}{\gamma\pi}\left[1 + \left(\frac{x - \mu}{\gamma}\right)^2\right]^{-1}.
\]
We write \(X \sim \text{Cauchy}(\mu, \gamma)\). (Specializing the standard \(t\)-density at \(\nu = 1\) with
location \(\mu\) and scale \(\gamma\) gives the prefactor
\(\Gamma(1)/[\gamma\sqrt{\pi}\,\Gamma(1/2)] = 1/(\gamma\pi)\), using \(\Gamma(1/2) = \sqrt{\pi}\).)
Consider the standard Cauchy distribution (\(\mu = 0, \gamma = 1\)). When we attempt to calculate its expected value:
\[
\mathbb{E}[X] = \frac{1}{\pi} \int_{-\infty}^{\infty} \frac{x}{1 + x^2}\, dx
\]
we find that the integral is not absolutely convergent. The integrand \(x/(1 + x^2)\) behaves like
\(1/x\) for large \(|x|\), and \(\int |x|/(1+x^2)\,dx\) diverges logarithmically (see
Improper Riemann Integrals).
Our definition of expected value requires absolute convergence, that is, the finiteness of \(\mathbb{E}[|X|]\). Even
the extended convention for the Lebesgue integral, which assigns the value \(+\infty\) or \(-\infty\) when exactly one
of the positive and negative parts \(X^+ = \max(X, 0)\) and \(X^- = \max(-X, 0)\) has infinite expectation, assigns
nothing here, since both parts have
infinite expectation. The mean of a Cauchy random variable is therefore undefined, even though the symmetric
Cauchy principal value \(\lim_{R \to \infty} \int_{-R}^{R} x/(1+x^2)\,dx\) equals zero. The variance, which
presupposes the mean, is undefined as well.
Why this matters. Since the mean and variance are undefined, the Law of Large Numbers
fails. Averaging \(n\) i.i.d. Cauchy variables does not make the sample mean \(\bar{X}_n\) settle down. That mean
follows the exact same Cauchy distribution as the individual observations. We state this without proof, since
its standard proof uses characteristic functions, which we do not develop. The
Convergence page returns to this example to explain why limit theorems
need conditions on moments.
Despite these theoretical challenges, the Cauchy distribution is highly useful. Bayesian modeling restricts it to
\(\mathbb{R}^+\) and uses the resulting Half-Cauchy distribution as a prior on scale parameters.
Definition: Half-Cauchy Distribution
The Half-Cauchy distribution with scale \(\gamma \gt 0\) is the Cauchy distribution
\(\text{Cauchy}(0, \gamma)\) restricted to \(x \geq 0\) and renormalized so its total mass is one. Its p.d.f. is
\[
f(x \mid \gamma) = \frac{2}{\pi \gamma}\left[1 + \left(\frac{x}{\gamma}\right)^2\right]^{-1}, \quad x \geq 0.
\]
(The factor \(2\) compensates for the loss of the negative half by the symmetry of the Cauchy density
about the origin.)
The Half-Cauchy is a common choice for scale parameters in hierarchical Bayesian priors. Its heavy tail allows
large values to remain probable while keeping a finite density at the origin. The resulting posterior also
tends to be less sensitive to the choice of hyperparameter \(\gamma\) than the Inverse-Gamma alternative is to
its own hyperparameters.
Laplace Distribution
The Cauchy distribution demonstrates that heavy tails can be extreme enough to prevent even the mean from existing. A
more moderate alternative, which retains finite moments of all orders while still placing more mass in the tails than
the normal distribution, is the Laplace distribution.
Definition: Laplace Distribution
The Laplace distribution (also called the double-sided exponential distribution) with location
\(\mu\) and scale \(b \gt 0\) has p.d.f.:
\[
f(y \mid \mu, b) = \frac{1}{2b}\exp\!\left(-\frac{|y - \mu|}{b}\right).
\]
Its moments are:
\[
\operatorname{mode}(Y) = \mathbb{E}[Y] = \mu, \quad \operatorname{Var}(Y) = 2b^2.
\]
Proof of the Laplace moments:
Substitute \(u = (y - \mu)/b\), so \(y = \mu + bu\) and \(dy = b\,du\). The density transforms to
\(f(y)\,dy = \tfrac{1}{2}e^{-|u|}\,du\), and \((y - \mu)^k = b^k u^k\). Both moments are integrals of a function of
\(Y\) against the density of \(Y\), which is the content of the
Expectation of a Function of a Random Variable,
and \(\mathbb{E}[Y]\) follows from \(\mathbb{E}[Y - \mu]\) by
Linearity of Expectation.
For the mean,
\[
\mathbb{E}[Y] - \mu = \mathbb{E}[Y - \mu] = b\int_{-\infty}^{\infty} u \cdot \tfrac{1}{2}e^{-|u|}\,du = 0,
\]
since \(u\,e^{-|u|}\) is odd and the integral is absolutely convergent (\(|u|\,e^{-|u|}\) integrates to
a finite value).
The variance is the second central moment in the sense of
Variance and Standard Deviation.
Since \(u^2 e^{-|u|}\) is even,
\[
\begin{align*}
\operatorname{Var}(Y)
&= \mathbb{E}[(Y - \mu)^2] \\\\
&= b^2 \int_{-\infty}^{\infty} u^2 \cdot \tfrac{1}{2} e^{-|u|}\,du \\\\
&= b^2 \int_{0}^{\infty} u^2 e^{-u}\,du \\\\
&= b^2 \,\Gamma(3) \\\\
&= 2b^2,
\end{align*}
\]
using \(\Gamma(n) = (n-1)!\) for positive integers \(n\). The mode is \(\mu\) because \(e^{-|y - \mu|/b}\) attains
its maximum exactly at \(y = \mu\).
Insight: Heavy Tails and Regularization
In machine learning, the choice of distribution often corresponds to the choice of regularization.
Under maximum-a-posteriori estimation, an i.i.d. zero-mean Gaussian prior on the
weights produces the \(L_2\) penalty of
Ridge,
and the analogous zero-mean Laplace prior produces the \(L_1\) penalty of
Lasso,
promoting sparsity. Replacing the Gaussian noise model of least squares with a Student's \(t\) noise model makes
the fitted regression more robust to outliers in the response.
The Student's \(t\), Cauchy, and Laplace distributions complete our toolkit of univariate distributions for robust
modeling. The step from single random variables to pairs of random variables brings in
covariance, the fundamental measure of linear dependence between them.