Gaussian Function
The normal (Gaussian) distribution is arguably the most important distribution
in all of statistics and machine learning. Before defining it, we study the underlying
Gaussian function and its integral properties. The key challenge is that the
Gaussian function has no elementary antiderivative, yet its improper integral over all of
\(\mathbb{R}\) can be evaluated exactly, which makes the normal distribution analytically tractable.
A Gaussian function is defined as:
\[
f(x) = e^{-x^2}
\]
and is often parametrized as
\[
f(x) = a e^{-\frac{(x-b)^2}{2c^2}} \tag{1}
\]
where \(a, b, c \in \mathbb{R}\), \(a \neq 0\), and \(c \neq 0\).
Since these functions have no elementary antiderivative, we represent their integral with the special function known as the
error function:
\[
\operatorname{erf}(z) = \frac{2}{\sqrt{\pi}} \int_{0}^z e^{-t^2}dt, \quad \operatorname{erf}: \mathbb{R} \to (-1, 1).
\]
Then,
\[
\int e^{-x^2} dx = \frac{\sqrt{\pi}}{2}\operatorname{erf}(x) + C.
\]
On the other hand, their improper integrals over \(\mathbb{R}\) can be evaluated exactly using the
Gaussian integral:
Theorem: Gaussian Integral
\[
\int_{-\infty} ^\infty e^{-x^2} dx = \sqrt{\pi}.
\]
Note: The generalized Gaussian integral is given by
\[
\int_{-\infty} ^\infty e^{-ax^2} dx = \sqrt{\frac{\pi}{a}} \quad a \gt 0 \tag{2}
\]
Proof:
Let \(I = \int_{-\infty} ^\infty e^{-x^2} dx\). Then
\[
\begin{align*}
I^2 &= \int_{-\infty} ^\infty \int_{-\infty} ^\infty e^{-u^2} e^{-v^2} du dv \\\\
&= \int_{-\infty} ^\infty \int_{-\infty} ^\infty e^{-(u^2 + v^2)} dudv.
\end{align*}
\]
Using polar coordinates \(u = r\cos \theta\) and \(v = r\sin \theta\), we have \(u^2 + v^2 = r^2\) and
\(dudv = r dr d\theta\).
The integrand is non-negative, so the passage to the double integral and the change to polar coordinates need no
integrability hypothesis. Thus
\[
\begin{align*}
I^2 &= \int_{0}^{2\pi} \int_{0}^{\infty} e^{-r^2}\,r\,dr\,d\theta \\\\
&= \left(\int_{0}^{2\pi} d\theta\right)\left(\int_{0}^{\infty} e^{-r^2}\,r\,dr\right).
\end{align*}
\]
The angular integral is \(2\pi\). For the radial integral, the antiderivative of \(e^{-r^2} r\) with respect to \(r\)
is \(-\tfrac{1}{2}e^{-r^2}\), so
\[
\begin{align*}
\int_{0}^{\infty} e^{-r^2}\,r\,dr
&= \left[-\tfrac{1}{2}e^{-r^2}\right]_{0}^{\infty} \\\\
&= 0 - \left(-\tfrac{1}{2}\right) \\\\
&= \tfrac{1}{2}.
\end{align*}
\]
Therefore \(I^2 = 2\pi \cdot \tfrac{1}{2} = \pi\), and since \(I \gt 0\),
\[
I = \sqrt{\pi}.
\]
This computation is the case \(a = 1\) of the generalized Gaussian integral. We now extend it to a scaling factor
\(a \gt 0\).
Let \(u = \sqrt{a}\,x\), so that \(x = \frac{u}{\sqrt{a}} \) and \(dx = \frac{1}{\sqrt{a}}du\). Substituting these into
(2), we obtain
\[
\begin{align*}
& \int_{-\infty} ^\infty e^{-a(\frac{u}{\sqrt{a}})^2} \frac{1}{\sqrt{a}}du \\\\
&= \frac{1}{\sqrt{a}} \int_{-\infty} ^\infty e^{-u^2}du \\\\
&= \frac{1}{\sqrt{a}} (\sqrt{\pi}) \\\\
&= \sqrt{\frac{\pi}{a}}.
\end{align*}
\]
Here, we use the parametrized Gaussian function (1). Let \(u = \frac{x-b}{c}\), which implies \(x = cu + b\) and
\(dx = c\,du\). The exponent becomes \(-\frac{(x-b)^2}{2c^2} = -\frac{u^2}{2}\). We must track the orientation of the
limits, since this depends on the sign of \(c\). If \(c \gt 0\), then \(u \to \pm\infty\) as \(x \to \pm\infty\), so the
limits are preserved:
\[
\int_{-\infty}^{\infty} a e^{-\frac{(x-b)^2}{2c^2}}\,dx = a c \int_{-\infty}^{\infty} e^{-\frac{u^2}{2}}\,du.
\]
If \(c \lt 0\), then \(u \to \mp\infty\) as \(x \to \pm\infty\), so the limits are reversed. Flipping them back introduces
a sign:
\[
\begin{align*}
\int_{-\infty}^{\infty} a e^{-\frac{(x-b)^2}{2c^2}}\,dx
&= a c \int_{+\infty}^{-\infty} e^{-\frac{u^2}{2}}\,du \\\\
&= -a c \int_{-\infty}^{\infty} e^{-\frac{u^2}{2}}\,du.
\end{align*}
\]
Since \(c = |c|\) when \(c \gt 0\) and \(-c = |c|\) when \(c \lt 0\), both cases combine into a single expression with
\(a|c|\). By (2) applied with parameter \(1/2\) in place of \(a\), we have
\(\int_{-\infty}^\infty e^{-u^2/2}\,du = \sqrt{2\pi}\), so
\[
\int_{-\infty}^{\infty} a e^{-\frac{(x-b)^2}{2c^2}}\,dx = a|c|\sqrt{2\pi}. \tag{3}
\]
Normal (Gaussian) Distribution
With the Gaussian integral established, we can now construct a proper probability density
function from the Gaussian function by choosing the parameters so that the total area under
the curve equals one.
Definition: Normal (Gaussian) Distribution
A random variable \(X\) has a normal (Gaussian) distribution with mean \(\mu \in \mathbb{R}\) and
variance \(\sigma^2 \gt 0\), where \(\sigma \gt 0\) denotes the positive square root, if its p.d.f. is:
\[
f(x) = \frac{1}{\sigma \sqrt{2\pi}}\exp\!\left(-\frac{(x - \mu)^2}{2\sigma^2}\right), \quad x \in \mathbb{R}.
\]
We write \(X \sim \mathcal{N}(\mu, \sigma^2)\).
The integral formula (3) shows that \(f\) integrates to \(1\). The density is the parametrized Gaussian (1)
with amplitude \(a = \frac{1}{\sigma\sqrt{2\pi}}\), location \(b = \mu\), and scale \(c = \sigma \gt 0\), so its integral
over \(\mathbb{R}\) equals \(a|c|\sqrt{2\pi} = ac\sqrt{2\pi} = 1\).
Note that the c.d.f. of the normal distribution does not have a closed form in elementary functions:
\[
F(x) = \int_{-\infty}^{x} \frac{1}{\sigma \sqrt{2\pi}}\exp\!\left(-\frac{(t - \mu)^2}{2\sigma^2}\right) dt.
\]
The above figure shows the normal p.d.f. curve. Here \(\mu\) is the center, and \(\sigma\) is the distance from the center
to the inflection point of the curve. One reason the normal distribution is so widely used in statistics and machine
learning is that its two parameters are exactly its mean and variance:
\[
\mathbb{E}[X] = \mu \quad \operatorname{Var}(X) = \sigma^2.
\]
In a normal distribution the measures of central tendency (mean, median, and mode) coincide.
Proof of \(\mathbb{E}[X] = \mu\) and \(\operatorname{Var}(X) = \sigma^2\):
Apply the substitution \(u = (x - \mu)/\sigma\), so \(x = \sigma u + \mu\) and \(dx = \sigma\,du\). For the mean,
\[
\begin{align*}
\mathbb{E}[X] &= \int_{-\infty}^{\infty} x \cdot \frac{1}{\sigma\sqrt{2\pi}}\exp\!\left(-\frac{(x-\mu)^2}{2\sigma^2}\right) dx \\\\
&= \frac{1}{\sqrt{2\pi}}\int_{-\infty}^{\infty} (\sigma u + \mu)\,e^{-u^2/2}\,du \\\\
&= \frac{\sigma}{\sqrt{2\pi}}\underbrace{\int_{-\infty}^{\infty} u\,e^{-u^2/2}\,du}_{= 0 \text{ (odd integrand)}}
+ \frac{\mu}{\sqrt{2\pi}}\underbrace{\int_{-\infty}^{\infty} e^{-u^2/2}\,du}_{= \sqrt{2\pi}} \\\\
&= 0 + \mu = \mu.
\end{align*}
\]
For the variance we start from the definition of
Variance and Standard Deviation and
evaluate \(\mathbb{E}\bigl[(X - \mu)^2\bigr]\) against the density of \(X\) itself, which is the content of the
Expectation of a Function of a Random Variable.
The same substitution gives \((x - \mu)^2 = \sigma^2 u^2\), so
\[
\begin{align*}
\operatorname{Var}(X) &= \mathbb{E}\bigl[(X - \mu)^2\bigr] \\\\
&= \int_{-\infty}^{\infty} (x - \mu)^2 \cdot \frac{1}{\sigma\sqrt{2\pi}} e^{-(x-\mu)^2/(2\sigma^2)}\,dx \\\\
&= \frac{\sigma^2}{\sqrt{2\pi}}\int_{-\infty}^{\infty} u^2\,e^{-u^2/2}\,du.
\end{align*}
\]
The remaining integral is evaluated by integration by parts with \(v = u\) and \(dw = u\,e^{-u^2/2}\,du\) (so
\(w = -e^{-u^2/2}\)):
\[
\begin{align*}
\int_{-\infty}^{\infty} u^2\,e^{-u^2/2}\,du
&= \bigl[-u\,e^{-u^2/2}\bigr]_{-\infty}^{\infty} + \int_{-\infty}^{\infty} e^{-u^2/2}\,du \\\\
&= 0 + \sqrt{2\pi},
\end{align*}
\]
the boundary term vanishing by exponential decay (quantified in the proof of the next theorem). Therefore
\(\operatorname{Var}(X) = (\sigma^2/\sqrt{2\pi})\cdot\sqrt{2\pi} = \sigma^2\).
The integration by parts in the variance computation is the first step of a recursion that reaches every even
moment. We record the general result for a centered normal variable.
Theorem: Moments of the Centered Normal Distribution
Let \(X \sim \mathcal{N}(0, \sigma^2)\) with \(\sigma \gt 0\). Then:
(a) Finiteness.
\(\mathbb{E}\bigl[|X|^m\bigr] \lt \infty\) for every integer \(m \geq 0\).
(b) Even moments.
For every integer \(k \geq 1\),
\[
\mathbb{E}\bigl[X^{2k}\bigr] = (2k - 1)(2k - 3) \cdots 3 \cdot 1 \cdot \sigma^{2k},
\]
the product running over the odd integers from \(1\) to \(2k - 1\). In particular,
\(\mathbb{E}[X^2] = \sigma^2\) and \(\mathbb{E}[X^4] = 3\sigma^4\).
Proof:
As in the variance computation, the
Expectation of a Function of a Random Variable
evaluates each moment against the density of \(X\), and the substitution \(u = x/\sigma\) turns it into
\[
\mathbb{E}\bigl[g(X)\bigr] = \frac{1}{\sqrt{2\pi}}\int_{-\infty}^{\infty} g(\sigma u)\,e^{-u^2/2}\,du
\]
for the continuous functions \(g(x) = |x|^m\) and \(g(x) = x^{2k}\), provided the integral converges absolutely.
For (a), fix \(m \geq 0\). The exponential series gives \(e^{y} \geq y^m/m!\) for \(y \geq 0\), so
\(e^{u^2/4} \geq u^{2m}/(4^m m!)\). Hence \(|u|^m e^{-u^2/4} \leq 4^m m!\,|u|^{-m} \leq 4^m m!\) for \(|u| \geq 1\),
while for \(|u| \lt 1\) the left side is at most \(1 \leq 4^m m!\). With \(C_m = 4^m m!\),
\[
\begin{align*}
|u|^m e^{-u^2/2}
&= \bigl(|u|^m e^{-u^2/4}\bigr)\, e^{-u^2/4} \\\\
&\leq C_m\, e^{-u^2/4}
\end{align*}
\]
for every \(u \in \mathbb{R}\). Taking \(1/4\) in place of \(a\) in (2), the function \(e^{-u^2/4}\) has integral
\(2\sqrt{\pi}\) over \(\mathbb{R}\). The integral for \(g(x) = |x|^m\) therefore converges, and
\(\mathbb{E}\bigl[|X|^m\bigr] \leq \sqrt{2}\,C_m \sigma^m\), which is finite.
For (b), write \(J_k(R) = \int_{-R}^{R} u^{2k} e^{-u^2/2}\,du\) for \(R \gt 0\) and
\(J_k = \int_{-\infty}^{\infty} u^{2k} e^{-u^2/2}\,du\), which is finite by the estimate in the proof of (a) and is the
limit of \(J_k(R)\) as \(R \to \infty\). Moreover, \(J_0 = \sqrt{2\pi}\), the integral already evaluated in deriving
(3). Fix \(k \geq 1\). Integrating by parts on \([-R, R]\) with \(v = u^{2k-1}\) and \(dw = u\,e^{-u^2/2}\,du\), so
that \(w = -e^{-u^2/2}\) as before, we obtain
\[
\begin{align*}
J_k(R)
&= \bigl[-u^{2k-1} e^{-u^2/2}\bigr]_{-R}^{R} + (2k - 1)\, J_{k-1}(R) \\\\
&= -2R^{2k-1} e^{-R^2/2} + (2k - 1)\, J_{k-1}(R).
\end{align*}
\]
The pointwise bound \(|u|^m e^{-u^2/4} \leq C_m\) from the proof of (a), taken with \(m = 2k - 1\), gives
\(2R^{2k-1} e^{-R^2/2} \leq 2C_{2k-1} e^{-R^2/4}\), so the boundary term tends to \(0\) as \(R \to \infty\). Letting
\(R \to \infty\) therefore yields \(J_k = (2k - 1) J_{k-1}\). Induction on \(k\) from \(J_0 = \sqrt{2\pi}\) gives
\(J_k = (2k - 1)(2k - 3) \cdots 3 \cdot 1 \cdot \sqrt{2\pi}\), and the substitution formula with \(g(x) = x^{2k}\)
turns this into \(\mathbb{E}\bigl[X^{2k}\bigr] = \sigma^{2k} J_k/\sqrt{2\pi}\), which completes the proof.
The simplest and most useful normal distribution is the one with zero mean and unit variance. We call it the
standard normal distribution denoted by
\[
Z \sim \mathcal{N}(0, 1).
\]
Any normally distributed random variable can be transformed into a standard normal random variable. If
\(X \sim \mathcal{N}(\mu, \sigma^2)\), then
\[
Z = \frac{X- \mu}{\sigma}\sim \mathcal{N}(0, 1).
\]
This process is called standardization, a special case of the linear transformation:
\[
Y = \alpha X + \beta \sim \mathcal{N}(\alpha\mu+\beta, \alpha^2\sigma^2) \quad (\alpha \neq 0).
\]
Standardization corresponds to \(\alpha = 1/\sigma\) and \(\beta = -\mu/\sigma\).
Linearity of Expectation and
the
Variance of a Linear Transformation
then give \(\mathbb{E}[Z] = (\mu - \mu)/\sigma = 0\) and \(\operatorname{Var}(Z) = \sigma^2/\sigma^2 = 1\).
Remark on notation. Up to this point we have written the density and c.d.f. of a single random variable
plainly as \(f(x)\) and \(F(x)\), since the variable in question was unambiguous. In the proof below we work with two
related variables \(X\) and \(Y = \alpha X + \beta\) simultaneously, so we must distinguish their densities and c.d.f.s. We
adopt the standard subscript convention \(f_X, F_X\) (resp. \(f_Y, F_Y\)) whenever multiple random variables are present
and confusion would otherwise be possible. We revert to plain \(f, F\) once the context fixes a single variable.
Proof of \(Y = \alpha X + \beta \sim \mathcal{N}(\alpha\mu + \beta, \alpha^2\sigma^2)\):
Assume \(\alpha \neq 0\). We derive the density of \(Y\) from its c.d.f. Consider first \(\alpha \gt 0\):
\[
F_Y(y) = P(\alpha X + \beta \leq y) = P\!\left(X \leq \frac{y - \beta}{\alpha}\right) = F_X\!\left(\frac{y - \beta}{\alpha}\right).
\]
Differentiating with respect to \(y\) and using the chain rule, we obtain
\(f_Y(y) = f_X\!\bigl((y - \beta)/\alpha\bigr)\cdot (1/\alpha)\). For \(\alpha \lt 0\), the inequality flips:
\(F_Y(y) = P(X \geq (y-\beta)/\alpha) = 1 - F_X\!\bigl((y-\beta)/\alpha\bigr)\), and differentiating gives
\(f_Y(y) = -f_X\!\bigl((y-\beta)/\alpha\bigr)\cdot(1/\alpha)\). Since \(\alpha \lt 0\) implies
\(-1/\alpha = 1/|\alpha|\), both cases combine to
\[
f_Y(y) = \frac{1}{|\alpha|}\,f_X\!\left(\frac{y - \beta}{\alpha}\right).
\]
Substituting the normal density,
\[
f_Y(y) = \frac{1}{|\alpha|\,\sigma\sqrt{2\pi}}\exp\!\left(-\frac{\bigl((y - \beta)/\alpha - \mu\bigr)^2}{2\sigma^2}\right).
\]
Rewriting the exponent,
\[
\frac{y - \beta}{\alpha} - \mu = \frac{y - \beta - \alpha\mu}{\alpha} = \frac{y - (\alpha\mu + \beta)}{\alpha},
\]
so its square is \((y - (\alpha\mu+\beta))^2 / \alpha^2\), and
\[
f_Y(y) = \frac{1}{|\alpha|\,\sigma\sqrt{2\pi}}\exp\!\left(-\frac{(y - (\alpha\mu + \beta))^2}{2\,\alpha^2\sigma^2}\right).
\]
This matches the \(\mathcal{N}(\alpha\mu + \beta, \alpha^2\sigma^2)\) density (with standard deviation
\(|\alpha|\sigma\) and variance \(\alpha^2\sigma^2\)). Therefore
\(Y \sim \mathcal{N}(\alpha\mu + \beta, \alpha^2\sigma^2)\).
The p.d.f. of \(Z\) is denoted by
\[
\phi(z) = \frac{1}{\sqrt{2\pi}}e^{-z^2/2}, \quad z \in \mathbb{R},
\]
and the c.d.f. of \(Z\) is written
\[
\Phi(z) = P(Z \leq z) = \int_{-\infty}^z \phi(u)\,du.
\]
Chi-Squared Distribution
If a single standard normal variable \(Z \sim \mathcal{N}(0,1)\) captures the behavior of one random measurement, what
distribution governs the sum of squared measurements? This question arises naturally in statistics. Whenever we
compute a sample variance, we are summing squared deviations. The answer is the chi-squared distribution,
which connects the normal distribution to the gamma distribution of the previous
page.
Definition: Chi-Squared Distribution
Let \(Z_1, Z_2, \ldots, Z_\nu\) be independent standard normal random variables,
\(Z_i \sim \mathcal{N}(0, 1)\). The random variable
\[
Q = Z_1^2 + Z_2^2 + \cdots + Z_\nu^2 = \sum_{i=1}^{\nu} Z_i^2
\]
is said to have a chi-squared distribution with \(\nu\) degrees of freedom, written
\(Q \sim \chi^2_\nu\). This construction requires \(\nu\) to be a positive integer. The density recorded below is
defined for every real \(\nu \gt 0\), and we take that density as the definition of \(\chi^2_\nu\) for non-integer
\(\nu\).
Proposition: Chi-Squared p.d.f.
The p.d.f. of \(Q \sim \chi^2_\nu\) is
\[
f(x) = \frac{1}{2^{\nu/2}\,\Gamma(\nu/2)}\, x^{\nu/2 - 1}\, e^{-x/2}, \quad x \gt 0.
\]
Equivalently, \(Q \sim \text{Gamma}\!\left(\alpha = \nu/2, \beta = 1/2\right)\) in the rate parametrization of the
Gamma Distribution.
Sketch. The proof has two ingredients. (i) For a single standard normal \(Z\), the change of variables
\(Y = Z^2\) (using the symmetry \(P(Z^2 \leq y) = 2\Phi(\sqrt{y}) - 1\) and differentiating) yields the density
\(f_Y(y) = (1/\sqrt{2\pi})\,y^{-1/2}e^{-y/2}\) for \(y \gt 0\), which matches \(\text{Gamma}(1/2, 1/2)\). (ii) The sum of
\(\nu\) independent \(\text{Gamma}(1/2, 1/2)\) random variables is \(\text{Gamma}(\nu/2, 1/2)\) (additivity of independent
gamma variables with a common rate parameter, proved via convolution of densities or moment generating functions).
Ingredient (i) is a one-variable computation. Ingredient (ii) requires the machinery of independence and joint
distributions, which we develop in
Limit Theorems & Product Measures.
The mean and variance follow immediately from the gamma parameters \(\alpha = \nu/2\) and \(\beta = 1/2\):
\[
\mathbb{E}[Q] = \frac{\alpha}{\beta} = \nu, \quad
\operatorname{Var}(Q) = \frac{\alpha}{\beta^2} = 2\nu.
\]
The chi-squared distribution plays a fundamental role in statistical inference. If \(X_1, \ldots, X_n\) are i.i.d.
\(\mathcal{N}(\mu, \sigma^2)\) with \(n \geq 2\) and we define the sample variance
\(s^2 = \frac{1}{n-1}\sum_{i=1}^{n}(X_i - \bar{X})^2\), then:
\[
\frac{(n-1)s^2}{\sigma^2} \sim \chi^2_{n-1}.
\]
We state this without proof. The result follows from Cochran's theorem, which decomposes the quadratic
form \(\sum_i (X_i - \bar{X})^2 / \sigma^2\) along orthogonal subspaces of \(\mathbb{R}^n\) and shows that the residual
component is a sum of \(n-1\) independent squared standard normals. Exactly one degree of freedom is "spent" on
estimating \(\mu\) by \(\bar{X}\). The proof requires multivariate normal theory not yet developed in this section. This
result is what allows us to construct confidence intervals for variances. Together with two further facts that we also
state without proof, that \(\bar{X} \sim \mathcal{N}(\mu, \sigma^2/n)\) exactly and that \(\bar{X}\) and \(s^2\) are
independent, it shows that \((\bar{X} - \mu)/(s/\sqrt{n})\) follows the
Student's \(t\)-distribution with \(n - 1\) degrees of freedom.
Insight: The Gamma → Chi-Squared → Student's \(t\) Chain
The relationship between these distributions forms a coherent chain. The gamma distribution is a
flexible family for non-negative continuous variables. Setting \(\alpha = \nu/2\) and \(\beta = 1/2\) specializes it to
the chi-squared distribution, which governs sums of squared independent standard normals. The
Student's \(t\)-distribution then arises as the ratio:
\[
T = \frac{Z}{\sqrt{Q/\nu}}, \quad Z \sim \mathcal{N}(0,1), \quad Q \sim \chi^2_\nu, \quad Z \perp Q
\]
where \(T \sim t_\nu\).
In practice, \(Z\) is the standardized signal and \(Q/\nu\) the estimated noise level. When we do not know the true
variance \(\sigma^2\) and must estimate it from data, replacing \(\sigma\) with \(s\) introduces extra uncertainty that
thickens the tails. The \(t\)-distribution's degrees of freedom parameter \(\nu = n - 1\) captures this precisely.
Central Limit Theorem
So far we have studied the normal distribution and its properties for a single random variable. In practice, however, we
almost always work with collections of observations. A fundamental question arises: if we average many independent
measurements, what can we say about the distribution of that average? The answer is provided by the
Central Limit Theorem, one of the most remarkable results in all of mathematics. The suitably
standardized average of independent, identically distributed measurements with finite, nonzero variance tends toward a
normal distribution, regardless of the underlying distribution of the individual measurements.
Now, we consider a set of random variables \(\{X_1, X_2, \ldots, X_n\}\). In random sampling, we assume
that the \(X_i\) have the same distribution and are mutually independent. Such a collection is called
independent and identically distributed (i.i.d.). In this case, a single sample mean
\(\bar{X} = \frac{1}{n}\sum_{i=1} ^n X_i\) varies from sample to sample, so in general it differs from the
population mean \(\mu\). Its expected value, however, is exactly \(\mu\) by
Linearity of Expectation:
\[
\begin{align*}
\mathbb{E}[\bar{X}] &= \mathbb{E}\!\left[\tfrac{1}{n}(X_1 + X_2 + \cdots +X_n)\right] \\\\
&= \tfrac{1}{n}\bigl[\mathbb{E}[X_1] + \mathbb{E}[X_2] + \cdots + \mathbb{E}[X_n]\bigr] \\\\
&= \tfrac{n\mu}{n} = \mu.
\end{align*}
\]
Thus the sample mean is an unbiased estimator of \(\mu\). We take for granted here that the variance of a
sum of independent random variables is the sum of their variances. The joint distributions needed to prove this are
developed later. Granting it, the variance of the sample mean is
\[
\begin{align*}
\operatorname{Var}(\bar{X}) &= \operatorname{Var}\!\left[\tfrac{1}{n}(X_1 + X_2 + \cdots +X_n)\right] \\\\
&= \tfrac{1}{n^2}\bigl[\operatorname{Var}(X_1) + \operatorname{Var}(X_2) + \cdots + \operatorname{Var}(X_n)\bigr] \\\\
&= \tfrac{n \sigma^2}{n^2} = \tfrac{\sigma^2}{n},
\end{align*}
\]
so \(\bar{X}\) concentrates around \(\mu\) with standard deviation \(\sigma/\sqrt{n}\). That concentration is the
foundation of the law of large numbers.
The naive sample-variance estimator \(\widetilde{s}^{\,2} = \tfrac{1}{n}\sum_{i=1}^n (X_i - \bar{X})^2\), however, is a
biased estimator of \(\sigma^2\). One can show that
\(\mathbb{E}[\widetilde{s}^{\,2}] = \tfrac{n-1}{n}\sigma^2\), which underestimates \(\sigma^2\) because the same data is
used both to center the deviations and to compute their average. To remove the bias, we adjust the denominator, assuming
\(n \geq 2\):
\[
s^2 = \frac{1}{n-1} \sum_{i=1}^n (X_i - \bar{X})^2,
\]
so that \(\mathbb{E}[s^2] = \sigma^2\).
To verify the bias factor, write each deviation about the population mean,
\(X_i - \bar{X} = (X_i - \mu) - (\bar{X} - \mu)\), and sum the squares:
\[
\sum_{i=1}^n (X_i - \bar{X})^2
= \sum_{i=1}^n (X_i - \mu)^2 - 2(\bar{X} - \mu)\sum_{i=1}^n (X_i - \mu) + n(\bar{X} - \mu)^2.
\]
Since \(\sum_{i=1}^n (X_i - \mu) = n(\bar{X} - \mu)\), the middle term equals \(-2n(\bar{X}-\mu)^2\), and the identity
collapses to
\[
\sum_{i=1}^n (X_i - \bar{X})^2 = \sum_{i=1}^n (X_i - \mu)^2 - n(\bar{X} - \mu)^2.
\]
Taking expectations and using \(\mathbb{E}[(X_i - \mu)^2] = \sigma^2\) together with
\(\mathbb{E}[(\bar{X} - \mu)^2] = \operatorname{Var}(\bar{X}) = \sigma^2/n\),
\[
\mathbb{E}\!\left[\sum_{i=1}^n (X_i - \bar{X})^2\right] = n\sigma^2 - n\cdot\frac{\sigma^2}{n} = (n-1)\sigma^2,
\]
which gives \(\mathbb{E}[\widetilde{s}^{\,2}] = \tfrac{n-1}{n}\sigma^2\) and hence \(\mathbb{E}[s^2] = \sigma^2\).
At this point we might ask how to make inferences about an unknown population distribution in practice. The
Central Limit Theorem (CLT) answers this. Whatever the population's underlying distribution,
provided it has finite, nonzero variance, the distribution of the standardized sample mean approaches the
standard normal distribution as \(n\) grows. This property is what makes the normal distribution the working tool
of statistical inference.
Theorem: Central Limit Theorem
Let \(X_1, X_2, \ldots\) be an infinite sequence of i.i.d. random variables with mean \(\mu\) and finite variance
\(\sigma^2 \gt 0\), and for each \(n\) let \(\bar{X}_n = \frac{1}{n}\sum_{i=1}^n X_i\). Then the distribution of
\[
Z_n = \frac{\bar{X}_n - \mu}{\frac{\sigma}{\sqrt{n}}}
\]
converges to the standard normal distribution as \(n \to \infty\).
Note: The standard deviation of \(\bar{X}_n\) is \(\sqrt{\frac{\sigma^2}{n}} = \frac{\sigma}{\sqrt{n}}\).
If the sample size \(n\) is large enough,
\[
\bar{X} \approx \mathcal{N}\!\left(\mu, \frac{\sigma^2}{n}\right)
\]
and
\[
\sum_{i =1} ^n X_i \approx \mathcal{N}(n\mu , n\sigma^2).
\]
Note: The rate at which this approximation becomes accurate depends on the underlying distribution. A common
rule of thumb is that mildly skewed distributions are well approximated by \(n \approx 30\), while highly skewed
distributions such as the exponential typically require \(n \gtrsim 100\) before the normal approximation is accurate in
practice.
The CLT explains the ubiquity of the normal distribution, and its extensions to independent summands that are not
identically distributed go further: many quantities built up from a large number of independent contributions, each of
which never deviates from its mean by more than a small fraction of the standard deviation of the total, are
approximately normally distributed. For a rigorous treatment of the i.i.d. case, with a proof that additionally assumes
the moment generating function exists near zero, see
Convergence.
Beyond the CLT, the Gaussian distribution is fundamental in information theory as the distribution that
maximizes entropy for a fixed mean and variance. That property makes it the
"most random" choice under these constraints.
Next, we introduce the Student's \(t\)-distribution,
which accounts for the additional uncertainty that arises when the population variance \(\sigma^2\)
is unknown and must be estimated from the sample.