Gamma Function
Before introducing the Gamma and Beta distributions, we need the special functions that serve as their building blocks.
The gamma function extends the factorial to non-integer (and even complex) arguments. This generalization
is essential. The normalizing constants of many continuous distributions involve factorials of non-integer parameters,
and the gamma function provides the mathematically rigorous way to handle them.
Definition: Gamma Function
The gamma function is defined as:
\[
\Gamma(z) = \int_{0}^{\infty} t^{z-1}e^{-t}\, dt, \quad z \gt 0.
\]
We work with real positive \(z\) throughout this page. The integral in fact converges for
all complex \(z\) with \(\operatorname{Re}(z) \gt 0\), and the gamma function extends to
a meromorphic function on the entire complex plane via analytic continuation. We do not
use the complex extension here.
Theorem: Gamma Function Extends the Factorial
For any positive integer \(n\),
\[
\Gamma (n+1) = n!
\]
Proof:
We first establish the recurrence \(\Gamma(n+1) = n\,\Gamma(n)\) for any \(n \gt 0\) by
integration by parts on
\[
\Gamma(n+1) = \int_0^\infty t^{n} e^{-t}\, dt.
\]
Take \(u = t^n\) (so \(du = n\,t^{n-1}\,dt\)) and \(dv = e^{-t}\,dt\) (so \(v = -e^{-t}\)). Then
\[
\begin{align*}
\Gamma(n+1) &= \bigl[-t^n e^{-t}\bigr]_0^\infty + \int_0^\infty n\,t^{n-1} e^{-t}\,dt \\\\
&= n\int_0^\infty t^{n-1} e^{-t}\,dt \\\\
&= n\,\Gamma(n),
\end{align*}
\]
using that the boundary term vanishes: \(t^n e^{-t} \to 0\) as \(t \to \infty\) (exponential
decay dominates polynomial growth) and \(t^n e^{-t} \to 0\) as \(t \to 0^+\) for any \(n \gt 0\).
We now prove \(\Gamma(n+1) = n!\) for positive integers \(n\) by induction on \(n\).
Base case (\(n = 1\)):
\[
\begin{align*}
\Gamma(2) &= \int_0^\infty t\,e^{-t}\,dt \\\\
&= \bigl[-t\,e^{-t}\bigr]_0^\infty + \int_0^\infty e^{-t}\,dt \\\\
&= 0 + 1 = 1 = 1!,
\end{align*}
\]
again by integration by parts with \(u = t\), \(dv = e^{-t}\,dt\).
Inductive step: Assume \(\Gamma(k+1) = k!\) for some integer \(k \geq 1\).
By the recurrence,
\[
\Gamma(k+2) = (k+1)\,\Gamma(k+1) = (k+1)\,k! = (k+1)!.
\]
By induction, \(\Gamma(n+1) = n!\) for all positive integers \(n\). Because
\(\Gamma(1) = \int_0^\infty e^{-t}\,dt = \bigl[-e^{-t}\bigr]_0^\infty = 1 = 0!\), we also have
\(\Gamma(n) = (n-1)!\) for every positive integer \(n\).
We now compute the most important non-integer value of the gamma function.
\[
\Gamma (\frac{1}{2}) = \int_{0} ^\infty t^{\frac{-1}{2}}e^{-t} dt
\]
Let \(t = u^2\), so \(dt =2udu\). Then
\[
\begin{align*}
\Gamma (\frac{1}{2}) &= \int_{0} ^\infty u^{2\cdot \frac{-1}{2}}e^{-u^2} 2udu \\\\
&= 2\int_{0} ^\infty e^{-u^2} du \\\\
&= \int_{-\infty} ^\infty e^{-u^2} du \\\\
&= \sqrt{\pi}
\end{align*}
\]
In the third line we used the symmetry of the even function \(e^{-u^2}\) to extend the integral
to the full real line, and the final equality is the famous Gaussian integral:
\[
\int_{-\infty} ^\infty e^{-x^2} dx = \sqrt{\pi}.
\]
This Gaussian Integral
is proved on the normal distribution page by passing to polar coordinates, so using it here
introduces no circularity.
Although the factorial \(n!\) is defined only for non-negative integers, the gamma function
extends it to all positive arguments via
\[
(z)! := \Gamma(z+1).
\]
The recurrence \(\Gamma(z+1) = z\,\Gamma(z)\) translates into the familiar factorial-style recursion
\((z)! = z\,(z-1)!\). The integration-by-parts argument in the proof above was carried out for an
arbitrary real \(n \gt 0\), so it establishes \(\Gamma(z+1) = z\,\Gamma(z)\) for every real \(z \gt 0\).
Combined with \(\Gamma(1/2) = \sqrt{\pi}\), the recursion yields:
\[
\begin{align*}
\Gamma(3/2) &= \tfrac{1}{2}\,\Gamma(1/2) = \tfrac{1}{2}\sqrt{\pi}, \\\\
\Gamma(5/2) &= \tfrac{3}{2}\,\Gamma(3/2) = \tfrac{3}{4}\sqrt{\pi}, \\\\
\Gamma(7/2) &= \tfrac{5}{2}\,\Gamma(5/2) = \tfrac{15}{8}\sqrt{\pi}.
\end{align*}
\]
Bonus: Volume of the \(n\)-dimensional ball of radius \(r\):
\[
V_n(r) = \frac{\pi^{n/2}}{\Gamma\!\left(\tfrac{n}{2} + 1\right)}\,r^n.
\]
(Recall that the sphere is the surface and the ball is the solid region it bounds.)
Gamma Distribution
With the gamma function in hand, we can now define a flexible family of continuous distributions
on \([0, \infty )\). The gamma distribution arises naturally as the distribution
of waiting times. If events occur independently at a constant average rate, the time until the
\(n\)-th event follows a gamma distribution with \(\alpha = n\).
Definition: Gamma Distribution
A random variable \(X\) has the gamma distribution with shape parameter
\(\alpha \gt 0\) and rate parameter \(\beta \gt 0\) if its p.d.f. is:
\[
f(x) = \frac{\beta^\alpha}{\Gamma(\alpha)}x^{\alpha-1}e^{-\beta x}, \quad x \gt 0,
\quad f(x) = 0 \text{ for } x \lt 0.
\]
We write \(X \sim \text{Gamma}(\alpha, \beta)\). The mean and variance are:
\[
\mathbb{E}[X] = \frac{\alpha}{\beta}, \quad \operatorname{Var}(X) = \frac{\alpha}{\beta^2}.
\]
That \(f\) integrates to \(1\) follows from the substitution \(t = \beta x\):
\[
\begin{align*}
\int_0^\infty \frac{\beta^\alpha}{\Gamma(\alpha)} x^{\alpha-1} e^{-\beta x}\,dx
&= \frac{\beta^\alpha}{\Gamma(\alpha)} \cdot \frac{1}{\beta^\alpha} \int_0^\infty t^{\alpha-1} e^{-t}\,dt \\\\
&= \frac{\Gamma(\alpha)}{\Gamma(\alpha)} = 1.
\end{align*}
\]
The density is stated on \(x \gt 0\) because \(x^{\alpha-1}\) is unbounded at the origin when
\(\alpha \lt 1\). The single point \(x = 0\) carries no probability, so leaving it unassigned changes
no integral.
The mean and variance of the gamma distribution:
Using the definition of the mean,
\[
\begin{align*}
\mathbb{E}[X] &= \int_{0}^\infty x f(x)dx \\\\
&= \frac{\beta^\alpha}{\Gamma(\alpha)} \int_{0}^\infty x^{\alpha}e^{-\beta x} dx
\end{align*}
\]
Let \(t = \beta x\) and so \(x = \frac{t}{\beta}\) and \(dx = \frac{1}{\beta}dt\).
Substituting these into the equation:
\[
\begin{align*}
\mathbb{E}[X] &= \frac{\beta^\alpha}{\Gamma(\alpha)} \int_{0}^\infty (\frac{t}{\beta})^{\alpha}e^{-t} \frac{1}{\beta}dt \\\\
&= \frac{\beta^\alpha}{\Gamma(\alpha) \beta^{\alpha +1}} \int_{0}^\infty t^{\alpha}e^{-t} dt \\\\
&= \frac{\Gamma (\alpha +1)}{\Gamma(\alpha) \beta}
\end{align*}
\]
Since \(\Gamma(\alpha +1) = \alpha \Gamma(\alpha)\),
\[
\begin{align*}
\mathbb{E}[X] &= \frac{\alpha \Gamma (\alpha)}{\Gamma(\alpha) \beta} \\\\
&= \frac{\alpha}{\beta}
\end{align*}
\]
For the variance we use the
Computational Identity for Variance
\(\operatorname{Var}(X) = \mathbb{E}[X^2] - (\mathbb{E}[X] )^2\). Its hypothesis, a finite second moment, is verified by the
computation of \(\mathbb{E}[X^2]\) below.
The Expectation of a
Function of a Random Variable lets us compute \(\mathbb{E}[X^2]\) from the density
directly, with the same substitution as above:
\[
\begin{align*}
\mathbb{E}[X^2] &= \int_{0}^\infty x^2 f(x)dx \\\\
&= \frac{\beta^\alpha}{\Gamma(\alpha)} \int_{0}^\infty x^{\alpha+1}e^{-\beta x} dx \\\\
&= \frac{\beta^\alpha}{\Gamma(\alpha) \beta^{\alpha +2}} \int_{0}^\infty t^{\alpha+1}e^{-t} dt \\\\
&= \frac{\Gamma (\alpha +2)}{\Gamma(\alpha) \beta^2}.
\end{align*}
\]
Since \(\Gamma(\alpha +2) = (\alpha +1)\Gamma (\alpha +1) = (\alpha +1 )\alpha \Gamma(\alpha)\),
\[
\begin{align*}
\mathbb{E}[X^2] &= \frac{(\alpha +1 )\alpha \Gamma(\alpha)}{\Gamma(\alpha) \beta^2} \\\\
&= \frac{\alpha(\alpha +1 )}{\beta^2}
\end{align*}
\]
Substitute the results:
\[
\begin{align*}
\operatorname{Var}(X) &= \mathbb{E}[X^2] - (\mathbb{E}[X] )^2 \\\\
&= \frac{\alpha(\alpha +1 )}{\beta^2} - (\frac{\alpha}{\beta})^2 \\\\
&= \frac{\alpha}{\beta^2}
\end{align*}
\]
Many references use the shape-scale parametrization
\(X \sim \text{Gamma}(k, \theta)\) with \(k = \alpha\) and \(\theta = 1/\beta\),
in which the density becomes
\[
f(x) = \frac{1}{\Gamma(k)\,\theta^k}\,x^{k-1}\,e^{-x/\theta}, \quad x \gt 0.
\]
We adopt the rate parametrization \((\alpha, \beta)\) throughout this page. Both
conventions are equivalent and are used interchangeably in the literature.
The Exponential distribution is a special case of the gamma distribution:
\[
\text{Exp}(\lambda) \equiv \text{Gamma}(1, \beta = \lambda).
\]
The gamma density then becomes
\[
f(x) = \lambda e^{-\lambda x} \quad x \geq 0.
\]
The exponential distribution models the waiting time in a process in which events occur continuously and
independently at a constant average rate \(\lambda\). For example, if a machine fails at a constant rate
of once every 20 years on average, then the time to failure follows the exponential distribution with
\(\lambda = \frac{1}{20}\).
The mean and variance of an exponential distribution are given by
\[
\mathbb{E}[X] = \frac{1}{\lambda} \quad \operatorname{Var}(X) = \frac{1}{\lambda^2}.
\]
This implies the mean is equal to the standard deviation.
The exponential c.d.f. is given by
\[
F(x) = \lambda \int_{0} ^x e^{-\lambda u} du = 1 - e^{-\lambda x}, \quad x \geq 0,
\]
and \(F(x) = 0\) for \(x \lt 0\).
More generally, when \(\alpha = n\) is a positive integer, \(\text{Gamma}(n, \beta)\) is known as the
Erlang distribution, which is the distribution of the sum of \(n\) independent
\(\text{Exp}(\beta)\) random variables. Making this statement rigorous requires the language of
independent random variables and product measures, which we develop in
Limit Theorems & Product Measures.
Beta Function
Just as the gamma function underpins the gamma distribution, the beta function
provides the normalizing constant for a distribution on the unit interval. The Beta-Gamma
identity below ties the beta function to the gamma function, and that connection is what makes the
beta distribution analytically tractable.
Definition: Beta Function
The beta function is defined as:
\[
B(a, b) = \int_{0}^1 t^{a-1}(1-t)^{b-1}\, dt, \quad a \gt 0, b \gt 0.
\]
As with the gamma function, the integral converges for all complex \(a, b\) with
\(\operatorname{Re}(a) \gt 0\) and \(\operatorname{Re}(b) \gt 0\). We restrict to the
real positive case throughout this page.
Theorem: Beta-Gamma Identity
The beta function can be represented by the gamma function:
\[
B(a, b) = \frac{\Gamma (a) \Gamma (b)}{\Gamma (a+b)}
\]
Proof:
Consider the product of two distinct gamma functions with inputs \(a, b \gt 0\).
\[
\begin{align*}
\Gamma (a) \Gamma(b) &= \int_{0}^\infty u^{a-1}e^{-u}du \cdot \int_{0}^\infty v^{b-1}e^{-v}dv \\\\
&= \int_{0}^\infty \int_{0}^\infty u^{a-1} v^{b-1} e^{-(u+v)} dudv
\end{align*}
\]
We perform the change of variables \((u, v) \mapsto (s, t)\) given by
\[
s = u + v, \quad t = \frac{u}{u+v},
\]
with inverse \(u = st\), \(v = (1-t)s\). The region \((u, v) \in (0,\infty)^2\) maps bijectively
onto \((s, t) \in (0, \infty) \times (0, 1)\). The Jacobian is
\[
\begin{align*}
\frac{\partial(u, v)}{\partial(s, t)}
&= \det\!\begin{pmatrix} \partial u/\partial s & \partial u/\partial t \\ \partial v/\partial s & \partial v/\partial t \end{pmatrix} \\\\
&= \det\!\begin{pmatrix} t & s \\ 1-t & -s \end{pmatrix} \\\\
&= -st - s(1-t) = -s,
\end{align*}
\]
so \(du\,dv = |{-s}|\,ds\,dt = s\,ds\,dt\).
The substitution gives
\[
\begin{align*}
\Gamma (a) \Gamma(b) &= \int_{0}^\infty \int_{0}^1 (st)^{a-1} ((1-t)s)^{b-1} e^{-s} \, s\,dt\,ds \\\\
&= \int_{0}^\infty s^{(a+b)-1} e^{-s} ds \cdot \int_{0}^1 t^{a-1} (1-t)^{b-1} dt \\\\
&= \Gamma (a+b) \cdot B(a, b),
\end{align*}
\]
where the second line uses \(s \cdot s^{a-1} \cdot s^{b-1} = s^{(a+b)-1}\). Both the passage to the
double integral and this factorization are legitimate because the integrand is non-negative
throughout, so no integrability hypothesis is needed. Therefore,
\[
B(a, b) = \frac{\Gamma (a) \Gamma (b)}{\Gamma (a+b)}.
\]
Beta Distribution
The beta function leads directly to a distribution on the unit interval \([0, 1]\).
This makes the beta distribution the natural choice for modeling quantities that
represent probabilities or proportions. Because its density has the same functional form as a binomial
likelihood viewed as a function of the success probability, the beta distribution also plays a central
role as a conjugate prior in
Bayesian inference.
Definition: Beta Distribution
A random variable \(X\) has a beta distribution on the unit interval
with parameters \(a \gt 0\) and \(b \gt 0\) if its p.d.f. is:
\[
f(x) = \frac{1}{B(a, b)}x^{a-1}(1-x)^{b-1}, \quad x \in (0, 1).
\]
We write \(X \sim \text{Beta}(a, b)\). The mean and variance are:
\[
\mathbb{E}[X] = \frac{a}{a+b}, \quad \operatorname{Var}(X) = \frac{ab}{(a+b)^2(a+b+1)}.
\]
We use the open interval \((0, 1)\) because the density diverges at \(x = 0\) when \(a \lt 1\)
(since \(x^{a-1} \to \infty\)) and at \(x = 1\) when \(b \lt 1\). The boundary points \(\{0, 1\}\)
have Lebesgue measure zero and contribute nothing to integrals, so probabilities such as
\(P(0 \leq X \leq 1) = 1\) remain unchanged.
Mean and Variance of the Beta Distribution:
We can rewrite the density using the Beta-Gamma identity:
\[
f(x) = \frac{\Gamma (a+b)}{\Gamma (a) \Gamma (b)}x^{a-1} (1-x)^{b-1} \quad x \in (0, 1).
\]
By definition,
\[
\begin{align*}
\mathbb{E}[X] &= \int_{0}^1 x f(x)dx \\\\
&= \int_{0}^1 x \frac{\Gamma (a+b)}{\Gamma (a) \Gamma (b)}x^{a-1} (1-x)^{b-1} dx \\\\
&= \frac{\Gamma (a+b)}{\Gamma (a) \Gamma (b)} \int_{0}^1 x^{a} (1-x)^{b-1} dx \\\\
&= \frac{\Gamma (a+b)}{\Gamma (a) \Gamma (b)} B(a+1, b) \\\\
&= \frac{\Gamma (a+b)}{\Gamma (a) \Gamma (b)} \frac{\Gamma (a+1) \Gamma (b)}{\Gamma (a+b+1)}
\end{align*}
\]
Since \(\Gamma (a+1) = a \Gamma (a)\) and similarly, \(\Gamma (a+b+1) = (a+b)\Gamma (a+b)\),
\[
\begin{align*}
\mathbb{E}[X] &= \frac{\Gamma (a+b)}{\Gamma (a) \Gamma (b)} \frac{a\Gamma (a) \Gamma (b)}{(a+b)\Gamma (a+b)} \\\\
&= \frac{a}{a+b}
\end{align*}
\]
For the variance we use the
Computational Identity for Variance
\(\operatorname{Var}(X) = \mathbb{E}[X^2] - (\mathbb{E}[X] )^2\). Its hypothesis, a finite second moment, is verified by the
computation of \(\mathbb{E}[X^2]\) below.
The second moment follows from the same route. The
Expectation of a Function
of a Random Variable gives
\[
\begin{align*}
\mathbb{E}[X^2] &= \int_{0}^1 x^2 f(x)dx \\\\
&= \frac{\Gamma (a+b)}{\Gamma (a) \Gamma (b)} \int_{0}^1 x^{a+1} (1-x)^{b-1} dx \\\\
&= \frac{\Gamma (a+b)}{\Gamma (a) \Gamma (b)} B(a+2, b) \\\\
&= \frac{\Gamma (a+b)}{\Gamma (a) \Gamma (b)} \frac{\Gamma (a+2) \Gamma (b)}{\Gamma (a+b+2)}
\end{align*}
\]
Since \(\Gamma (a+2) = (a+1) \Gamma (a+1) = (a+1) a \Gamma (a)\), and similarly, \(\Gamma (a+b+2) = (a+b+1)(a+b)\Gamma (a+b)\),
\[
\begin{align*}
\mathbb{E}[X^2] &= \frac{\Gamma (a+b)}{\Gamma (a) \Gamma (b)} \frac{(a+1)a\Gamma (a) \Gamma (b)}{(a+b+1)(a+b)\Gamma (a+b)} \\\\
&= \frac{a(a+1)}{(a+b)(a+b+1)}.
\end{align*}
\]
Substitute the results:
\[
\begin{align*}
\operatorname{Var}(X) &= \mathbb{E}[X^2] - (\mathbb{E}[X] )^2 \\\\
&= \frac{a(a+1)}{(a+b)(a+b+1)} - (\frac{a}{a+b})^2 \\\\
&= \frac{a(a+1)(a+b) - a^2 (a+b+1) }{(a+b)^2(a+b+1)} \\\\
&= \frac{ab}{(a+b)^2(a+b+1)}
\end{align*}
\]
The beta family includes several important special cases. The simplest is the uniform
distribution on \([0, 1]\), which arises when both parameters equal one.
\[
U[0, 1] \equiv \text{Beta}(1, 1), \quad \text{so } f(x) = 1 \text{ on } [0, 1].
\]
The uniform law on a general interval \([c, d]\) is not itself a beta distribution, but rather the image
of \(U[0, 1]\) under the affine map \(x \mapsto c + (d-c)x\). Its p.d.f. is
\[
f(x) = \frac{1}{d - c} \quad x \in [c, d],
\]
and \(f(x) = 0\) outside \([c, d]\).
The uniform distribution represents situations in which every value in \([c, d]\) is equally
likely. For example, it is widely used for generating random
numbers in programming languages.
The mean and variance of the uniform distribution are given by
\[
\mathbb{E}[X] = \frac{c+d}{2} \quad \operatorname{Var}(X) = \frac{(d-c)^2}{12}
\]
and its c.d.f. is given by
\[
F(x) = \frac{x-c}{d-c} \quad x \in [c, d].
\]
For \(x \lt c\) we have \(F(x) = 0\), and \(F(x) = 1\) for \(x \gt d\).
Mean and Variance of the Uniform Distribution:
The definition of the mean gives
\[
\begin{align*}
\mathbb{E}[X] &= \int_{c}^d x \frac{1}{d-c}dx \\\\
&= \frac{1}{d-c} (\frac{d^2 - c^2}{2}) \\\\
&= \frac{c+d}{2}
\end{align*}
\]
The Expectation of a
Function of a Random Variable gives the second moment,
\[
\begin{align*}
\mathbb{E}[X^2] &= \int_{c}^d x^2 \frac{1}{d-c}dx \\\\
&= \frac{1}{d-c} (\frac{d^3 - c^3}{3}) \\\\
&= \frac{c^2 +cd + d^2}{3}
\end{align*}
\]
which is finite, so the
Computational Identity for Variance
gives
\[
\begin{align*}
\operatorname{Var}(X) &= \mathbb{E}[X^2] - (\mathbb{E}[X] )^2 \\\\
&= \frac{c^2 +cd + d^2}{3} - (\frac{c+d}{2})^2 \\\\
&= \frac{4(c^2 + cd +d^2)-3(c^2 + 2cd + d^2)}{12} \\\\
&= \frac{(d-c)^2}{12}.
\end{align*}
\]
Insight: Why Beta and Gamma? (Conjugate Priors)
In Bayesian inference, a conjugate prior is a prior distribution that, when
multiplied by the likelihood, results in a posterior distribution of the same functional form (family) as the prior.
- Beta distribution is the conjugate prior for Bernoulli and Binomial likelihoods.
This makes it the standard choice for modeling uncertainty about a probability (for example, click-through rates).
- Gamma distribution is the conjugate prior for the rate of a Poisson likelihood and for
the precision (inverse variance) of a Gaussian with known mean.
Most common distributions in machine learning belong to the
exponential family, which guarantees the existence of
a conjugate prior, so updating our beliefs reduces to simple algebraic additions to parameters. When
that prior's normalizing constant is known in closed form, as for the beta and gamma priors above,
no integration is required.
Interactive Visualization
The demo below has three tabs. The first two plot the gamma and beta densities as we move the
parameter sliders, and they draw this page's derivations onto the plot. The dashed line marks the
mean (\(\alpha/\beta\) for the gamma, \(a/(a+b)\) for the beta), the dotted line marks the mode,
and the shaded band spans mean \(\pm \sigma\), so we can watch \(\operatorname{Var}(X) = \alpha/\beta^2\)
shrink as \(\beta\) grows. A toggle overlays the CDF.
The special-case buttons jump to particular members of the two families: the exponential
(\(\alpha = 1\)) and Erlang cases discussed above, the chi-squared distribution (a gamma with
\(\beta = \tfrac{1}{2}\)), the uniform \(\text{Beta}(1,1)\), and the arcsine distribution
\(\text{Beta}(\tfrac{1}{2}, \tfrac{1}{2})\). For \(\alpha \lt 1\) in the gamma, or \(a \lt 1\) or
\(b \lt 1\) in the beta, the density genuinely diverges at the boundary, and the demo says so rather
than hiding it.
On the beta tab, the Bayesian update toggle brings the conjugate-prior discussion above
to life. We choose a prior \(\text{Beta}(a, b)\) and dial in \(k\) successes out of \(n\) trials, and the
posterior \(\text{Beta}(a + k,\, b + n - k)\) is drawn over the prior. No integration is required, only
parameter arithmetic, and the posterior mean is pulled from the prior mean toward the sample
proportion \(k/n\). We develop this machinery properly in Bayesian inference later in this section.
The third tab plots the gamma function itself: the smooth curve through the factorial points
\((n, (n-1)!)\), with \(\Gamma(\tfrac{1}{2}) = \sqrt{\pi}\) marked. Both results were proved at the top
of this page, and the plot makes them visible as geometry.
With the Gamma and Beta distributions established, we now turn to the single most important distribution in all of statistics:
the normal (Gaussian) distribution.