Null Hypothesis Significance Test
Once we have a statistical model (or hypothesis), we need to assess whether it is plausible given our data
\(\mathcal{D}\). On the MLE page, we developed
maximum likelihood estimation as a method for fitting parameters to data. MLE answers the question "what is the
best estimate?" but it does not answer a complementary question: "is the effect we observe real, or could it be due to chance?"
Hypothesis testing provides a principled framework for making such decisions under uncertainty.
Although Bayesian inference can replace many frequentist techniques and is
especially popular in modern machine learning, frequentist methods remain valuable. They are often simpler to compute, more
standardized, and provide complementary insights. Here, we introduce the
null hypothesis significance test (NHST).
Definition: Hypotheses
A hypothesis test involves two competing statements:
- Null Hypothesis \(H_0\):
The default assumption (for example, "the treatment has no effect").
- Alternative Hypothesis \(H_1\):
The claim we wish to support (for example, "the treatment has a positive effect").
Hypothesis testing can be viewed as a binary classification problem. Given data \(\mathcal{D}\), we decide
between \(H_0\) and \(H_1\).
Our reasoning follows the logic of proof by contradiction. If the observed data would be extremely unlikely under \(H_0\), we
reject the null hypothesis in favor of \(H_1\). However, rejecting \(H_0\) does not prove \(H_1\) is true, and
failing to reject \(H_0\) does not prove \(H_0\) is true. It only means the evidence is insufficient. Because
our conclusion can be wrong, we must account for two types of error:
Definition: Type I and Type II Errors
- Type I error (false positive):
Rejecting \(H_0\) when it is actually true.
- Type II error (false negative):
Failing to reject \(H_0\) when \(H_1\) is actually true.
The Type I error rate \(\alpha\) is called the significance level of the test.
It represents the probability of mistakenly rejecting \(H_0\) when it is true, and is typically
set to 0.05 or 0.01 in practice.
To decide whether to reject \(H_0\), we compute a test statistic \(T(\mathcal{D})\), a function of the data that
summarizes the evidence against \(H_0\). We then compare it to the distribution of \(T(\tilde{\mathcal{D}})\) under hypothetical
datasets \(\tilde{\mathcal{D}}\) drawn assuming \(H_0\) is true.
Definition: p-Value
The p-value is the probability, under \(H_0\), of obtaining a test statistic at least as extreme as
the one observed:
\[
p = P\!\left(T(\tilde{\mathcal{D}}) \geq T(\mathcal{D}) \middle| \tilde{\mathcal{D}} \sim H_0\right).
\]
The form above is the upper-tailed p-value, appropriate when \(H_1\) predicts larger values of \(T\).
The lower-tailed variant replaces \(\geq\) with \(\leq\), and the two-sided version uses
\(P(|T(\tilde{\mathcal{D}})| \geq |T(\mathcal{D})|)\) when the test statistic is symmetric under \(H_0\).
The choice depends on the alternative hypothesis.
If \(p \lt \alpha\), the observed result is deemed unlikely under \(H_0\), and we reject the null hypothesis.
It is essential to interpret p-values correctly. A p-value of 0.05 does not mean that \(H_1\) is true with
probability 0.95. The p-value measures the compatibility of the data with \(H_0\), not the probability that \(H_0\) is
true or false.
While NHST provides a systematic framework, it has well-known limitations:
- Statistical significance does not imply practical significance.
Even if \(H_0\) is rejected, the effect size may be too small to matter.
- p-values depend on sample size. With very large datasets,
even tiny, practically irrelevant differences may yield small p-values.
- Bayesian approaches offer an
alternative by directly computing the probability of hypotheses given the data,
rather than relying on fixed significance thresholds.
To make the NHST framework concrete, we now work through a specific example
using the t-test.
t-Tests
The population standard deviation \(\sigma\) is typically unknown in practice. We therefore replace it with the sample standard
deviation \(s\) and use the Student's t-distribution
instead of the normal distribution. The resulting procedure is called a t-test.
Distributional assumption. If \(n \geq 2\) and \(X_1, \ldots, X_n\) are i.i.d. \(\mathcal{N}(\mu, \sigma^2)\),
then the standardized sample mean \((\bar{X} - \mu)/(s/\sqrt{n})\) follows a \(t_{n-1}\) distribution exactly. To see this,
note that \(Z := \sqrt{n}(\bar{X} - \mu)/\sigma \sim \mathcal{N}(0,1)\) and \(Q := (n-1)s^2/\sigma^2 \sim \chi^2_{n-1}\), and
that \(Z\) and \(Q\) are independent. We take these three facts for granted here. All of them follow from the orthogonal
decomposition of a Gaussian sample into its mean and its residuals (Cochran's theorem), which we do not develop. The ratio
\(Z/\sqrt{Q/(n-1)} = (\bar{X} - \mu)/(s/\sqrt{n})\) then matches the
standard t-distribution definition. For
non-normal populations, the result holds only asymptotically via the Central Limit Theorem.
Example:
Suppose we are analyzing the test scores of students in a school. Historically, the average test score is 70. A researcher
believes that a new teaching method has improved scores. To test this, we collect a sample of 30 students' scores after
using the new method.
- \(H_0\): The new method has no effect, meaning the true mean is still 70.
- \(H_1\): The new method increases the average score, meaning the mean is greater than 70.
We collected a sample of \(n = 30\) students with the following observed statistics:
- Sample mean: \(\bar{x} = 75.20\).
- Sample standard deviation: \(s = 9.00\).
where
\[
s = \sqrt{\frac{1}{n-1}\sum_{i=1}^n (x_i - \bar{x})^2}.
\]
Since the population standard deviation \(\sigma\) is unknown, we use a one-sample t-test. With
\(\mu_0 = 70\) the mean under \(H_0\), the test statistic is given by:
\[
t = \frac{\bar{x} - \mu_0}{s / \sqrt{n}} = \frac{75.20 - 70}{9.00 / \sqrt{30}} \approx 3.16.
\]
Under \(H_0\), and assuming the scores are i.i.d. normal, the test statistic follows a
Student's t-distribution with \(n - 1 = 29\) degrees
of freedom, and the p-value is the upper-tail probability beyond the computed value.
Then we have \(p \approx 0.0018\) by numerical computation (the area in the upper tail of the \(t_{29}\) distribution). Set
the significance level \(\alpha = 0.05\). Since \(p \lt 0.05\), we reject \(H_0\). Therefore, there is strong statistical
evidence that the new teaching method increases students' test scores.
We can only say that the data we observed (test scores) are very unlikely under the assumption that the true mean is still
70. It does not mean that:
- the new teaching method definitely increases test scores.
- the probability that \(H_0\) is true is 0.0018.
- the effect is practically significant.
Confidence Intervals vs Credible Intervals
Hypothesis testing gives a binary answer, reject or fail to reject. It says nothing about how close our estimate might
be to the true parameter. In practice, we often want a range of plausible values for \(\theta\). Both frequentist and Bayesian
statistics provide such intervals, but they differ fundamentally in interpretation.
Definition: Confidence Interval (Frequentist)
For a level \(\alpha \in (0,1)\), a \(100(1 - \alpha)\%\) confidence interval for \(\theta\) is a
random interval \([L(\mathcal{D}),\, U(\mathcal{D})]\) satisfying
\[
P_\theta\!\left(\theta \in [L(\mathcal{D}), U(\mathcal{D})]\right) \geq 1 - \alpha
\quad \text{for all } \theta \in \Theta,
\]
where the probability is taken with respect to the sampling distribution of \(\mathcal{D}\) under parameter \(\theta\). In
frequency terms, if we were to repeat the experiment many times and construct such an interval each time, then in the long
run at least \(100(1 - \alpha)\%\) of those intervals would contain the true parameter \(\theta\).
One subtlety is critical. A 95% CI does not mean "there is a 95% probability that \(\theta\) lies in this
interval." In frequentist statistics, \(\theta\) is a fixed constant. It either lies in the interval or it does
not. The probability statement refers to the procedure, not to any single interval.
In Bayesian statistics, the interpretation is reversed. The data are fixed (since
they are observed) and the parameter is treated as a random variable with a posterior distribution.
Definition: Credible Interval (Bayesian)
For a level \(\alpha \in (0,1)\), a \(100(1 - \alpha)\%\) credible interval for \(\theta\) is an
interval \(C_\alpha(\mathcal{D}) = [L, U]\) satisfying
\[
P\!\left(\theta \in [L, U] \middle| \mathcal{D}\right) = 1 - \alpha,
\]
where the probability is taken with respect to the posterior distribution \(p(\theta \mid \mathcal{D})\). For continuous
posteriors, this condition does not uniquely determine the interval. Infinitely many such intervals exist. A common choice is
the equal-tailed (or central) credible interval given by the \(\alpha/2\) and
\(1 - \alpha/2\) posterior quantiles.
Unlike a confidence interval, a credible interval directly states that, given the data and the prior, the parameter falls within
the interval with the stated probability.
Example:
Suppose we toss a coin \(n = 100\) times and observe 60 heads. Now we want to estimate the probability of getting heads.
First, we try the frequentist approach. The point estimate for the probability of heads \(p\) is
\(\hat{p} = \frac{60}{100} = 0.6\), and the standard error (SE) for a proportion is given by:
\[
\begin{align*}
\operatorname{SE} &= \sqrt{\frac{\hat{p}(1-\hat{p})}{n}} \\\\
&= \sqrt{\frac{0.6 \times 0.4}{100}} \\\\
&\approx 0.049.
\end{align*}
\]
For a 95% CI using normal approximation, the critical value is \(z_{0.025} \approx 1.96\). Then the CI
is given by
\[
\begin{align*}
\operatorname{CI} &= [\hat{p}-z_{0.025} \times \operatorname{SE}, \quad \hat{p}+z_{0.025}\times \operatorname{SE}] \\\\
&\approx [0.504, 0.696].
\end{align*}
\]
If we repeated the experiment (tossing the coin 100 times) many times and computed a 95% CI each time, about 95% of the
resulting intervals would cover the true \(p\).
Note. A z-score is any value that has been standardized to represent the number of standard
deviations away from the mean. The critical value is a specific z-score used as a threshold in hypothesis
testing or confidence interval calculations. In our case, \(z_{0.025} \approx 1.96\) is the critical value that separates the
central 95% of the distribution from the outer 5% (2.5% in each tail).
In the Bayesian approach, we assume a uniform prior for the probability \(p\), which is equivalent to a
Beta distribution:
\[
p \sim \operatorname{Beta}(1, 1).
\]
With 60 heads and 40 tails, the likelihood is given by a binomial distribution. In the Bayesian framework, the
posterior distribution is:
\[
p \sim \operatorname{Beta}(1+60, 1+40) = \operatorname{Beta}(61, 41).
\]
Note. The Beta distribution is a
conjugate prior for the binomial likelihood.
Because the likelihood for coin tosses is binomial, a Beta prior yields a Beta posterior.
A 95% credible interval (CrI) is typically obtained by finding the 2.5th and 97.5th percentiles of the posterior
distribution. These percentiles can be computed using the inverse cumulative distribution function for the Beta
distribution. For example,
\[
\begin{align*}
\operatorname{CrI} &= [\operatorname{invBeta}(0.025, 61, 41), \quad \operatorname{invBeta}(0.975, 61, 41)] \\\\
&\approx [0.50, 0.69].
\end{align*}
\]
Given the observed data and the chosen prior, there is a 95% probability that \(p\) falls between 0.50 and 0.69. This
interval directly reflects our uncertainty about \(p\) after seeing the data.
The credible interval is often considered more intuitive, because it directly answers the question "What is the probability that
the parameter falls within this interval, given the data and our prior beliefs?" Frequentist methods, on the other hand, provide
guarantees on long-run performance without the need for a prior, which can be an advantage in settings where subjective beliefs
are hard to justify.
Bootstrap
The confidence intervals derived above rely on distributional assumptions. The usual assumption is that the sampling
distribution is approximately normal by the Central Limit Theorem. These analytical approximations may be unreliable when the
sample size is small, when the estimator is a complex function of the data, or when the underlying distribution is far from
normal. The bootstrap method provides a powerful non-parametric alternative, although its own
justification is also asymptotic. This procedure is motivated by the
Glivenko-Cantelli theorem, which
guarantees that the empirical cdf converges uniformly to the true cdf almost surely as the sample size grows.
The idea is conceptually simple. We treat the observed sample as a proxy for the population and resample from it with replacement
to generate many bootstrap samples. By computing the estimator on each bootstrap sample, we obtain an empirical
approximation of the estimator's sampling distribution. This approach provides a flexible, non-parametric way to assess
uncertainty and construct confidence intervals.
Example:
We have a small dataset of 10 people's heights (in cm):
\[
x = [160,165,170,175,180,185,190,195,200,205].
\]
The sample mean is \(\hat{\mu} = 182.5\). Since our sample size is small, we might not be able to assume the
sampling distribution of the mean is normal. Instead of assuming normality, we use bootstrap resampling to
approximate it empirically.
Here, we randomly draw 10 values from the original sample with replacement. Some values may be repeated, and others may be
missing in a given resample. For example, we might obtain:
\[
\begin{align*}
&x_1 = [165,175,175,190,185,160,200,195,180,185], \\\\
&\hat{\mu} = 181.0.
\end{align*}
\]
Then we repeat this process many times (for example, 10,000 times). Each time, we create a new resampled dataset, compute its
mean, and store it. Once we have 10,000 bootstrap sample means, we can use them to estimate the confidence interval.
Connections to Machine Learning
Hypothesis testing and confidence intervals appear throughout machine learning practice. When comparing two models,
paired t-tests or bootstrap confidence intervals on the performance difference help
determine whether one model is genuinely better or the gap is within sampling noise. The
confusion matrix and ROC analysis extend hypothesis testing ideas to
classifier evaluation. The bootstrap is particularly valuable in deep learning, where the sampling distribution of complex
metrics (for example, BLEU scores or F1) has no closed-form expression. In
Bayesian inference, credible intervals provide a principled alternative that
directly quantifies parameter uncertainty.
With the tools for estimation (MLE) and inference (hypothesis testing, confidence intervals) established, we can now
apply them to one of the most important models in statistics. On the linear regression page,
we develop linear regression as the intersection of least-squares optimization and maximum likelihood estimation
under Gaussian noise.