III - Probability & Statistics

The Mathematics of Uncertainty and Inference

Probability and Statistics are the mathematics of reasoning when the answer is not certain. Probability runs in one direction: given a model of the world, it predicts how the data should behave. Statistics runs the other way: given the data, it asks which model produced them. The second direction is harder, and it splits into two traditions that this section develops side by side. One treats the unknown quantity as fixed, asks which value the data fit best, and asks how often such a procedure would mislead us over repeated samples. The other treats the unknown quantity as uncertain in its own right, gives it a distribution, and lets the data reshape that distribution. Much of probabilistic machine learning, variational methods included, sits on the second view.

Within the Compass, this section is where the discrete and the continuous meet. On one side are the counting arguments of Section IV (Discrete Mathematics & Algorithms), where there are finitely many outcomes. On the other are the integrals and limits of Section II (Calculus to Optimization & Analysis), where quantities vary continuously. A sum over outcomes and an integral over a continuum look like different operations, but with the right notion of measurement they turn out to be the same one. The matrices of Section I (Linear Algebra to Algebraic Foundations) run throughout, recording how quantities vary together, and everything here eventually feeds the inference that drives Section V (Machine Learning).

The section's longest work is making all of this rigorous. Using the measure theory of Section II (Calculus to Optimization & Analysis), it gives exact meaning to a density, to the ratio of one probability to another, and to conditioning on information already seen. On that base it builds a second strand: chance that unfolds in continuous time. It starts from Brownian motion, the model of noise added up over time, develops a calculus that works along paths too rough for ordinary derivatives, and arrives at equations that describe how a whole probability distribution moves over time. The strand ends with a result that ties it to current practice. A single path of distributions solves the equation of motion of a whole family of random dynamics, from one with no noise at all to ones with as much noise as we like. That is the common ground beneath the generative models that turn noise into data.

Uncertainty changes what checking can mean. A conclusion drawn from finite data is never certain, so the useful question is how much a given amount of data can support, and with what confidence. Generative models raise a sharper version of the same question. They produce samples no one can check one by one, so the deeper test is whether the process as a whole reaches the right distribution. In its ideal form, with the exact field the model is meant to learn, that is a statement about how distributions move, and this section is where it is studied. Not every link in that chain is proved here yet. Where a step is taken on trust, the page says so plainly, and says which step it is.