V - Machine Learning

The Mathematics of Synthesis and Intelligence

Machine Learning is the practice of learning a rule from data or experience: fitting a model well enough that it says something useful about cases it has never seen. The field moves fast, and its tools come and go. The mathematics underneath them is far more stable than the libraries built on top. This section covers the core ideas, from regression and classification to neural networks and how they are trained, modern deep architectures, learning by trial and reward, and models that generate new data, while keeping the mathematical structure of each method in view. A particular architecture may be obsolete in a few years. Understanding why it works lets you adapt to whatever replaces it, and even build the replacement.

Within the Compass, this section is where everything else comes together. The algebra of Section I (Linear Algebra to Algebraic Foundations), the optimization and geometry of Section II (Calculus to Optimization & Analysis), the inference of Section III (Probability & Statistics), and the algorithms of Section IV (Discrete Mathematics & Algorithms) all meet here, on real problems. But this is not the end of the road. Each topic here is better seen as a vantage point than as a destination: a place from which the earlier tools can be watched working together, and from which the mathematics still ahead comes into view.

That is why the curriculum keeps circling back. An idea often appears here first as an application and is only later built properly from below, and then the earlier page reads differently. The clearest case is the principle that a good architecture respects the structure of its data. It can now be followed along two paths. One runs through representation theory to networks that respect rotations by construction. The other runs through simplicial complexes to networks on discrete domains, which pass messages along edges and triangles and can detect holes in the complex built from the data. Generative models make the same trip. Diffusion and flow models are met here first and are later underwritten by the continuous-time probability of Section III (Probability & Statistics), with the steps taken on trust named there. A third path describes learning systems in the language of category theory from Section IV (Discrete Mathematics & Algorithms), and it now has its first pages: a layer as a map with parameters, and backpropagation as composition running in reverse. They work in the simplest setting, built only from what the Compass already owns, and the larger picture is still forming.

Trust is hardest to earn here. A trained model is a large function found by optimization, and no one can read it line by line. Mathematics offers partial ways in. A guarantee can be built into the architecture, as with a symmetry that holds by design. A training objective can be stated exactly, so that it is clear what it rewards and what it ignores. And a method can be taken apart into the steps that rest on proof and the steps that rest on experiment. Security shows the same tension from another angle. A system of noisy linear equations looks like a regression problem, yet once its arithmetic wraps around a fixed modulus in high dimension, it is believed hard enough to build encryption on.