Intro to Machine Learning

The Law of Intelligence Learning Paradigms Basic Task Categories Standard Process of ML

The Law of Intelligence

Artificial intelligence (AI) has become one of the most influential technologies of our time, powering applications from search engines to self-driving cars. Before diving into its technical details, it is worth stepping back and asking: what is intelligence itself, and what does it mean to replicate it artificially?

Consider an analogy from physics. The laws of motion and aerodynamics govern both natural and human-made flight. We accept without hesitation that birds can fly, and we trust airplanes to carry us safely across continents because we understand the shared physical principles. Similarly, if we could uncover the fundamental laws of intelligence, we might someday build machines that "think" with the same confidence we have in machines that fly.

The catch is that we have no such laws yet. Absent them, "artificial intelligence" describes our aspirations as much as our achievements. The term covers everything from genuinely powerful pattern-matching systems to speculative claims that outrun the current science. What we can speak about precisely is a more modest and more rigorous subject: machine learning (ML), the mathematical and algorithmic study of how systems improve at specific tasks by processing data. ML draws on well-developed frameworks such as Bayesian decision theory, information theory, and statistical learning theory, and it is the focus of this section.

A widely cited formal definition comes from computer scientist Tom M. Mitchell:

A computer program is said to learn from experience \(E\), with respect to some class of tasks \(T\) and performance measure \(P\), if its performance at tasks in \(T\), as measured by \(P\), improves with experience \(E\). (T. Mitchell, Machine Learning, McGraw Hill, 1997)

A standard mathematical formalization of this definition is the optimization problem behind empirical risk minimization, in which the objective is the average loss over the dataset:

Definition: Learning

Given a hypothesis space \(\mathcal{H}\) of candidate models parameterized by \(\theta \in \Theta\), an objective function \(J(\theta; \mathcal{D})\) measuring prediction quality (typically based on a loss function), and a dataset \(\mathcal{D} = \{(x_i, y_i)\}_{i=1}^{N}\), the learner seeks optimal parameters \(\theta^*\) that minimize \(J\): \[ \theta^* = \arg\min_{\theta \in \Theta} J(\theta; \mathcal{D}). \]

Every component of this formulation draws on material developed earlier. The parameter space \(\Theta\) lives in a vector space, the optimization is driven by gradient-based methods, the objective is often built from a likelihood, as in maximum likelihood estimation, or from a posterior, and the algorithmic procedure has a well-defined computational complexity.

One branch of machine learning, called deep learning, utilizes large neural networks to perform complex tasks such as:

One application of deep learning is the development of Large Language Models (LLMs). These models are typically built from deep neural networks with the transformer architecture and are trained on massive text datasets. LLMs have demonstrated strong performance on language understanding, text generation, translation, and tasks that require multi-step inference.

Beyond digital text processing, ML also extends into Physical AI and autonomous systems. Real-world interaction is inherently stochastic. Sensor noise, unobserved physical properties, and environmental variability mean that even classical robotic systems incorporate probabilistic methods such as Kalman filtering, which already treats the system state as a probability distribution. Physical AI extends this probabilistic treatment from the state to the learned models themselves, so that a system can estimate its own epistemic uncertainty, the uncertainty that stems from limited data and knowledge rather than from sensor noise. Such estimates offer a principled basis for real-time safety and risk management, complementing rather than replacing classical safety mechanisms.

With this broad motivation in hand, we now turn to the distinctions that organize the field, beginning with the difference between supervised and unsupervised learning.

Learning Paradigms

Machine learning encompasses various approaches. The first distinction among them is the presence or absence of labeled data, which determines the mathematical formulation of the learning problem.

Definition: Supervised Learning

The model is trained on a labeled dataset \(\mathcal{D} = \{(x_i, y_i)\}_{i=1}^{N}\), where each input \(x_i \in \mathbb{R}^D\) is paired with an output label \(y_i \in \mathcal{Y}\). The goal is to learn a mapping \(f: \mathbb{R}^D \to \mathcal{Y}\) that generalizes well to unseen data. From a probabilistic perspective, we model the conditional distribution \(p(y \mid x; \theta)\) and either estimate \(\theta\) by maximum likelihood estimation (MLE) or place a posterior distribution over \(\theta\) by Bayesian inference. Common applications include image classification, spam detection, and medical diagnosis.

Definition: Unsupervised Learning

Here the dataset consists only of inputs \(\mathcal{D} = \{x_i\}_{i=1}^{N}\) with no target labels. The model attempts to uncover intrinsic structures or underlying patterns in two primary ways:

  • Structural Analysis: Identifying discrete groupings (clustering) or finding low-dimensional latent representations (dimensionality reduction) that capture the data's dominant variance.
  • Generative Modeling: Explicitly modeling the data-generating distribution \(p(x; \theta)\). This allows the system not only to synthesize new samples but also to assess how plausible a new input is under the learned distribution, a signal used for out-of-distribution detection in Physical AI.
Applications include customer segmentation, anomaly detection, and synthetic data generation.

While supervised and unsupervised learning form the mathematical foundation of data modeling, a third major paradigm exists where data is acquired through action:

Definition: Reinforcement Learning (RL)

Unlike learning from a static dataset, an agent interacts with a dynamic environment to maximize the expected cumulative reward it receives as scalar signals. This is a fundamentally different formulation where the model must balance exploration of the unknown with exploitation of known rewards. In Physical AI, RL can incorporate uncertainty-aware constraints, allowing robots to recognize "out-of-distribution" states and trigger safety aborts to prevent hardware damage. (See: Reinforcement Learning)

Modern machine learning often bridges these three major paradigms (supervised, unsupervised, and reinforcement learning) through hybrid approaches:

Basic Task Categories

The learning paradigms above (supervised, unsupervised, reinforcement) describe how a model learns. Orthogonal to this is the question of what the model predicts. Machine learning tasks are broadly categorized based on the nature of the output space:

From a unified probabilistic perspective, these categories can be viewed as different ways of modeling the data-generating process. Whether we are predicting a continuous value (Regression), assigning a label (Classification), or discovering hidden manifolds (Dimensionality Reduction), we are essentially seeking the underlying mathematical structure that governs the observed data.

These categories are not mutually exclusive. In practice, they frequently overlap. For instance, a VAE simultaneously performs non-linear dimensionality reduction and generative modeling. Modern architectures often integrate multiple paradigms to handle complex, high-dimensional data.

Each of these categories is explored in dedicated pages within this section, and all of them share the workflow outlined next.

Standard Process of ML

Regardless of whether we are performing regression, classification, or clustering, every machine learning project follows a common pipeline. Understanding this pipeline matters because each step introduces its own mathematical and engineering considerations. These considerations range from the statistical properties of the data to the convergence guarantees of the optimizer and the deployment realities of distribution shift.

The Machine Learning Pipeline

  1. Problem Definition.
    Articulate the problem precisely and decide whether machine learning is the right tool. This means specifying the input space \(\mathcal{X}\), the output space \(\mathcal{Y}\), and the performance criterion. ML failures often trace back to vague problem definitions. A model can only optimize what we measure, so the choice of metric quietly determines what behavior the system will learn. This is a manifestation of Goodhart's law, commonly paraphrased as the statement that when a measure becomes a target, it ceases to be a good measure.
  2. Data Collection and Curation.
    Gather data that is relevant, diverse, and representative of the deployment distribution. For most of the 2010s the dominant strategy was simply to collect more data on the assumption that bigger was better. Since around 2023, the emphasis has shifted. Data quality has come to rival data quantity as the binding constraint on model performance. Carefully curated, deduplicated, and filtered datasets can match or outperform much larger noisy ones, and licensing, provenance, and contamination have become first-class engineering concerns. The data must be sufficiently rich to capture the underlying distribution \(p(x, y)\) we wish to model, yet no richer than what we can verify and trust.
  3. Data Preprocessing.
    Clean and prepare the data by handling missing values, encoding categorical variables, and normalizing features. Feature scaling, for instance, typically speeds up gradient descent because it tends to improve the condition number of the Hessian of the loss (for least squares with design matrix \(X\), the matrix \(X^\top X\)). For text and image data, preprocessing also includes tokenization, augmentation, and standardization choices that often matter as much as the model architecture.
  4. Data Splitting and Contamination Control.
    Divide the dataset into training, validation, and test sets. This separation is essential for estimating generalization performance. Cross-validation techniques refine it by rotating the validation role across several folds of the training data. In the foundation-model era, an additional concern has emerged: contamination. The test set may overlap with the corpus on which a pretrained model was already trained, and such overlap inflates evaluation scores without genuine generalization. Guarding against contamination requires explicit overlap checks and the use of held-out or freshly constructed evaluation sets.
  5. Model Selection.
    Choose a hypothesis class \(\mathcal{H}\) appropriate to the problem. In classical ML this means picking a linear, tree-based, or kernel-based model family or a neural-network architecture, and confronting the fundamental bias-variance tradeoff. More expressive models fit training data better (lower bias) but may generalize poorly (higher variance). When pretrained models are available, model selection often reduces to choosing among pretrained foundation models (a vision transformer, a language model, a diffusion backbone) and a transfer strategy: full fine-tuning, parameter-efficient methods like LoRA or adapters, or zero-shot prompting.
  6. Training or Fine-Tuning.
    Optimize the model parameters \(\theta\) by minimizing an empirical loss on the training set. For neural networks, this involves backpropagation, an efficient application of the chain rule, combined with stochastic gradient descent or its variants (Adam, the related AdamW, and alternatives such as Lion). When starting from a pretrained model, training is usually called fine-tuning and uses smaller learning rates, fewer steps, and often updates only a fraction of the original parameters. Specialized training paradigms extend this basic gradient-descent loop in domain-specific ways. They include supervised fine-tuning (SFT), RLHF, direct preference optimization (DPO), and knowledge distillation.
  7. Evaluation.
    Assess the model's performance on the validation set using metrics appropriate to the task. Classification uses accuracy, precision, recall, and F1, regression uses mean squared error, and generative tasks use BLEU, ROUGE, and pairwise human preference. Beyond task accuracy, evaluation also includes behavioral and safety assessments: robustness to adversarial inputs, calibration of uncertainty estimates, fairness across subgroups, and red-team probing for harmful outputs. A model that scores high on task accuracy but fails safety evaluation is not yet ready for deployment.
  8. Hyperparameter Tuning.
    Optimize the hyperparameters that govern the learning process, such as the regularization strength \(\lambda\), learning rate, batch size, network depth, and dropout rate. Standard search methods are grid search, random search, and Bayesian optimization. Hyperparameter tuning sits in an outer loop around training and is computationally expensive. Principled methods such as successive halving and population-based training can reduce the cost substantially.
  9. Testing.
    Evaluate the final model on the held-out test set to obtain an unbiased estimate of generalization performance. The test set must remain untouched throughout hyperparameter tuning and model selection. Every peek at the test set during development is a small leak that biases the final estimate. Statistical rigor at this stage is what separates a trustworthy result from an over-tuned one.
  10. Deployment and Monitoring.
    Integrate the model into a production environment and monitor its behavior continuously. In practice, the data distribution drifts over time (distribution shift), requiring periodic retraining. For safety-critical systems, this phase relies on out-of-distribution (OOD) detection. If a model's estimated epistemic uncertainty exceeds a learned threshold, the system recognizes it is operating outside its confident regime and can trigger fallback policies or immediate aborts.
  11. Inference-Time Compute.
    This addition to the pipeline was popularized around 2024 by reasoning-focused language models. Rather than spending compute only at training time, the system spends additional compute at inference. It generates multiple candidate solutions, performs internal search or reflection, or extends its chain of reasoning (chain-of-thought), and then selects the best output. The approach recasts model performance as a function of three axes (training data, model parameters, and inference compute) rather than two.

This pipeline reflects the workflow of a classical ML project built from scratch. In foundation-model practice, the expensive early stages of large-scale data collection and pretraining are typically replaced by transfer from a pretrained model. The mathematical principles examined throughout this section apply across both regimes. They include convergence, generalization, regularization, and optimization geometry.

The ML pipeline brings together the material of every earlier section. Data representation relies on Linear Algebra to Algebraic Foundations. Each data point is a vector, each dataset is a matrix, and transformations like PCA are eigenvalue problems. Optimization is governed by Calculus to Optimization & Analysis. Gradient descent, convexity, and convergence rates determine whether training succeeds. Generalization is a question of Probability & Statistics, which spans everything from the bias-variance tradeoff to the law of large numbers that motivates empirical risk minimization. Finally, the computational feasibility of each algorithm depends on Discrete Mathematics & Algorithms. In the pages ahead, we explore each of these connections in depth.

Looking further ahead, these foundations converge in a recurring viewpoint: Geometric Deep Learning (GDL). GDL unifies Convolutional Neural Networks (translation symmetry), Graph Neural Networks (permutation symmetry on nodes), and equivariant networks for 3D data (rotation and translation symmetry under Lie groups such as \(SO(3)\) and \(SE(3)\)). The unifying principle is that the architecture of a neural network should respect the symmetries of the data it operates on. Lie groups, smooth manifolds, group representations, and the graph Laplacian supply the mathematical machinery this principle requires, and the earlier sections develop them in preparation for the GDL pages later in this section. A second viewpoint, Categorical Deep Learning, asks which compositional structure a network's layers must preserve.