Intro to Geometric Deep Learning

Introduction Geometric Deep Learning: A Unifying Framework Graph Neural Networks: From Spectral to Spatial Beyond Permutation: Continuous Group Equivariance Interactive Demo

Introduction

On the page on deep neural networks, we surveyed the major architectural innovations of modern deep learning, namely convolutional networks, residual connections, attention, and the Transformer. That page closed its survey with an observation that runs through all of them. Each architecture is built around a symmetry of its input domain. Convolutional networks commute with translations of the input image (exactly so for stride one on an unbounded or periodic grid), and self-attention commutes with permutations of the input tokens. The shared principle is that the architecture encodes the symmetry, and architectures that respect their domain's symmetry tend to generalize from far less data than architectures that do not.

That observation is the starting point of Geometric Deep Learning (GDL): a unifying framework which views CNNs, Transformers, graph neural networks, and equivariant architectures for 3D data as instances of a single design principle, namely equivariance under the symmetry group of the input domain.

The word "geometric" here is meant in its deep mathematical sense. A geometry is most naturally understood as the study of what remains invariant under a chosen group of transformations. Different choices of group give different geometries, such as a grid under translation, a set under permutation, a graph under node relabelling, or a smooth manifold under rotation. An architecture is "geometric" when it respects the symmetries of its domain.

The present page introduces this framework, then develops its most directly accessible instance, the graph neural network (GNN), which generalizes message passing from regular grids and complete graphs to arbitrary graphs. We then survey how the framework extends beyond permutation symmetry to the continuous symmetries of three-dimensional space. This is the territory of equivariant neural networks for molecular structure, point clouds, and rigid-body manipulation.

The mathematical machinery for the continuous-symmetry case is substantial. It draws on Lie group theory, on representation theory of compact groups, and on the differential geometry of smooth manifolds. The present page stays at the level of design principle, showing why those foundations are needed and what kind of architecture they support.

Geometric Deep Learning: A Unifying Framework

The GDL framework, articulated by Bronstein, Bruna, Cohen, and Veličković, organizes deep architectures by the symmetry group of their input domain and the way each layer is required to commute with that group's action. Concretely, let \(\mathcal{X}\) be a domain with symmetry group \(G\) acting on it, and let \(G\) act on \(\mathcal{Y}\) as well. A layer \(f : \mathcal{X} \to \mathcal{Y}\) is said to be \(G\)-equivariant if \[ f(g \cdot x) = g \cdot f(x) \quad \text{for every } g \in G,\ x \in \mathcal{X}, \] and \(G\)-invariant if the right-hand side is replaced by \(f(x)\) itself.

Equivariance preserves symmetry information through the network, whereas invariance discards it. In practice, deep architectures stack equivariant layers and finish with an invariant pooling step, so that the output is invariant to the chosen symmetry while intermediate representations remain sensitive to it.

The major architectures we have already surveyed all fit this pattern. The table below organizes them by the domain they act on and the symmetry group encoded in the architecture.

Remark: Architectures Organized by Symmetry

Architecture Domain Symmetry group Equivariance type
MLP \(\mathbb{R}^d\) trivial none
CNN grid (image / volume) translation \(\mathbb{Z}^d\) translation-equivariant
Transformer sequence / set symmetric group \(S_n\) permutation-equivariant
GNN graph (features \(\mathbf{X}\) + adjacency \(\mathbf{A}\)) symmetric group \(S_n\) joint permutation of features and topology: \((\mathbf{X}, \mathbf{A}) \mapsto (P\mathbf{X}, P\mathbf{A}P^\top)\)
Steerable / \(SE(3)\)-equivariant NN 3D point cloud, molecular structure \(SO(3)\), \(SE(3)\) continuous group equivariance

The table makes the unification explicit. The MLP, alone among the architectures listed, encodes no symmetry and is the baseline against which all the others are GDL instances. CNNs, Transformers, GNNs, and \(SE(3)\)-equivariant networks all share the same template. The template identifies the symmetry group of the data domain and constrains each layer to commute with its action. The differences between the architectures are which group, and how the group acts.

The Transformer's place in this table deserves a closer look, because it is in some respects the cleanest example of GDL design philosophy already at industrial scale. Self-attention computes \[ \mathrm{Attn}(\mathbf{Q}, \mathbf{K}, \mathbf{V}) = \operatorname{softmax}\!\left( \tfrac{\mathbf{Q}\mathbf{K}^\top}{\sqrt{d_k}}\right)\mathbf{V}, \] and, because \(\mathbf{Q}\), \(\mathbf{K}\), and \(\mathbf{V}\) are all read off from the same matrix of tokens, permuting the rows of that matrix permutes the rows of the output identically. The layer therefore commutes with the symmetric group \(S_n\) acting on the token axis.

A Transformer without positional encoding is therefore a permutation-equivariant set processor. Positional encoding deliberately breaks this symmetry to inject sequence order back into the architecture. From the GDL viewpoint, a Transformer-based large language model is built on a fundamentally GDL-shaped backbone. Equivariance was an architectural choice, not an afterthought, and it is thought to be part of why the architecture scales.

Graph neural networks generalize this picture in a direction the Transformer does not cover. A Transformer treats every pair of tokens as connected, so its connectivity is fixed by the architecture rather than given by the data. Most real data carries structure of its own. Only adjacent atoms in a molecule interact directly, only cited pairs in a citation network are linked, and connectivity in a road network is constrained by physical layout. A GNN takes the connectivity as part of its input. That input is a pair \((\mathbf{X}, \mathbf{A})\) of node features and adjacency matrix, and the GNN is designed to be equivariant under the joint action of \(S_n\) on both: \[ f(P\mathbf{X},\, P\mathbf{A}P^\top) = P\,f(\mathbf{X}, \mathbf{A}) \quad \text{for every } P \in S_n. \]

Both the Transformer (without positional encoding) and the GNN are \(S_n\)-equivariant, and the difference lies in what \(P\) acts on. The Transformer permutes a feature matrix, whereas the GNN permutes a feature-adjacency pair, with the adjacency matrix conjugated so that the graph structure is permuted consistently with the node relabelling. If we restrict attention to one fixed graph, the permutations preserving \(\mathbf{A}\) are exactly the graph isomorphisms from that graph onto itself, and they form its automorphism group. Those are the only relabellings that permute the features while leaving the graph itself fixed. The architecture, by contrast, is built to handle any graph, which is why a single trained GNN generalizes to graphs unseen at training time.

Graph Neural Networks: From Spectral to Spatial

Two essentially different approaches to constructing graph-based architectures emerged historically. The spectral approach, which starts from signal processing on graphs, defines convolution on a graph through the Graph Fourier Transform, parameterizes filters in the spectral domain, and recovers tractable architectures by polynomial approximation. The spatial approach defines a layer directly as local message passing between adjacent vertices. The two viewpoints are connected, since a widely used spatial architecture, the GCN, is mathematically a first-order Chebyshev approximation of a spectral filter. We treat them in turn.

The Spectral Side: Filters in the Frequency Domain

The graph Laplacian \(\mathbf{L} = \mathbf{D} - \mathbf{A}\) is real symmetric and positive semi-definite, so it admits an eigendecomposition \(\mathbf{L} = \mathbf{V}\boldsymbol{\Lambda}\mathbf{V}^\top\) with real non-negative eigenvalues. The eigenvectors play the role of "graph frequencies," and the Graph Fourier Transform \(\hat{\mathbf{f}} = \mathbf{V}^\top \mathbf{f}\) decomposes a graph signal into its frequency components. A spectral graph convolution applies a filter \(g(\boldsymbol{\Lambda})\) to the spectrum: \[ \mathbf{f}_{\mathrm{out}} = \mathbf{V}\, g(\boldsymbol{\Lambda})\, \mathbf{V}^\top \mathbf{f}_{\mathrm{in}}. \]

Computing this directly requires the full eigendecomposition of \(\mathbf{L}\), which is prohibitive for large graphs. The standard remedy approximates the filter by a low-degree polynomial in \(\mathbf{L}\), typically built from Chebyshev polynomials of the rescaled Laplacian. The resulting architectures are efficient and localized, since a degree-\(K\) polynomial filter respects \(K\)-hop neighborhoods. The first-order case of this construction is the Graph Convolutional Network (GCN), the widely used architecture named above.

Our graph Laplacian page gives the full development: the Chebyshev approximation, the rescaling \(\tilde{\mathbf{L}} = (2/\lambda_{\max})\mathcal{L}_{\mathrm{sym}} - \mathbf{I}\) of the symmetric normalized Laplacian \(\mathcal{L}_{\mathrm{sym}} = \mathbf{D}^{-1/2}\mathbf{L}\mathbf{D}^{-1/2}\), which is defined when no vertex is isolated and whose largest eigenvalue is \(\lambda_{\max}\), and the GCN derivation. Here we recall only what we need to connect the spectral and spatial perspectives.

The Spatial Side: Message Passing

The spatial perspective begins from a different place. Rather than diagonalizing the Laplacian, we ask directly: what local computation respects graph structure? The answer takes the form of a message passing layer. At each layer \(k\), every vertex \(v\) updates its hidden state \(\mathbf{h}_v^{(k)}\) by aggregating messages from its neighbors and combining them with its own current state:

Definition: Message Passing Layer

A message passing layer updates vertex features by \[ \mathbf{h}_v^{(k+1)} = \mathrm{UPDATE}\!\left(\mathbf{h}_v^{(k)},\ \mathrm{AGGREGATE}\big(\{\!\!\{ \mathbf{h}_u^{(k)} : u \in \mathcal{N}(v) \}\!\!\}\big)\right), \] where \(\mathcal{N}(v)\) is the neighborhood of \(v\), \(\{\!\!\{\cdot\}\!\!\}\) denotes a multiset, \(\mathrm{AGGREGATE}\) is a permutation-invariant function over the neighborhood (sum, mean, max, attention-weighted sum), and \(\mathrm{UPDATE}\) is a learnable transformation (typically a small MLP).

The choice of \(\mathrm{AGGREGATE}\) and \(\mathrm{UPDATE}\) determines the architecture. The GCN uses a normalized-degree weighted sum derived from \(\mathcal{L}_{\mathrm{sym}}\). GraphSAGE uses sampled neighborhoods with a learnable pooling function, which addresses scalability on very large graphs. The Graph Attention Network (GAT) computes neighborhood weights by an attention mechanism and so recovers much of the flexibility of self-attention while restricting interactions to graph edges.

The Graph Isomorphism Network (GIN) replaces the aggregation with a sum followed by an MLP, a choice motivated by its connection to the 1-Weisfeiler-Lehman (1-WL) graph isomorphism test. With injective aggregation and update maps, GIN matches the distinguishing power of 1-WL, and its sum aggregation separates multisets of neighbor features that mean- or max-pooling cannot.

The crucial structural property of every message passing layer is automatic permutation equivariance. Relabelling the vertices of the graph produces a correspondingly relabelled output, because \(\mathrm{AGGREGATE}\) operates on a multiset of messages and is by definition invariant under their order. This property places GNNs in the GDL framework, since the architecture is constrained by its very construction to satisfy \(f(P\mathbf{X}, P\mathbf{A}P^\top) = Pf(\mathbf{X}, \mathbf{A})\) for every \(P \in S_n\).

Comparison with CNNs and Transformers

CNNs, Transformers, and GNNs admit a unified description as message passing on graphs, distinguished only by which graph and which aggregation:

The continuum from CNN through Transformer to GNN is therefore one of increasing data-dependence in the connectivity structure: regular grid → complete graph → arbitrary graph. The Transformer's success at scale suggests that the complete graph is often a useful default when no informative connectivity is given. The GNN is the natural choice when connectivity carries information that should be respected.

This equivariance with respect to graph structure is exactly the property that has made GNNs the architecture of choice where data is intrinsically relational. Examples include molecular property prediction (where atoms and bonds form a graph), recommendation systems and social network analysis (where the user-item or user-user graph carries the predictive signal), knowledge graph completion, and traffic and network science. Specific benchmark architectures evolve quickly, but the underlying design principle of encoding the graph's symmetry into the layer has remained stable across a decade of empirical development.

Beyond Permutation: Continuous Group Equivariance

The GDL principle generalizes naturally beyond the discrete symmetries we have discussed so far. Consider a molecule represented not as an abstract graph but as a set of atoms with positions in three-dimensional space. Rotating or translating the entire molecule produces the same molecule, with its energy, stability, and chemical identity unchanged. A learned function predicting any of these properties should therefore be invariant under the group \(SE(3)\) of rigid motions, which combines rotations \(SO(3)\) with three-dimensional translations. A learned function predicting a vector-valued geometric quantity (a force, an orientation) should instead be equivariant under \(SE(3)\). Rotating the input then rotates the output by the same rotation.

Implicit in this distinction is a subtlety we have so far suppressed in our notation. The equivariance condition \(f(g \cdot x) = g \cdot f(x)\) writes the action of \(g\) the same way on both sides, but the actual action depends on the type of the quantity being acted on. A rotation \(R \in SO(3)\) acts trivially on a scalar (energy, charge density), as the standard \(3 \times 3\) rotation matrix on a vector (force, dipole moment), and via more elaborate formulas on higher-order tensors.

The systematic study of how a single group acts in different ways on different feature types is the representation theory of the group, and equivariant networks for 3D data are organized precisely by which representation each layer-feature transforms under. The same shorthand \(g \cdot x\) can therefore stand for genuinely distinct linear maps, and a careful theory must distinguish them.

Both \(SO(3)\) and \(SE(3)\) are continuous groups, and concretely they are matrix Lie groups. Building neural network layers that commute with their action requires more than the multiset symmetrization that powers GNNs.

A common approach decomposes vertex features into irreducible representations of the rotation group (scalars, vectors, higher-order tensors corresponding to the spherical harmonics), constrains messages to transform appropriately under rotation, and combines them via Clebsch-Gordan products, the tensor product decompositions that respect the irreducible representation structure. The result is a graph-like architecture whose messages are themselves geometric objects rather than scalar features. Such an architecture knows, by its construction, how to handle the rotation of its input.

The mathematical machinery this requires includes irreducible representations of compact Lie groups, spherical harmonics, the Peter-Weyl theorem, and equivariant tensor product decompositions. The systematic construction of equivariant neural networks built on this machinery is the subject of our page on equivariant neural networks. Here we record only the application-side outcome.

Encoding continuous symmetry into the architecture has produced concrete advances in structural biology. There, AlphaFold's prediction of three-dimensional protein structure from amino-acid sequence, recognized by the 2024 Nobel Prize in Chemistry, relies on attention mechanisms designed to respect the rigid-motion symmetries of residues in 3D space. Equivariant architectures are also a leading family in machine-learning interatomic potentials for molecular dynamics, and an active research direction in robotic manipulation, where the geometry of physical configuration spaces naturally calls for \(SE(3)\)-equivariant policies.

Why the Mathematics Comes Next

Three threads converge in equivariant neural networks for 3D data. Lie groups provide the symmetry groups themselves and their algebraic structure. Representation theory decomposes feature spaces into irreducible building blocks compatible with the symmetry. The differential geometry of smooth manifolds supplies the natural setting for "data on a curved domain" and the analogue of the graph Laplacian (the Laplace-Beltrami operator) on such domains. The manifold series and the representation theory pages develop these threads, and the Peter-Weyl theorem connects them back to harmonic analysis on Hilbert spaces. The GDL viewpoint makes the destination concrete. Each of these mathematical developments pays off in a specific class of equivariant architecture.

Interactive Demo

The demo above places the central claim of this page side by side. All three panels run the same layer template \[ \mathbf{h}^{(k+1)} = \tanh\!\big( w_k\, \hat{\mathbf{A}}\, \mathbf{h}^{(k)} + b_k \mathbf{1} \big), \] with the GNN and the Transformer differing only in the structure matrix \(\hat{\mathbf{A}}\). The GNN uses the fixed, sparse, degree-normalized adjacency \(\hat{\mathbf{A}} = \tilde{\mathbf{D}}^{-1/2}(\mathbf{A} + \mathbf{I})\tilde{\mathbf{D}}^{-1/2}\) of the displayed graph. The Transformer computes a dense attention matrix \(\hat{\mathbf{A}}(\mathbf{h})_{ij} = \operatorname{softmax}_j(h_i h_j / \tau)\) from the current features, with every pair of vertices connected. The MLP replaces \(w_k \hat{\mathbf{A}}\) with a dense matrix \(\mathbf{W}_k\) of free parameters that mixes node coordinates by their index, and gives each vertex its own bias in place of \(b_k \mathbf{1}\). It uses no structure matrix at all, which is why its panel draws no edges.

The GNN and Transformer panels share the same scalar weights \((w_k, b_k)\), so whatever differs between those two panels is attributable to the structure matrix alone. Node colors show the scalar feature at each vertex (blue negative, red positive). All weights are untrained and drawn once from a fixed seed.

The Layer slider scrubs a single forward pass through all three panels simultaneously. Layer 0 is the input, and each successive layer shows the features after one application of the template. Beneath each graph, a heatmap displays the structure matrix in effect for the displayed transition. The GNN's heatmap never changes, because its \(\hat{\mathbf{A}}\) is fixed by the graph. The Transformer's is recomputed from the features at every layer, so it visibly changes as we scrub. The MLP's shows its own \(\mathbf{W}_k\), whose entries bear no relation to the drawn edges because the architecture has no notion of them. The Depth slider changes the number of layers, which redraws the whole weight stack from a new fixed seed.

The Shuffle Node Labels button is where the GDL claim becomes testable. Shuffling draws a random permutation \(P\) and applies it jointly to the node features and the adjacency matrix in all three panels at once. The equivariance condition \(f(P\mathbf{h}, P\mathbf{A}P^\top) = Pf(\mathbf{h}, \mathbf{A})\) is then evaluated on the original, unshuffled data with the same fixed weights, and each panel reports its own error.

For the GNN and the Transformer, the error sits at floating-point zero. For the MLP, the same shuffle produces an order-one discrepancy, flagged in red. The dense weight matrix mixes node coordinates in a way that depends on their indexing. The contrast is not a matter of better or worse training. It is built into the architectures by construction, before any training has occurred.

The attention used here is genuine scaled dot-product attention, specialized to scalar features with identity query and key maps. The specialization loses nothing structural. At feature dimension one, the query and key maps contribute only a single scalar multiplier on \(h_i h_j\), which is absorbed here into the temperature \(\tau = 0.15\). The value map is likewise absorbed into the shared weight \(w_k\).

The full mechanism, with learned projections, multiple heads, and positional encoding, is computed exactly, number by number, at toy dimension in the decoder-block demo. The present demo isolates the one structural fact that survives every such refinement, namely that attention weights computed from the features themselves transform covariantly under \(S_n\).