Introduction
On the page on deep neural networks, we surveyed
the major architectural innovations of modern deep learning, namely convolutional networks,
residual connections, attention, and the
Transformer. That page closed its survey with an observation that runs through all of them.
Each architecture is built around a symmetry of its input domain. Convolutional
networks commute with translations of the input image (exactly so for stride one on an
unbounded or periodic grid), and self-attention commutes with permutations of the input
tokens. The shared principle is that the architecture encodes the symmetry,
and architectures that respect their domain's symmetry tend to generalize from far less data
than architectures that do not.
That observation is the starting point of Geometric Deep Learning (GDL): a
unifying framework which views CNNs, Transformers, graph neural networks, and equivariant
architectures for 3D data as instances of a single design principle, namely equivariance
under the symmetry group of the input domain.
The word "geometric" here is meant in its deep mathematical sense. A geometry is most
naturally understood as the study of what remains invariant under a chosen group
of transformations. Different choices of group give different geometries, such as a grid
under translation, a set under permutation, a graph under node relabelling, or a smooth
manifold under rotation. An architecture is "geometric" when it respects the symmetries
of its domain.
The present page introduces this framework, then develops its most directly accessible
instance, the graph neural network (GNN), which generalizes message passing
from regular grids and complete graphs to arbitrary graphs. We then survey how the framework
extends beyond permutation symmetry to the continuous symmetries of three-dimensional space.
This is the territory of equivariant neural networks for molecular structure, point clouds,
and rigid-body manipulation.
The mathematical machinery for the continuous-symmetry case is substantial. It draws on
Lie group theory, on
representation theory of compact groups, and on the differential geometry of smooth
manifolds. The present page stays at the level of design principle, showing why those
foundations are needed and what kind of architecture they support.
Geometric Deep Learning: A Unifying Framework
The GDL framework, articulated by Bronstein, Bruna, Cohen, and
Veličković, organizes deep architectures by the symmetry group of their
input domain and the way each layer is required to commute with that group's action.
Concretely, let \(\mathcal{X}\) be a domain with symmetry group \(G\) acting on it, and
let \(G\) act on \(\mathcal{Y}\) as well. A layer \(f : \mathcal{X} \to \mathcal{Y}\)
is said to be \(G\)-equivariant if
\[
f(g \cdot x) = g \cdot f(x) \quad \text{for every } g \in G,\ x \in \mathcal{X},
\]
and \(G\)-invariant if the right-hand side is replaced by \(f(x)\) itself.
Equivariance preserves symmetry information through the network, whereas invariance discards
it. In practice, deep architectures stack equivariant layers and finish with an invariant
pooling step, so that the output is invariant to the chosen symmetry while intermediate
representations remain sensitive to it.
The major architectures we have already surveyed all fit this pattern. The table below
organizes them by the domain they act on and the symmetry group encoded in the architecture.
Remark: Architectures Organized by Symmetry
| Architecture |
Domain |
Symmetry group |
Equivariance type |
| MLP |
\(\mathbb{R}^d\) |
trivial |
none |
| CNN |
grid (image / volume) |
translation \(\mathbb{Z}^d\) |
translation-equivariant |
| Transformer |
sequence / set |
symmetric group \(S_n\) |
permutation-equivariant |
| GNN |
graph (features \(\mathbf{X}\) + adjacency \(\mathbf{A}\)) |
symmetric group \(S_n\) |
joint permutation of features and topology: \((\mathbf{X}, \mathbf{A}) \mapsto (P\mathbf{X}, P\mathbf{A}P^\top)\) |
| Steerable / \(SE(3)\)-equivariant NN |
3D point cloud, molecular structure |
\(SO(3)\), \(SE(3)\) |
continuous group equivariance |
The table makes the unification explicit. The MLP, alone among the architectures listed,
encodes no symmetry and is the baseline against which all the others are GDL instances.
CNNs, Transformers, GNNs, and \(SE(3)\)-equivariant networks all share the same template.
The template identifies the symmetry group of the data domain and constrains each layer to
commute with its action. The differences between the architectures are which group,
and how the group acts.
The Transformer's place in this table deserves a closer look, because it is in some respects
the cleanest example of GDL design philosophy already at industrial scale.
Self-attention
computes
\[
\mathrm{Attn}(\mathbf{Q}, \mathbf{K}, \mathbf{V}) = \operatorname{softmax}\!\left(
\tfrac{\mathbf{Q}\mathbf{K}^\top}{\sqrt{d_k}}\right)\mathbf{V},
\]
and, because \(\mathbf{Q}\), \(\mathbf{K}\), and \(\mathbf{V}\) are all read off from the
same matrix of tokens, permuting the rows of that matrix permutes the rows of the output
identically. The layer therefore commutes with the
symmetric group
\(S_n\) acting on the token axis.
A Transformer without
positional encoding is therefore a
permutation-equivariant set processor. Positional encoding deliberately breaks
this symmetry to inject sequence order back into the architecture. From the GDL viewpoint,
a Transformer-based large language model is built on a fundamentally GDL-shaped backbone.
Equivariance was an architectural choice, not an afterthought, and it is thought to be part
of why the architecture scales.
Graph neural networks generalize this picture in a direction the Transformer does not cover.
A Transformer treats every pair of tokens as connected, so its connectivity is fixed by the
architecture rather than given by the data. Most real data carries structure of its own.
Only adjacent atoms in a molecule interact directly, only cited pairs in a citation network
are linked, and connectivity in a road network is constrained by physical layout. A GNN
takes the connectivity as part of its input. That input is a pair
\((\mathbf{X}, \mathbf{A})\) of node features and adjacency matrix, and the GNN is designed
to be equivariant under the joint action of \(S_n\) on both:
\[
f(P\mathbf{X},\, P\mathbf{A}P^\top) = P\,f(\mathbf{X}, \mathbf{A})
\quad \text{for every } P \in S_n.
\]
Both the Transformer (without positional encoding) and the GNN are \(S_n\)-equivariant, and
the difference lies in what \(P\) acts on. The Transformer permutes a feature matrix,
whereas the GNN permutes a feature-adjacency pair, with the adjacency matrix conjugated so
that the graph structure is permuted consistently with the node relabelling. If we restrict
attention to one fixed graph, the permutations preserving \(\mathbf{A}\) are exactly the
graph isomorphisms
from that graph onto itself, and they form its automorphism group. Those are the only
relabellings that permute the features while leaving the graph itself fixed. The
architecture, by contrast, is built to handle any graph, which is why a single
trained GNN generalizes to graphs unseen at training time.
Graph Neural Networks: From Spectral to Spatial
Two essentially different approaches to constructing graph-based architectures emerged
historically. The spectral approach, which starts from signal processing
on graphs, defines convolution on a graph through the
Graph Fourier Transform,
parameterizes filters in the spectral domain, and recovers tractable architectures by
polynomial approximation. The spatial approach defines a layer directly as local message passing between adjacent vertices.
The two viewpoints are connected, since a widely used spatial architecture, the
GCN, is mathematically a first-order Chebyshev approximation of a spectral filter. We
treat them in turn.
The Spectral Side: Filters in the Frequency Domain
The graph Laplacian
\(\mathbf{L} = \mathbf{D} - \mathbf{A}\) is real symmetric and
positive semi-definite,
so it admits an
eigendecomposition
\(\mathbf{L} = \mathbf{V}\boldsymbol{\Lambda}\mathbf{V}^\top\) with real non-negative
eigenvalues. The eigenvectors play the role of "graph frequencies," and the
Graph Fourier Transform
\(\hat{\mathbf{f}} = \mathbf{V}^\top \mathbf{f}\) decomposes a graph signal into its
frequency components. A spectral graph convolution applies a filter
\(g(\boldsymbol{\Lambda})\) to the spectrum:
\[
\mathbf{f}_{\mathrm{out}}
= \mathbf{V}\, g(\boldsymbol{\Lambda})\, \mathbf{V}^\top \mathbf{f}_{\mathrm{in}}.
\]
Computing this directly requires the full eigendecomposition of \(\mathbf{L}\), which is
prohibitive for large graphs. The standard remedy approximates the filter by a low-degree
polynomial in \(\mathbf{L}\), typically built from Chebyshev polynomials of the rescaled
Laplacian. The resulting architectures are efficient and localized, since a
degree-\(K\) polynomial filter respects \(K\)-hop neighborhoods. The first-order case of
this construction is the Graph Convolutional Network (GCN), the widely used
architecture named above.
Our graph Laplacian page
gives the full development: the Chebyshev approximation, the rescaling
\(\tilde{\mathbf{L}} = (2/\lambda_{\max})\mathcal{L}_{\mathrm{sym}} - \mathbf{I}\) of
the symmetric normalized Laplacian
\(\mathcal{L}_{\mathrm{sym}} = \mathbf{D}^{-1/2}\mathbf{L}\mathbf{D}^{-1/2}\), which
is defined when no vertex is isolated and whose largest eigenvalue is
\(\lambda_{\max}\), and the GCN derivation. Here we recall only what we need to
connect the spectral and spatial perspectives.
The Spatial Side: Message Passing
The spatial perspective begins from a different place. Rather than diagonalizing the
Laplacian, we ask directly: what local computation respects graph structure?
The answer takes the form of a message passing layer. At each layer
\(k\), every vertex \(v\) updates its hidden state \(\mathbf{h}_v^{(k)}\) by aggregating
messages from its neighbors and combining them with its own current state:
Definition: Message Passing Layer
A message passing layer updates vertex features by
\[
\mathbf{h}_v^{(k+1)}
= \mathrm{UPDATE}\!\left(\mathbf{h}_v^{(k)},\ \mathrm{AGGREGATE}\big(\{\!\!\{ \mathbf{h}_u^{(k)} : u \in \mathcal{N}(v) \}\!\!\}\big)\right),
\]
where \(\mathcal{N}(v)\) is the
neighborhood
of \(v\), \(\{\!\!\{\cdot\}\!\!\}\) denotes a multiset, \(\mathrm{AGGREGATE}\) is a
permutation-invariant function over the neighborhood (sum, mean, max, attention-weighted
sum), and \(\mathrm{UPDATE}\) is a learnable transformation (typically a small MLP).
The choice of \(\mathrm{AGGREGATE}\) and \(\mathrm{UPDATE}\) determines the architecture.
The GCN uses a normalized-degree weighted sum derived from \(\mathcal{L}_{\mathrm{sym}}\).
GraphSAGE uses sampled neighborhoods with a learnable pooling function, which addresses
scalability on very large graphs. The Graph Attention Network (GAT) computes neighborhood
weights by an attention mechanism and so recovers much of the flexibility of self-attention
while restricting interactions to graph edges.
The Graph Isomorphism Network (GIN) replaces the aggregation with a sum followed by an MLP,
a choice motivated by its connection to the 1-Weisfeiler-Lehman (1-WL)
graph isomorphism test. With injective aggregation and update maps, GIN matches the
distinguishing power of 1-WL, and its sum aggregation separates multisets of neighbor
features that mean- or max-pooling cannot.
The crucial structural property of every message passing layer is automatic
permutation equivariance. Relabelling the vertices of the graph produces a
correspondingly relabelled output, because \(\mathrm{AGGREGATE}\) operates on a multiset of
messages and is by definition invariant under their order. This property places GNNs in the
GDL framework, since the architecture is constrained by its very construction to satisfy
\(f(P\mathbf{X}, P\mathbf{A}P^\top) = Pf(\mathbf{X}, \mathbf{A})\) for every \(P \in S_n\).
Comparison with CNNs and Transformers
CNNs, Transformers, and GNNs admit a unified description as message passing on graphs,
distinguished only by which graph and which aggregation:
-
A CNN performs message passing on a regular grid graph with fixed local
neighborhoods and a learned weight per edge offset. The same weights are applied at
every position, which is what makes the layer translation-equivariant.
-
A Transformer performs message passing on the
complete graph over the input tokens, with attention-derived edge weights.
The layer is permutation-equivariant in the strong sense that all permutations of
\(S_n\) commute with it.
-
A GNN performs message passing on an input graph supplied as
part of the data rather than fixed by the architecture, with \(S_n\) acting
jointly on the feature-adjacency pair as formalized above.
The continuum from CNN through Transformer to GNN is therefore one of increasing
data-dependence in the connectivity structure: regular grid → complete graph
→ arbitrary graph. The Transformer's success at scale suggests that the complete graph
is often a useful default when no informative connectivity is given. The GNN is the natural
choice when connectivity carries information that should be respected.
This equivariance with respect to graph structure is exactly the property that has made GNNs
the architecture of choice where data is intrinsically relational. Examples include
molecular property prediction (where atoms and bonds form a graph), recommendation systems
and social network analysis (where the user-item or user-user graph carries the predictive
signal), knowledge graph completion, and traffic and network science. Specific benchmark
architectures evolve quickly, but the underlying design principle of encoding the graph's
symmetry into the layer has remained stable across a decade of empirical development.
Beyond Permutation: Continuous Group Equivariance
The GDL principle generalizes naturally beyond the discrete symmetries we have discussed
so far. Consider a molecule represented not as an abstract graph but as a set of atoms
with positions in three-dimensional space. Rotating or translating the entire molecule
produces the same molecule, with its energy, stability, and chemical identity unchanged.
A learned function predicting any of these properties should therefore be invariant under
the group \(SE(3)\) of
rigid motions,
which combines rotations \(SO(3)\) with three-dimensional translations. A learned
function predicting a vector-valued geometric quantity (a force, an orientation) should
instead be equivariant under \(SE(3)\). Rotating the input then rotates the
output by the same rotation.
Implicit in this distinction is a subtlety we have so far suppressed in our notation. The
equivariance condition \(f(g \cdot x) = g \cdot f(x)\) writes the action of \(g\) the same
way on both sides, but the actual action depends on the type of the
quantity being acted on. A rotation \(R \in SO(3)\) acts trivially on a scalar (energy,
charge density), as the standard \(3 \times 3\) rotation matrix on a vector (force, dipole
moment), and via more elaborate formulas on higher-order tensors.
The systematic study of how a single group acts in different ways on different feature types
is the representation theory of the group, and equivariant networks for 3D
data are organized precisely by which representation each layer-feature transforms under.
The same shorthand \(g \cdot x\) can therefore stand for genuinely distinct linear maps, and
a careful theory must distinguish them.
Both \(SO(3)\) and \(SE(3)\) are continuous groups, and concretely they are
matrix Lie groups. Building
neural network layers that commute with their action requires more than the multiset
symmetrization that powers GNNs.
A common approach decomposes vertex features into
irreducible representations of the rotation group (scalars, vectors,
higher-order tensors corresponding to the spherical harmonics), constrains messages to
transform appropriately under rotation, and combines them via Clebsch-Gordan products, the
tensor product decompositions that respect the irreducible representation structure. The
result is a graph-like architecture whose messages are themselves geometric objects rather
than scalar features. Such an architecture knows, by its construction, how to
handle the rotation of its input.
The mathematical machinery this requires includes irreducible representations of compact Lie
groups, spherical harmonics, the
Peter-Weyl theorem, and
equivariant tensor product decompositions. The systematic construction of equivariant neural
networks built on this machinery is the subject of our page on
equivariant neural networks. Here we
record only the application-side outcome.
Encoding continuous symmetry into the architecture has produced concrete advances in
structural biology. There, AlphaFold's prediction of three-dimensional
protein structure from amino-acid sequence, recognized by the 2024 Nobel Prize in
Chemistry, relies on attention mechanisms designed to respect the rigid-motion symmetries
of residues in 3D space. Equivariant architectures are also a leading family in
machine-learning interatomic potentials for molecular dynamics, and an active research
direction in robotic manipulation, where the geometry of physical configuration spaces
naturally calls for \(SE(3)\)-equivariant policies.
Why the Mathematics Comes Next
Three threads converge in equivariant neural networks for 3D data. Lie groups
provide the symmetry groups themselves and their algebraic structure.
Representation theory decomposes feature spaces into irreducible building
blocks compatible with the symmetry. The
differential geometry of smooth manifolds supplies the natural setting for
"data on a curved domain" and the analogue of the graph Laplacian (the Laplace-Beltrami
operator) on such domains. The
manifold series
and the representation theory pages develop these threads, and the
Peter-Weyl theorem
connects them back to harmonic analysis on Hilbert spaces. The GDL viewpoint makes the
destination concrete. Each of these mathematical developments pays off in a specific
class of equivariant architecture.
Interactive Demo
The demo above places the central claim of this page side by side. All three panels run the
same layer template
\[
\mathbf{h}^{(k+1)} = \tanh\!\big( w_k\, \hat{\mathbf{A}}\, \mathbf{h}^{(k)} + b_k \mathbf{1} \big),
\]
with the GNN and the Transformer differing only in the
structure matrix \(\hat{\mathbf{A}}\). The GNN uses the fixed, sparse,
degree-normalized adjacency \(\hat{\mathbf{A}} = \tilde{\mathbf{D}}^{-1/2}(\mathbf{A} + \mathbf{I})\tilde{\mathbf{D}}^{-1/2}\)
of the displayed graph. The Transformer computes a dense attention
matrix \(\hat{\mathbf{A}}(\mathbf{h})_{ij} = \operatorname{softmax}_j(h_i h_j / \tau)\)
from the current features, with every pair of vertices connected. The
MLP replaces \(w_k \hat{\mathbf{A}}\) with a dense matrix
\(\mathbf{W}_k\) of free parameters that mixes node coordinates by their index, and gives
each vertex its own bias in place of \(b_k \mathbf{1}\). It uses no structure matrix at
all, which is why its panel draws no edges.
The GNN and Transformer panels share the same scalar weights \((w_k, b_k)\), so
whatever differs between those two panels is attributable to the structure matrix alone.
Node colors show the scalar feature at each vertex (blue negative, red positive). All
weights are untrained and drawn once from a fixed seed.
The Layer slider scrubs a single forward pass through all three panels
simultaneously. Layer 0 is the input, and each successive layer shows the features after one
application of the template. Beneath each graph, a heatmap displays the structure matrix in
effect for the displayed transition. The GNN's heatmap never changes, because its
\(\hat{\mathbf{A}}\) is fixed by the graph. The Transformer's is recomputed from the
features at every layer, so it visibly changes as we scrub. The MLP's shows its own
\(\mathbf{W}_k\), whose entries bear no relation to the drawn edges because the architecture
has no notion of them. The Depth slider changes the number of layers, which
redraws the whole weight stack from a new fixed seed.
The Shuffle Node Labels button is where the GDL claim becomes testable.
Shuffling draws a random permutation \(P\) and applies it jointly to the node features and
the adjacency matrix in all three panels at once. The equivariance condition
\(f(P\mathbf{h}, P\mathbf{A}P^\top) = Pf(\mathbf{h}, \mathbf{A})\) is then evaluated on the
original, unshuffled data with the same fixed weights, and each panel reports its own error.
For the GNN and the Transformer, the error sits at floating-point zero. For the MLP, the
same shuffle produces an order-one discrepancy, flagged in red. The dense weight matrix
mixes node coordinates in a way that depends on their indexing. The contrast is not a matter
of better or worse training. It is built into the architectures by construction, before any
training has occurred.
The attention used here is genuine
scaled dot-product attention,
specialized to scalar features with identity query and key maps. The specialization loses
nothing structural. At feature dimension one, the query and key maps contribute only a
single scalar multiplier on \(h_i h_j\), which is absorbed here into the temperature
\(\tau = 0.15\). The value map is likewise absorbed into the shared weight \(w_k\).
The full mechanism, with learned projections, multiple heads, and positional encoding, is
computed exactly, number by number, at toy dimension in the
decoder-block demo. The present demo
isolates the one structural fact that survives every such refinement, namely that attention
weights computed from the features themselves transform covariantly under \(S_n\).