The Lie Group-Lie Algebra Correspondence
Over the past three pages, we have built a dictionary between Lie groups and Lie
algebras. A
matrix Lie group
\(G\) determines a
Lie algebra
\(\mathfrak{g} = T_I G\), a vector space equipped with a
Lie bracket
\([X, Y] = XY - YX\). The
matrix exponential
\(\exp : \mathfrak{g} \to G\) sends the algebra to the group, converting linear
(infinitesimal) data into nonlinear (finite) group elements.
We now address the central question: to what extent does the Lie algebra
determine the Lie group?
The Exponential Map as a Local Diffeomorphism
The exponential map \(\exp : \mathfrak{g} \to G\) is, in general, neither injective
nor surjective globally. For instance, in \(SO(2)\), the map
\(\theta \mapsto \exp(\theta J)\) wraps \(\mathbb{R}\) infinitely many times around
the circle, so it is not injective. However, near the origin, the exponential map is
well-behaved:
Theorem: Local Diffeomorphism
Let \(G\) be a matrix Lie group with Lie algebra \(\mathfrak{g}\). The exponential
map \(\exp : \mathfrak{g} \to G\) is a diffeomorphism from a
neighborhood of \(0 \in \mathfrak{g}\) onto a neighborhood of \(I \in G\). That
is, every group element in that neighborhood of \(I\) has a unique logarithm in
that neighborhood of \(0\).
The proof uses the inverse function theorem on manifolds. The derivative of \(\exp\) at
\(0\) is the identity map \(\mathfrak{g} \to \mathfrak{g}\) (since
\(\frac{d}{dt}\big|_{t=0} \exp(tA) = A\)), and the identity map is invertible. Hence
\(\exp\) is a local diffeomorphism near \(0\). The full details require the machinery of
smooth manifolds. For now,
we state the theorem and draw its consequences.
The local diffeomorphism property means that the Lie algebra \(\mathfrak{g}\) faithfully
encodes the local structure of \(G\), the structure in a neighborhood
of the identity. Two Lie groups with isomorphic Lie algebras are
locally isomorphic, in the sense that they look identical near their
respective identities. They may, however, differ globally in their topology.
We will see a dramatic example of this phenomenon in the next section.
Lie Algebra Homomorphisms
The correspondence between groups and algebras extends to their maps. Just as a
group homomorphism
preserves the group operation, a Lie algebra homomorphism preserves the bracket.
Definition: Lie Algebra Homomorphism
A Lie algebra homomorphism is a linear map
\(\varphi : \mathfrak{g} \to \mathfrak{h}\) between Lie algebras that preserves
the bracket:
\[
\varphi([X, Y]) = [\varphi(X), \varphi(Y)]
\quad \text{for all } X, Y \in \mathfrak{g}.
\]
A bijective Lie algebra homomorphism is a Lie algebra isomorphism.
The key theorem of the Lie correspondence is that group homomorphisms automatically
induce Lie algebra homomorphisms. The passage from group to algebra is
functorial.
Theorem: Group Homomorphisms Induce Lie Algebra Homomorphisms
Let \(\Phi : G \to H\) be a Lie group homomorphism (that is, a continuous group
homomorphism between matrix Lie groups). Then the derivative at the
identity,
\[
d\Phi_I : \mathfrak{g} \to \mathfrak{h}, \quad
d\Phi_I(A) = \left.\frac{d}{dt}\right|_{t=0} \Phi(\exp(tA)),
\]
is a Lie algebra homomorphism:
\(d\Phi_I([A, B]) = [d\Phi_I(A),\, d\Phi_I(B)]\).
Proof sketch:
Since \(\Phi\) is a group homomorphism, it intertwines the exponential maps:
\[
\Phi(\exp(tA)) = \exp(t \cdot d\Phi_I(A))
\]
for all \(A \in \mathfrak{g}\) and \(t \in \mathbb{R}\). (This follows from the
fact that \(t \mapsto \Phi(\exp(tA))\) is a one-parameter subgroup of \(H\) with
initial velocity \(d\Phi_I(A)\), so by the
one-parameter subgroup theorem,
it equals \(\exp(t \cdot d\Phi_I(A))\).)
Now, \([A, B] \in \mathfrak{g}\) by
closure under the commutator,
and differentiating \(\exp(tA)\,B\,\exp(-tA)\) at \(t = 0\) by the product rule
gives \([A, B] = \frac{d}{dt}\big|_{t=0} \exp(tA)\,B\,\exp(-tA)\).
Applying \(d\Phi_I\) and using the intertwining property:
\[
\begin{align*}
d\Phi_I([A, B])
&= \left.\frac{d}{dt}\right|_{t=0}
\Phi(\exp(tA))\, d\Phi_I(B)\, \Phi(\exp(-tA)) \\\\
&= \left.\frac{d}{dt}\right|_{t=0}
\exp(t\,d\Phi_I(A))\, d\Phi_I(B)\, \exp(-t\,d\Phi_I(A)) \\\\
&= [d\Phi_I(A),\, d\Phi_I(B)].
\end{align*}
\]
That \(d\Phi_I\) is linear, and that it commutes with the \(t\)-derivative in
the first step, both rest on the smoothness of \(\Phi\), which we take for granted
here.
Example: The Determinant and the Trace
The determinant
\(\det : GL(n, \mathbb{R}) \to \mathbb{R} \setminus \{0\}\) is a Lie group
homomorphism. Its derivative at the identity follows from
property (d)
of the matrix exponential:
\[
\begin{align*}
d(\det)_I(A)
&= \left.\frac{d}{dt}\right|_{t=0} \det(\exp(tA)) \\\\
&= \left.\frac{d}{dt}\right|_{t=0} e^{t\,\mathrm{tr}(A)} \\\\
&= \mathrm{tr}(A).
\end{align*}
\]
So the induced Lie algebra homomorphism
\(\mathrm{tr} : \mathfrak{gl}(n, \mathbb{R}) \to \mathbb{R}\) is the trace.
Since \(\mathbb{R}\) is abelian (its Lie bracket is zero), the homomorphism
condition \(\mathrm{tr}([A, B]) = [\mathrm{tr}(A), \mathrm{tr}(B)] = 0\)
reduces to \(\mathrm{tr}(AB - BA) = 0\), which indeed holds because
\(\mathrm{tr}(AB) = \mathrm{tr}(BA)\).
The
kernel
of \(\det\) is \(SL(n)\). Correspondingly, the kernel of \(\mathrm{tr}\) is
\(\mathfrak{sl}(n)\). Group-level and algebra-level kernels correspond perfectly.
The Baker-Campbell-Hausdorff Formula
We now arrive at the deepest result of the Lie correspondence. The group multiplication
near the identity is entirely determined by the Lie bracket. This is
formalized by the Baker-Campbell-Hausdorff (BCH) formula.
The smallness condition on \(\|X\|, \|Y\|\) is what makes the right-hand side
meaningful. Even when \(X, Y\) are arbitrary elements of \(\mathfrak{g}\), the product
\(\exp(X)\exp(Y)\) is always an element of \(G\). To obtain \(Z\) by inverting
\(\exp\), however, we need \(\exp(X)\exp(Y)\) to lie in the neighborhood of \(I\)
on which the
local diffeomorphism
inverts \(\exp\).
For \(X\) and \(Y\) small enough, the group multiplication, a curved and nonlinear
operation, is recovered from the Lie bracket through the series above. The full series
involves increasingly complex nested brackets and is determined by a universal
explicit formula (the Dynkin formula). We do not prove the BCH formula (one standard
proof uses formal power series in non-commuting variables). Instead, we focus on its
meaning and consequences.
The key message. The group multiplication
\((g, h) \mapsto gh\) is a nonlinear operation on a curved manifold. The BCH
formula shows that, near the identity, this nonlinear operation is entirely encoded
by the Lie bracket, a bilinear operation on a vector space. Every coefficient
in the BCH series is built from iterated brackets and nothing else. This is why the Lie
algebra, with its bracket, determines the local structure of the group.
Special cases. The BCH formula illuminates the results of the
preceding pages:
(a) Commuting case. If \([X, Y] = 0\), then all bracket terms in the
BCH series vanish, and \(Z = X + Y\). For small \(X\) and \(Y\), this recovers
\(\exp(X)\exp(Y) = \exp(X + Y)\), which
property (b)
of the matrix exponential gives for every commuting pair.
(b) First-order approximation. Keeping only the first bracket term:
\[
\exp(X)\exp(Y) \approx \exp\!\left(X + Y + \tfrac{1}{2}[X, Y]\right)
\]
for small \(X, Y\). The failure of \(\exp\) to be a homomorphism is measured, to
leading order, by \(\frac{1}{2}[X, Y]\). This makes precise the statement from
The Matrix Exponential
that the commutator is the "first correction term."
(c) Abelian groups. If \(\mathfrak{g}\) is abelian (\([X, Y] = 0\)
for all \(X, Y\)), then case (a) gives \(\exp(X)\exp(Y) = \exp(X + Y)\) for every pair
of small elements, and property (b) of the matrix exponential extends it to every
pair, so the exponential map is a group homomorphism from \((\mathfrak{g}, +)\) to
\(G\). This is exactly what happens for \(SO(2)\). The exponential
\(\theta \mapsto \exp(\theta J)\) is a homomorphism because
\(\mathfrak{so}(2) \cong \mathbb{R}\) is abelian.
The Double Cover: \(SU(2)\) and \(SO(3)\)
The BCH formula tells us that isomorphic Lie algebras produce locally isomorphic Lie
groups. But "locally isomorphic" does not mean "isomorphic." Two groups can share
the same infinitesimal structure while differing in their global topology. The most
important example of this phenomenon is the relationship between
\(SU(2)\)
and
\(SO(3)\).
The Lie Algebra Isomorphism \(\mathfrak{su}(2) \cong \mathfrak{so}(3)\)
The Lie algebra \(\mathfrak{su}(2)\) consists of \(2 \times 2\) traceless
skew-Hermitian matrices. A standard basis is:
\[
\begin{align*}
F_1 &= \begin{pmatrix} 0 & -i \\ -i & 0 \end{pmatrix}, \\\\
F_2 &= \begin{pmatrix} 0 & -1 \\ 1 & 0 \end{pmatrix}, \\\\
F_3 &= \begin{pmatrix} -i & 0 \\ 0 & i \end{pmatrix}.
\end{align*}
\]
(These are \(-i\) times the Pauli matrices \(\sigma_1, \sigma_2, \sigma_3\).) Each
\(F_k\) is skew-Hermitian (\(F_k^* = -F_k\)) and traceless by inspection, and
multiplying the matrices gives
\[
\begin{align*}
[F_1, F_2] &= 2F_3, \\\\
[F_2, F_3] &= 2F_1, \\\\
[F_3, F_1] &= 2F_2.
\end{align*}
\]
Setting \(\tilde{E}_k = \frac{1}{2}F_k\) for \(k = 1, 2, 3\) gives a basis satisfying
\[
\begin{align*}
[\tilde{E}_1, \tilde{E}_2] &= \tilde{E}_3, \\\\
[\tilde{E}_2, \tilde{E}_3] &= \tilde{E}_1, \\\\
[\tilde{E}_3, \tilde{E}_1] &= \tilde{E}_2.
\end{align*}
\]
The bracket relations are exactly those of the basis
\(\{E_1, E_2, E_3\}\) of \(\mathfrak{so}(3)\) computed in
Lie Algebras and the Lie Bracket.
The linear map \(\tilde{E}_k \mapsto E_k\) is therefore a Lie algebra isomorphism:
\[
\mathfrak{su}(2) \cong \mathfrak{so}(3).
\]
Both are 3-dimensional real Lie algebras with structure constants given by the
Levi-Civita symbol. At the Lie algebra level, they are indistinguishable.
The Groups are Not Isomorphic
Despite their identical Lie algebras, \(SU(2)\) and \(SO(3)\) are
not isomorphic as groups. An element \(U \in SU(2)\) with \(U^2 = I\)
is unitary, hence diagonalizable, with eigenvalues \(\pm 1\) whose product is
\(\det U = 1\), so \(U = \pm I\). Thus \(-I\) is the only element of order \(2\) in
\(SU(2)\), while \(SO(3)\) contains several, for example \(\mathrm{diag}(1, -1, -1)\)
and \(\mathrm{diag}(-1, 1, -1)\), and an isomorphism preserves the orders of elements.
The difference is also topological:
\(SU(2)\) is homeomorphic to \(S^3\). Every element of \(SU(2)\) has
the form
\(\begin{pmatrix} \alpha & -\bar{\beta} \\ \beta & \bar{\alpha} \end{pmatrix}\) with
\(|\alpha|^2 + |\beta|^2 = 1\), which identifies \(SU(2)\) with the unit sphere in
\(\mathbb{C}^2 \cong \mathbb{R}^4\). Since the
sphere \(S^3\) is simply connected,
so is \(SU(2)\). Every closed loop in it can be continuously shrunk to a point.
\(SO(3)\) is homeomorphic to \(\mathbb{RP}^3\). The real projective
space \(\mathbb{RP}^3 = S^3 / \{x \sim -x\}\) is obtained from \(S^3\) by identifying
antipodal points. This space is not
simply connected. The
fundamental group
of \(SO(3)\) is \(\pi_1(SO(3)) \cong \mathbb{Z}/2\mathbb{Z}\). There exist loops in
\(SO(3)\) (a rotation by \(2\pi\) about any axis) that cannot be continuously deformed
to the identity, but traversing such a loop twice (rotation by \(4\pi\)) yields
a contractible loop. The homeomorphism, the value of the fundamental group, and the
behavior of the \(2\pi\) loop are all derived from the double cover constructed below,
in the proposition on the
fundamental groups of \(SU(2)\) and \(SO(3)\)
and the discussion after it. This page constructs the double cover and establishes its
algebraic structure and its continuity.
The Double Cover Homomorphism
The Lie algebra isomorphism \(\mathfrak{su}(2) \cong \mathfrak{so}(3)\) lifts to a
group homomorphism, but not to an isomorphism. There is a surjective two-to-one map
\(\Phi : SU(2) \to SO(3)\), called the double cover.
Theorem (The Double Cover \(SU(2) \to SO(3)\))
The conjugation action \(\Phi : SU(2) \to SO(3)\), \(U \mapsto \Phi_U\),
constructed below, is a surjective group homomorphism that is two-to-one,
with kernel
\[
\ker(\Phi) = \{I, -I\} \cong \mathbb{Z}/2\mathbb{Z}.
\]
Consequently \(SU(2)/\{I, -I\} \cong SO(3)\) as groups. Each rotation \(R \in SO(3)\) is
the image of exactly one pair \(\{U, -U\} \subset SU(2)\).
Construction:
The group \(SU(2)\) acts on its own Lie algebra \(\mathfrak{su}(2)\) by
conjugation: for \(U \in SU(2)\) and \(X \in \mathfrak{su}(2)\), define
\[
\Phi_U(X) = U X U^*.
\]
Since conjugation preserves skew-Hermiticity and tracelessness, \(\Phi_U\) maps
\(\mathfrak{su}(2)\) to itself.
The bilinear form
\(\langle X, Y \rangle = -\frac{1}{2}\,\mathrm{tr}(XY)\) is a genuine (positive-definite)
inner product on \(\mathfrak{su}(2)\). It is real-valued and symmetric, since for
\(X, Y \in \mathfrak{su}(2)\) the conjugate of \(\mathrm{tr}(XY)\) is
\(\mathrm{tr}((XY)^*) = \mathrm{tr}((-Y)(-X)) = \mathrm{tr}(YX) = \mathrm{tr}(XY)\).
For \(X \in \mathfrak{su}(2)\) we have
\(X^* = -X\), hence
\[
\langle X, X \rangle = -\tfrac{1}{2}\,\mathrm{tr}(X^2)
= \tfrac{1}{2}\,\mathrm{tr}(X^* X) = \tfrac{1}{2}\,\|X\|_F^2 \geq 0,
\]
with equality if and only if \(X = 0\), where \(\|\cdot\|_F\) denotes the Frobenius norm.
Moreover, \(\Phi_U\) preserves this inner product:
\[
\langle \Phi_U(X), \Phi_U(Y) \rangle
= -\tfrac{1}{2}\,\mathrm{tr}(UXU^* \cdot UYU^*)
= -\tfrac{1}{2}\,\mathrm{tr}(XY) = \langle X, Y \rangle.
\]
The basis \(\{\tilde{E}_1, \tilde{E}_2, \tilde{E}_3\}\) is orthogonal for this inner
product, with \(\langle \tilde{E}_k, \tilde{E}_k \rangle = \tfrac{1}{4}\) for each
\(k\), so the coordinate map \(\mathfrak{su}(2) \to \mathbb{R}^3\) it defines
multiplies every length by \(2\). The matrix of \(\Phi_U\) in this basis is
therefore orthogonal, and we regard \(\Phi_U\) as an isometry of \(\mathbb{R}^3\).
Its entries \(4\langle \tilde{E}_j, \Phi_U(\tilde{E}_k) \rangle\) are real
polynomials in the real and imaginary parts of the entries of \(U\), so
\(U \mapsto \Phi_U\) is continuous, indeed smooth. Each \(\Phi_U\) also preserves
orientation. Its determinant is \(\pm 1\), depends continuously on \(U\), and equals
\(1\) at \(U = I\), and \(SU(2)\) is connected, so it equals \(1\) throughout. The
group \(SU(2)\) is connected because for \(U \in SU(2)\) the conditions \(U^{-1} = U^*\) and \(\det U = 1\) force
\(U = \begin{pmatrix} a & -\overline{b} \\ b & \overline{a} \end{pmatrix}\) with
\(|a|^2 + |b|^2 = 1\), which identifies \(SU(2)\) with the unit sphere in
\(\mathbb{C}^2\), a path-connected set. Hence \(\Phi_U \in SO(3)\).
The map \(\Phi : SU(2) \to SO(3)\), \(U \mapsto \Phi_U\), is a group homomorphism
(since conjugation respects composition), and it is continuous, so it is a Lie group
homomorphism in the sense of the
induced homomorphism theorem.
Regard \(d\Phi_I(A)\), through the basis \(\tilde{E}\), as a linear map of
\(\mathfrak{su}(2)\). Since \(\exp(tA)^* = \exp(-tA)\) for
\(A \in \mathfrak{su}(2)\), we get
\(d\Phi_I(A)(Y) = \frac{d}{dt}\big|_{t=0} \exp(tA)\,Y\,\exp(-tA) = [A, Y]\). The
basis \(\tilde{E}_k\) has the same bracket relations as \(E_k\). For instance
\([\tilde{E}_1, \tilde{E}_1] = 0\), \([\tilde{E}_1, \tilde{E}_2] = \tilde{E}_3\),
and \([\tilde{E}_1, \tilde{E}_3] = -\tilde{E}_2\), so the columns of the matrix of
\(Y \mapsto [\tilde{E}_1, Y]\) in the basis \(\tilde{E}\) are \((0, 0, 0)^\top\),
\((0, 0, 1)^\top\), and \((0, -1, 0)^\top\), which is the matrix \(E_1\), and the
same computation gives \(E_2\) and \(E_3\) (compare the example of \(\mathrm{ad}\)
for \(\mathfrak{so}(3)\) below). Hence
\(d\Phi_I : \mathfrak{su}(2) \to \mathfrak{so}(3)\) is the Lie algebra isomorphism
\(\tilde{E}_k \mapsto E_k\) established above. By the identity
\(\Phi(\exp(tA)) = \exp(t \cdot d\Phi_I(A))\) from the proof of that theorem, each
\(\exp(B)\) with \(B \in \mathfrak{so}(3)\) equals \(\Phi(\exp(A))\) for the \(A\)
with \(d\Phi_I(A) = B\), so the image \(\Phi(SU(2))\) contains
\(\exp(\mathfrak{so}(3))\).
Surjectivity of \(\Phi\) now reduces to surjectivity of the exponential map
\(\exp : \mathfrak{so}(3) \to SO(3)\).
Rodrigues' formula
expresses every rotation about a given axis as
\(\exp(\theta\,\hat{\mathbf{n}}_\times)\), and
Euler's rotation theorem
states that every element of \(SO(3)\) is such a rotation. Hence
\(\exp(\mathfrak{so}(3)) = SO(3)\), and \(\Phi\) is surjective.
The kernel and the two-to-one structure:
We have \(\Phi_U = \mathrm{Id}\) if and only if
\(UXU^* = X\) for all \(X \in \mathfrak{su}(2)\), which is equivalent to saying
\(U\) commutes with each basis matrix \(F_1, F_2, F_3\).
We extract \(U = \lambda I\) by a direct calculation. Since \(U\) commutes with
\(F_3 = \mathrm{diag}(-i, i)\), a diagonal matrix with distinct eigenvalues,
\(U\) must itself be diagonal, say \(U = \mathrm{diag}(a, b)\). The commutation
\(U F_1 = F_1 U\) with
\(F_1 = \bigl(\begin{smallmatrix} 0 & -i \\ -i & 0 \end{smallmatrix}\bigr)\) gives
the off-diagonal condition \(-ia = -ib\), hence \(a = b\), so
\(U = \lambda I\). The \(SU(2)\) conditions \(U^* U = I\) and \(\det(U) = 1\) force
\(\lambda^2 = 1\) with \(|\lambda| = 1\), so \(\lambda = \pm 1\):
\[
\ker(\Phi) = \{I, -I\} \cong \mathbb{Z}/2\mathbb{Z}.
\]
By the
First Isomorphism Theorem
and the surjectivity established above,
\[
SU(2) / \{I, -I\} \cong SO(3).
\]
Equivalently, the fiber of \(\Phi\) over any \(R \in SO(3)\) is a single coset
\(U \cdot \{I, -I\} = \{U, -U\}\). The map is therefore exactly two-to-one.
As a set, and as a group, \(SO(3)\) is thus obtained from \(SU(2) \cong S^3\) by
identifying each element with its negative, the same identification that produces the
projective space \(\mathbb{RP}^3\) from \(S^3\). That the induced bijection from
\(\mathbb{RP}^3\) onto \(SO(3)\) is a homeomorphism is proved alongside the fundamental
group computation, as noted above.
Quaternions and 3D Graphics
The double cover \(SU(2) \to SO(3)\) is the mathematical foundation of
quaternion rotations in computer graphics and game engines. The
group \(SU(2)\) is isomorphic to the group of unit quaternions
\(\{q \in \mathbb{H} : |q| = 1\}\). Under this isomorphism the double cover
corresponds to the action of a unit quaternion \(q\) on \(\mathbb{R}^3\) by
\(\mathbf{v} \mapsto q \mathbf{v} \bar{q}\), where \(\mathbf{v}\) is identified
with a purely imaginary quaternion. The sign ambiguity, in which \(q\) and
\(-q\) yield the same rotation, is precisely the
\(\mathbb{Z}/2\mathbb{Z}\) kernel.
Quaternion representations of rotations are preferred over Euler angles in many
applications because they avoid gimbal lock, the degeneracy of
Euler angle parameterizations, and they admit smooth interpolation via
SLERP (Spherical Linear Interpolation) on \(S^3\). The
mathematical reason is that the covering map \(S^3 \to SO(3)\) is a local
diffeomorphism at every point, whereas an Euler-angle parameterization has
configurations where its differential drops rank. The shortest great-circle arc of
\(S^3\) between two unit quaternions is unique whenever they are not antipodal,
and replacing \(q\) by \(-q\), which gives the same rotation, avoids the
antipodal case.
The Adjoint Representations
The double cover construction above used a specific instance of a general operation:
the group \(G\) acting on its own Lie algebra \(\mathfrak{g}\) by conjugation. This
operation, the adjoint representation, is the most natural example
of a group representation and the bridge to the systematic study of representation
theory.
The Adjoint Representation of the Group
For each \(g \in G\), the conjugation map
\(C_g : G \to G\), \(C_g(h) = ghg^{-1}\), is a group automorphism. Recall that the
normal subgroups
of \(G\) are those satisfying \(gH = Hg\) for every \(g \in G\), equivalently
\(gHg^{-1} = H\). The adjoint
representation is the infinitesimal version. Instead of conjugating subgroups,
we conjugate Lie algebra elements.
Definition: The Adjoint Representation \(\mathrm{Ad}\)
Let \(G\) be a matrix Lie group with Lie algebra \(\mathfrak{g}\). For each
\(g \in G\), define the linear map
\[
\mathrm{Ad}(g) : \mathfrak{g} \to \mathfrak{g}, \quad
\mathrm{Ad}(g)(X) = gXg^{-1}.
\]
The map \(\mathrm{Ad} : G \to GL(\mathfrak{g})\), \(g \mapsto \mathrm{Ad}(g)\),
is a group homomorphism, the adjoint representation of \(G\).
Here \(GL(\mathfrak{g})\) denotes the group of invertible linear maps of
\(\mathfrak{g}\) to itself.
Verification:
\(\mathrm{Ad}(g)\) maps \(\mathfrak{g}\) to \(\mathfrak{g}\):
If \(X \in \mathfrak{g}\), then \(\exp(t \cdot gXg^{-1}) = g\exp(tX)g^{-1} \in G\)
for all \(t\) (since \(\exp(tX) \in G\) and \(G\) is closed under conjugation).
Hence \(gXg^{-1} \in \mathfrak{g}\).
\(\mathrm{Ad}(g)\) is linear:
\(\mathrm{Ad}(g)(aX + bY) = g(aX + bY)g^{-1} = a\,gXg^{-1} + b\,gYg^{-1} = a\,\mathrm{Ad}(g)(X) + b\,\mathrm{Ad}(g)(Y)\).
\(\mathrm{Ad}\) is a group homomorphism:
\(\mathrm{Ad}(g_1 g_2)(X) = (g_1 g_2)X(g_1 g_2)^{-1} = g_1(g_2 X g_2^{-1})g_1^{-1} = \mathrm{Ad}(g_1)(\mathrm{Ad}(g_2)(X))\),
so \(\mathrm{Ad}(g_1 g_2) = \mathrm{Ad}(g_1) \circ \mathrm{Ad}(g_2)\). Also,
\(\mathrm{Ad}(I)(X) = X\), so \(\mathrm{Ad}(I) = \mathrm{Id}\).
The double cover \(\Phi : SU(2) \to SO(3)\) constructed in the previous section is
precisely \(\mathrm{Ad} : SU(2) \to GL(\mathfrak{su}(2)) \cong GL(3, \mathbb{R})\),
whose image is \(SO(3)\). This illustrates the adjoint representation's
geometric content, namely how the group rotates its own infinitesimal generators.
The Adjoint Representation of the Lie Algebra
Differentiating a group homomorphism yields a Lie algebra homomorphism. Applied to
\(\mathrm{Ad} : G \to GL(\mathfrak{g})\), this general principle of the Lie
correspondence gives a map from \(\mathfrak{g}\) to \(\mathfrak{gl}(\mathfrak{g})\).
Definition: The Adjoint Representation \(\mathrm{ad}\)
The derivative of \(\mathrm{Ad}\) at the identity defines the
adjoint representation of the Lie algebra, a map
\(\mathrm{ad} : \mathfrak{g} \to \mathfrak{gl}(\mathfrak{g})\) into the Lie
algebra of all linear maps \(\mathfrak{g} \to \mathfrak{g}\) under the
commutator, given by
\[
\begin{align*}
\mathrm{ad}(X)(Y)
&= \left.\frac{d}{dt}\right|_{t=0} \mathrm{Ad}(\exp(tX))(Y) \\\\
&= \left.\frac{d}{dt}\right|_{t=0} \exp(tX)\,Y\,\exp(-tX).
\end{align*}
\]
By the product rule, this derivative equals the commutator, which
lies in \(\mathfrak{g}\) by
closure under the commutator:
\[
\mathrm{ad}(X)(Y) = XY - YX = [X, Y].
\]
The adjoint representation of the Lie algebra is simply the Lie bracket itself,
viewed as a linear map \(Y \mapsto [X, Y]\) for each fixed \(X\).
Theorem: \(\mathrm{ad}\) is a Lie Algebra Homomorphism
The map \(\mathrm{ad} : \mathfrak{g} \to \mathfrak{gl}(\mathfrak{g})\) is a
Lie algebra homomorphism:
\[
\mathrm{ad}([X, Y]) = [\mathrm{ad}(X),\, \mathrm{ad}(Y)]
= \mathrm{ad}(X) \circ \mathrm{ad}(Y) - \mathrm{ad}(Y) \circ \mathrm{ad}(X).
\]
Proof:
We must show that for all \(Z \in \mathfrak{g}\):
\[
\mathrm{ad}([X, Y])(Z) = \mathrm{ad}(X)(\mathrm{ad}(Y)(Z))
- \mathrm{ad}(Y)(\mathrm{ad}(X)(Z)).
\]
The left-hand side is \([[X, Y], Z]\). The right-hand side is
\([X, [Y, Z]] - [Y, [X, Z]]\). The claim is therefore:
\[
[[X, Y], Z] = [X, [Y, Z]] - [Y, [X, Z]].
\]
This claim is precisely the
Jacobi identity
rearranged. The standard form
\([X, [Y, Z]] + [Y, [Z, X]] + [Z, [X, Y]] = 0\) can be rewritten as
\(-[Z, [X, Y]] = [X, [Y, Z]] + [Y, [Z, X]]\), that is,
\([[X, Y], Z] = [X, [Y, Z]] - [Y, [X, Z]]\) (using antisymmetry
\([Z, X] = -[X, Z]\)).
The statement "\(\mathrm{ad}\) is a Lie algebra homomorphism" is equivalent
to the Jacobi identity. Put differently, the Jacobi identity says that for each fixed
\(X\) the map \([X, \,\cdot\,]\) is a derivation, one that satisfies a "product rule"
with respect to the bracket.
Conjugation by the Exponential
The group and algebra adjoints are joined by a single formula relating conjugation by
\(e^{X}\) to the exponential of \(\mathrm{ad}(X)\). It is the identity we will invoke
whenever a representation is conjugated by an exponential, and it makes the group-level
conjugation \(Y \mapsto e^{X} Y e^{-X}\) computable as a series of iterated brackets.
Theorem: Conjugation Is the Exponential of \(\mathrm{ad}\)
Take \(G = GL(n, \mathbb{C})\), whose Lie algebra is \(M_n(\mathbb{C})\), so
that both \(\mathrm{Ad}\) and \(\mathrm{ad}\) are defined on it. For any
\(X \in M_n(\mathbb{C})\), the map \(\mathrm{ad}(X) : M_n(\mathbb{C}) \to
M_n(\mathbb{C})\) is given by \(\mathrm{ad}(X)(Y) = [X, Y]\). Then for any
\(Y \in M_n(\mathbb{C})\),
\[
e^{X}\, Y\, e^{-X} = \mathrm{Ad}(e^{X})(Y) = e^{\mathrm{ad}(X)}(Y),
\]
where
\[
e^{\mathrm{ad}(X)}(Y) = Y + [X, Y] + \frac{1}{2!}\bigl[X, [X, Y]\bigr]
+ \frac{1}{3!}\Bigl[X, \bigl[X, [X, Y]\bigr]\Bigr] + \cdots.
\]
Proof:
The first equality is the definition of \(\mathrm{Ad}(e^{X})\). Conjugation
of \(Y\) by the group element \(e^{X}\) is
\(\mathrm{Ad}(e^{X})(Y) = e^{X} Y e^{-X}\).
For the second equality, fix \(X\) and consider the family of operators \(t \mapsto
\mathrm{Ad}(e^{tX})\) on the space \(M_n(\mathbb{C})\). Because \(\mathrm{Ad}\) is a
group homomorphism and \(t \mapsto e^{tX}\) is a
one-parameter subgroup,
\[
\mathrm{Ad}(e^{(s+t)X}) = \mathrm{Ad}(e^{sX})\,\mathrm{Ad}(e^{tX}),
\]
so \(t \mapsto \mathrm{Ad}(e^{tX})\) is a one-parameter subgroup of
\(GL(M_n(\mathbb{C}))\). Its generator at \(t = 0\) is, by the very definition of the
Lie algebra adjoint,
\[
\left.\frac{d}{dt}\right|_{t=0} \mathrm{Ad}(e^{tX}) = \mathrm{ad}(X).
\]
A one-parameter subgroup is the exponential of its generator, so
\(\mathrm{Ad}(e^{tX}) = e^{t\,\mathrm{ad}(X)}\). Setting \(t = 1\) gives
\(\mathrm{Ad}(e^{X}) = e^{\mathrm{ad}(X)}\), which is the claimed identity. Expanding the
operator exponential \(e^{\mathrm{ad}(X)} = \sum_{k \geq 0} \frac{1}{k!}\,
\mathrm{ad}(X)^k\) and applying it to \(Y\) produces the nested-bracket series, since
\(\mathrm{ad}(X)^k(Y)\) is the \(k\)-fold bracket \([X, [X, \ldots, [X, Y]\ldots]]\).
The formula turns a conjugation, a product of two exponentials around \(Y\), into a
single series of iterated brackets. Its payoff is in representation theory. When a
representation \(\pi\) is conjugated by \(e^{\pi(X)}\), the result is
\(e^{\mathrm{ad}(\pi(X))}(\pi(Y))\). For the small algebras whose brackets close after
a step or two this collapses to a short, exact expression.
Example: \(\mathrm{ad}\) for \(\mathfrak{so}(3)\)
Using the basis \(\{E_1, E_2, E_3\}\) of \(\mathfrak{so}(3)\) and its bracket
relations, we can write \(\mathrm{ad}(E_i)\) as a \(3 \times 3\) matrix with respect
to that basis.
For \(\mathrm{ad}(E_1)\) we have \(\mathrm{ad}(E_1)(E_1) = [E_1, E_1] = 0\),
\(\mathrm{ad}(E_1)(E_2) = [E_1, E_2] = E_3\), and
\(\mathrm{ad}(E_1)(E_3) = [E_1, E_3] = -E_2\). So, reading off the matrix of
\(\mathrm{ad}(E_1)\) with respect to the ordered basis
\((E_1, E_2, E_3)\):
\[
[\mathrm{ad}(E_1)] = \begin{pmatrix} 0 & 0 & 0 \\ 0 & 0 & -1 \\ 0 & 1 & 0 \end{pmatrix}.
\]
Similarly:
\[
\begin{align*}
[\mathrm{ad}(E_2)] &= \begin{pmatrix} 0 & 0 & 1 \\ 0 & 0 & 0 \\ -1 & 0 & 0 \end{pmatrix}, \\\\
[\mathrm{ad}(E_3)] &= \begin{pmatrix} 0 & -1 & 0 \\ 1 & 0 & 0 \\ 0 & 0 & 0 \end{pmatrix}.
\end{align*}
\]
The matrix \([\mathrm{ad}(E_i)]\) has the same numerical entries as the
matrix \(E_i\) itself. This is not an identity-map assertion. The operator
\(\mathrm{ad}(E_i)\) lives in \(\mathfrak{gl}(\mathfrak{so}(3))\), a space of linear
operators on a 3-dimensional vector space, while \(E_i \in \mathfrak{so}(3)\) is a
matrix acting on \(\mathbb{R}^3\). The two spaces of endomorphisms are distinct. What
the coincidence reflects is the
cross-product isomorphism
\((\mathbb{R}^3, \times) \cong (\mathfrak{so}(3), [\,\cdot\,,\,\cdot\,])\).
Concretely, the hat map identifies a vector
\(\boldsymbol{\omega} \in \mathbb{R}^3\) with the matrix
\(\hat{\boldsymbol{\omega}}_\times \in \mathfrak{so}(3)\) that implements
\(\mathbf{v} \mapsto \boldsymbol{\omega} \times \mathbf{v}\) on \(\mathbb{R}^3\).
Under this identification, the bracket
\(\mathrm{ad}(\hat{\boldsymbol{\omega}}_\times)(\hat{\mathbf{v}}_\times)
= [\hat{\boldsymbol{\omega}}_\times, \hat{\mathbf{v}}_\times]\) corresponds to
the vector \(\boldsymbol{\omega} \times \mathbf{v}\), so \(\mathrm{ad}\) on
\(\mathfrak{so}(3)\) becomes the cross-product map on \(\mathbb{R}^3\). The
matrix representing this map in the standard basis of \(\mathbb{R}^3\) is
precisely \(\hat{\boldsymbol{\omega}}_\times\), by definition of the hat map.
The two appearances of \(E_i\) (once as an element of \(\mathfrak{so}(3)\),
once as the matrix of \(\mathrm{ad}(E_i)\)) are thus the hat map applied to the
same standard basis vector \(\mathbf{e}_i \in \mathbb{R}^3\) on both sides.
The coincidence is an exceptional feature of dimension three. The dimensions of
\(\mathfrak{so}(3)\) and of the vector space it acts on both equal \(3\), so
\(\mathrm{ad}\) can be realised as a linear map \(\mathbb{R}^3 \to \mathbb{R}^3\).
No analogous coincidence occurs for \(\mathfrak{so}(n)\) with \(n \neq 3\),
since \(\dim \mathfrak{so}(n) = n(n-1)/2 \neq n\) in general.
Connections and Outlook
Lie Algebras in Deep Learning
In equivariant neural networks, layers must commute with the
action of a symmetry group \(G\) on their inputs and outputs, a condition called
equivariance. For instance, a network processing 3D point clouds
should produce the same output however the input is rotated (invariance under
\(SO(3)\), the special case of equivariance in which \(SO(3)\) acts trivially
on the output).
Enforcing equivariance at the group level requires checking the
constraint for every group element, an uncountable family of conditions. The Lie
correspondence provides a shortcut. For connected groups, equivariance of a linear
layer between representations of \(G\) is equivalent to
infinitesimal equivariance under the Lie algebra
\(\mathfrak{g}\). This reduces the problem to finitely many linear constraints
(one for each basis element of \(\mathfrak{g}\)), which can be incorporated
directly into the network architecture. Some equivariant architectures accordingly
impose the constraints at the level of \(\mathfrak{g}\) rather than through
explicit group elements.
Summary of the Four-Page Arc
We have developed the basic correspondence between matrix Lie groups and their Lie
algebras. Let us trace the logical arc:
Matrix Lie Groups defined the classical
groups as closed subgroups of \(GL(n)\), with Cartan's theorem guaranteeing smooth
manifold structure.
The Matrix Exponential provided
the bridge between linear data and nonlinear group elements, establishing one-parameter
subgroups as the "straight lines" in the group.
Lie Algebras and the Lie Bracket
formalized the tangent space at the identity as a Lie algebra, a vector space with a
bracket encoding non-commutativity.
The present page established the Lie correspondence: the algebra
determines the group locally (BCH formula), group homomorphisms induce algebra
homomorphisms, and the adjoint representations describe the group's action on its own
infinitesimal generators.
Where We Go from Here
Representation Theory. The adjoint representations \(\mathrm{Ad}\)
and \(\mathrm{ad}\) are examples of
group and algebra representations, homomorphisms from \(G\) to
\(GL(V)\), or from \(\mathfrak{g}\) to \(\mathfrak{gl}(V)\), for some vector space
\(V\). Representation theory studies these systematically: which vector spaces can
\(G\) act on? When can a representation be decomposed into simpler pieces (irreducible
representations)? For compact matrix Lie groups, the pages that follow show that every
finite-dimensional representation splits into irreducible pieces and that Schur's lemma
governs the maps between those pieces. For \(\mathfrak{sl}(2,\mathbb{C})\) they classify
the finite-dimensional irreducible representations completely.
Smooth Manifolds. The
tangent space \(T_I G\), defined here concretely as velocity vectors of curves in
\(G \subseteq M_n(\mathbb{C})\), will be generalized to the tangent space \(T_p M\) at
any point of an abstract smooth manifold. For compact groups such as \(SO(n)\) and
\(SU(n)\), the group exponential coincides with the Riemannian exponential map at
\(I\) of a bi-invariant metric.
Equivariant Neural Networks. The
equivariant networks
page builds rotation-equivariant layers for three-dimensional data from the
representation theory of \(SO(3)\): irreducible representations, Schur's lemma, and
the Clebsch-Gordan decomposition. The Lie algebra \(\mathfrak{so}(3)\) returns there
through its complexification \(\mathfrak{sl}(2,\mathbb{C})\), the setting in which the
Clebsch-Gordan theorem is stated.