From Fourier to Compact Groups
Fourier analysis on the circle rests on a single structural fact. The exponentials
\(e^{in\theta}\) form an
orthonormal basis
of the square-integrable functions on \(S^1\). Every periodic signal decomposes into these
elementary frequencies, and the decomposition is orthogonal, complete, and norm-preserving. The
exponentials are not an arbitrary choice. The circle is a group under addition of angles, and each
\(e^{in\theta}\) is a one-dimensional representation of that group, since the map
\(\theta \mapsto e^{in\theta}\) sends the group operation to multiplication of complex numbers.
Classical Fourier analysis is, in this light, the harmonic analysis of the abelian group \(S^1\),
and the basis functions are exactly its irreducible representations.
The aim of this page is to identify the analogue of this basis for a general compact group, where
the group need not be abelian and its irreducible representations need not be one-dimensional. The
replacement for the exponentials will be the matrix entries of the irreducible
representations, and the replacement for the orthonormal basis of \(L^2(S^1)\) will be a basis of
\(L^2(K)\) assembled from those entries. The central results are the orthogonality of characters and
the completeness theorem of Peter and Weyl, which together promote the circle's Fourier expansion to
a noncommutative harmonic analysis valid on every compact matrix Lie group.
Throughout, \(K\) denotes a compact matrix Lie group, that is, a compact closed subgroup of
\(GL(n, \mathbb{C})\). Such a group is by
Cartan's closed subgroup theorem
a smooth embedded submanifold of \(GL(n, \mathbb{C})\) whose group operations are smooth, hence a
Lie group
in the sense of manifold theory, and every construction of that theory applies to it. Being a subset
of \(M_n(\mathbb{C})\) it is metrizable, so it is also compact Hausdorff.
Integration over \(K\) is taken against the normalized
Haar integral,
written \(\int_K f(x)\, dx\). It is the integral against the
Haar volume form
\(\omega_K\), the unique left-invariant top-degree form on \(K\) of total integral \(1\). The
hypothesis of that result, a left-invariant orientation, is met here. The group admits a
left-invariant global frame,
and the top-degree form taking the value \(1\) on that frame at every point is left-invariant and
nowhere vanishing, so it
determines an orientation,
left-invariant because the form is. Normalization means \(\int_K 1\, dx = 1\), and left invariance
means that \(\int_K f(yx)\, dx = \int_K f(x)\, dx\) for every fixed \(y \in K\).
The integral is right invariant as well. Write \(R_y\) for
right translation
by \(y \in K\), a diffeomorphism of \(K\) with inverse \(R_{y^{-1}}\). Since
\(R_y \circ L_g = L_g \circ R_y\) and \(L_g^{*}\omega_K = \omega_K\), the pullback
\(R_y^{*}\omega_K\) is again left-invariant. Top-degree alternating forms on a \(k\)-dimensional
space form a
one-dimensional space,
here with \(k = \dim K\), so \(R_y^{*}\omega_K = c\,\omega_K\) for some function \(c\) on \(K\), and
pulling back by \(L_g\) turns the last identity into \((c \circ L_g)\,\omega_K = c\,\omega_K\),
forcing \(c\) to be constant. That constant is nonzero because \(\omega_K\) vanishes nowhere, so
\(R_y\) is orientation-preserving when \(c \gt 0\) and orientation-reversing when \(c \lt 0\), and
part (d) of the
properties of the integral
gives
\[
c = \int_K R_y^{*}\omega_K = \operatorname{sgn}(c) \int_K \omega_K = \operatorname{sgn}(c),
\]
so \(c = \pm 1\). Applying the same part to \(f\omega_K\), whose pullback is
\(c\,(f \circ R_y)\,\omega_K\),
\[
c \int_K f(xy)\, dx = \int_K R_y^{*}(f\omega_K) = \operatorname{sgn}(c) \int_K f(x)\, dx = c \int_K f(x)\, dx,
\]
and dividing by \(c\) leaves \(\int_K f(xy)\, dx = \int_K f(x)\, dx\). Conjugation
\(x \mapsto yxy^{-1}\) is a left translation followed by a right translation, so the integral is
invariant under it too. We use these invariances repeatedly and without further comment.
The Haar integral also determines a measure, which is what gives meaning to \(L^2(K)\). The integral is
positive, since for continuous \(f \geq 0\) not identically zero the form \(f\omega_K\) is
everywhere non-negative and strictly positive somewhere, so part (c) of the
properties of the integral
makes \(\int_K f\, dx \gt 0\). Boundedness follows by applying positivity to
\(\sup_K |f| - \operatorname{Re}(\alpha f)\), with \(\alpha\) a unimodular scalar chosen so that
\(\alpha \int_K f\, dx = \left| \int_K f\, dx \right|\), which yields
\(\left| \int_K f\, dx \right| \leq \sup_K |f|\). Because \(K\) is compact, \(C_0(K) = C(K)\), so
the
representation of positive functionals by measures
presents the Haar integral as integration against a unique regular Borel measure \(\mu_K\) on \(K\),
with \(\mu_K(K) = \int_K 1\, dx = 1\). We write \(L^2(K)\) for the
\(L^2\) space
of this measure and keep the notation \(\int_K f(x)\, dx\) for integration against it. That \(C(K)\)
is dense in \(L^2(K)\) is a standard fact about regular Borel measures on compact spaces, which we
take as given.
Characters and Class Functions
The objects that play the role of frequencies are built from traces of representations. We work
throughout with finite-dimensional
representations
of \(K\) on complex vector spaces. Recall that every such representation of a compact group may be
taken to be unitary with respect to an inner product obtained by
averaging an arbitrary one over the group.
Definition: Character of a Representation
Let \((\Pi, V)\) be a finite-dimensional representation of a compact matrix Lie group \(K\). The
character of \(\Pi\) is the function \(\chi_\Pi : K \to \mathbb{C}\) defined by
\[
\chi_\Pi(x) = \operatorname{tr}\big(\Pi(x)\big).
\]
In particular \(\chi_\Pi(e) = \operatorname{tr}(I) = \dim V\), so the character records the dimension of
the representation at the identity.
The character discards all information about \(\Pi\) except what survives the trace, yet it
retains enough to detect the representation completely. Two features make it the right object.
First, it is unchanged under isomorphism of representations, because the trace is invariant under
conjugation. If \(\Pi' = S\Pi S^{-1}\) then
\(\operatorname{tr}(\Pi'(x)) = \operatorname{tr}(S\Pi(x)S^{-1}) = \operatorname{tr}(\Pi(x))\).
Second, and for the same reason, the character is constant on conjugacy classes of \(K\). This
second property is fundamental enough to name.
Definition: Class Function
A function \(f : K \to \mathbb{C}\) is a class function if it is constant on conjugacy
classes, that is, if
\[
f(yxy^{-1}) = f(x) \quad \text{for all } x, y \in K.
\]
Equivalently, \(f\) is invariant under the conjugation action of \(K\) on itself.
Every character is a class function. The verification is immediate from the multiplicativity of \(\Pi\) and the
cyclic invariance of the trace: for any \(x, y \in K\),
\[
\chi_\Pi(yxy^{-1}) = \operatorname{tr}\big(\Pi(y)\Pi(x)\Pi(y)^{-1}\big) = \operatorname{tr}\big(\Pi(x)\big) = \chi_\Pi(x),
\]
where the middle equality uses \(\operatorname{tr}(ABA^{-1}) = \operatorname{tr}(B)\). The continuous
class functions form a subspace of \(L^2(K)\), and the orthogonality theorem will show that the
characters of the irreducible representations sit inside this subspace as an orthonormal system.
The Peter-Weyl theorem will then show that this system is complete. It is to the space of class
functions what the exponentials \(e^{in\theta}\) are to \(L^2(S^1)\). On the circle the two
pictures coincide, because every element is its own conjugacy class and so every function is a
class function. That degeneration is one we return to once the general theorem is in hand.
Orthonormality of Characters
We equip \(L^2(K)\) with the inner product
\[
\langle f, g \rangle = \int_K \overline{f(x)}\, g(x)\, dx,
\]
linear in the second argument. The orthogonality theorem states that the characters of the
irreducible representations form an orthonormal set with respect to this product. Distinct
irreducibles have orthogonal characters, and each irreducible character has unit norm. This is the
exact analogue of the relation \(\langle e^{im\theta}, e^{in\theta} \rangle = \delta_{mn}\)
underlying ordinary Fourier series, now without any assumption of commutativity.
Theorem: Orthonormality of Irreducible Characters
Let \((\Pi, V)\) and \((\Sigma, W)\) be irreducible representations of a compact matrix Lie group \(K\).
Then
\[
\langle \chi_\Pi, \chi_\Sigma \rangle = \int_K \overline{\chi_\Pi(x)}\, \chi_\Sigma(x)\, dx =
\begin{cases} 1 & \text{if } \Pi \cong \Sigma, \\\\ 0 & \text{if } \Pi \not\cong \Sigma. \end{cases}
\]
The proof rests on two preliminary facts. The first turns the Haar integral of a representation
into a concrete geometric object, an orthogonal projection. The second computes the trace of a
tensor product. After establishing these we assemble the orthogonality relation by recognizing the
integral above as the dimension of a space of intertwining maps, which Schur's lemma then
evaluates.
Averaging as a Projection
For any finite-dimensional representation \((\Pi, V)\) of \(K\), write
\[
V^K = \{\, v \in V : \Pi(x)v = v \text{ for all } x \in K \,\}
\]
for the subspace of vectors fixed by every group element. Averaging the operators \(\Pi(x)\) over the group
produces a projection onto this subspace.
Proposition: The Averaging Projection
Let \((\Pi, V)\) be a finite-dimensional representation of a compact matrix Lie group \(K\), and define
the operator
\[
P = \int_K \Pi(x)\, dx \in \operatorname{End}(V),
\]
the integral taken entrywise in any basis. Then \(P\) is a projection of \(V\) onto \(V^K\).
It maps \(V\) into \(V^K\), and \(Pv = v\) for every \(v \in V^K\). Consequently
\(\operatorname{tr}(P) = \dim V^K\).
Proof
\(P\) maps into \(V^K\).
Fix \(y \in K\) and \(v \in V\). Using the linearity of \(\Pi(y)\), the homomorphism property
\(\Pi(y)\Pi(x) = \Pi(yx)\), and the left-invariance of the Haar integral,
\[
\Pi(y)\, Pv = \Pi(y) \int_K \Pi(x)v\, dx = \int_K \Pi(yx)v\, dx = \int_K \Pi(x)v\, dx = Pv.
\]
The third equality substitutes \(x \mapsto yx\), under which \(dx\) is invariant. Since this holds for every
\(y \in K\), the vector \(Pv\) is fixed by the representation, so \(Pv \in V^K\).
\(P\) fixes \(V^K\) pointwise.
If \(v \in V^K\) then \(\Pi(x)v = v\) for all \(x\), so
\[
Pv = \int_K \Pi(x)v\, dx = \int_K v\, dx = \left( \int_K 1\, dx \right) v = v,
\]
using the normalization \(\int_K 1\, dx = 1\). The two properties together say that \(P\) lands
in \(V^K\) and restricts to the identity there, so \(P^2 = P\) and the image of \(P\) is exactly
\(V^K\). Consequently \(V = V^K \oplus \ker P\), because every \(v\) splits as \(Pv + (v - Pv)\)
with \(P(v - Pv) = 0\), while a vector lying in both summands is at once fixed and killed by
\(P\). In a basis adapted to this splitting the matrix of \(P\) is \(\operatorname{diag}(I, 0)\)
with an identity block of size \(\dim V^K\), so \(\operatorname{tr}(P) = \dim V^K\).
When \(\Pi\) is irreducible and is not the trivial representation, the only vector fixed by all of
\(K\) is the origin. The subspace \(V^K\) is invariant, it is not all of \(V\) because \(\Pi\) is
not trivial, and irreducibility leaves only \(V^K = \{0\}\). The proposition then records that
\(\int_K \Pi(x)\, dx = 0\).
The Trace of a Tensor Product
The orthogonality integral pairs two representations, and the product of their traces will be read as the trace
of a single representation on the
tensor product
space. The underlying linear-algebra identity is the following, an immediate consequence of the
characteristic property of the tensor product,
which guarantees that \(A \otimes B\) is the well-defined operator acting by
\((A \otimes B)(v \otimes w) = Av \otimes Bw\). That property is stated there for real vector spaces,
and both its statement and its proof use only the vector space operations, so it transfers verbatim
to complex ones.
The tensor-trace identity
For operators \(A\) on \(V\) and \(B\) on \(W\), one has \(\operatorname{tr}(A \otimes B) = \operatorname{tr}(A)\,\operatorname{tr}(B)\).
To see this, choose bases \(\{v_j\}\) of \(V\) and \(\{w_l\}\) of \(W\), so that \(\{v_j \otimes w_l\}\) is a
basis of \(V \otimes W\). The matrix of \(A \otimes B\) in this basis has diagonal entries \(A_{jj} B_{ll}\),
whence
\[
\operatorname{tr}(A \otimes B) = \sum_{j, l} A_{jj} B_{ll} = \left(\sum_j A_{jj}\right)\!\left(\sum_l B_{ll}\right) = \operatorname{tr}(A)\,\operatorname{tr}(B).
\]
Assembling the Orthogonality Relation
Two further ingredients connect the conjugate character to a dual representation. Fix an orthonormal
basis for the invariant inner product recalled above, so that \(\Pi(x)\) is unitary and
\(\Pi(x)^{-1}\) is its conjugate transpose. Then \([\Pi(x^{-1})]^\top = \overline{\Pi(x)}\), the
entrywise conjugate, so the matrix of the
dual representation,
which acts by \(\Pi^*(x) = [\Pi(x^{-1})]^\top\), is the entrywise conjugate of the matrix of
\(\Pi(x)\). Taking traces,
\[
\overline{\chi_\Pi(x)} = \overline{\operatorname{tr}(\Pi(x))} = \operatorname{tr}\big(\overline{\Pi(x)}\big) = \operatorname{tr}\big(\Pi^*(x)\big) = \chi_{\Pi^*}(x).
\]
The conjugate of a character is therefore the character of the dual representation. We can now prove
the theorem.
Proof of Orthonormality
From the pairing to a tensor character.
Combining the conjugate-character identity with the tensor-trace identity, the integrand of
\(\langle \chi_\Pi, \chi_\Sigma \rangle\) becomes the character of the tensor product \(\Pi^* \otimes \Sigma\):
\[
\overline{\chi_\Pi(x)}\, \chi_\Sigma(x) = \operatorname{tr}\big(\Pi^*(x)\big)\, \operatorname{tr}\big(\Sigma(x)\big) = \operatorname{tr}\big((\Pi^* \otimes \Sigma)(x)\big).
\]
Integrating gives the dimension of an invariant subspace.
Integrating over \(K\) and exchanging trace with the entrywise Haar integral, then applying the averaging
projection to the representation \(\Pi^* \otimes \Sigma\) on \(V^* \otimes W\),
\[
\langle \chi_\Pi, \chi_\Sigma \rangle = \int_K \operatorname{tr}\big((\Pi^* \otimes \Sigma)(x)\big)\, dx = \operatorname{tr}\!\left( \int_K (\Pi^* \otimes \Sigma)(x)\, dx \right) = \operatorname{tr}(P) = \dim\big( (V^* \otimes W)^K \big),
\]
where \(P\) is the projection onto the fixed subspace \((V^* \otimes W)^K\).
Identifying the fixed subspace with intertwiners.
The space \(V^* \otimes W\) is identified with \(\operatorname{Hom}(V, W)\), the linear maps
from \(V\) to \(W\), by sending \(\varphi \otimes w\) to the map \(v \mapsto \varphi(v)\, w\).
This is a linear isomorphism, since it carries the basis assembled from a basis of \(V^*\) and
one of \(W\) to the matrix units of \(\operatorname{Hom}(V, W)\). Under it the action
\(\Pi^*(x) \otimes \Sigma(x)\) of \(x \in K\) becomes \(T \mapsto \Sigma(x)\, T\, \Pi(x)^{-1}\),
because \(\Pi^*(x)\varphi = \varphi \circ \Pi(x)^{-1}\). A map is fixed by every \(x\) precisely
when \(\Sigma(x) T = T \Pi(x)\) for all \(x\), that is, precisely when \(T\) is an intertwining
map between \(\Pi\) and \(\Sigma\). Hence
\[
(V^* \otimes W)^K \cong \operatorname{Hom}_K(V, W),
\]
the space of intertwiners.
Schur's lemma evaluates the dimension.
By
Schur's lemma
in its first part, a nonzero intertwiner between irreducibles is an isomorphism, so
\(\operatorname{Hom}_K(V, W) = 0\) when \(\Pi \not\cong \Sigma\). When \(\Pi \cong \Sigma\), fix
an isomorphism \(S\) between them. The third part of the same lemma makes every nonzero
intertwiner a scalar multiple of \(S\), so \(\operatorname{Hom}_K(V, W) = \mathbb{C}S\) is
one-dimensional. Therefore
\[
\langle \chi_\Pi, \chi_\Sigma \rangle = \dim \operatorname{Hom}_K(V, W) = \begin{cases} 1 & \Pi \cong \Sigma, \\\\ 0 & \Pi \not\cong \Sigma, \end{cases}
\]
which is the assertion of the theorem.
The orthonormality relation already separates representations. Two irreducibles with the same
character are isomorphic, because \(\langle \chi_\Pi, \chi_\Sigma \rangle = 1\) forces
\(\Pi \cong \Sigma\). What remains is the deeper question of completeness. Do the characters
exhaust the class functions, leaving no direction in \(L^2(K)\) orthogonal to all of them? That is
the content of the Peter-Weyl theorem, to which we now turn.
Matrix Entries and the Peter-Weyl Theorem
Completeness cannot be proved at the level of characters alone. A character is a single function
extracted from a representation by the trace, and the trace collapses too much. The route to
completeness passes instead through the full collection of matrix coefficients of a
representation, of which the character is the special case formed by summing the diagonal. We
isolate these coefficients, prove that they are dense in \(L^2(K)\), and only then descend to
characters by averaging over conjugacy.
Definition: Matrix Entries of a Representation
Let \((\Pi, V)\) be a finite-dimensional representation of \(K\) and fix a basis \(\{v_j\}\) of \(V\). The
matrix entries of \(\Pi\) are the functions \(K \to \mathbb{C}\) given by
\[
x \longmapsto \big(\Pi(x)\big)_{jk},
\]
the \((j,k)\) entry of the matrix of \(\Pi(x)\) in the basis \(\{v_j\}\). More generally, and independently of
any basis, a function of the form
\[
f(x) = \operatorname{tr}\big(\Pi(x) A\big), \quad A \in \operatorname{End}(V),
\]
is called a matrix entry for \(\Pi\). Choosing \(A\) to be the matrix unit with a single \(1\)
in position \((k,j)\) recovers the coordinate function \((\Pi(x))_{jk}\), and a general \(A\)
produces a linear combination of coordinate functions. Taking \(A = I\) gives the character
\(\chi_\Pi\).
The basis-free description \(f(x) = \operatorname{tr}(\Pi(x)A)\) is the one we use, because it transforms cleanly
under the group operations. Of these, the single fact the completeness proof relies on is how an integral of
conjugated operators behaves on an irreducible representation.
Conjugation-averaging on an irreducible
If \((\Pi, V)\) is irreducible, then for every operator \(A\) on \(V\),
\[
\int_K \Pi(y)\, A\, \Pi(y)^{-1}\, dy = c\, I \quad \text{for some scalar } c,
\]
and taking traces on both sides fixes the scalar as \(c = \operatorname{tr}(A)/\dim V\). Here
the integral commutes with the trace, and
\(\operatorname{tr}(\Pi(y)A\Pi(y)^{-1}) = \operatorname{tr}(A)\) is constant in \(y\). To see
that the integral is a scalar, let \(B\) denote the operator on the left. For any \(x \in K\),
the left-invariance of the Haar integral gives
\[
\Pi(x)\, B\, \Pi(x)^{-1} = \int_K \Pi(xy)\, A\, \Pi(xy)^{-1}\, dy = \int_K \Pi(y)\, A\, \Pi(y)^{-1}\, dy = B,
\]
so \(B\) commutes with every \(\Pi(x)\). By
Schur's lemma
a self-intertwiner of an irreducible complex representation is a scalar, hence \(B = cI\).
The Algebra of Matrix Entries
Let \(\mathcal{A} \subseteq C(K)\) be the set of all functions expressible as a finite linear combination of
matrix entries, ranging over all finite-dimensional representations of \(K\). The Peter-Weyl theorem asserts that
\(\mathcal{A}\) is uniformly dense in \(C(K)\), and therefore dense in \(L^2(K)\). We prove this by verifying that
\(\mathcal{A}\) satisfies the hypotheses of the
Stone-Weierstrass theorem
on the compact Hausdorff space \(K\): that it is a self-adjoint subalgebra containing the constants and separating
points. The uniform closure of such an algebra is then all of \(C(K)\).
Theorem: Peter-Weyl (Density of Matrix Entries)
Let \(K\) be a compact matrix Lie group. The space \(\mathcal{A}\) of finite linear combinations of matrix
entries of finite-dimensional representations of \(K\) is dense in \(C(K)\) in the uniform norm, and hence
dense in \(L^2(K)\).
Proof
\(\mathcal{A}\) is an algebra.
By definition \(\mathcal{A}\) is the span of the matrix entries, hence a linear subspace. For
products, the tensor-trace identity gives, for representations \(\Pi, \Sigma\) and operators
\(A, B\),
\[
\operatorname{tr}\big(\Pi(x)A\big)\, \operatorname{tr}\big(\Sigma(x)B\big) = \operatorname{tr}\big((\Pi \otimes \Sigma)(x)\, (A \otimes B)\big),
\]
since \((\Pi(x)A) \otimes (\Sigma(x)B) = (\Pi \otimes \Sigma)(x)\,(A \otimes B)\). The
right-hand side is a matrix entry of the tensor product representation \(\Pi \otimes \Sigma\),
so the product of two elements of \(\mathcal{A}\) lies again in \(\mathcal{A}\) and
\(\mathcal{A}\) is closed under multiplication.
\(\mathcal{A}\) contains the constants.
The trivial representation \(x \mapsto 1\) on \(\mathbb{C}\) has the constant function \(1\) as its matrix
entry, so every constant lies in \(\mathcal{A}\).
\(\mathcal{A}\) is self-adjoint.
The complex conjugate of a matrix entry is a matrix entry of the dual representation. With
\(\Pi\) unitary in an orthonormal basis, as arranged above, the matrix of \(\Pi^*(x)\) is the
entrywise conjugate of that of \(\Pi(x)\), so
\[
\overline{\operatorname{tr}\big(\Pi(x)A\big)} = \operatorname{tr}\big(\overline{\Pi(x)}\,\overline{A}\big) = \operatorname{tr}\big(\Pi^*(x)\,\overline{A}\big),
\]
a matrix entry for \(\Pi^*\). Hence \(\mathcal{A}\) is closed under complex conjugation.
\(\mathcal{A}\) separates points.
This is the one step that uses the
hypothesis that \(K\) is a matrix Lie group. By definition such a group comes with a faithful
finite-dimensional representation \(\Pi_0\), namely the inclusion
\(K \hookrightarrow GL(n, \mathbb{C})\) itself. If \(x \neq y\) in \(K\), then
\(\Pi_0(x) \neq \Pi_0(y)\) by faithfulness, so some matrix entry \((\Pi_0)_{jk}\) takes
different values at \(x\) and \(y\). Thus \(\mathcal{A}\) separates the points of \(K\).
Conclusion via Stone-Weierstrass.
Multiplication and conjugation are continuous for the uniform norm, so the uniform closure
\(\overline{\mathcal{A}}\) is again a subalgebra closed under conjugation, and it contains the
constants and separates points because \(\mathcal{A}\) does. The space \(K\) is compact
Hausdorff, so the
Stone-Weierstrass theorem
gives \(\overline{\mathcal{A}} = C(K)\). Finally \(\|g\|_2 \leq \|g\|_\infty\) because
\(\mu_K(K) = 1\), so uniform approximation implies approximation in \(L^2(K)\), and \(C(K)\) is
dense there. Hence \(\mathcal{A}\) is dense in \(L^2(K)\) as well.
From Matrix Entries to Characters
Density of matrix entries is the Peter-Weyl theorem in its primary form. Completeness of
characters is the statement that a continuous class function orthogonal to every irreducible
character must vanish, and it follows by projecting matrix entries onto the class functions. The
bridge is an averaging over conjugacy that sends each matrix entry to a linear combination of
characters.
Corollary: Completeness of Characters
Let \(K\) be a compact matrix Lie group. If \(f\) is a continuous class function on \(K\) satisfying
\(\langle \chi_\Pi, f \rangle = 0\) for every irreducible representation \(\Pi\), then
\(f = 0\). Equivalently, no nonzero continuous class function on \(K\) is orthogonal to every
irreducible character.
Proof
Conjugation-averaging sends matrix entries to characters.
For a function \(g \in \mathcal{A}\), define its average over conjugacy classes,
\[
(\mathcal{C}g)(x) = \int_K g\big(y^{-1} x y\big)\, dy.
\]
The result is a class function, by the invariance of the Haar integral under \(y \mapsto zy\). Applying this
to a single matrix entry \(g(x) = \operatorname{tr}(\Pi(x) A)\) and using the homomorphism property,
\[
(\mathcal{C}g)(x) = \int_K \operatorname{tr}\big(\Pi(y^{-1}) \Pi(x) \Pi(y) A\big)\, dy = \operatorname{tr}\!\left( \Pi(x) \int_K \Pi(y)\, A\, \Pi(y)^{-1}\, dy \right),
\]
where the second equality uses the cyclic invariance of the trace and
\(\Pi(y^{-1}) = \Pi(y)^{-1}\). If \(\Pi\) is irreducible, the conjugation-averaging identity
evaluates the inner integral as \((\operatorname{tr} A / \dim V)\, I\), so that \(\mathcal{C}g\)
is the multiple \((\operatorname{tr} A / \dim V)\, \chi_\Pi\) of a single irreducible character.
A general \(\Pi\) reduces to that case. Being a finite-dimensional representation of a compact
matrix Lie group it is
completely reducible,
which in its internal form
means that \(V = U_1 \oplus \cdots \oplus U_r\) with every \(U_i\) invariant and irreducible,
the restriction \(\Pi_i\) of \(\Pi\) to \(U_i\) being a representation in its own right. The
projection \(p_i\) onto \(U_i\) along the remaining summands commutes with every \(\Pi(x)\),
because \(\Pi(x)\) carries each summand into itself. Writing \(A_i\) for the operator
\(p_i A|_{U_i}\) on \(U_i\) and using \(\sum_i p_i = I\) together with cyclicity of the trace,
\[
\operatorname{tr}\big(\Pi(x) A\big) = \sum_{i=1}^{r} \operatorname{tr}\big(\Pi(x)\, p_i A p_i\big) = \sum_{i=1}^{r} \operatorname{tr}\big(\Pi_i(x)\, A_i\big).
\]
Every summand is a matrix entry for an irreducible representation, so the previous case applies
to it. Since \(\mathcal{C}\) is linear and every element of \(\mathcal{A}\) is a finite
combination of matrix entries, conjugation-averaging carries \(\mathcal{A}\) into the span of
the irreducible characters.
A class function orthogonal to all characters is orthogonal to its own approximants.
Suppose \(f\) is a continuous class function with \(\langle \chi_\Pi, f \rangle = 0\) for every irreducible
\(\Pi\). By the density of \(\mathcal{A}\), choose \(g_n \in \mathcal{A}\) converging to \(f\) in \(L^2(K)\).
Each conjugation-average \(\mathcal{C}g_n\) is a linear combination of characters, so
\(\langle \mathcal{C}g_n, f \rangle = 0\). Because \(f\) is itself a class function, averaging the integrand of
\(\langle g_n, f \rangle\) over conjugacy does not change its value:
\[
\langle g_n, f \rangle = \int_K \overline{g_n(x)}\, f(x)\, dx = \int_K \overline{(\mathcal{C}g_n)(x)}\, f(x)\, dx = \langle \mathcal{C}g_n, f \rangle = 0.
\]
The middle equality is where the class-function property of \(f\) enters. Written out,
\(\langle \mathcal{C}g_n, f \rangle\) is an integral over \(K \times K\) of
\(\overline{g_n(y^{-1}xy)}\, f(x)\), a continuous and therefore bounded function, measurable for
the product \(\sigma\)-algebra because \(K\) is a compact metric space and hence second
countable. The measure \(\mu_K \otimes \mu_K\) being finite,
Fubini's theorem
allows integrating in \(x\) first. For fixed \(y\) the substitution \(x \mapsto y x y^{-1}\),
under which \(dx\) is invariant, turns \(\int_K \overline{g_n(y^{-1}xy)}\, f(x)\, dx\) into
\(\int_K \overline{g_n(x)}\, f(y x y^{-1})\, dx\), which equals \(\langle g_n, f \rangle\)
because \(f\) is a class function. Averaging over \(y\) leaves that value unchanged.
Passing to the limit.
Since \(g_n \to f\) in \(L^2(K)\), we have
\(\langle g_n, f \rangle \to \langle f, f \rangle = \|f\|_2^2\). But every
\(\langle g_n, f \rangle = 0\), so \(\|f\|_2 = 0\) and \(f = 0\) almost everywhere. Being
continuous, \(f\) is identically zero. The irreducible characters therefore leave no nonzero
class function orthogonal to all of them, which is their completeness.
The Peter-Weyl Decomposition and the Window to GDL
The density of matrix entries upgrades at once to an orthonormal basis. The space \(L^2(K)\) is a
Hilbert space,
carrying the inner product \(\langle f, g \rangle = \int_K \overline{f}\, g\, dx\) and complete by
the
Riesz-Fischer theorem.
We call a family in a Hilbert space an orthonormal basis when its members are orthonormal and their
finite linear combinations are dense, which is the condition of the
definition for sequences
with the countability of the family dropped. The orthogonality relations for the matrix entries
supply the first half of that condition and the density theorem supplies the second, and together
they yield the decomposition that is the harmonic-analytic core of the theory.
Theorem: The Peter-Weyl Decomposition of \(L^2(K)\)
Let \(K\) be a compact matrix Lie group, and let \(\widehat{K}\) be a set of representatives of
the isomorphism classes of irreducible unitary representations. For each \(\Pi \in \widehat{K}\)
of dimension \(d_\Pi\), fix an orthonormal basis of its space for a \(K\)-invariant inner
product, giving matrix entries \((\Pi(x))_{jk}\). Then the rescaled matrix entries
\[
\sqrt{d_\Pi}\,(\Pi(x))_{jk}, \quad \Pi \in \widehat{K}, \quad 1 \leq j, k \leq d_\Pi,
\]
form an orthonormal basis of \(L^2(K)\). Equivalently, \(L^2(K)\) decomposes as the Hilbert-space direct sum
\[
L^2(K) = \bigoplus_{\Pi \in \widehat{K}} \mathcal{E}_\Pi,
\]
where \(\mathcal{E}_\Pi\) is the \(d_\Pi^2\)-dimensional space spanned by the entries of
\(\Pi\). The summands are mutually orthogonal and their finite sums are dense, which is what the
direct sum asserts here.
Proof
Orthogonality and normalization.
The relations come from the same mechanism that gave the character relations, applied to the
full coefficients rather than their traces. Let \((\Pi, V)\) and \((\Sigma, W)\) be irreducible
members of \(\widehat{K}\), with the chosen orthonormal bases, and for a linear map
\(T : V \to W\) set
\[
T^{\#} = \int_K \Sigma(y)\, T\, \Pi(y)^{-1}\, dy,
\]
the integral taken entrywise. Left invariance gives
\(\Sigma(x)\, T^{\#}\, \Pi(x)^{-1} = T^{\#}\) for every \(x \in K\), so \(T^{\#}\) is an
intertwining map. Unitarity in these bases turns \((\Pi(y)^{-1})_{kj}\) into
\(\overline{\Pi(y)_{jk}}\). Let \(E_{mk}\) carry the \(k\)-th basis vector of \(V\) to the
\(m\)-th of \(W\) and the others to zero. Then
\[
\big(E_{mk}^{\#}\big)_{lj} = \int_K \Sigma(y)_{lm}\, \overline{\Pi(y)_{jk}}\, dy = \big\langle (\Pi)_{jk}, (\Sigma)_{lm} \big\rangle.
\]
If \(\Pi \not\cong \Sigma\), the first part of Schur's lemma forces \(E_{mk}^{\#} = 0\), so
entries of inequivalent irreducibles are orthogonal in \(L^2(K)\). If \(\Sigma = \Pi\), the
conjugation-averaging identity evaluates \(E_{mk}^{\#}\) as
\((\operatorname{tr} E_{mk}/d_\Pi)\, I = (\delta_{km}/d_\Pi)\, I\), whose \((l, j)\) entry is
\(\delta_{km}\delta_{jl}/d_\Pi\). Within a single irreducible \(\Pi\) we therefore have
\[
\big\langle (\Pi)_{jk}, (\Pi)_{lm} \big\rangle = \frac{1}{d_\Pi}\, \delta_{jl}\, \delta_{km},
\]
and rescaling each entry by \(\sqrt{d_\Pi}\) produces an orthonormal family. All these pairings
are finite because \(K\) is compact, so every continuous function lies in \(L^2(K)\) and the
Cauchy-Schwarz inequality
controls them.
Completeness.
By the density theorem of the previous section, the span of all matrix entries is dense in \(L^2(K)\). An
orthonormal family whose span is dense is an orthonormal basis. Hence the rescaled entries form an orthonormal
basis, and grouping them by representation gives the stated orthogonal decomposition into the finite-dimensional
blocks \(\mathcal{E}_\Pi\).
Every \(f \in L^2(K)\) thus expands in matrix entries, with coefficients obtained by integrating
\(f\) against their conjugates. These are the noncommutative Fourier coefficients of \(f\). Only
countably many of them are nonzero for a fixed \(f\), since for any finite subfamily \(\{e_i\}\) of
the rescaled entries the expansion of
\(0 \leq \|f - \sum_i \langle e_i, f \rangle e_i\|_2^2 = \|f\|_2^2 - \sum_i |\langle e_i, f \rangle|^2\)
leaves fewer than \(n^2\|f\|_2^2\) coefficients of modulus above \(1/n\). Enumerate the nonzero ones
as a sequence and let \(M\) be the closed span of the corresponding entries, a Hilbert space because
a
closed subset of a complete space is complete,
and separable because the rational combinations of that sequence are dense in it. Then \(f \in M\).
Writing \(f = m + m^{\perp}\) as in the
projection theorem,
every entry of the enumerated sequence lies in \(M\) and is therefore orthogonal to \(m^{\perp}\),
while every remaining entry is orthogonal to \(M\) and has coefficient zero against \(f\), hence is
orthogonal to \(m^{\perp}\) as well. Being orthogonal to a basis, \(m^{\perp} = 0\), and
Parseval's identity
applied inside \(M\) gives both the \(L^2\)-convergent expansion of \(f\) and the energy identity
\(\|f\|_2^2 = \sum_i |\langle e_i, f \rangle|^2\). This is the precise sense in which the Peter-Weyl
theorem is the harmonic analysis of a compact group. It furnishes the basis, the transform, and the
energy identity, with the irreducible representations playing the role of frequencies.
The Circle Recovered
Take \(K = U(1)\), the circle realized as a compact matrix Lie group inside \(GL(1, \mathbb{C})\).
It is abelian, so its
irreducible complex representations are one-dimensional,
and they are exactly \(\Pi_n(e^{i\theta}) = e^{in\theta}\) for \(n \in \mathbb{Z}\), a
classification we take for granted here. Each is its own single matrix entry, and the factor
\(d_\Pi = 1\) makes the rescaling trivial, so the Peter-Weyl decomposition states that
\(\{e^{in\theta}\}_{n \in \mathbb{Z}}\) is an orthonormal basis of \(L^2(U(1))\). Its normalized
Haar measure is \(d\theta/2\pi\), and taking \(L = \pi\) in the
completeness of the trigonometric system
gives the family \((2\pi)^{-1/2}e^{-in\theta}\) in \(L^2\) of the unnormalized measure, which is the
same family after the relabelling \(n \mapsto -n\) and the rescaling that normalizes the measure.
The general theorem is the noncommutative extension of that one fact. Where the circle offered
scalar exponentials indexed by integers, a nonabelian compact group offers matrix-valued entries
indexed by its irreducible representations.
The window to geometric deep learning
The Peter-Weyl decomposition is the analytic foundation for learning architectures that respect
a continuous symmetry. Functions on the group itself expand in the matrix entries of its
irreducible representations, which is the basis this theorem provides. Data on a homogeneous
space \(K/H\) is reached through the same basis, since functions on \(K/H\) are the functions on
\(K\) invariant under right translation by \(H\), and the expansion of such a function retains
only the entries carrying that invariance. Signals on the sphere under the rotation group, with
\(S^2 = SO(3)/SO(2)\), are the standard instance. The coefficients in each block
\(\mathcal{E}_\Pi\) transform among themselves under translation and never mix across blocks,
because each block is spanned by the entries of a single irreducible representation.
A feature indexed by a fixed irreducible is precisely a quantity that transforms in a
prescribed, irreducible way under rotation. Convolution against a function on the group becomes,
in this basis, a block-diagonal operation acting independently on each irreducible component.
The mathematical content underlying such architectures is the decomposition proved here, namely
that the irreducible blocks exhaust the function space and are mutually orthogonal. The deeper
development, in which these blocks become the typed features of equivariant networks on
homogeneous spaces, belongs to the geometric deep learning track and is taken up there.