Peter-Weyl Theorem

From Fourier to Compact Groups Orthonormality of Characters Matrix Entries and the Peter-Weyl Theorem The Peter-Weyl Decomposition and the Window to GDL

From Fourier to Compact Groups

Fourier analysis on the circle rests on a single structural fact. The exponentials \(e^{in\theta}\) form an orthonormal basis of the square-integrable functions on \(S^1\). Every periodic signal decomposes into these elementary frequencies, and the decomposition is orthogonal, complete, and norm-preserving. The exponentials are not an arbitrary choice. The circle is a group under addition of angles, and each \(e^{in\theta}\) is a one-dimensional representation of that group, since the map \(\theta \mapsto e^{in\theta}\) sends the group operation to multiplication of complex numbers. Classical Fourier analysis is, in this light, the harmonic analysis of the abelian group \(S^1\), and the basis functions are exactly its irreducible representations.

The aim of this page is to identify the analogue of this basis for a general compact group, where the group need not be abelian and its irreducible representations need not be one-dimensional. The replacement for the exponentials will be the matrix entries of the irreducible representations, and the replacement for the orthonormal basis of \(L^2(S^1)\) will be a basis of \(L^2(K)\) assembled from those entries. The central results are the orthogonality of characters and the completeness theorem of Peter and Weyl, which together promote the circle's Fourier expansion to a noncommutative harmonic analysis valid on every compact matrix Lie group.

Throughout, \(K\) denotes a compact matrix Lie group, that is, a compact closed subgroup of \(GL(n, \mathbb{C})\). Such a group is by Cartan's closed subgroup theorem a smooth embedded submanifold of \(GL(n, \mathbb{C})\) whose group operations are smooth, hence a Lie group in the sense of manifold theory, and every construction of that theory applies to it. Being a subset of \(M_n(\mathbb{C})\) it is metrizable, so it is also compact Hausdorff.

Integration over \(K\) is taken against the normalized Haar integral, written \(\int_K f(x)\, dx\). It is the integral against the Haar volume form \(\omega_K\), the unique left-invariant top-degree form on \(K\) of total integral \(1\). The hypothesis of that result, a left-invariant orientation, is met here. The group admits a left-invariant global frame, and the top-degree form taking the value \(1\) on that frame at every point is left-invariant and nowhere vanishing, so it determines an orientation, left-invariant because the form is. Normalization means \(\int_K 1\, dx = 1\), and left invariance means that \(\int_K f(yx)\, dx = \int_K f(x)\, dx\) for every fixed \(y \in K\).

The integral is right invariant as well. Write \(R_y\) for right translation by \(y \in K\), a diffeomorphism of \(K\) with inverse \(R_{y^{-1}}\). Since \(R_y \circ L_g = L_g \circ R_y\) and \(L_g^{*}\omega_K = \omega_K\), the pullback \(R_y^{*}\omega_K\) is again left-invariant. Top-degree alternating forms on a \(k\)-dimensional space form a one-dimensional space, here with \(k = \dim K\), so \(R_y^{*}\omega_K = c\,\omega_K\) for some function \(c\) on \(K\), and pulling back by \(L_g\) turns the last identity into \((c \circ L_g)\,\omega_K = c\,\omega_K\), forcing \(c\) to be constant. That constant is nonzero because \(\omega_K\) vanishes nowhere, so \(R_y\) is orientation-preserving when \(c \gt 0\) and orientation-reversing when \(c \lt 0\), and part (d) of the properties of the integral gives \[ c = \int_K R_y^{*}\omega_K = \operatorname{sgn}(c) \int_K \omega_K = \operatorname{sgn}(c), \] so \(c = \pm 1\). Applying the same part to \(f\omega_K\), whose pullback is \(c\,(f \circ R_y)\,\omega_K\), \[ c \int_K f(xy)\, dx = \int_K R_y^{*}(f\omega_K) = \operatorname{sgn}(c) \int_K f(x)\, dx = c \int_K f(x)\, dx, \] and dividing by \(c\) leaves \(\int_K f(xy)\, dx = \int_K f(x)\, dx\). Conjugation \(x \mapsto yxy^{-1}\) is a left translation followed by a right translation, so the integral is invariant under it too. We use these invariances repeatedly and without further comment.

The Haar integral also determines a measure, which is what gives meaning to \(L^2(K)\). The integral is positive, since for continuous \(f \geq 0\) not identically zero the form \(f\omega_K\) is everywhere non-negative and strictly positive somewhere, so part (c) of the properties of the integral makes \(\int_K f\, dx \gt 0\). Boundedness follows by applying positivity to \(\sup_K |f| - \operatorname{Re}(\alpha f)\), with \(\alpha\) a unimodular scalar chosen so that \(\alpha \int_K f\, dx = \left| \int_K f\, dx \right|\), which yields \(\left| \int_K f\, dx \right| \leq \sup_K |f|\). Because \(K\) is compact, \(C_0(K) = C(K)\), so the representation of positive functionals by measures presents the Haar integral as integration against a unique regular Borel measure \(\mu_K\) on \(K\), with \(\mu_K(K) = \int_K 1\, dx = 1\). We write \(L^2(K)\) for the \(L^2\) space of this measure and keep the notation \(\int_K f(x)\, dx\) for integration against it. That \(C(K)\) is dense in \(L^2(K)\) is a standard fact about regular Borel measures on compact spaces, which we take as given.

Characters and Class Functions

The objects that play the role of frequencies are built from traces of representations. We work throughout with finite-dimensional representations of \(K\) on complex vector spaces. Recall that every such representation of a compact group may be taken to be unitary with respect to an inner product obtained by averaging an arbitrary one over the group.

Definition: Character of a Representation

Let \((\Pi, V)\) be a finite-dimensional representation of a compact matrix Lie group \(K\). The character of \(\Pi\) is the function \(\chi_\Pi : K \to \mathbb{C}\) defined by \[ \chi_\Pi(x) = \operatorname{tr}\big(\Pi(x)\big). \] In particular \(\chi_\Pi(e) = \operatorname{tr}(I) = \dim V\), so the character records the dimension of the representation at the identity.

The character discards all information about \(\Pi\) except what survives the trace, yet it retains enough to detect the representation completely. Two features make it the right object. First, it is unchanged under isomorphism of representations, because the trace is invariant under conjugation. If \(\Pi' = S\Pi S^{-1}\) then \(\operatorname{tr}(\Pi'(x)) = \operatorname{tr}(S\Pi(x)S^{-1}) = \operatorname{tr}(\Pi(x))\). Second, and for the same reason, the character is constant on conjugacy classes of \(K\). This second property is fundamental enough to name.

Definition: Class Function

A function \(f : K \to \mathbb{C}\) is a class function if it is constant on conjugacy classes, that is, if \[ f(yxy^{-1}) = f(x) \quad \text{for all } x, y \in K. \] Equivalently, \(f\) is invariant under the conjugation action of \(K\) on itself.

Every character is a class function. The verification is immediate from the multiplicativity of \(\Pi\) and the cyclic invariance of the trace: for any \(x, y \in K\), \[ \chi_\Pi(yxy^{-1}) = \operatorname{tr}\big(\Pi(y)\Pi(x)\Pi(y)^{-1}\big) = \operatorname{tr}\big(\Pi(x)\big) = \chi_\Pi(x), \] where the middle equality uses \(\operatorname{tr}(ABA^{-1}) = \operatorname{tr}(B)\). The continuous class functions form a subspace of \(L^2(K)\), and the orthogonality theorem will show that the characters of the irreducible representations sit inside this subspace as an orthonormal system. The Peter-Weyl theorem will then show that this system is complete. It is to the space of class functions what the exponentials \(e^{in\theta}\) are to \(L^2(S^1)\). On the circle the two pictures coincide, because every element is its own conjugacy class and so every function is a class function. That degeneration is one we return to once the general theorem is in hand.

Orthonormality of Characters

We equip \(L^2(K)\) with the inner product \[ \langle f, g \rangle = \int_K \overline{f(x)}\, g(x)\, dx, \] linear in the second argument. The orthogonality theorem states that the characters of the irreducible representations form an orthonormal set with respect to this product. Distinct irreducibles have orthogonal characters, and each irreducible character has unit norm. This is the exact analogue of the relation \(\langle e^{im\theta}, e^{in\theta} \rangle = \delta_{mn}\) underlying ordinary Fourier series, now without any assumption of commutativity.

Theorem: Orthonormality of Irreducible Characters

Let \((\Pi, V)\) and \((\Sigma, W)\) be irreducible representations of a compact matrix Lie group \(K\). Then \[ \langle \chi_\Pi, \chi_\Sigma \rangle = \int_K \overline{\chi_\Pi(x)}\, \chi_\Sigma(x)\, dx = \begin{cases} 1 & \text{if } \Pi \cong \Sigma, \\\\ 0 & \text{if } \Pi \not\cong \Sigma. \end{cases} \]

The proof rests on two preliminary facts. The first turns the Haar integral of a representation into a concrete geometric object, an orthogonal projection. The second computes the trace of a tensor product. After establishing these we assemble the orthogonality relation by recognizing the integral above as the dimension of a space of intertwining maps, which Schur's lemma then evaluates.

Averaging as a Projection

For any finite-dimensional representation \((\Pi, V)\) of \(K\), write \[ V^K = \{\, v \in V : \Pi(x)v = v \text{ for all } x \in K \,\} \] for the subspace of vectors fixed by every group element. Averaging the operators \(\Pi(x)\) over the group produces a projection onto this subspace.

Proposition: The Averaging Projection

Let \((\Pi, V)\) be a finite-dimensional representation of a compact matrix Lie group \(K\), and define the operator \[ P = \int_K \Pi(x)\, dx \in \operatorname{End}(V), \] the integral taken entrywise in any basis. Then \(P\) is a projection of \(V\) onto \(V^K\). It maps \(V\) into \(V^K\), and \(Pv = v\) for every \(v \in V^K\). Consequently \(\operatorname{tr}(P) = \dim V^K\).

Proof

\(P\) maps into \(V^K\).
Fix \(y \in K\) and \(v \in V\). Using the linearity of \(\Pi(y)\), the homomorphism property \(\Pi(y)\Pi(x) = \Pi(yx)\), and the left-invariance of the Haar integral, \[ \Pi(y)\, Pv = \Pi(y) \int_K \Pi(x)v\, dx = \int_K \Pi(yx)v\, dx = \int_K \Pi(x)v\, dx = Pv. \] The third equality substitutes \(x \mapsto yx\), under which \(dx\) is invariant. Since this holds for every \(y \in K\), the vector \(Pv\) is fixed by the representation, so \(Pv \in V^K\).

\(P\) fixes \(V^K\) pointwise.
If \(v \in V^K\) then \(\Pi(x)v = v\) for all \(x\), so \[ Pv = \int_K \Pi(x)v\, dx = \int_K v\, dx = \left( \int_K 1\, dx \right) v = v, \] using the normalization \(\int_K 1\, dx = 1\). The two properties together say that \(P\) lands in \(V^K\) and restricts to the identity there, so \(P^2 = P\) and the image of \(P\) is exactly \(V^K\). Consequently \(V = V^K \oplus \ker P\), because every \(v\) splits as \(Pv + (v - Pv)\) with \(P(v - Pv) = 0\), while a vector lying in both summands is at once fixed and killed by \(P\). In a basis adapted to this splitting the matrix of \(P\) is \(\operatorname{diag}(I, 0)\) with an identity block of size \(\dim V^K\), so \(\operatorname{tr}(P) = \dim V^K\).

When \(\Pi\) is irreducible and is not the trivial representation, the only vector fixed by all of \(K\) is the origin. The subspace \(V^K\) is invariant, it is not all of \(V\) because \(\Pi\) is not trivial, and irreducibility leaves only \(V^K = \{0\}\). The proposition then records that \(\int_K \Pi(x)\, dx = 0\).

The Trace of a Tensor Product

The orthogonality integral pairs two representations, and the product of their traces will be read as the trace of a single representation on the tensor product space. The underlying linear-algebra identity is the following, an immediate consequence of the characteristic property of the tensor product, which guarantees that \(A \otimes B\) is the well-defined operator acting by \((A \otimes B)(v \otimes w) = Av \otimes Bw\). That property is stated there for real vector spaces, and both its statement and its proof use only the vector space operations, so it transfers verbatim to complex ones.

The tensor-trace identity

For operators \(A\) on \(V\) and \(B\) on \(W\), one has \(\operatorname{tr}(A \otimes B) = \operatorname{tr}(A)\,\operatorname{tr}(B)\). To see this, choose bases \(\{v_j\}\) of \(V\) and \(\{w_l\}\) of \(W\), so that \(\{v_j \otimes w_l\}\) is a basis of \(V \otimes W\). The matrix of \(A \otimes B\) in this basis has diagonal entries \(A_{jj} B_{ll}\), whence \[ \operatorname{tr}(A \otimes B) = \sum_{j, l} A_{jj} B_{ll} = \left(\sum_j A_{jj}\right)\!\left(\sum_l B_{ll}\right) = \operatorname{tr}(A)\,\operatorname{tr}(B). \]

Assembling the Orthogonality Relation

Two further ingredients connect the conjugate character to a dual representation. Fix an orthonormal basis for the invariant inner product recalled above, so that \(\Pi(x)\) is unitary and \(\Pi(x)^{-1}\) is its conjugate transpose. Then \([\Pi(x^{-1})]^\top = \overline{\Pi(x)}\), the entrywise conjugate, so the matrix of the dual representation, which acts by \(\Pi^*(x) = [\Pi(x^{-1})]^\top\), is the entrywise conjugate of the matrix of \(\Pi(x)\). Taking traces, \[ \overline{\chi_\Pi(x)} = \overline{\operatorname{tr}(\Pi(x))} = \operatorname{tr}\big(\overline{\Pi(x)}\big) = \operatorname{tr}\big(\Pi^*(x)\big) = \chi_{\Pi^*}(x). \] The conjugate of a character is therefore the character of the dual representation. We can now prove the theorem.

Proof of Orthonormality

From the pairing to a tensor character.
Combining the conjugate-character identity with the tensor-trace identity, the integrand of \(\langle \chi_\Pi, \chi_\Sigma \rangle\) becomes the character of the tensor product \(\Pi^* \otimes \Sigma\): \[ \overline{\chi_\Pi(x)}\, \chi_\Sigma(x) = \operatorname{tr}\big(\Pi^*(x)\big)\, \operatorname{tr}\big(\Sigma(x)\big) = \operatorname{tr}\big((\Pi^* \otimes \Sigma)(x)\big). \]

Integrating gives the dimension of an invariant subspace.
Integrating over \(K\) and exchanging trace with the entrywise Haar integral, then applying the averaging projection to the representation \(\Pi^* \otimes \Sigma\) on \(V^* \otimes W\), \[ \langle \chi_\Pi, \chi_\Sigma \rangle = \int_K \operatorname{tr}\big((\Pi^* \otimes \Sigma)(x)\big)\, dx = \operatorname{tr}\!\left( \int_K (\Pi^* \otimes \Sigma)(x)\, dx \right) = \operatorname{tr}(P) = \dim\big( (V^* \otimes W)^K \big), \] where \(P\) is the projection onto the fixed subspace \((V^* \otimes W)^K\).

Identifying the fixed subspace with intertwiners.
The space \(V^* \otimes W\) is identified with \(\operatorname{Hom}(V, W)\), the linear maps from \(V\) to \(W\), by sending \(\varphi \otimes w\) to the map \(v \mapsto \varphi(v)\, w\). This is a linear isomorphism, since it carries the basis assembled from a basis of \(V^*\) and one of \(W\) to the matrix units of \(\operatorname{Hom}(V, W)\). Under it the action \(\Pi^*(x) \otimes \Sigma(x)\) of \(x \in K\) becomes \(T \mapsto \Sigma(x)\, T\, \Pi(x)^{-1}\), because \(\Pi^*(x)\varphi = \varphi \circ \Pi(x)^{-1}\). A map is fixed by every \(x\) precisely when \(\Sigma(x) T = T \Pi(x)\) for all \(x\), that is, precisely when \(T\) is an intertwining map between \(\Pi\) and \(\Sigma\). Hence \[ (V^* \otimes W)^K \cong \operatorname{Hom}_K(V, W), \] the space of intertwiners.

Schur's lemma evaluates the dimension.
By Schur's lemma in its first part, a nonzero intertwiner between irreducibles is an isomorphism, so \(\operatorname{Hom}_K(V, W) = 0\) when \(\Pi \not\cong \Sigma\). When \(\Pi \cong \Sigma\), fix an isomorphism \(S\) between them. The third part of the same lemma makes every nonzero intertwiner a scalar multiple of \(S\), so \(\operatorname{Hom}_K(V, W) = \mathbb{C}S\) is one-dimensional. Therefore \[ \langle \chi_\Pi, \chi_\Sigma \rangle = \dim \operatorname{Hom}_K(V, W) = \begin{cases} 1 & \Pi \cong \Sigma, \\\\ 0 & \Pi \not\cong \Sigma, \end{cases} \] which is the assertion of the theorem.

The orthonormality relation already separates representations. Two irreducibles with the same character are isomorphic, because \(\langle \chi_\Pi, \chi_\Sigma \rangle = 1\) forces \(\Pi \cong \Sigma\). What remains is the deeper question of completeness. Do the characters exhaust the class functions, leaving no direction in \(L^2(K)\) orthogonal to all of them? That is the content of the Peter-Weyl theorem, to which we now turn.

Matrix Entries and the Peter-Weyl Theorem

Completeness cannot be proved at the level of characters alone. A character is a single function extracted from a representation by the trace, and the trace collapses too much. The route to completeness passes instead through the full collection of matrix coefficients of a representation, of which the character is the special case formed by summing the diagonal. We isolate these coefficients, prove that they are dense in \(L^2(K)\), and only then descend to characters by averaging over conjugacy.

Definition: Matrix Entries of a Representation

Let \((\Pi, V)\) be a finite-dimensional representation of \(K\) and fix a basis \(\{v_j\}\) of \(V\). The matrix entries of \(\Pi\) are the functions \(K \to \mathbb{C}\) given by \[ x \longmapsto \big(\Pi(x)\big)_{jk}, \] the \((j,k)\) entry of the matrix of \(\Pi(x)\) in the basis \(\{v_j\}\). More generally, and independently of any basis, a function of the form \[ f(x) = \operatorname{tr}\big(\Pi(x) A\big), \quad A \in \operatorname{End}(V), \] is called a matrix entry for \(\Pi\). Choosing \(A\) to be the matrix unit with a single \(1\) in position \((k,j)\) recovers the coordinate function \((\Pi(x))_{jk}\), and a general \(A\) produces a linear combination of coordinate functions. Taking \(A = I\) gives the character \(\chi_\Pi\).

The basis-free description \(f(x) = \operatorname{tr}(\Pi(x)A)\) is the one we use, because it transforms cleanly under the group operations. Of these, the single fact the completeness proof relies on is how an integral of conjugated operators behaves on an irreducible representation.

Conjugation-averaging on an irreducible

If \((\Pi, V)\) is irreducible, then for every operator \(A\) on \(V\), \[ \int_K \Pi(y)\, A\, \Pi(y)^{-1}\, dy = c\, I \quad \text{for some scalar } c, \] and taking traces on both sides fixes the scalar as \(c = \operatorname{tr}(A)/\dim V\). Here the integral commutes with the trace, and \(\operatorname{tr}(\Pi(y)A\Pi(y)^{-1}) = \operatorname{tr}(A)\) is constant in \(y\). To see that the integral is a scalar, let \(B\) denote the operator on the left. For any \(x \in K\), the left-invariance of the Haar integral gives \[ \Pi(x)\, B\, \Pi(x)^{-1} = \int_K \Pi(xy)\, A\, \Pi(xy)^{-1}\, dy = \int_K \Pi(y)\, A\, \Pi(y)^{-1}\, dy = B, \] so \(B\) commutes with every \(\Pi(x)\). By Schur's lemma a self-intertwiner of an irreducible complex representation is a scalar, hence \(B = cI\).

The Algebra of Matrix Entries

Let \(\mathcal{A} \subseteq C(K)\) be the set of all functions expressible as a finite linear combination of matrix entries, ranging over all finite-dimensional representations of \(K\). The Peter-Weyl theorem asserts that \(\mathcal{A}\) is uniformly dense in \(C(K)\), and therefore dense in \(L^2(K)\). We prove this by verifying that \(\mathcal{A}\) satisfies the hypotheses of the Stone-Weierstrass theorem on the compact Hausdorff space \(K\): that it is a self-adjoint subalgebra containing the constants and separating points. The uniform closure of such an algebra is then all of \(C(K)\).

Theorem: Peter-Weyl (Density of Matrix Entries)

Let \(K\) be a compact matrix Lie group. The space \(\mathcal{A}\) of finite linear combinations of matrix entries of finite-dimensional representations of \(K\) is dense in \(C(K)\) in the uniform norm, and hence dense in \(L^2(K)\).

Proof

\(\mathcal{A}\) is an algebra.
By definition \(\mathcal{A}\) is the span of the matrix entries, hence a linear subspace. For products, the tensor-trace identity gives, for representations \(\Pi, \Sigma\) and operators \(A, B\), \[ \operatorname{tr}\big(\Pi(x)A\big)\, \operatorname{tr}\big(\Sigma(x)B\big) = \operatorname{tr}\big((\Pi \otimes \Sigma)(x)\, (A \otimes B)\big), \] since \((\Pi(x)A) \otimes (\Sigma(x)B) = (\Pi \otimes \Sigma)(x)\,(A \otimes B)\). The right-hand side is a matrix entry of the tensor product representation \(\Pi \otimes \Sigma\), so the product of two elements of \(\mathcal{A}\) lies again in \(\mathcal{A}\) and \(\mathcal{A}\) is closed under multiplication.

\(\mathcal{A}\) contains the constants.
The trivial representation \(x \mapsto 1\) on \(\mathbb{C}\) has the constant function \(1\) as its matrix entry, so every constant lies in \(\mathcal{A}\).

\(\mathcal{A}\) is self-adjoint.
The complex conjugate of a matrix entry is a matrix entry of the dual representation. With \(\Pi\) unitary in an orthonormal basis, as arranged above, the matrix of \(\Pi^*(x)\) is the entrywise conjugate of that of \(\Pi(x)\), so \[ \overline{\operatorname{tr}\big(\Pi(x)A\big)} = \operatorname{tr}\big(\overline{\Pi(x)}\,\overline{A}\big) = \operatorname{tr}\big(\Pi^*(x)\,\overline{A}\big), \] a matrix entry for \(\Pi^*\). Hence \(\mathcal{A}\) is closed under complex conjugation.

\(\mathcal{A}\) separates points.
This is the one step that uses the hypothesis that \(K\) is a matrix Lie group. By definition such a group comes with a faithful finite-dimensional representation \(\Pi_0\), namely the inclusion \(K \hookrightarrow GL(n, \mathbb{C})\) itself. If \(x \neq y\) in \(K\), then \(\Pi_0(x) \neq \Pi_0(y)\) by faithfulness, so some matrix entry \((\Pi_0)_{jk}\) takes different values at \(x\) and \(y\). Thus \(\mathcal{A}\) separates the points of \(K\).

Conclusion via Stone-Weierstrass.
Multiplication and conjugation are continuous for the uniform norm, so the uniform closure \(\overline{\mathcal{A}}\) is again a subalgebra closed under conjugation, and it contains the constants and separates points because \(\mathcal{A}\) does. The space \(K\) is compact Hausdorff, so the Stone-Weierstrass theorem gives \(\overline{\mathcal{A}} = C(K)\). Finally \(\|g\|_2 \leq \|g\|_\infty\) because \(\mu_K(K) = 1\), so uniform approximation implies approximation in \(L^2(K)\), and \(C(K)\) is dense there. Hence \(\mathcal{A}\) is dense in \(L^2(K)\) as well.

From Matrix Entries to Characters

Density of matrix entries is the Peter-Weyl theorem in its primary form. Completeness of characters is the statement that a continuous class function orthogonal to every irreducible character must vanish, and it follows by projecting matrix entries onto the class functions. The bridge is an averaging over conjugacy that sends each matrix entry to a linear combination of characters.

Corollary: Completeness of Characters

Let \(K\) be a compact matrix Lie group. If \(f\) is a continuous class function on \(K\) satisfying \(\langle \chi_\Pi, f \rangle = 0\) for every irreducible representation \(\Pi\), then \(f = 0\). Equivalently, no nonzero continuous class function on \(K\) is orthogonal to every irreducible character.

Proof

Conjugation-averaging sends matrix entries to characters.
For a function \(g \in \mathcal{A}\), define its average over conjugacy classes, \[ (\mathcal{C}g)(x) = \int_K g\big(y^{-1} x y\big)\, dy. \] The result is a class function, by the invariance of the Haar integral under \(y \mapsto zy\). Applying this to a single matrix entry \(g(x) = \operatorname{tr}(\Pi(x) A)\) and using the homomorphism property, \[ (\mathcal{C}g)(x) = \int_K \operatorname{tr}\big(\Pi(y^{-1}) \Pi(x) \Pi(y) A\big)\, dy = \operatorname{tr}\!\left( \Pi(x) \int_K \Pi(y)\, A\, \Pi(y)^{-1}\, dy \right), \] where the second equality uses the cyclic invariance of the trace and \(\Pi(y^{-1}) = \Pi(y)^{-1}\). If \(\Pi\) is irreducible, the conjugation-averaging identity evaluates the inner integral as \((\operatorname{tr} A / \dim V)\, I\), so that \(\mathcal{C}g\) is the multiple \((\operatorname{tr} A / \dim V)\, \chi_\Pi\) of a single irreducible character.

A general \(\Pi\) reduces to that case. Being a finite-dimensional representation of a compact matrix Lie group it is completely reducible, which in its internal form means that \(V = U_1 \oplus \cdots \oplus U_r\) with every \(U_i\) invariant and irreducible, the restriction \(\Pi_i\) of \(\Pi\) to \(U_i\) being a representation in its own right. The projection \(p_i\) onto \(U_i\) along the remaining summands commutes with every \(\Pi(x)\), because \(\Pi(x)\) carries each summand into itself. Writing \(A_i\) for the operator \(p_i A|_{U_i}\) on \(U_i\) and using \(\sum_i p_i = I\) together with cyclicity of the trace, \[ \operatorname{tr}\big(\Pi(x) A\big) = \sum_{i=1}^{r} \operatorname{tr}\big(\Pi(x)\, p_i A p_i\big) = \sum_{i=1}^{r} \operatorname{tr}\big(\Pi_i(x)\, A_i\big). \] Every summand is a matrix entry for an irreducible representation, so the previous case applies to it. Since \(\mathcal{C}\) is linear and every element of \(\mathcal{A}\) is a finite combination of matrix entries, conjugation-averaging carries \(\mathcal{A}\) into the span of the irreducible characters.

A class function orthogonal to all characters is orthogonal to its own approximants.
Suppose \(f\) is a continuous class function with \(\langle \chi_\Pi, f \rangle = 0\) for every irreducible \(\Pi\). By the density of \(\mathcal{A}\), choose \(g_n \in \mathcal{A}\) converging to \(f\) in \(L^2(K)\). Each conjugation-average \(\mathcal{C}g_n\) is a linear combination of characters, so \(\langle \mathcal{C}g_n, f \rangle = 0\). Because \(f\) is itself a class function, averaging the integrand of \(\langle g_n, f \rangle\) over conjugacy does not change its value: \[ \langle g_n, f \rangle = \int_K \overline{g_n(x)}\, f(x)\, dx = \int_K \overline{(\mathcal{C}g_n)(x)}\, f(x)\, dx = \langle \mathcal{C}g_n, f \rangle = 0. \] The middle equality is where the class-function property of \(f\) enters. Written out, \(\langle \mathcal{C}g_n, f \rangle\) is an integral over \(K \times K\) of \(\overline{g_n(y^{-1}xy)}\, f(x)\), a continuous and therefore bounded function, measurable for the product \(\sigma\)-algebra because \(K\) is a compact metric space and hence second countable. The measure \(\mu_K \otimes \mu_K\) being finite, Fubini's theorem allows integrating in \(x\) first. For fixed \(y\) the substitution \(x \mapsto y x y^{-1}\), under which \(dx\) is invariant, turns \(\int_K \overline{g_n(y^{-1}xy)}\, f(x)\, dx\) into \(\int_K \overline{g_n(x)}\, f(y x y^{-1})\, dx\), which equals \(\langle g_n, f \rangle\) because \(f\) is a class function. Averaging over \(y\) leaves that value unchanged.

Passing to the limit.
Since \(g_n \to f\) in \(L^2(K)\), we have \(\langle g_n, f \rangle \to \langle f, f \rangle = \|f\|_2^2\). But every \(\langle g_n, f \rangle = 0\), so \(\|f\|_2 = 0\) and \(f = 0\) almost everywhere. Being continuous, \(f\) is identically zero. The irreducible characters therefore leave no nonzero class function orthogonal to all of them, which is their completeness.

The Peter-Weyl Decomposition and the Window to GDL

The density of matrix entries upgrades at once to an orthonormal basis. The space \(L^2(K)\) is a Hilbert space, carrying the inner product \(\langle f, g \rangle = \int_K \overline{f}\, g\, dx\) and complete by the Riesz-Fischer theorem. We call a family in a Hilbert space an orthonormal basis when its members are orthonormal and their finite linear combinations are dense, which is the condition of the definition for sequences with the countability of the family dropped. The orthogonality relations for the matrix entries supply the first half of that condition and the density theorem supplies the second, and together they yield the decomposition that is the harmonic-analytic core of the theory.

Theorem: The Peter-Weyl Decomposition of \(L^2(K)\)

Let \(K\) be a compact matrix Lie group, and let \(\widehat{K}\) be a set of representatives of the isomorphism classes of irreducible unitary representations. For each \(\Pi \in \widehat{K}\) of dimension \(d_\Pi\), fix an orthonormal basis of its space for a \(K\)-invariant inner product, giving matrix entries \((\Pi(x))_{jk}\). Then the rescaled matrix entries \[ \sqrt{d_\Pi}\,(\Pi(x))_{jk}, \quad \Pi \in \widehat{K}, \quad 1 \leq j, k \leq d_\Pi, \] form an orthonormal basis of \(L^2(K)\). Equivalently, \(L^2(K)\) decomposes as the Hilbert-space direct sum \[ L^2(K) = \bigoplus_{\Pi \in \widehat{K}} \mathcal{E}_\Pi, \] where \(\mathcal{E}_\Pi\) is the \(d_\Pi^2\)-dimensional space spanned by the entries of \(\Pi\). The summands are mutually orthogonal and their finite sums are dense, which is what the direct sum asserts here.

Proof

Orthogonality and normalization.
The relations come from the same mechanism that gave the character relations, applied to the full coefficients rather than their traces. Let \((\Pi, V)\) and \((\Sigma, W)\) be irreducible members of \(\widehat{K}\), with the chosen orthonormal bases, and for a linear map \(T : V \to W\) set \[ T^{\#} = \int_K \Sigma(y)\, T\, \Pi(y)^{-1}\, dy, \] the integral taken entrywise. Left invariance gives \(\Sigma(x)\, T^{\#}\, \Pi(x)^{-1} = T^{\#}\) for every \(x \in K\), so \(T^{\#}\) is an intertwining map. Unitarity in these bases turns \((\Pi(y)^{-1})_{kj}\) into \(\overline{\Pi(y)_{jk}}\). Let \(E_{mk}\) carry the \(k\)-th basis vector of \(V\) to the \(m\)-th of \(W\) and the others to zero. Then \[ \big(E_{mk}^{\#}\big)_{lj} = \int_K \Sigma(y)_{lm}\, \overline{\Pi(y)_{jk}}\, dy = \big\langle (\Pi)_{jk}, (\Sigma)_{lm} \big\rangle. \] If \(\Pi \not\cong \Sigma\), the first part of Schur's lemma forces \(E_{mk}^{\#} = 0\), so entries of inequivalent irreducibles are orthogonal in \(L^2(K)\). If \(\Sigma = \Pi\), the conjugation-averaging identity evaluates \(E_{mk}^{\#}\) as \((\operatorname{tr} E_{mk}/d_\Pi)\, I = (\delta_{km}/d_\Pi)\, I\), whose \((l, j)\) entry is \(\delta_{km}\delta_{jl}/d_\Pi\). Within a single irreducible \(\Pi\) we therefore have \[ \big\langle (\Pi)_{jk}, (\Pi)_{lm} \big\rangle = \frac{1}{d_\Pi}\, \delta_{jl}\, \delta_{km}, \] and rescaling each entry by \(\sqrt{d_\Pi}\) produces an orthonormal family. All these pairings are finite because \(K\) is compact, so every continuous function lies in \(L^2(K)\) and the Cauchy-Schwarz inequality controls them.

Completeness.
By the density theorem of the previous section, the span of all matrix entries is dense in \(L^2(K)\). An orthonormal family whose span is dense is an orthonormal basis. Hence the rescaled entries form an orthonormal basis, and grouping them by representation gives the stated orthogonal decomposition into the finite-dimensional blocks \(\mathcal{E}_\Pi\).

Every \(f \in L^2(K)\) thus expands in matrix entries, with coefficients obtained by integrating \(f\) against their conjugates. These are the noncommutative Fourier coefficients of \(f\). Only countably many of them are nonzero for a fixed \(f\), since for any finite subfamily \(\{e_i\}\) of the rescaled entries the expansion of \(0 \leq \|f - \sum_i \langle e_i, f \rangle e_i\|_2^2 = \|f\|_2^2 - \sum_i |\langle e_i, f \rangle|^2\) leaves fewer than \(n^2\|f\|_2^2\) coefficients of modulus above \(1/n\). Enumerate the nonzero ones as a sequence and let \(M\) be the closed span of the corresponding entries, a Hilbert space because a closed subset of a complete space is complete, and separable because the rational combinations of that sequence are dense in it. Then \(f \in M\). Writing \(f = m + m^{\perp}\) as in the projection theorem, every entry of the enumerated sequence lies in \(M\) and is therefore orthogonal to \(m^{\perp}\), while every remaining entry is orthogonal to \(M\) and has coefficient zero against \(f\), hence is orthogonal to \(m^{\perp}\) as well. Being orthogonal to a basis, \(m^{\perp} = 0\), and Parseval's identity applied inside \(M\) gives both the \(L^2\)-convergent expansion of \(f\) and the energy identity \(\|f\|_2^2 = \sum_i |\langle e_i, f \rangle|^2\). This is the precise sense in which the Peter-Weyl theorem is the harmonic analysis of a compact group. It furnishes the basis, the transform, and the energy identity, with the irreducible representations playing the role of frequencies.

The Circle Recovered

Take \(K = U(1)\), the circle realized as a compact matrix Lie group inside \(GL(1, \mathbb{C})\). It is abelian, so its irreducible complex representations are one-dimensional, and they are exactly \(\Pi_n(e^{i\theta}) = e^{in\theta}\) for \(n \in \mathbb{Z}\), a classification we take for granted here. Each is its own single matrix entry, and the factor \(d_\Pi = 1\) makes the rescaling trivial, so the Peter-Weyl decomposition states that \(\{e^{in\theta}\}_{n \in \mathbb{Z}}\) is an orthonormal basis of \(L^2(U(1))\). Its normalized Haar measure is \(d\theta/2\pi\), and taking \(L = \pi\) in the completeness of the trigonometric system gives the family \((2\pi)^{-1/2}e^{-in\theta}\) in \(L^2\) of the unnormalized measure, which is the same family after the relabelling \(n \mapsto -n\) and the rescaling that normalizes the measure. The general theorem is the noncommutative extension of that one fact. Where the circle offered scalar exponentials indexed by integers, a nonabelian compact group offers matrix-valued entries indexed by its irreducible representations.

The window to geometric deep learning

The Peter-Weyl decomposition is the analytic foundation for learning architectures that respect a continuous symmetry. Functions on the group itself expand in the matrix entries of its irreducible representations, which is the basis this theorem provides. Data on a homogeneous space \(K/H\) is reached through the same basis, since functions on \(K/H\) are the functions on \(K\) invariant under right translation by \(H\), and the expansion of such a function retains only the entries carrying that invariance. Signals on the sphere under the rotation group, with \(S^2 = SO(3)/SO(2)\), are the standard instance. The coefficients in each block \(\mathcal{E}_\Pi\) transform among themselves under translation and never mix across blocks, because each block is spanned by the entries of a single irreducible representation.

A feature indexed by a fixed irreducible is precisely a quantity that transforms in a prescribed, irreducible way under rotation. Convolution against a function on the group becomes, in this basis, a block-diagonal operation acting independently on each irreducible component. The mathematical content underlying such architectures is the decomposition proved here, namely that the irreducible blocks exhaust the function space and are mutually orthogonal. The deeper development, in which these blocks become the typed features of equivariant networks on homogeneous spaces, belongs to the geometric deep learning track and is taken up there.