Orthonormal bases and orthogonal projections

This lesson supplies the inner-product facts used in Representations and complete reducibility, Proposition 2.1 and its following paragraph. It belongs to the linear-algebra foundation. It does not require the functional-analysis course.

Proofs, explanations, examples, solutions and programme integration produced by OpenAI Codex — GPT-6 Astra, Ultra effort, October 2026. Self-checked by the writing AI; no human or independent review is claimed. These are standard results, not discoveries.

Prerequisites and sources

We use real and complex arithmetic, complex conjugation, positive square roots of positive real numbers, finite induction, vector-space axioms and finite sums. The programme’s earlier lesson From bases to projections proves that every subspace of a finite-dimensional vector space has a finite basis and that a basis extends to the whole space. No completeness, limit of vectors, choice of an infinite basis, or projection theorem is assumed in the proofs of Theorems 1–2. Exercise 3 is an optional comparison using square-summable sequences and convergence of geometric series.

John M. Erdman’s Functional Analysis and Operator Algebras: An Introduction, version 4 October 2015, Chapter 1, states Gram–Schmidt orthonormalization and the orthogonal-decomposition corollary (source label 0001514). The English and Indonesian source passages are part of the programme’s Erdman edition. Those particular passages contain statements without proofs. This bridge supplies the elementary proofs and connects them to their use in representation theory. It uses Erdman’s convention, linear in the first variable. The author’s source edition and the programme edition retain his authorship; he has not reviewed or endorsed this bridge.

This exposition is offered under Creative Commons Attribution–ShareAlike 4.0, preserving the terms of the Erdman material it develops. The linked basis-extension lesson retains its own CC BY-SA 2.5 licence; its text is not copied into this lesson. Exact source identities and the bounded comparison with other programme proofs are recorded in SOURCE_REVIEW.json.

Inner products and orthogonality

Let 𝔽\mathbb F be ℝ\mathbb R or ℂ\mathbb C. An inner product on an 𝔽\mathbb F-vector space VV is a map ⟨,⟩:V×V→𝔽\langle\ ,\ \rangle:V\times V\to\mathbb F such that

⟨ax+by,z⟩=a⟨x,z⟩+b⟨y,z⟩,⟨y,x⟩=⟨x,y⟩¯, \langle ax+by,z\rangle=a\langle x,z\rangle+b\langle y,z\rangle, \qquad \langle y,x\rangle=\overline{\langle x,y\rangle},

and ⟨x,x⟩\langle x,x\rangle is a positive real number for every nonzero xx. Here a,b∈𝔽a,b\in\mathbb F; conjugation does nothing over ℝ\mathbb R. The two identities imply

⟨x,ay+bz⟩=a¯⟨x,y⟩+b¯⟨x,z⟩. \langle x,ay+bz\rangle=\overline a\langle x,y\rangle+\overline b\langle x,z\rangle.

They also give ⟨0,x⟩=⟨x,0⟩=0\langle0,x\rangle=\langle x,0\rangle=0. Put ∥x∥=⟨x,x⟩\|x\|=\sqrt{\langle x,x\rangle}. In particular ∥ax∥=|a|∥x∥\|ax\|=|a|\|x\|, and ∥x∥=0\|x\|=0 exactly when x=0x=0. We will need only these facts and direct expansions, not the triangle inequality.

Vectors are orthogonal when their inner product is zero. For a subspace UU, define

U⟂={v∈V:⟨v,u⟩=0 for every u∈U}. U^\perp=\{v\in V:\langle v,u\rangle=0\text{ for every }u\in U\}.

Linearity in the first variable makes U⟂U^\perp a subspace. Conjugate symmetry makes orthogonality symmetric. A family (ej)(e_j) is orthonormal when ⟨ei,ej⟩\langle e_i,e_j\rangle is one for i=ji=j and zero otherwise. For a finite linear combination, taking its inner product with eie_i recovers its coefficient. Thus every orthonormal family is linearly independent.

Gram–Schmidt with the coefficients in the correct order

Theorem 1. Let v1,v2,…v_1,v_2,\ldots be a linearly independent finite or countably infinite sequence in a real or complex inner-product space. There is an orthonormal sequence of the same length such that, for every available positive integer jj,

span⁡(e1,…,ej)=span⁡(v1,…,vj). \operatorname{span}(e_1,\ldots,e_j)=\operatorname{span}(v_1,\ldots,v_j).

No completeness or finite dimension of the ambient space is needed for this statement. For an empty sequence, take the empty sequence.

Proof. Suppose the vectors through ej−1e_{j-1} have been constructed, with the asserted span equality. Set

wj=vj−∑i=1j−1⟨vj,ei⟩ei. w_j=v_j-\sum_{i=1}^{j-1}\langle v_j,e_i\rangle e_i.

For k<jk<j, linearity and orthonormality give

⟨wj,ek⟩=⟨vj,ek⟩−∑i=1j−1⟨vj,ei⟩⟨ei,ek⟩=0. \langle w_j,e_k\rangle =\langle v_j,e_k\rangle-\sum_{i=1}^{j-1}\langle v_j,e_i\rangle\langle e_i,e_k\rangle=0.

If wj=0w_j=0, then vjv_j belongs to the span of the previous eie_i, hence of the previous viv_i. This contradicts independence. Therefore ∥wj∥>0\|w_j\|>0, and we may define

ej=wj∥wj∥. e_j=\frac{w_j}{\|w_j\|}.

This vector has inner product one with itself and zero with each previous vector. It belongs to the span of v1,…,vjv_1,\ldots,v_j. Conversely,

vj=∥wj∥ej+∑i=1j−1⟨vj,ei⟩ei v_j=\|w_j\|e_j+\sum_{i=1}^{j-1}\langle v_j,e_i\rangle e_i

belongs to the span of e1,…,eje_1,\ldots,e_j. Together with the induction hypothesis this proves both inclusions of spans. At j=1j=1 the sums are empty, so the same argument starts the induction. In the countably infinite case this construction defines each term by a finite calculation; it makes no assertion about convergence of infinite linear combinations. ▫\square

Corollary 1. Every finite-dimensional real or complex inner-product space has an orthonormal basis, and so does every subspace of it.

Proof. Apply Theorem 1 to a finite basis, using the earlier basis-extension lesson for a subspace’s finite basis. Span equality gives a spanning orthonormal family. In dimension zero the empty family is the required basis. ▫\square

Orthogonal projection and decomposition

Theorem 2. Let UU be a finite-dimensional subspace of any real or complex inner-product space VV. Then every v∈Vv\in V has a unique expression

v=u+w,u∈U,w∈U⟂. v=u+w,\qquad u\in U,\quad w\in U^\perp.

Thus V=U⊕U⟂V=U\oplus U^\perp. In particular this holds for every subspace of a finite-dimensional VV. If (e1,…,em)(e_1,\ldots,e_m) is any orthonormal basis of UU, the orthogonal projection is

Pv=∑i=1m⟨v,ei⟩ei. Pv=\sum_{i=1}^m\langle v,e_i\rangle e_i.

The map is independent of the chosen orthonormal basis. It is linear, satisfies P2=PP^2=P, has image UU and kernel U⟂U^\perp, and obeys

⟨Px,y⟩=⟨x,Py⟩,∥Pv∥2+∥v−Pv∥2=∥v∥2. \langle Px,y\rangle=\langle x,Py\rangle, \qquad \|Pv\|^2+\|v-Pv\|^2=\|v\|^2.

Proof. The restriction of the inner product to UU is positive definite, so Corollary 1 supplies its orthonormal basis. The formula plainly takes values in UU. For every kk, expansion gives

⟨v−Pv,ek⟩=⟨v,ek⟩−∑i⟨v,ei⟩⟨ei,ek⟩=0. \langle v-Pv,e_k\rangle =\langle v,e_k\rangle-\sum_i\langle v,e_i\rangle\langle e_i,e_k\rangle=0.

Conjugate-linearity in the second variable then gives v−Pv∈U⟂v-Pv\in U^\perp. If z∈U∩U⟂z\in U\cap U^\perp, taking its inner product with itself gives z=0z=0. Subtracting two decompositions therefore proves uniqueness. This also proves independence from the basis, since every basis formula produces a decomposition with the same two specified subspaces.

Linearity of PP follows from linearity of each coefficient in vv. For u∈Uu\in U, the unique decomposition is u=u+0u=u+0, so Pu=uPu=u. Consequently the image is exactly UU and P2=PP^2=P. Its formula vanishes on U⟂U^\perp; conversely Pv=0Pv=0 in the decomposition implies v∈U⟂v\in U^\perp. This identifies its kernel.

Orthogonality gives

⟨Px,y⟩=⟨Px,Py⟩=⟨x,Py⟩. \langle Px,y\rangle=\langle Px,Py\rangle=\langle x,Py\rangle.

Expanding the inner product of v=Pv+(v−Pv)v=Pv+(v-Pv) with itself makes the two cross terms zero and proves the squared-length identity. It gives ∥Pv∥≤∥v∥\|Pv\|\le\|v\| as well. No norm-completeness assertion was used. If U=0U=0, the empty sum gives P=0P=0; if U=VU=V, then P=1VP=1_V. These statements include V=0V=0. ▫\square

Corollary 2 (nearest vector). For every v∈Vv\in V, PvPv is the unique vector of UU minimizing distance to vv.

Proof. For u∈Uu\in U, the vectors v−Pvv-Pv and Pv−uPv-u are orthogonal. Expanding their squared length gives

∥v−u∥2=∥v−Pv∥2+∥Pv−u∥2. \|v-u\|^2=\|v-Pv\|^2+\|Pv-u\|^2.

The last summand is nonnegative and vanishes exactly for u=Pvu=Pv. ▫\square

The two uses in representation theory

Let a finite group act linearly on a finite-dimensional complex vector space. Choose any basis and define an initial inner product by (x,y)0=∑jxjyj¯(x,y)_0=\sum_j x_j\overline{y_j} in its unique coordinates. It is linear first, Hermitian, and positive definite because a nonzero vector has a nonzero coordinate. The representation-theory lesson constructs the invariant average

⟨x,y⟩G=1|G|∑g∈G(gx,gy)0. \langle x,y\rangle_G=\frac1{|G|}\sum_{g\in G}(gx,gy)_0.

Each group element acts invertibly, so every term on the diagonal is positive when x≠0x\ne0. Right multiplication permutes the group elements, which proves invariance. Corollary 1 supplies the orthonormal basis used at the end of Proposition 2.1. If MM is the matrix of a group element in that basis, invariance gives M*M=IM^*M=I: the inner products of its columns are those of the original basis vectors, with complex conjugation on the second argument (equivalently their conjugates give the usual matrix entries of M*MM^*M).

For an invariant subspace UU, a vector v∈U⟂v\in U^\perp, an element g∈Gg\in G, and u∈Uu\in U, invariance gives

⟨gv,u⟩G=⟨v,g−1u⟩G=0, \langle gv,u\rangle_G=\langle v,g^{-1}u\rangle_G=0,

since g−1u∈Ug^{-1}u\in U. Thus U⟂U^\perp is invariant. Theorem 2 now supplies the decomposition used immediately after Proposition 2.1. The orthogonal projection also commutes with the action: applying gg to both parts of the unique decomposition preserves their subspaces, so P(gv)=gPvP(gv)=gPv.

These arguments close these two particular linear-algebra uses. They do not prove every other prerequisite of the representation-theory course. Averaging over a general field has its own characteristic condition, and the earlier basis/projection bridge supplies its different, purely algebraic starting projection.

A complex example

In ℂ2\mathbb C^2 use ⟨x,y⟩=x1y1¯+x2y2¯\langle x,y\rangle=x_1\overline{y_1}+x_2\overline{y_2}. Starting with v1=(1,i)v_1=(1,i), v2=(0,1)v_2=(0,1), Theorem 1 gives

e1=(1,i)2,⟨v2,e1⟩=−i2,w2=(i/2,1/2),e2=(i,1)2. e_1=\frac{(1,i)}{\sqrt2},\quad \langle v_2,e_1\rangle=-\frac{i}{\sqrt2},\quad w_2=(i/2,1/2),\quad e_2=\frac{(i,1)}{\sqrt2}.

The two vectors have squared length one and ⟨e1,e2⟩=(−i+i)/2=0\langle e_1,e_2\rangle=(-i+i)/2=0. For U=ℂ(1,i)U=\mathbb C(1,i), Theorem 2 gives

P=12(1−ii1). P=\frac12\begin{pmatrix}1&-i\\i&1\end{pmatrix}.

For example P(0,1)=(−i/2,1/2)P(0,1)=(-i/2,1/2) and (0,1)−P(0,1)=(i/2,1/2)(0,1)-P(0,1)=(i/2,1/2). Reversing the arguments in the coefficient would not give this decomposition. In complex spaces, the coefficient order is substantive.

Exercises with solutions

Exercise 1. Explain why a positive semidefinite form is not enough for Gram–Schmidt as stated.

Solution. On the one-dimensional space ℂ\mathbb C, the zero form is positive semidefinite. The singleton (1)(1) is linearly independent but its squared length is zero, so it cannot be normalized to length one. The proof needs positive definiteness precisely when it divides by ∥wj∥\|w_j\|.

Exercise 2. For the complex matrix PP above, verify directly that P2=P=P*P^2=P=P^*, find its image and kernel, and decompose (1,0)(1,0).

Solution. Write A=2PA=2P. Multiplication gives

A2=(2−2i2i2)=2A,A*=A. A^2=\begin{pmatrix}2&-2i\\2i&2\end{pmatrix}=2A, \qquad A^*=A.

Thus P2=P=P*P^2=P=P^*. Its image is spanned by (1,i)(1,i), since its first column is half that vector and its second is −i-i times the first. Its kernel is the line x−iy=0x-iy=0, spanned by (i,1)(i,1). Finally

(1,0)=(1/2,i/2)+(1/2,−i/2) (1,0)=(1/2,i/2)+(1/2,-i/2)

has its first term in the image and its second in the kernel. Their inner product is zero.

Exercise 3. Why does Theorem 2 keep the assumption that UU is finite-dimensional, even when the ambient space is allowed to be infinite-dimensional?

Solution. In the usual sequence space ℓ2(ℕ)\ell^2(\mathbb N), let UU consist of sequences with only finitely many nonzero entries. If x∈U⟂x\in U^\perp, taking inner products with each coordinate vector gives xn=0x_n=0 for every nn, so U⟂=0U^\perp=0. But the sequence (2−n)n≥1(2^{-n})_{n\ge1} belongs to ℓ2\ell^2, because ∑n≥14−n=1/3\sum_{n\ge1}4^{-n}=1/3, and does not belong to UU. Therefore ℓ2≠U⊕U⟂\ell^2\ne U\oplus U^\perp. This does not contradict the later projection theorem for closed subspaces of a Hilbert space; this UU is not closed, since its truncations converge to that same sequence.

For the broader result, see Hilbert spaces and compact operators, Theorems 2.1–2.2. That proof uses completeness and closedness to obtain a nearest vector. It remains the programme’s general Hilbert-space treatment; the elementary proof here does not replace or narrow it.