Reading guide · Proof index

Analysis with vector spaces

Jiří Lebl, Basic Analysis I–II, version 6.3. Free author edition of this section. Selection and attribution · Notation.

L8.2.4: Source Euclidean proof plus P9.3 for its full stated normed-space generality.

Proof.

As we said we only prove the proposition for euclidean spaces, so suppose that X=RnX = \R^n and the norm is the standard euclidean norm. The general case is left as an exercise.
Let {e1,e2,…,en}\{ e_1,e_2,\ldots,e_n \} be the standard basis of Rn.\R^n\text{.} Write x∈Rn,x \in \R^n\text{,} with ∥x∥=1,\snorm{x} = 1\text{,} as
x=∑k=1nck ek.\begin{equation*} x = \sum_{k=1}^n c_k \, e_k . \end{equation*}
Since ek⋅eℓ=0e_k \cdot e_\ell = 0 whenever k≠ℓk \neq \ell and ek⋅ek=1,e_k \cdot e_k = 1\text{,} we have ck=x⋅ek.c_k = x \cdot e_k\text{.} By Cauchy–Schwarz,
∣ck∣=∣x⋅ek∣≤∥x∥ ∥ek∥=1.\begin{equation*} \sabs{c_k} = \sabs{ x \cdot e_k } \leq \snorm{x} \, \snorm{e_k} = 1 . \end{equation*}
Then
∥Ax∥=∥∑k=1nck Aek∥≤∑k=1n∣ck∣ ∥Aek∥≤∑k=1n∥Aek∥.\begin{equation*} \snorm{Ax} = \norm{\sum_{k=1}^n c_k \, Ae_k} \leq \sum_{k=1}^n \sabs{c_k} \, \snorm{Ae_k} \leq \sum_{k=1}^n \snorm{Ae_k} . \end{equation*}
The right-hand side does not depend on x.x\text{.} We found a finite upper bound for ∥Ax∥\snorm{Ax} independent of x,x\text{,} so ∥A∥<∞.\snorm{A} < \infty\text{.}
Take normed vector spaces XX and Y,Y\text{,} and A∈L(X,Y)A \in L(X,Y) with ∥A∥<∞.\snorm{A} < \infty\text{.} For v,w∈X,v,w \in X\text{,}
∥Av−Aw∥=∥A(v−w)∥≤∥A∥ ∥v−w∥.\begin{equation*} \snorm{Av - Aw} = \bnorm{A(v-w)} \leq \snorm{A} \, \snorm{v-w} . \end{equation*}
As ∥A∥<∞,\snorm{A} < \infty\text{,} the inequality above says that AA is Lipschitz with constant ∥A∥.\snorm{A}\text{.}

L8.2.5: Operator norm sum, scalar and composition inequalities and norm axioms.

Proof.

First, since all the spaces are finite-dimensional, then all the operator norms are finite, and the statements make sense to begin with.
For i, let x∈Xx \in X be arbitrary. Then
∥(A+B)x∥=∥Ax+Bx∥≤∥Ax∥+∥Bx∥≤∥A∥ ∥x∥+∥B∥ ∥x∥=(∥A∥+∥B∥)∥x∥.\begin{equation*} \bnorm{(A+B)x} = \snorm{Ax+Bx} \leq \snorm{Ax}+\snorm{Bx} \leq \snorm{A} \, \snorm{x}+\snorm{B} \,\snorm{x} = \bigl(\snorm{A}+\snorm{B}\bigr) \snorm{x} . \end{equation*}
So ∥A+B∥≤∥A∥+∥B∥.\snorm{A+B} \leq \snorm{A}+\snorm{B}\text{.} Similarly,
∥(cA)x∥=∣c∣ ∥Ax∥≤(∣c∣ ∥A∥)∥x∥.\begin{equation*} \bnorm{(cA)x} = \sabs{c} \, \snorm{Ax} \leq \bigl(\sabs{c} \,\snorm{A}\bigr) \snorm{x} . \end{equation*}
Thus ∥cA∥≤∣c∣ ∥A∥.\snorm{cA} \leq \sabs{c} \, \snorm{A}\text{.} Next,
∣c∣ ∥Ax∥=∥cAx∥≤∥cA∥ ∥x∥.\begin{equation*} \sabs{c} \, \snorm{Ax} = \snorm{cAx} \leq \snorm{cA} \, \snorm{x} . \end{equation*}
Hence ∣c∣ ∥A∥≤∥cA∥.\sabs{c} \, \snorm{A} \leq \snorm{cA}\text{.}
For ii, write
∥BAx∥≤∥B∥ ∥Ax∥≤∥B∥ ∥A∥ ∥x∥.\begin{equation*} \snorm{BAx} \leq \snorm{B} \, \snorm{Ax} \leq \snorm{B} \, \snorm{A} \, \snorm{x} . \qedhere \end{equation*}

L8.2.6: Invertible neighbourhood and continuity of inversion in every finite-dimensional normed space.

Proof.

Let us prove i. We know something about A−1A^{-1} and A−B;A-B\text{;} they are linear operators. So apply them to a vector:
A−1(A−B)x=x−A−1Bx.\begin{equation*} A^{-1}(A-B)x = x-A^{-1}Bx . \end{equation*}
Therefore,
∥x∥=∥A−1(A−B)x+A−1Bx∥≤∥A−1∥ ∥A−B∥ ∥x∥+∥A−1∥ ∥Bx∥.\begin{equation*} \begin{split} \snorm{x} & = \bnorm{A^{-1} (A-B)x + A^{-1}Bx} \\ & \leq \snorm{A^{-1}}\,\snorm{A-B}\, \snorm{x} + \snorm{A^{-1}}\,\snorm{Bx} . \end{split} \end{equation*}
Assume x≠0x \neq 0 and so ∥x∥≠0.\snorm{x} \neq 0\text{.} Using (8.2), we obtain
∥x∥<∥x∥+∥A−1∥ ∥Bx∥.\begin{equation*} \snorm{x} < \snorm{x} + \snorm{A^{-1}} \, \snorm{Bx} . \end{equation*}
Thus ∥Bx∥≠0\snorm{Bx} \neq 0 for all x≠0,x \neq 0\text{,} and consequently Bx≠0Bx \neq 0 for all x≠0.x \neq 0\text{.} So BB is one-to-one; if Bx=By,Bx = By\text{,} then B(x−y)=0,B(x-y) = 0\text{,} so x=y.x=y\text{.} As BB is a one-to-one linear mapping from XX to X,X\text{,} which is finite-dimensional, it is also onto by Proposition 8.1.18. Therefore, BB is invertible. It follows that, in particular, GL(X)GL(X) is open.
Let us prove ii. We must show that the inverse is continuous. Fix an A∈GL(X).A \in GL(X)\text{.} Let BB be near A,A\text{,} specifically ∥A−B∥<12∥A−1∥.\snorm{A-B} < \frac{1}{2 \snorm{A^{-1}}}\text{.} Then (8.2) is satisfied and BB is invertible. A similar computation as above (using B−1yB^{-1}y instead of xx) gives
∥B−1y∥≤∥A−1∥ ∥A−B∥ ∥B−1y∥+∥A−1∥ ∥y∥≤12∥B−1y∥+∥A−1∥ ∥y∥,\begin{equation*} \snorm{B^{-1}y} \leq \snorm{A^{-1}} \, \snorm{A-B} \, \snorm{B^{-1}y} + \snorm{A^{-1}} \, \snorm{y} \leq \frac{1}{2} \snorm{B^{-1}y} + \snorm{A^{-1}}\,\snorm{y} , \end{equation*}
or
∥B−1y∥≤2∥A−1∥ ∥y∥.\begin{equation*} \snorm{B^{-1}y} \leq 2\snorm{A^{-1}}\,\snorm{y} . \end{equation*}
So ∥B−1∥≤2∥A−1∥.\snorm{B^{-1}} \leq 2 \snorm{A^{-1}} \text{.}
Now
A−1(A−B)B−1=A−1(AB−1−I)=B−1−A−1,\begin{equation*} A^{-1}(A-B)B^{-1} = A^{-1}(AB^{-1}-I) = B^{-1}-A^{-1} , \end{equation*}
and
∥B−1−A−1∥=∥A−1(A−B)B−1∥≤∥A−1∥ ∥A−B∥ ∥B−1∥≤2∥A−1∥2∥A−B∥.\begin{equation*} \snorm{B^{-1}-A^{-1}} = \bnorm{A^{-1}(A-B)B^{-1}} \leq \snorm{A^{-1}}\,\snorm{A-B}\,\snorm{B^{-1}} \leq 2\snorm{A^{-1}}^2 \snorm{A-B} . \end{equation*}
Therefore, as BB tends to A,A\text{,} ∥B−1−A−1∥\snorm{B^{-1}-A^{-1}} tends to 0, and so the inverse operation is a continuous function at A.A\text{.}

L8.2.7: Matrix-coordinate/operator-norm topology equivalence; full preceding Frobenius estimates.

Subsection 8.2.2 Matrices

Once we fix a basis in a finite-dimensional vector space X,X\text{,} we can represent a vector of XX as an nn-tuple of numbers—a vector in Rn.\R^n\text{.} The same can be done with L(X,Y),L(X,Y)\text{,} bringing us to matrices, which are a convenient way to represent finite-dimensional linear transformations. Suppose {x1,x2,…,xn}\{ x_1, x_2, \ldots, x_n \} and {y1,y2,…,ym}\{ y_1, y_2, \ldots, y_m \} are bases for vector spaces XX and YY respectively. A linear operator is determined by its values on the basis. Given A∈L(X,Y),A \in L(X,Y)\text{,} AxjA x_j is an element of Y.Y\text{.} Define the numbers ai,ja_{i,j} via
Axj=∑i=1mai,j yi,\begin{equation*} A x_j = \sum_{i=1}^m a_{i,j} \, y_i ,\tag{8.3} \end{equation*}
and write them as a matrix, which we, by slight abuse of notation, also call A,A\text{,}
A=[a1,1a1,2⋯a1,na2,1a2,2⋯a2,n⋮⋮⋱⋮am,1am,2⋯am,n].\begin{equation*} A = \begin{bmatrix} a_{1,1} & a_{1,2} & \cdots & a_{1,n} \\ a_{2,1} & a_{2,2} & \cdots & a_{2,n} \\ \vdots & \vdots & \ddots & \vdots \\ a_{m,1} & a_{m,2} & \cdots & a_{m,n} \end{bmatrix} . \end{equation*}
We sometimes write AA as [ai,j].[a_{i,j}]\text{.} We say AA is an mm-by-nn matrix. The jjth column of the matrix contains precisely the coefficients that represent AxjA x_j in terms of the basis {y1,y2,…,ym}.\{ y_1,y_2,\ldots,y_m \}\text{.} Given the numbers ai,j,a_{i,j}\text{,} then via the formula (8.3), we find the corresponding linear operator, as it is determined by the action on a basis. Hence, once we fix bases on XX and Y,Y\text{,} we have a one-to-one correspondence between L(X,Y)L(X,Y) and the mm-by-nn matrices. When
z=∑j=1nzj xj,\begin{equation*} z = \sum_{j=1}^n z_j \, x_j , \end{equation*}
then
Az=∑j=1nzj Axj=∑j=1nzj(∑i=1mai,j yi)=∑i=1m(∑j=1nai,j zj)yi,\begin{equation*} A z = \sum_{j=1}^n z_j \, A x_j = \sum_{j=1}^n z_j \left( \sum_{i=1}^m a_{i,j}\, y_i \right) = \sum_{i=1}^m \left(\sum_{j=1}^n a_{i,j}\, z_j \right) y_i , \end{equation*}
which gives rise to the familiar rule for matrix multiplication, thinking of zz as a column vector, that is, an nn-by-1 matrix. More generally, if BB is an nn-by-rr matrix with entries bj,k,b_{j,k}\text{,} then the matrix for C=ABC = AB is an mm-by-rr matrix whose (i,k)(i,k)th entry ci,kc_{i,k} is
ci,k=∑j=1nai,j bj,k.\begin{equation*} c_{i,k} = \sum_{j=1}^n a_{i,j}\,b_{j,k} . \end{equation*}
A way to remember it is if you order the indices as we do—row, column—and put the elements in the same order as the matrices, then the “middle index” is “summed-out.”
There is a one-to-one correspondence between matrices and linear operators in L(X,Y),L(X,Y)\text{,} once we fix bases in XX and Y.Y\text{.} If we choose different bases, we get different matrices. This is an important distinction. The operator AA acts on elements of X,X\text{,} while the matrix is something that works with nn-tuples of numbers, that is, vectors of Rn.\R^n\text{.} By convention, we use standard bases in Rn\R^n unless otherwise specified, and we identify L(Rn,Rm)L(\R^n,\R^m) with the set of mm-by-nn matrices.
A linear mapping changing one basis to another is represented by a square matrix in which the columns represent vectors of the second basis in terms of the first basis. We call such a linear mapping a change of basis. So for two choices of a basis in an nn-dimensional vector space, there is a linear mapping (a change of basis) taking one basis to the other, and this corresponds to an nn-by-nn matrix which does the corresponding operation on Rn.\R^n\text{.}
Suppose X=Rn,X=\R^n\text{,} Y=Rm,Y=\R^m\text{,} and all the bases are just the standard bases. Using the Cauchy–Schwarz inequality, with c=(c1,c2,…,cn)∈Rn,c=(c_1,c_2,\ldots,c_n) \in \R^n\text{,} compute
∥Ac∥2=∑i=1m(∑j=1nai,j cj)2≤∑i=1m((∑j=1n(ai,j)2)(∑j=1n(cj)2))=(∑i=1m∑j=1n(ai,j)2)∥c∥2.\begin{equation*} \snorm{Ac}^2 = \sum_{i=1}^m { \left(\sum_{j=1}^n a_{i,j} \, c_j \right)}^2 \leq \sum_{i=1}^m \left( \left(\sum_{j=1}^n {(a_{i,j})}^2 \right) \left(\sum_{j=1}^n {(c_j)}^2 \right) \right) = \left( \sum_{i=1}^m \sum_{j=1}^n {(a_{i,j})}^2 \right) \snorm{c}^2 . \end{equation*}
In other words, we have a bound on the operator norm (note that equality rarely happens)
∥A∥≤∑i=1m∑j=1n(ai,j)2.\begin{equation*} \snorm{A} \leq \sqrt{\sum_{i=1}^m \sum_{j=1}^n {(a_{i,j})}^2} . \end{equation*}
The right-hand side is the euclidean norm on Rnm,\R^{nm}\text{,} the space of all the entries of the matrix. If the entries go to zero, then ∥A∥\snorm{A} goes to zero. Conversely,
∑i=1m∑j=1n(ai,j)2=∑j=1n∥Aej∥2≤∑j=1n∥A∥2=n∥A∥2.\begin{equation*} \sum_{i=1}^m \sum_{j=1}^n {(a_{i,j})}^2 = \sum_{j=1}^n \snorm{A e_j}^2 \leq \sum_{j=1}^n \snorm{A}^2 = n \snorm{A}^2 . \end{equation*}
So if the operator norm of AA goes to zero, so do the entries. In particular, if AA is fixed and BB is changing, then the entries of BB go to the entries of AA if and only if BB goes to AA in operator norm (∥A−B∥\snorm{A-B} goes to zero). We have proved:

L8.2.8: All seven determinant properties; continuity uses P10.1.

Proof.

We go through the proof quickly, as you have likely seen it before. Item i is trivial. For ii, note that each term in the definition of the determinant contains exactly one factor from each column. Item iii follows as switching two columns is switching the two corresponding numbers in every element in Sn.S_n\text{.} Hence, all the signs are changed. Item iv follows because if two columns are equal, and we switch them, we get the same matrix back. So item iii says the determinant must be 0. Item v follows because the product in each term in the definition includes one element from the zero column. Item vi follows as det⁡\det is a polynomial in the entries of the matrix and hence continuous (as a function of the entries of the matrix). A function defined on matrices is continuous in the operator norm if and only if it is continuous as a function of the entries (Proposition 8.2.7). Finally, item vii is a direct computation.

L8.2.9: Determinant multiplication, transpose invariance and invertibility criterion.

Proof.

Let b1,b2,…,bnb_1,b_2,\ldots,b_n be the columns of B.B\text{.} Then
AB=[Ab1Ab2⋯Abn].\begin{equation*} AB = [ Ab_1 \quad Ab_2 \quad \cdots \quad Ab_n ] . \end{equation*}
That is, the columns of ABAB are Ab1,Ab2,…,Abn.Ab_1,Ab_2,\ldots,Ab_n\text{.}
Let bj,kb_{j,k} denote the elements of BB and aja_j the columns of A.A\text{.} By linearity of the determinant,
det⁡(AB)=det⁡([Ab1Ab2⋯Abn])=det⁡([∑j=1nbj,1ajAb2⋯Abn])=∑j=1nbj,1det⁡([ajAb2⋯Abn])=∑1≤j1,j2,…,jn≤nbj1,1bj2,2⋯bjn,ndet⁡([aj1aj2⋯ajn])=(∑(j1,j2,…,jn)∈Snbj1,1bj2,2⋯bjn,nsgn⁡(j1,j2,…,jn))det⁡([a1a2⋯an]).\begin{equation*} \begin{split} \det(AB) & = \det \bigl([ Ab_1 \quad Ab_2 \quad \cdots \quad Ab_n ] \bigr) = \det \left(\left[ \sum_{j=1}^n b_{j,1} a_j \quad Ab_2 \quad \cdots \quad Ab_n \right]\right) \\ & = \sum_{j=1}^n b_{j,1} \det \bigl([ a_j \quad Ab_2 \quad \cdots \quad Ab_n ]\bigr) \\ & = \sum_{1 \leq j_1,j_2,\ldots,j_n \leq n} b_{j_1,1} b_{j_2,2} \cdots b_{j_n,n} \det \bigl([ a_{j_1} \quad a_{j_2} \quad \cdots \quad a_{j_n} ]\bigr) \\ & = \left( \sum_{(j_1,j_2,\ldots,j_n) \in S_n} b_{j_1,1} b_{j_2,2} \cdots b_{j_n,n} \operatorname{sgn}(j_1,j_2,\ldots,j_n) \right) \det \bigl([ a_{1} \quad a_{2} \quad \cdots \quad a_{n} ]\bigr) . \end{split} \end{equation*}
In the last equality, we sum over the elements of SnS_n instead of all nn-tuples for integers between 1 and n,n\text{,} because when two columns in the determinant are the same, then the determinant is zero. Reordering the columns to the original ordering obtains the sgn.
The conclusion that det⁡(AB)=det⁡(A)det⁡(B)\det(AB) = \det(A)\det(B) follows by recognizing that the expression in parentheses above is the determinant of B.B\text{.} We obtain this by plugging in A=I.A=I\text{.} The expression we get for the determinant of BB has rows and columns swapped, so as a bonus, we have also just proved that the determinant of a matrix and its transpose are equal.
Let us prove the “Furthermore.” If AA is invertible, then A−1A=I.A^{-1}A = I\text{.} Consequently det⁡(A−1)det⁡(A)=det⁡(A−1A)=det⁡(I)=1.\det(A^{-1})\det(A) = \det(A^{-1}A) = \det(I) = 1\text{.} If AA is not invertible, then it is not one-to-one, and so AA takes some nonzero vector to zero. In other words, the columns of AA are linearly dependent. Suppose
∑k=1nγk ak=0,\begin{equation*} \sum_{k=1}^n \gamma_k\, a_k = 0 , \end{equation*}
where not all γk\gamma_k are equal to 0. Without loss of generality, suppose γ1≠0.\gamma_1\neq 0\text{.} Take
B≔[γ100⋯0γ210⋯0γ301⋯0⋮⋮⋮⋱⋮γn00⋯1].\begin{equation*} B \coloneqq \begin{bmatrix} \gamma_1 & 0 & 0 & \cdots & 0 \\ \gamma_2 & 1 & 0 & \cdots & 0 \\ \gamma_3 & 0 & 1 & \cdots & 0 \\ \vdots & \vdots & \vdots & \ddots & \vdots \\ \gamma_n & 0 & 0 & \cdots & 1 \end{bmatrix} . \end{equation*}
Using the definition of the determinant (there is only a single permutation σ\sigma for which ∏i=1nbi,σi\prod_{i=1}^n b_{i,\sigma_i} is nonzero) we find det⁡(B)=γ1≠0.\det(B) = \gamma_1 \neq 0\text{.} Then det⁡(AB)=det⁡(A)det⁡(B)=γ1det⁡(A).\det(AB) = \det(A)\det(B) = \gamma_1\det(A)\text{.} The first column of ABAB is zero, and hence det⁡(AB)=0.\det(AB) = 0\text{.} We conclude det⁡(A)=0.\det(A) = 0\text{.}