L8.2.4: Source Euclidean proof plus P9.3 for its full stated normed-space generality.
Proposition8.2.4.
Let X and Y be normed vector spaces with X finite-dimensional, and let A∈L(X,Y). Then ∥A∥<∞, and A is uniformly continuous (Lipschitz with constant ∥A∥).
Proof.
As we said we only prove the proposition for euclidean spaces, so suppose that X=Rn and the norm is the standard euclidean norm. The general case is left as an exercise.
Let {e1,e2,…,en} be the standard basis of Rn. Write x∈Rn, with ∥x∥=1, as
x=k=1∑nckek.
Since ek⋅eℓ=0 whenever k=ℓ and ek⋅ek=1, we have ck=x⋅ek. By Cauchy–Schwarz,
Thus ∥Bx∥=0 for all x=0, and consequently Bx=0 for all x=0. So B is one-to-one; if Bx=By, then B(x−y)=0, so x=y. As B is a one-to-one linear mapping from X to X, which is finite-dimensional, it is also onto by Proposition 8.1.18. Therefore, B is invertible. It follows that, in particular, GL(X) is open.
Let us prove ii. We must show that the inverse is continuous. Fix an A∈GL(X). Let B be near A, specifically ∥A−B∥<2∥A−1∥1. Then (8.2) is satisfied and B is invertible. A similar computation as above (using B−1y instead of x) gives
Therefore, as B tends to A,∥B−1−A−1∥ tends to 0, and so the inverse operation is a continuous function at A.
L8.2.7: Matrix-coordinate/operator-norm topology equivalence; full preceding Frobenius estimates.
Subsection8.2.2Matrices
Once we fix a basis in a finite-dimensional vector space X, we can represent a vector of X as an n-tuple of numbers—a vector in Rn. The same can be done with L(X,Y), bringing us to matrices, which are a convenient way to represent finite-dimensional linear transformations. Suppose {x1,x2,…,xn} and {y1,y2,…,ym} are bases for vector spaces X and Y respectively. A linear operator is determined by its values on the basis. Given A∈L(X,Y),Axj is an element of Y. Define the numbers ai,j via
Axj=i=1∑mai,jyi,(8.3)
and write them as a matrix, which we, by slight abuse of notation, also call A,
We sometimes write A as [ai,j]. We say A is an m-by-n matrix. The jth column of the matrix contains precisely the coefficients that represent Axj in terms of the basis {y1,y2,…,ym}. Given the numbers ai,j, then via the formula (8.3), we find the corresponding linear operator, as it is determined by the action on a basis. Hence, once we fix bases on X and Y, we have a one-to-one correspondence between L(X,Y) and the m-by-n matrices. When
which gives rise to the familiar rule for matrix multiplication, thinking of z as a column vector, that is, an n-by-1 matrix. More generally, if B is an n-by-r matrix with entries bj,k, then the matrix for C=AB is an m-by-r matrix whose (i,k)th entry ci,k is
ci,k=j=1∑nai,jbj,k.
A way to remember it is if you order the indices as we do—row, column—and put the elements in the same order as the matrices, then the “middle index” is “summed-out.”
There is a one-to-one correspondence between matrices and linear operators in L(X,Y), once we fix bases in X and Y. If we choose different bases, we get different matrices. This is an important distinction. The operator A acts on elements of X, while the matrix is something that works with n-tuples of numbers, that is, vectors of Rn. By convention, we use standard bases in Rn unless otherwise specified, and we identify L(Rn,Rm) with the set of m-by-n matrices.
A linear mapping changing one basis to another is represented by a square matrix in which the columns represent vectors of the second basis in terms of the first basis. We call such a linear mapping a change of basis. So for two choices of a basis in an n-dimensional vector space, there is a linear mapping (a change of basis) taking one basis to the other, and this corresponds to an n-by-n matrix which does the corresponding operation on Rn.
Suppose X=Rn,Y=Rm, and all the bases are just the standard bases. Using the Cauchy–Schwarz inequality, with c=(c1,c2,…,cn)∈Rn, compute
In other words, we have a bound on the operator norm (note that equality rarely happens)
∥A∥≤i=1∑mj=1∑n(ai,j)2.
The right-hand side is the euclidean norm on Rnm, the space of all the entries of the matrix. If the entries go to zero, then ∥A∥ goes to zero. Conversely,
So if the operator norm of A goes to zero, so do the entries. In particular, if A is fixed and B is changing, then the entries of B go to the entries of A if and only if B goes to A in operator norm (∥A−B∥ goes to zero). We have proved:
Proposition8.2.7.
The topology (the set of open sets) on L(Rn,Rm) is the same whether we consider L(Rn,Rm) as a metric space using the operator norm, or the euclidean metric of Rnm.
In particular, let S be a metric space and let π:L(Rn,Rm)→Rnm identify an operator with the nm-tuple of entries of the corresponding matrix. Then f:S→L(Rn,Rm) is continuous if and only if π∘f:S→Rnm is continuous. Similarly for g:L(Rn,Rm)→S and g∘π−1:Rnm→S.
L8.2.8: All seven determinant properties; continuity uses P10.1.
Proposition8.2.8.
det(I)=1.
For every j=1,2,…,n, the function xj↦det([x1x2⋯xn]) is linear.
If two columns of a matrix are interchanged, then the determinant changes sign.
If two columns of A are equal, then det(A)=0.
If a column is zero, then det(A)=0.
A↦det(A) is a continuous function on L(Rn).
det([acbd])=ad−bc, and det([a])=a.
Proof.
We go through the proof quickly, as you have likely seen it before. Item i is trivial. For ii, note that each term in the definition of the determinant contains exactly one factor from each column. Item iii follows as switching two columns is switching the two corresponding numbers in every element in Sn. Hence, all the signs are changed. Item iv follows because if two columns are equal, and we switch them, we get the same matrix back. So item iii says the determinant must be 0. Item v follows because the product in each term in the definition includes one element from the zero column. Item vi follows as det is a polynomial in the entries of the matrix and hence continuous (as a function of the entries of the matrix). A function defined on matrices is continuous in the operator norm if and only if it is continuous as a function of the entries (Proposition 8.2.7). Finally, item vii is a direct computation.
L8.2.9: Determinant multiplication, transpose invariance and invertibility criterion.
Proposition8.2.9.
If A and B are n-by-n matrices, then det(AB)=det(A)det(B). Furthermore, A is invertible if and only if det(A)=0 and in this case, det(A−1)=det(A)1.
Proof.
Let b1,b2,…,bn be the columns of B. Then
AB=[Ab1Ab2⋯Abn].
That is, the columns of AB are Ab1,Ab2,…,Abn.
Let bj,k denote the elements of B and aj the columns of A. By linearity of the determinant,
In the last equality, we sum over the elements of Sn instead of all n-tuples for integers between 1 and n, because when two columns in the determinant are the same, then the determinant is zero. Reordering the columns to the original ordering obtains the sgn.
The conclusion that det(AB)=det(A)det(B) follows by recognizing that the expression in parentheses above is the determinant of B. We obtain this by plugging in A=I. The expression we get for the determinant of B has rows and columns swapped, so as a bonus, we have also just proved that the determinant of a matrix and its transpose are equal.
Let us prove the “Furthermore.” If A is invertible, then A−1A=I. Consequently det(A−1)det(A)=det(A−1A)=det(I)=1. If A is not invertible, then it is not one-to-one, and so A takes some nonzero vector to zero. In other words, the columns of A are linearly dependent. Suppose
k=1∑nγkak=0,
where not all γk are equal to 0. Without loss of generality, suppose γ1=0. Take
B:=γ1γ2γ3⋮γn010⋮0001⋮0⋯⋯⋯⋱⋯000⋮1.
Using the definition of the determinant (there is only a single permutation σ for which ∏i=1nbi,σi is nonzero) we find det(B)=γ1=0. Then det(AB)=det(A)det(B)=γ1det(A). The first column of AB is zero, and hence det(AB)=0. We conclude det(A)=0.