Geometry of Linear Maps
These pairs of pictures contrast the geometric action of the nonlinear maps and
with the linear maps and .
Each of the four pictures shows the domain on the left mapped to the codomain on the right. Arrows trace where each map sends , , , , and .
The nonlinear maps distort the domain in transforming it into the range. For instance, is further from than it is from —this map spreads the domain out unevenly so that a domain interval near is spread apart more than is a domain interval near . The linear maps are nicer, more regular, in that for each map all of the domain spreads by the same factor. The map on the left spreads all intervals apart to be twice as wide while on the right keeps intervals the same length but reverses their orientation, as with the rising interval from to being transformed to the falling interval from to .
The only linear maps from to are multiplications by a scalar but in higher dimensions more can happen. For instance, this linear transformation of rotates vectors counterclockwise.
The transformation of that projects vectors into the -plane is also not simply a rescaling.
Despite this additional variety, even in higher dimensions linear maps behave nicely. Consider a linear and use the standard bases to represent it by a matrix . Recall from Theorem V.2.7 that factors into where and are nonsingular and is a partial-identity matrix. Recall also that nonsingular matrices factor into elementary matrices , which are matrices that come from the identity after one Gaussian row operation, so each matrix is one of these three kinds
with , . So if we understand the geometric effect of a linear map described by a partial-identity matrix and the effect of the linear maps described by the elementary matrices then we will in some sense completely understand the effect of any linear map. (The pictures below stick to transformations of for ease of drawing but the principles extend for maps from any to any .)
The geometric effect of the linear transformation represented by a partial-identity matrix is projection.
The geometric effect of the matrices is to stretch vectors by a factor of along the -th axis. This map stretches by a factor of along the -axis.
If or if then the -th component goes the other way, here to the left.
Either of these stretches is a dilation.
A transformation represented by a matrix interchanges the -th and -th axes. This is reflection about the line .
Permutations involving more than two axes decompose into a combination of swaps of pairs of axes; see Exercise 7.
The remaining matrices have the form . For instance performs .
In the picture below, the vector with the first component of is affected less than the vector with the first component of . The vector is mapped to a that is only higher than while is higher than .
Any vector with a first component of would be affected in the same way as : it would slide up by . And any vector with a first component of would slide up , as was . That is, the transformation represented by affects vectors depending on their -th component.
Another way to see this point is to consider the action of this map on the unit square. In the next picture, vectors with a first component of , such as the origin, are not pushed vertically at all but vectors with a positive first component slide up. Here, all vectors with a first component of , the entire right side of the square, slide to the same extent. In general, vectors on the same vertical line slide by the same amount, by twice their first component. The resulting shape has the same base and height as the square (and thus the same area) but the right angle corners are gone.
For contrast, the next picture shows the effect of the map represented by . Here vectors are affected according to their second component: slides horizontally by twice .
In general, for any , the sliding happens so that vectors with the same -th component are slid by the same amount. This kind of map is a shear.
With that we understand the geometric effect of the four types of matrices on the right-hand side of and so in some sense we understand the action of any matrix . Thus, even in higher dimensions the geometry of linear maps is easy: it is built by putting together a number of components, each of which acts in a simple way.
We will apply this understanding in two ways. The first way is to prove something general about the geometry of linear maps. Recall that under a linear map, the image of a subspace is a subspace and thus the linear transformation represented by maps lines through the origin to lines through the origin. (The dimension of the image space cannot be greater than the dimension of the domain space, so a line can’t map onto, say, a plane.) We will show that maps any line—not just one through the origin— to a line. The proof is simple: the partial-identity projection and the elementary ’s each turn a line input into a line output; verifying the four cases is Exercise 5. Therefore their composition also preserves lines.
The second way that we will apply the geometric understanding of linear maps is to elucidate a point from Calculus. Below is a picture of the action of the one-variable real function . As with the nonlinear functions pictured earlier, the geometric effect of this map is irregular in that at different domain points it has different effects; for example as the input goes from to , the associated output at first decreases, then pauses for an instant, and then increases.
But in Calculus we focus less on the map overall and more on the local effect of the map. Below we look closely at what this map does near . The derivative is so that near we have . That is, in a neighborhood of , in carrying the domain over this map causes it to grow by a factor of —it is, locally, approximately, a dilation. The picture below shows this as a small interval in the domain carried over to an interval in the codomain that is three times as wide.
In higher dimensions the core idea is the same but more can happen. For a function and a point , the derivative is defined to be the linear map that best approximates how changes near . So the geometry described above directly applies to the derivative.
We close by remarking how this point of view makes clear an often misunderstood result about derivatives, the Chain Rule. Recall that, under suitable conditions on the two functions, the derivative of the composition is this.
For instance the derivative of is .
Where does this come from? Consider .
The first map dilates the neighborhood of by a factor of
and the second map follows that by dilating a neighborhood of by a factor of
and when combined, the composition dilates by the product of the two. In higher dimensions the map expressing how a function changes near a point is a linear map, and is represented by a matrix. The Chain Rule multiplies the matrices.
Exercises
Exercise 1 Supplied answer
Use the decomposition to find the combination of dilations, flips, skews, and projections that produces the map represented with respect to the standard bases by this matrix.
Answer. This Gaussian reduction
gives the reduced echelon form of the matrix. Now the two column operations of taking times the first column and adding it to the second, and then of swapping columns two and three produce this partial identity.
All of that translates into matrix terms as: where
and
the given matrix factors as .
Exercise 2 Supplied answer
What combination of dilations, flips, skews, and projections produces a rotation counterclockwise by radians?
Answer. We will first represent the map with a matrix , perform the row operations and, if needed, column operations to reduce it to a partial-identity matrix. We will then translate that into a factorization . Substituting into the general matrix
gives this representation.
Gauss’s Method is routine.
That translates to a matrix equation in this way.
Taking inverses to solve for yields this factorization.
Exercise 3 Supplied answer
If a map is nonsingular then to get from its representation to the identity matrix we do not need any column operations, so that in the matrix is the identity. An example of a nonsingular map is the transformation that rotates vectors clockwise by radians.
Find the matrix representing this map with respect to the standard bases.
Use Gauss-Jordan to reduce to the identity, without column operations.
Translate that to a matrix equation .
Solve the matrix equation for .
Describe as a combination of dilations, flips, skews, and projections (the identity is a trivial projection).
Answer.
Recall that rotation counterclockwise by radians is represented with respect to the standard basis in this way.
A clockwise angle is the negative of a counterclockwise one.
This Gauss-Jordan reduction
produces the identity matrix. Thus we do not need column-swapping operations to end with a partial-identity.
In matrix multiplication the reduction is
(note that composition of the Gaussian operations is from right to left).
Taking inverses
gives the desired factorization of . The partial identity is .
Reading the composition from right to left (and ignoring the identity matrices as trivial) gives that has the same effect as first performing this skew
followed by a dilation that multiplies all first components by (this is a shrink in that is less than ) and all second components by , followed by another skew.
For an example we start with the unit vector whose angle with the -axis is and apply the components of in turn.
We can easily verify that the resulting vector has unit length and forms an angle with the -axis of , which is indeed a rotation clockwise of radians since .
Exercise 4 Supplied answer
Show that any linear transformation of is a map that multiplies by a scalar .
Answer. Represent it with respect to the standard bases . That produces a matrix. The only entry is the scalar .
Exercise 5 Supplied answer
Show that linear maps preserve the linear structures of a space.
Show that for any linear map from to , the image of any line is a line. The image may be a degenerate line, that is, a single point.
Show that the image of any linear surface is a linear surface. This generalizes the result that under a linear map the image of a subspace is a subspace.
Linear maps preserve other linear ideas. Show that linear maps preserve “betweeness”: if the point is between and then the image of is between the image of and the image of .
Answer.
A line is a subset of of the form . The image of a point on that line is , and the set of such vectors, as ranges over the reals, is a line (albeit, degenerate if ).
This is an obvious extension of the prior argument.
If the point is between the points and then the line from to has in it. That is, there is a such that (where is the endpoint of , etc.). Now, as in the argument of the first item, linearity shows that .
Exercise 6 Supplied answer
Use a picture like the one that appears in the discussion of the Chain Rule to answer: if a function has an inverse, what’s the relationship between how the function —locally, approximately —dilates space, and how its inverse dilates space (assuming, of course, that it has an inverse)?
Exercise 7 Supplied answer
Show that any permutation, any reordering, of the numbers , …, , the map
can be done with a composition of maps, each of which only swaps a single pair of coordinates. Hint: you can use induction on . (Remark: in the fourth chapter we will show this and we will also show that the parity of the number of swaps used is determined by . That is, although a particular permutation could be expressed in two different ways with two different numbers of swaps, either both ways use an even number of swaps, or both use an odd number.)
Answer. We can show this by induction on the number of components in the vector. In the base case the only permutation is the trivial one, and the map
is expressible as a composition of swaps—as zero swaps. For the inductive step we assume that the map induced by any permutation of fewer than numbers can be expressed with swaps only, and we consider the map induced by a permutation of numbers.
Consider the number such that . The map
will, when followed by the swap of the -th and -th components, give the map . Now, the inductive hypothesis gives that is achievable as a composition of swaps.