Original English by Jim Hefferon — 34 validated sections. The original mathematics and supplied answers below are preserved. This is a partial-book reading edition, not the complete book or an Everyday-English rewrite.

Source, reuse and conversion details

Source revision df2262e089a02651c127f1dd12649c4622ee1383; CC BY-SA 2.5 option, with original component credits retained. This is not an Everyday-English rewrite. The complete active source topic is included. The original end-of-file marker and comments remain in the editable source. Cross-section links require the bound Change of Basis reader for offline use. Fifteen bounded source-context rules are available; this does not claim complete prerequisite closure.

AI-assisted conversion and source checks; no human review is claimed. Current rebuild runtime is documented in the credit below; earlier intermediate work is not reattributed.

Notes about the original source and supplied answers

These six bounded source findings are separate from the unchanged text, formulas and diagrams. Opening them may reveal answers. One is a later source qualification, not an error correction. They are not an exhaustive correctness audit or human review.

  1. Source note 1: Exercise 1 prints the row-reduction product as P and a Q that swaps columns 1 and 2, then claims H=PBQ. Exact rational replay disproves that product. The described column operations give B=P H (D S23); solving for H requires the inverses of those reduction products. Original matrices remain unchanged.
  2. Source note 2: Both Exercise 2 rotation representations are juxtaposed with their matrices without equals signs. The mirror preserves those exact expressions; it does not silently insert a sign.
  3. Source note 3: Exercise 3 describes the general counterclockwise angle theta but labels the representation t_{pi/4}, while the matrix uses theta. Use a general theta label or explicitly specialize the matrix; this mirror preserves the original mismatch.
  4. Source note 4: The next Exercise 3 representation writes t_{-pi/4}, with literal letters p and i rather than the TeX command for pi. MathML conversion preserves this source error.
  5. Source note 5: Exercise 6 evaluates the inverse derivative at x instead of at f(x). With differentiability and a nonzero derivative, the relation is (f^{-1})'(f(x))=1/f'(x). For the invertible real function f(x)=x+x^3, the inverse derivative at 2 is 1/4, not 1/f'(2)=1/13. The original formula remains unchanged.
  6. Source note 6: The body states that any line maps to a line without mentioning collapse to a point. Exercise 5 and its answer explicitly allow a degenerate line and give the condition h(w)=0. Carry that qualification with modular extracts of the earlier narrative. This is a later source clarification, not a newly invented correction.

Wide formulas scroll horizontally. Focus a mathematical region and use the arrow keys.

Source-preserving rebuild, navigation, source packaging and current deterministic checks: OpenAI Codex — GPT-6 Astra, Ultra effort. Jim Hefferon remains the author of the mathematics. Earlier intermediate-conversion runtime identity is not established by its retained receipts and is not reassigned to this rebuild. No human review or exhaustive proof certification is claimed.

Geometry of Linear Maps

These pairs of pictures contrast the geometric action of the nonlinear maps f 1 ( x ) = e x and f 2 ( x ) = x 2

The nonlinear exponential map sends five marked real-number inputs, minus two through two, from the left number line to positive outputs on the right. The arrows spread unevenly. The square map sends minus two and two to four, minus one and one to one, and zero to zero. Arrows connect the left domain and right codomain number lines.

with the linear maps h 1 ( x ) = 2 x and h 2 ( x ) = − x .

The linear map x to 2x doubles the distance of every marked input from zero. Five arrows connect corresponding positions on the two number lines. The linear map x to minus x reverses the marked number-line positions and preserves distances. All five connecting arrows cross at their midpoint.

Each of the four pictures shows the domain ℝ on the left mapped to the codomain ℝ on the right. Arrows trace where each map sends x = 0 , x = 1 , x = 2 , x = − 1 , and x = − 2 .

The nonlinear maps distort the domain in transforming it into the range. For instance, f 1 ( 1 ) is further from f 1 ( 2 ) than it is from f 1 ( 0 ) —this map spreads the domain out unevenly so that a domain interval near x = 2 is spread apart more than is a domain interval near x = 0 . The linear maps are nicer, more regular, in that for each map all of the domain spreads by the same factor. The map  h 1 on the left spreads all intervals apart to be twice as wide while on the right  h 2 keeps intervals the same length but reverses their orientation, as with the rising interval from 1 to 2 being transformed to the falling interval from − 1 to  − 2 .

The only linear maps from ℝ to ℝ are multiplications by a scalar but in higher dimensions more can happen. For instance, this linear transformation of ℝ 2 rotates vectors counterclockwise.

A plane vector is rotated counterclockwise. The central formula maps (x,y) to (x cos theta minus y sin theta, x sin theta plus y cos theta).

The transformation of ℝ 3 that projects vectors into the x z -plane is also not simply a rescaling.

A vector in three-dimensional space projects onto the xz-plane: (x,y,z) maps to (x,0,z). The dashed plane and the two coordinate diagrams are retained.

Despite this additional variety, even in higher dimensions linear maps behave nicely. Consider a linear h : ℝ n → ℝ m and use the standard bases to represent it by a matrix H . Recall from Theorem V.2.7 that H factors into H = P B Q where P and Q are nonsingular and B is a partial-identity matrix. Recall also that nonsingular matrices factor into elementary matrices P B Q = T n T n − 1 ⋯ T s B T s − 1 ⋯ T 1 , which are matrices that come from the identity I after one Gaussian row operation, so each T matrix is one of these three kinds

I ⟶ k ρ i ( M i ( k ) I ⟶ ρ i ↔ ρ j ( P i , j I ⟶ k ρ i + ρ j ( C i , j ( k )

with i ≠ j , k ≠ 0 . So if we understand the geometric effect of a linear map described by a partial-identity matrix and the effect of the linear maps described by the elementary matrices then we will in some sense completely understand the effect of any linear map. (The pictures below stick to transformations of ℝ 2 for ease of drawing but the principles extend for maps from any ℝ n to any ℝ m .)

The geometric effect of the linear transformation represented by a partial-identity matrix is projection.

( x y z ) → ( 1 0 0 0 1 0 0 0 0 ) ( x y 0 )

The geometric effect of the M i ( k ) matrices is to stretch vectors by a factor of k along the i -th axis. This map stretches by a factor of 3 along the x -axis.

Horizontal dilation multiplies the x coordinate by three and leaves y unchanged. The original vector and its horizontally stretched image appear on matching axes.

If 0 ≤ k < 1 or if k < 0 then the i -th component goes the other way, here to the left.

Horizontal multiplication by minus two reverses and doubles the x coordinate while leaving y unchanged. The image lies left of the vertical axis.

Either of these stretches is a dilation.

A transformation represented by a P i , j matrix interchanges the i -th and j -th axes. This is reflection about the line x i = x j .

Swapping x and y reflects a plane vector across the dashed line y=x. A pale copy of the input vector provides comparison in the image axes.

Permutations involving more than two axes decompose into a combination of swaps of pairs of axes; see Exercise 7.

The remaining matrices have the form C i , j ( k ) . For instance C 1 , 2 ( 2 ) performs 2 ρ 1 + ρ 2 .

( x y ) → ( 1 0 2 1 ) ( x 2 x + y )

In the picture below, the vector u → with the first component of 1 is affected less than the vector v → with the first component of 2 . The vector u → is mapped to a h ( u → ) that is only 2 higher than u → while h ( v → ) is 4 higher than v → .

Vertical shear maps (x,y) to (x,2x+y). Input vectors u and v with different x coordinates acquire different vertical displacements; their images are labelled h(u) and h(v).

Any vector with a first component of 1 would be affected in the same way as u → : it would slide up by 2 . And any vector with a first component of 2 would slide up 4 , as was v → . That is, the transformation represented by C i , j ( k ) affects vectors depending on their i -th component.

Another way to see this point is to consider the action of this map on the unit square. In the next picture, vectors with a first component of 0 , such as the origin, are not pushed vertically at all but vectors with a positive first component slide up. Here, all vectors with a first component of 1 , the entire right side of the square, slide to the same extent. In general, vectors on the same vertical line slide by the same amount, by twice their first component. The resulting shape has the same base and height as the square (and thus the same area) but the right angle corners are gone.

Vertical shear (x,y) to (x,2x+y) takes a shaded unit square to a parallelogram with the same base and height.

For contrast, the next picture shows the effect of the map represented by C 2 , 1 ( 2 ) . Here vectors are affected according to their second component: ( x y ) slides horizontally by twice y .

Horizontal shear (x,y) to (x+2y,y) takes a shaded unit square to a parallelogram. Points with equal y slide equally.

In general, for any C i , j ( k ) , the sliding happens so that vectors with the same i -th component are slid by the same amount. This kind of map is a shear.

With that we understand the geometric effect of the four types of matrices on the right-hand side of H = T n T n − 1 ⋯ T j B T j − 1 ⋯ T 1 and so in some sense we understand the action of any matrix  H . Thus, even in higher dimensions the geometry of linear maps is easy: it is built by putting together a number of components, each of which acts in a simple way.

We will apply this understanding in two ways. The first way is to prove something general about the geometry of linear maps. Recall that under a linear map, the image of a subspace is a subspace and thus the linear transformation h represented by H maps lines through the origin to lines through the origin. (The dimension of the image space cannot be greater than the dimension of the domain space, so a line can’t map onto, say, a plane.) We will show that h maps any line—not just one through the origin— to a line. The proof is simple: the partial-identity projection B and the elementary T i ’s each turn a line input into a line output; verifying the four cases is Exercise 5. Therefore their composition also preserves lines.

The second way that we will apply the geometric understanding of linear maps is to elucidate a point from Calculus. Below is a picture of the action of the one-variable real function y ( x ) = x 2 + x . As with the nonlinear functions pictured earlier, the geometric effect of this map is irregular in that at different domain points it has different effects; for example as the input  x goes from 2 to − 2 , the associated output  f ( x ) at first decreases, then pauses for an instant, and then increases.

The nonlinear map x to x squared plus x sends marked points on a left number line to a right number line; some different inputs have the same output.

But in Calculus we focus less on the map overall and more on the local effect of the map. Below we look closely at what this map does near x = 1 . The derivative is d y / d x = 2 x + 1 so that near x = 1 we have Δ y ≈ 3 ⋅ Δ x . That is, in a neighborhood of x = 1 , in carrying the domain over this map causes it to grow by a factor of 3 —it is, locally, approximately, a dilation. The picture below shows this as a small interval in the domain ( 1 − Δ x . . 1 + Δ x ) carried over to an interval in the codomain ( 2 − Δ y . . 2 + Δ y ) that is three times as wide.

A small interval around x=1 maps to an interval around y=2. The original braces and arrow illustrate local dilation by the derivative three.

In higher dimensions the core idea is the same but more can happen. For a function y : ℝ n → ℝ m and a point x → ∈ ℝ n , the derivative is defined to be the linear map h : ℝ n → ℝ m that best approximates how y changes near y ( x → ) . So the geometry described above directly applies to the derivative.

We close by remarking how this point of view makes clear an often misunderstood result about derivatives, the Chain Rule. Recall that, under suitable conditions on the two functions, the derivative of the composition is this.

d ( g ∘ f ) d x ( x ) = d g d x ( f ( x ) ) ⋅ d f d x ( x )

For instance the derivative of sin ⁡ ( x 2 + 3 x ) is cos ⁡ ( x 2 + 3 x ) ⋅ ( 2 x + 3 ) .

Where does this come from? Consider f , g : ℝ → ℝ .

Three number lines show x, f(x), and g(f(x)). Successive arrows and neighborhood braces illustrate composition of local dilations in the chain rule.

The first map f dilates the neighborhood of x by a factor of

d f d x ( x )

and the second map g follows that by dilating a neighborhood of f ( x ) by a factor of

d g d x ( f ( x ) )

and when combined, the composition dilates by the product of the two. In higher dimensions the map expressing how a function changes near a point is a linear map, and is represented by a matrix. The Chain Rule multiplies the matrices.

Exercises

  1. Exercise 1 Supplied answer

    Use the H = P B Q decomposition to find the combination of dilations, flips, skews, and projections that produces the map h : ℝ 3 → ℝ 3 represented with respect to the standard bases by this matrix.

    H = ( 1 2 1 3 6 0 1 2 2 )

    Back to Exercise 1

    Answer. This Gaussian reduction

    ⟶ − ρ 1 + ρ 3 − 3 ρ 1 + ρ 2 ( ( 1 2 1 0 0 − 3 0 0 1 ) ⟶ ( 1 / 3 ) ρ 2 + ρ 3 ( ( 1 2 1 0 0 − 3 0 0 0 ) ⟶ ( − 1 / 3 ) ρ 2 ( ( 1 2 1 0 0 1 0 0 0 ) ⟶ − ρ 2 + ρ 1 ( ( 1 2 0 0 0 1 0 0 0 )

    gives the reduced echelon form of the matrix. Now the two column operations of taking − 2 times the first column and adding it to the second, and then of swapping columns two and three produce this partial identity.

    B = ( 1 0 0 0 1 0 0 0 0 )

    All of that translates into matrix terms as: where

    P = ( 1 − 1 0 0 1 0 0 0 1 ) ( 1 0 0 0 − 1 / 3 0 0 0 1 ) ( 1 0 0 0 1 0 0 1 / 3 1 ) ( 1 0 0 0 1 0 − 1 0 1 ) ( 1 0 0 − 3 1 0 0 0 1 )

    and

    Q = ( 1 − 2 0 0 1 0 0 0 1 ) ( 0 1 0 1 0 0 0 0 1 )

    the given matrix factors as P B Q .

  2. Exercise 2 Supplied answer

    What combination of dilations, flips, skews, and projections produces a rotation counterclockwise by 2 π / 3 radians?

    Back to Exercise 2

    Answer. We will first represent the map with a matrix H , perform the row operations and, if needed, column operations to reduce it to a partial-identity matrix. We will then translate that into a factorization H = P B Q . Substituting into the general matrix

    Rep ℰ 2 , ℰ 2 ( r θ ) ( cos ⁡ θ − sin ⁡ θ sin ⁡ θ cos ⁡ θ )

    gives this representation.

    Rep ℰ 2 , ℰ 2 ( r 2 π / 3 ) ( − 1 / 2 − 3 / 2 3 / 2 − 1 / 2 )

    Gauss’s Method is routine.

    ⟶ 3 ρ 1 + ρ 2 ( ( − 1 / 2 − 3 / 2 0 − 2 ) ⟶ ( − 1 / 2 ) ρ 2 − 2 ρ 1 ( ( 1 3 0 1 ) ⟶ − 3 ρ 2 + ρ 1 ( ( 1 0 0 1 )

    That translates to a matrix equation in this way.

    ( 1 − 3 0 1 ) ( − 2 0 0 − 1 / 2 ) ( 1 0 3 1 ) ( − 1 / 2 − 3 / 2 3 / 2 − 1 / 2 ) = I

    Taking inverses to solve for H yields this factorization.

    ( − 1 / 2 − 3 / 2 3 / 2 − 1 / 2 ) = ( 1 0 − 3 1 ) ( − 1 / 2 0 0 − 2 ) ( 1 3 0 1 ) I

  3. Exercise 3 Supplied answer

    If a map is nonsingular then to get from its representation to the identity matrix we do not need any column operations, so that in H = P B Q the matrix Q is the identity. An example of a nonsingular map is the transformation t − π / 4 : ℝ 2 → ℝ 2 that rotates vectors clockwise by π / 4  radians.

    1. Find the matrix H representing this map with respect to the standard bases.

    2. Use Gauss-Jordan to reduce H to the identity, without column operations.

    3. Translate that to a matrix equation T j T j − 1 ⋯ T 1 H = I .

    4. Solve the matrix equation for H .

    5. Describe H as a combination of dilations, flips, skews, and projections (the identity is a trivial projection).

    Back to Exercise 3

    Answer.

    1. Recall that rotation counterclockwise by θ  radians is represented with respect to the standard basis in this way.

      Rep ℰ 2 , ℰ 2 ( t π / 4 ) = ( cos ⁡ θ − sin ⁡ θ sin ⁡ θ cos ⁡ θ )

      A clockwise angle is the negative of a counterclockwise one.

      Rep ℰ 2 , ℰ 2 ( t − p i / 4 ) = ( cos ⁡ ( − π / 4 ) − sin ⁡ ( − π / 4 ) sin ⁡ ( − π / 4 ) cos ⁡ ( − π / 4 ) ) = ( 2 / 2 2 / 2 − 2 / 2 2 / 2 )

    2. This Gauss-Jordan reduction

      ⟶ ρ 1 + ρ 2 ( ( 2 / 2 2 / 2 0 2 ) ⟶ ( 1 / 2 ) ρ 2 ( 2 / 2 ) ρ 1 ( ( 1 1 0 1 ) ⟶ − ρ 2 + ρ 1 ( ( 1 0 0 1 )

      produces the identity matrix. Thus we do not need column-swapping operations to end with a partial-identity.

    3. In matrix multiplication the reduction is

      ( 1 − 1 0 1 ) ( 2 / 2 0 0 1 / 2 ) ( 1 0 1 1 ) H = I

      (note that composition of the Gaussian operations is from right to left).

    4. Taking inverses

      H = ( 1 0 − 1 1 ) ( 2 / 2 0 0 2 ) ( 1 1 0 1 ) ⏟ P I

      gives the desired factorization of H . The partial identity is I .

    5. Reading the composition from right to left (and ignoring the identity matrices as trivial) gives that H has the same effect as first performing this skew

      Supplied-answer-only horizontal shear maps (x,y) to (x+y,y), with input vectors u and v and their labelled images.

      followed by a dilation that multiplies all first components by 2 / 2 (this is a shrink in that 2 / 2 ≈ 0.707 is less than 1 ) and all second components by 2 , followed by another skew.

      Supplied-answer-only vertical shear maps (x,y) to (x,-x+y); the image of v lies below the horizontal axis.

      For an example we start with the unit vector whose angle with the x -axis is π / 6 and apply the components of H in turn.

      Supplied-answer-only three-stage rotation construction. Starting at (square root of three over two, one half), apply the shear x+y, dilation by square root of two over two horizontally and square root of two vertically, then shear -x+y. All intermediate coordinate labels are preserved.

      We can easily verify that the resulting vector has unit length and forms an angle with the x -axis of − π / 12 , which is indeed a rotation clockwise of π / 4 radians since ( π / 6 ) − ( π / 4 ) = − π / 12 .

  4. Exercise 4 Supplied answer

    Show that any linear transformation of ℝ 1 is a map h k that multiplies by a scalar x ↦ k x .

    Back to Exercise 4

    Answer. Represent it with respect to the standard bases ℰ 1 , ℰ 1 . That produces a 1 × 1 matrix. The only entry is the scalar  k .

  5. Exercise 5 Supplied answer

    Show that linear maps preserve the linear structures of a space.

    1. Show that for any linear map from ℝ n to ℝ m , the image of any line is a line. The image may be a degenerate line, that is, a single point.

    2. Show that the image of any linear surface is a linear surface. This generalizes the result that under a linear map the image of a subspace is a subspace.

    3. Linear maps preserve other linear ideas. Show that linear maps preserve “betweeness”: if the point B is between A and C then the image of B is between the image of A and the image of C .

    Back to Exercise 5

    Answer.

    1. A line is a subset of ℝ n of the form { v → = u → + t ⋅ w → ∣ t ∈ ℝ } . The image of a point on that line is h ( v → ) = h ( u → + t ⋅ w → ) = h ( u → ) + t ⋅ h ( w → ) , and the set of such vectors, as t ranges over the reals, is a line (albeit, degenerate if h ( w → ) = 0 → ).

    2. This is an obvious extension of the prior argument.

    3. If the point  B is between the points  A and  C then the line from A to C has B in it. That is, there is a t ∈ ( 0 . . 1 ) such that b → = a → + t ⋅ ( c → − a → ) (where B is the endpoint of b → , etc.). Now, as in the argument of the first item, linearity shows that h ( b → ) = h ( a → ) + t ⋅ h ( c → − a → ) .

  6. Exercise 6 Supplied answer

    Use a picture like the one that appears in the discussion of the Chain Rule to answer: if a function f : ℝ → ℝ has an inverse, what’s the relationship between how the function —locally, approximately —dilates space, and how its inverse dilates space (assuming, of course, that it has an inverse)?

    Back to Exercise 6

    Answer. The two are inverse. For instance, for a fixed x ∈ ℝ , if f ′ ( x ) = k (with k ≠ 0 ) then ( f − 1 ) ′ ( x ) = 1 / k .

    Supplied-answer-only inverse-function diagram. Neighborhoods around x, f(x), and f inverse of f(x) are connected on three number lines. The first and last intervals have equal widths.

  7. Exercise 7 Supplied answer

    Show that any permutation, any reordering, p of the numbers 1 , …, n , the map

    ( x 1 x 2 ⋮ x n ) ↦ ( x p ( 1 ) x p ( 2 ) ⋮ x p ( n ) )

    can be done with a composition of maps, each of which only swaps a single pair of coordinates. Hint: you can use induction on n . (Remark: in the fourth chapter we will show this and we will also show that the parity of the number of swaps used is determined by p . That is, although a particular permutation could be expressed in two different ways with two different numbers of swaps, either both ways use an even number of swaps, or both use an odd number.)

    Back to Exercise 7

    Answer. We can show this by induction on the number of components in the vector. In the n = 1 base case the only permutation is the trivial one, and the map

    ( x 1 ) ↦ ( x 1 )

    is expressible as a composition of swaps—as zero swaps. For the inductive step we assume that the map induced by any permutation of fewer than n numbers can be expressed with swaps only, and we consider the map induced by a permutation p of n numbers.

    ( x 1 x 2 ⋮ x n ) ↦ ( x p ( 1 ) x p ( 2 ) ⋮ x p ( n ) )

    Consider the number  i such that p ( i ) = n . The map

    ( x 1 x 2 ⋮ x i ⋮ x n ) ⟼ p ^ ( x p ( 1 ) x p ( 2 ) ⋮ x p ( n ) ⋮ x n )

    will, when followed by the swap of the i -th and n -th components, give the map  p . Now, the inductive hypothesis gives that p ^ is achievable as a composition of swaps.