Original English by Jim Hefferon — 34 validated sections. The original mathematics and supplied answers below are preserved. This is a partial-book reading edition, not the complete book or an Everyday-English rewrite.

Source, reuse and conversion details

Source revision df2262e089a02651c127f1dd12649c4622ee1383; CC BY-SA 2.5 option, with original component credits retained. This is not an Everyday-English rewrite. The complete original source section is included. Applicable definitions, statements, prior-question data and their original diagrams travel separately with modular extracts; supplied answers never become question context.

Conversion and source notes: OpenAI Codex — GPT-6, Ultra effort. Original author and source credits are retained.

Notes about the original source and its answers

These twenty-two findings are separate from the unchanged original text and formulas. Opening them may reveal answers. This is not an exhaustive mathematical correctness audit or human review. An answer container is not certification of a complete, correct worked solution.

  1. Source note 1: The expanded first derivative omits a factor of two on v_1s_1+v_2s_2. The final stationary-point formula is correct, but the intermediate equality is not.
  2. Source note 2: The second derivative is positive whenever the direction vector is nonzero, not only when neither component is zero. The one-zero-component case is not a trivial question.
  3. Source note 3: The displayed calculation divides by v dot v although the answer also covers v=0. The zero case requires a separate nonzero line generator.
  4. Source note 4: The closest-point comparison follows from the right angle and the Pythagorean theorem; the ordinary triangle inequality alone does not justify this hypotenuse comparison.
  5. Source note 5: A comma is missing between the second and third vectors in this displayed ordered orthonormal basis.
  6. Source note 6: The sentence repeats the word can.
  7. Source note 7: The orthogonal-vector exercise permits zero vectors, but its projection formulas use each vector as a line generator and need nonzero denominators.
  8. Source note 8: The last projection subscript uses v_k instead of the bound orthogonal vector kappa_k.
  9. Source note 9: These answer formulas use v_k or v_2 where the corresponding bound orthogonal vector is kappa_k or kappa_2.
  10. Source note 10: A displayed dot product takes a scalar fraction as its second argument before multiplying by kappa_i; the dot product should instead enclose the entire scalar-vector product.
  11. Source note 11: The orthogonal-basis calculation uses beta subscripts from the prior nonorthogonal basis while the explicit vectors are the current kappa basis.
  12. Source note 12: The squared norm in the Bessel example is 30, not 50.
  13. Source note 13: The first proof sentence invokes the null space of orthogonal projection before this lemma establishes that projection along M perp is well-defined. This is a proof-order dependency gap, not a false conclusion.
  14. Source note 14: An equals sign is missing between B_N and the displayed basis sequence.
  15. Source note 15: The phrase all the the repeats the word the.
  16. Source note 16: An equals sign is missing after M perp before its set description.
  17. Source note 17: The follow-up projection answer uses the wrong plane matrix and incorrectly says (1,2,-1) belongs to S. Its correct orthogonal projection onto the stated S is (0,2,0).
  18. Source note 18: As in the earlier line case, the closest-point proof needs the right-triangle Pythagorean comparison; the ordinary triangle inequality alone does not imply this conclusion.
  19. Source note 19: An equals sign is missing between the named projection and its matrix-vector expression.
  20. Source note 20: The final scalar coefficient is joined to w_n by dotprod instead of scalar multiplication.
  21. Source note 21: An equals sign is missing between the null-space symbol and its set description.
  22. Source note 22: The sentence To finish we taking a basis has a verb-form error.

Wide formulas, diagrams and tables scroll horizontally. Focus a region and use the arrow keys.

Source-preserving rebuild, navigation, source packaging and current deterministic checks: OpenAI Codex — GPT-6 Astra, Ultra effort. Jim Hefferon remains the author of the mathematics. Earlier intermediate-conversion runtime identity is not established by its retained receipts and is not reassigned to this rebuild. No human review or exhaustive proof certification is claimed.

Projection

This section is optional. It is a prerequisite only for the final two sections of Chapter Five, and some Topics.

We have described projection from ℝ 3 into its x y -plane subspace as a shadow map. This shows why but it also shows that some shadows fall upward.

Projection onto the xy-plane: the vector (1,2,2) points above the plane; its shadow has the same first two coordinates. A vertical downward arrow joins the vector tip to the shadow tip. Projection onto the xy-plane: the vector (1,2,minus 1) points below the plane. A vertical upward arrow joins its tip to the planar shadow tip.

So perhaps a better description is: the projection of v → is the vector p → in the plane with the property that someone standing on p → and looking straight up or down—that is, looking orthogonally to the plane— sees the tip of v → . In this section we will generalize this to other projections, orthogonal and non-orthogonal.

Orthogonal Projection Into a Line

We first consider orthogonal projection of a vector v → into a line ℓ . This shows a figure walking out on the line to a point p → such that the tip of v → is directly above them, where “above” does not mean parallel to the y -axis but instead means orthogonal to the line.

A person stands on the sloping line at the orthogonal projection of a vector. The dotted line of sight is perpendicular to the sloping line, not parallel to the vertical coordinate axis.

Since the line is the span of some vector ℓ = { c ⋅ s → ∣ c ∈ ℝ } , we have a coefficient c p → with the property that v → − c p → s → is orthogonal to c p → s → .

The vector v decomposes into its line projection c_p times s and the residual v minus c_p times s. The projection runs along the sloping line; the residual from its tip to the tip of v is perpendicular to that line.

To solve for this coefficient, observe that because v → − c p → s → is orthogonal to a scalar multiple of s → , it must be orthogonal to s → itself. Then ( v → − c p → s → ) ⋅ s → = 0 gives that c p → = v → ⋅ s → / s → ⋅ s → .

Definition 1.1 The orthogonal projection of v → into the line spanned by a nonzero s → is this vector.

proj [ s → ] ( v → ) = v → ⋅ s → s → ⋅ s → ⋅ s →

(That says ‘spanned by s → ’ instead the more formal ‘span of the set { s → } ’. This more casual phrase is common.)

Example 1.2 To orthogonally project the vector ( 2 3 ) into the line y = 2 x , first pick a direction vector for the line.

s → = ( 1 2 )

The calculation is easy.

The vector (2,3) projects orthogonally onto the line y=2x. Its projection is (8/5,16/5), shown by an arrow on the steeper line; a dotted segment joins the two tips.

( 2 3 ) ⋅ ( 1 2 ) ( 1 2 ) ⋅ ( 1 2 ) ⋅ ( 1 2 ) = 8 5 ⋅ ( 1 2 ) = ( 8 / 5 16 / 5 )

Example 1.3 In ℝ 3 , the orthogonal projection of a general vector

( x y z )

into the y -axis is

( x y z ) ⋅ ( 0 1 0 ) ( 0 1 0 ) ⋅ ( 0 1 0 ) ⋅ ( 0 1 0 ) = ( 0 y 0 )

which matches our intuitive expectation.

The picture above showing the figure walking out on the line until v → ’s tip is overhead is one way to think of the orthogonal projection of a vector into a line. We finish this subsection with two other ways.

Example 1.4 A railroad car left on an east-west track without its brake is pushed by a wind blowing toward the northeast at fifteen miles per hour; what speed will the car reach?

An unbraked railroad car on an east-west track is pushed by northeast-directed wind. Five wind arrows are drawn around and across the car; the track constrains its motion.

For the wind we use a vector of length 15 that points toward the northeast.

v → = ( 15 1 / 2 15 1 / 2 )

The car is only affected by the part of the wind blowing in the east-west direction—the part of v → in the direction of the x -axis is this (the picture has the same perspective as the railroad car picture above).

North and east directions accompany a northeast wind arrow and its horizontal eastward component. The railcar example uses a wind speed of 15 miles per hour.

p → = ( 15 1 / 2 0 )

So the car will reach a velocity of 15 1 / 2 miles per hour toward the east.

Thus, another way to think of the picture that precedes the definition is that it shows v → as decomposed into two parts, the part p → with the line, and the part that is orthogonal to the line (shown above on the north-south axis). These two are non-interacting in the sense that the east-west car is not at all affected by the north-south part of the wind (see Exercise 1.10). So we can think of the orthogonal projection of v → into the line spanned by s → as the part of v → that lies in the direction of s → .

Still another useful way to think of orthogonal projection into a line is to have the person stand on the vector, not the line. This person holds a rope looped over the line. As they pull, the loop slides on the line.

A person stands at the original vector tip and pulls a rope looped over the sloping line. When tight, the rope is perpendicular to the line; the light-gray arrow gives the closest point on the line.

When it is tight, the rope is orthogonal to the line. That is, we can think of the projection p → as being the vector in the line that is closest to v → (see Exercise 1.17).

Example 1.5 A submarine is tracking a ship moving along the line y = 3 x + 2 . Torpedo range is one-half mile. If the sub stays where it is, at the origin on the chart below, will the ship pass within range?

North and east axes show a ship track along the dashed affine line y=3x+2. The submarine is at the origin. The line does not pass through the origin; the adjacent example shifts the map before applying projection.

The formula for projection into a line does not immediately apply because the line doesn’t pass through the origin, and so isn’t the span of any s → . To adjust for this, we start by shifting the entire map down two units. Now the line is y = 3 x , a subspace. We project to get the point p → on the line closest to

v → = ( 0 − 2 )

the sub’s shifted position.

p → = ( 0 − 2 ) ⋅ ( 1 3 ) ( 1 3 ) ⋅ ( 1 3 ) ⋅ ( 1 3 ) = ( − 3 / 5 − 9 / 5 )

The distance between v → and p → is about 0.63  miles. The ship will never be in range.

Exercises

  1. Exercise 1.6 Worked answer

    Recommended. Project the first vector orthogonally into the line spanned by the second vector.

    1. ( 2 1 ) , ( 3 − 2 )

    2. ( 2 1 ) , ( 3 0 )

    3. ( 1 1 4 ) , ( 1 2 − 1 )

    4. ( 1 1 4 ) , ( 3 3 12 )

    Back to Exercise 1.6

    Answer.

    1. ( 2 1 ) ⋅ ( 3 − 2 ) ( 3 − 2 ) ⋅ ( 3 − 2 ) ⋅ ( 3 − 2 ) = 4 13 ⋅ ( 3 − 2 ) = ( 12 / 13 − 8 / 13 )

    2. ( 2 1 ) ⋅ ( 3 0 ) ( 3 0 ) ⋅ ( 3 0 ) ⋅ ( 3 0 ) = 2 3 ⋅ ( 3 0 ) = ( 2 0 )

    3. ( 1 1 4 ) ⋅ ( 1 2 − 1 ) ( 1 2 − 1 ) ⋅ ( 1 2 − 1 ) ⋅ ( 1 2 − 1 ) = − 1 6 ⋅ ( 1 2 − 1 ) = ( − 1 / 6 − 1 / 3 1 / 6 )

    4. ( 1 1 4 ) ⋅ ( 3 3 12 ) ( 3 3 12 ) ⋅ ( 3 3 12 ) ⋅ ( 3 3 12 ) = 1 3 ⋅ ( 3 3 12 ) = ( 1 1 4 )

  2. Exercise 1.7 Worked answer

    Recommended. Project the vector orthogonally into the line.

    1. ( 2 − 1 4 ) ,   { c ( − 3 1 − 3 ) ∣ c ∈ ℝ }

    2. ( − 1 − 1 ) , the line y = 3 x

    Back to Exercise 1.7

    Answer.

    1. ( 2 − 1 4 ) ⋅ ( − 3 1 − 3 ) ( − 3 1 − 3 ) ⋅ ( − 3 1 − 3 ) ⋅ ( − 3 1 − 3 ) = − 19 19 ⋅ ( − 3 1 − 3 ) = ( 3 − 1 3 )

    2. Writing the line as

      { c ⋅ ( 1 3 ) ∣ c ∈ ℝ }

      gives this projection.

      ( − 1 − 1 ) ⋅ ( 1 3 ) ( 1 3 ) ⋅ ( 1 3 ) ⋅ ( 1 3 ) = − 4 10 ⋅ ( 1 3 ) = ( − 2 / 5 − 6 / 5 )

  3. Exercise 1.8 Worked answer

    Although pictures guided our development of Definition 1.1, we are not restricted to spaces that we can draw. In ℝ 4 project this vector into this line.

    v → = ( 1 2 1 3 ) ℓ = { c ⋅ ( − 1 1 − 1 1 ) ∣ c ∈ ℝ }

    Back to Exercise 1.8

    Answer. ( 1 2 1 3 ) ⋅ ( − 1 1 − 1 1 ) ( − 1 1 − 1 1 ) ⋅ ( − 1 1 − 1 1 ) ⋅ ( − 1 1 − 1 1 ) = 3 4 ⋅ ( − 1 1 − 1 1 ) = ( − 3 / 4 3 / 4 − 3 / 4 3 / 4 )

  4. Exercise 1.9 Worked answer

    Recommended. Definition 1.1 uses two vectors s → and v → . Consider the transformation of ℝ 2 resulting from fixing

    s → = ( 3 1 )

    and projecting v → into the line that is the span of s → . Apply it to these vectors.

    1. ( 1 2 )

    2. ( 0 4 )

    Show that in general the projection transformation is this.

    ( x 1 x 2 ) ↦ ( ( 9 x 1 + 3 x 2 ) / 10 ( 3 x 1 + x 2 ) / 10 )

    Express the action of this transformation with a matrix.

    Back to Exercise 1.9

    Answer.

    1. ( 1 2 ) ⋅ ( 3 1 ) ( 3 1 ) ⋅ ( 3 1 ) ⋅ ( 3 1 ) = 1 2 ⋅ ( 3 1 ) = ( 3 / 2 1 / 2 )

    2. ( 0 4 ) ⋅ ( 3 1 ) ( 3 1 ) ⋅ ( 3 1 ) ⋅ ( 3 1 ) = 2 5 ⋅ ( 3 1 ) = ( 6 / 5 2 / 5 )

    In general the projection is this.

    ( x 1 x 2 ) ⋅ ( 3 1 ) ( 3 1 ) ⋅ ( 3 1 ) ⋅ ( 3 1 ) = 3 x 1 + x 2 10 ⋅ ( 3 1 ) = ( ( 9 x 1 + 3 x 2 ) / 10 ( 3 x 1 + x 2 ) / 10 )

    The appropriate matrix is this.

    ( 9 / 10 3 / 10 3 / 10 1 / 10 )

  5. Exercise 1.10 Worked answer

    Example 1.4 suggests that projection breaks v → into two parts, proj [ s → ] ( v → ) and v → − proj [ s → ] ( v → ) , that are non-interacting. Recall that the two are orthogonal. Show that any two nonzero orthogonal vectors make up a linearly independent set.

    Back to Exercise 1.10

    Answer. Suppose that v → 1 and v → 2 are nonzero and orthogonal. Consider the linear relationship c 1 v → 1 + c 2 v → 2 = 0 → . Take the dot product of both sides of the equation with v → 1 to get that

    v → 1 ⋅ ( c 1 v → 1 + c 2 v → 2 ) = c 1 ⋅ ( v → 1 ⋅ v → 1 ) + c 2 ⋅ ( v → 1 ⋅ v → 2 ) = c 1 ⋅ ( v → 1 ⋅ v → 1 ) + c 2 ⋅ 0 = c 1 ⋅ ( v → 1 ⋅ v → 1 )

    is equal to v → 1 ⋅ 0 → = 0 → . With the assumption that v → 1 is nonzero, this gives that c 1 is zero. Showing that c 2 is zero is similar.

  6. Exercise 1.11 Worked answer

    1. What is the orthogonal projection of v → into a line if v → is a member of that line?

    2. Show that if v → is not a member of the line then the set { v → , v → − proj [ s → ] ( v → ) } is linearly independent.

    Back to Exercise 1.11

    Answer.

    1. If the vector v → is in the line then the orthogonal projection is v → . To verify this by calculation, note that since v → is in the line we have that v → = c v → ⋅ s → for some scalar c v → .

      v → ⋅ s → s → ⋅ s → ⋅ s → = c v → ⋅ s → ⋅ s → s → ⋅ s → ⋅ s → = c v → ⋅ s → ⋅ s → s → ⋅ s → ⋅ s → = c v → ⋅ 1 ⋅ s → = v →

      (Remark. If we assume that v → is nonzero then we can simplify the above by taking s → to be v → .)

    2. Write c p → s → for the projection proj [ s → ] ( v → ) . Note that, by the assumption that v → is not in the line, both v → and v → − c p → s → are nonzero. Note also that if c p → is zero then we are actually considering the one-element set { v → } , and with v → nonzero, this set is necessarily linearly independent. Therefore, we are left considering the case that c p → is nonzero.

      Setting up a linear relationship

      a 1 ( v → ) + a 2 ( v → − c p → s → ) = 0 →

      leads to the equation ( a 1 + a 2 ) ⋅ v → = a 2 c p → ⋅ s → . Because v → isn’t in the line, the scalars a 1 + a 2 and a 2 c p → must both be zero. We handled the c p → = 0 case above, so the remaining case is that a 2 = 0 , and this gives that a 1 = 0 also. Hence the set is linearly independent.

  7. Exercise 1.12 Worked answer

    Definition 1.1 requires that s → be nonzero. Why? What is the right definition of the orthogonal projection of a vector into the (degenerate) line spanned by the zero vector?

    Back to Exercise 1.12

    Answer. If s → is the zero vector then the expression

    proj [ s → ] ( v → ) = v → ⋅ s → s → ⋅ s → ⋅ s →

    contains a division by zero, and so is undefined. As for the right definition, for the projection to lie in the span of the zero vector, it must be defined to be 0 → .

  8. Exercise 1.13 Worked answer

    Are all vectors the projection of some other vector into some line?

    Back to Exercise 1.13

    Answer. Any vector in ℝ n is the projection of some other into a line, provided that the dimension n is greater than one. (Clearly, any vector is the projection of itself into a line containing itself; the question is to produce some vector other than v → that projects to v → .)

    Suppose that v → ∈ ℝ n with n > 1 . If v → ≠ 0 → then we consider the line ℓ = { c v → ∣ c ∈ ℝ } and if v → = 0 → we take ℓ to be any (non-degenerate) line at all (actually, we needn’t distinguish between these two cases—see the prior exercise). Let v 1 , … , v n be the components of v → ; since n > 1 , there are at least two. If some v i is zero then the vector w → = e → i is perpendicular to v → . If none of the components is zero then the vector w → whose components are v 2 , − v 1 , 0 , … , 0 is perpendicular to v → . In either case, observe that v → + w → does not equal v → , and that v → is the projection of v → + w → into ℓ .

    ( v → + w → ) ⋅ v → v → ⋅ v → ⋅ v → = ( v → ⋅ v → v → ⋅ v → + w → ⋅ v →   v → ⋅ v → ) ⋅ v → = v → ⋅ v → v → ⋅ v → ⋅ v → = v →

    We can dispose of the remaining n = 0 and n = 1 cases. The dimension n = 0 case is the trivial vector space, here there is only one vector and so it cannot be expressed as the projection of a different vector. In the dimension n = 1 case there is only one (non-degenerate) line, and every vector is in it, hence every vector is the projection only of itself.

  9. Exercise 1.14 Worked answer

    Show that the projection of v → into the line spanned by s → has length equal to the absolute value of the number v → ⋅ s → divided by the length of the vector s → .

    Back to Exercise 1.14

    Answer. The proof is a calculation. Recall that for any vector u → , the length is determined by | u → | 2 = u → ⋅ u → . So this is the square of the length.

    | v → ⋅ s → s → ⋅ s → ⋅ s → | 2 = ( v → ⋅ s → s → ⋅ s → ⋅ s → ) ⋅ ( v → ⋅ s → s → ⋅ s → ⋅ s → ) = ( v → ⋅ s → s → ⋅ s → ) 2 ⋅ ( s → ⋅ s → ) = ( v → ⋅ s → ) 2 ( s → ⋅ s → ) 2 ⋅ ( s → ⋅ s → ) = ( v → ⋅ s → ) 2 | s → | 2

  10. Exercise 1.15 Worked answer

    Find the formula for the distance from a point to a line.

    Back to Exercise 1.15

    Answer. Because the projection of v → into the line spanned by s → is this,

    v → ⋅ s → s → ⋅ s → ⋅ s →

    the distance squared from the point to the line is this.

    | v → − v → ⋅ s → s → ⋅ s → ⋅ s → | 2 = ( v → − v → ⋅ s → s → ⋅ s → ⋅ s → ) ⋅ ( v → − v → ⋅ s → s → ⋅ s → ⋅ s → ) = v → ⋅ v → − v → ⋅ ( v → ⋅ s → s → ⋅ s → ⋅ s → ) − ( v → ⋅ s → s → ⋅ s → ⋅ s → ) ⋅ v → + ( v → ⋅ s → s → ⋅ s → ⋅ s → ) ⋅ ( v → ⋅ s → s → ⋅ s → ⋅ s → ) = v → ⋅ v → − 2 ⋅ ( v → ⋅ s → s → ⋅ s → ) ⋅ v → ⋅ s → + ( v → ⋅ s → s → ⋅ s → ) 2 ⋅ ( s → ⋅ s → ) = ( v → ⋅ v → ) ⋅ ( s → ⋅ s → ) − 2 ⋅ ( v → ⋅ s → ) 2 + ( v → ⋅ s → ) 2 s → ⋅ s → = ( v → ⋅ v → ) ( s → ⋅ s → ) − ( v → ⋅ s → ) 2 s → ⋅ s →

  11. Exercise 1.16 Worked answer

    Find the scalar c such that the point ( c s 1 , c s 2 ) is a minimum distance from the point ( v 1 , v 2 ) by using Calculus (i.e., consider the distance function, set the first derivative equal to zero, and solve). Generalize to ℝ n .

    Back to Exercise 1.16

    Answer. Because square root is a strictly increasing function, we can minimize d ( c ) = ( c s 1 − v 1 ) 2 + ( c s 2 − v 2 ) 2 instead of the square root of d . The derivative is d d / d c = 2 ( c s 1 − v 1 ) ⋅ s 1 + 2 ( c s 2 − v 2 ) ⋅ s 2 . Setting it equal to zero 2 ( c s 1 − v 1 ) ⋅ s 1 + 2 ( c s 2 − v 2 ) ⋅ s 2 = c ⋅ ( 2 s 1 2 + 2 s 2 2 ) − ( v 1 s 1 + v 2 s 2 ) = 0 gives the only critical point.

    c = v 1 s 1 + v 2 s 2 s 1 2 + s 2 2 = v → ⋅ s → s → ⋅ s →

    Now the second derivative with respect to c

    d 2 d d c 2 = 2 s 1 2 + 2 s 2 2

    is strictly positive (as long as neither s 1 nor s 2 is zero, in which case the question is trivial) and so the critical point is a minimum.

    The generalization to ℝ n is straightforward. Consider d n ( c ) = ( c s 1 − v 1 ) 2 + ⋯ + ( c s n − v n ) 2 , take the derivative, etc.

  12. Exercise 1.17 Worked answer

    Recommended. Let p → be the orthogonal projection of v → ∈ ℝ n onto a line  ℓ . Show that p → is the point in the line closest to  v → .

    Back to Exercise 1.17

    Answer. Suppose w → ∈ ℓ . Because this is orthogonal projection, the two vectors v → − p → and p → − w → are at a right angle. The Triangle Inequality applies and the hypotenuse v → − w → is therefore at least as long as v → − p → .

  13. Exercise 1.18 Worked answer

    Prove that the orthogonal projection of a vector into a line has length less than or equal to that of the vector.

    Back to Exercise 1.18

    Answer. For any vector u → , the length squared is given by | u → | 2 = u → ⋅ u → . So this is the square of the projection’s length.

    | v → ⋅ s → s → ⋅ s → ⋅ s → | 2 = ( v → ⋅ s → s → ⋅ s → ⋅ s → ) ⋅ ( v → ⋅ s → s → ⋅ s → ⋅ s → ) = ( v → ⋅ s → s → ⋅ s → ) 2 ⋅ ( s → ⋅ s → ) = ( v → ⋅ s → ) 2 ( s → ⋅ s → ) 2 ⋅ ( s → ⋅ s → ) = ( v → ⋅ s → ) 2 | s → | 2

    Thus the length is | v → ⋅ s → | / | s → | , the absolute value of v → ⋅ s → divided by the length of  s → . The Cauchy-Schwarz inequality, | v → ⋅ s → | ≤ | v → | ⋅ | s → | , gives that | v → ⋅ s → | / | s → | ≤ | v → | ⋅ | s → | / | s → | = | v → | .

  14. Exercise 1.19 Worked answer

    Recommended. Show that the definition of orthogonal projection into a line does not depend on the spanning vector: if s → is a nonzero multiple of q → then ( v → ⋅ s → / s → ⋅ s → ) ⋅ s → equals ( v → ⋅ q → / q → ⋅ q → ) ⋅ q → .

    Back to Exercise 1.19

    Answer. Write c s → for q → , and calculate: ( v → ⋅ c s → / c s → ⋅ c s → ) ⋅ c s → = ( v → ⋅ s → / s → ⋅ s → ) ⋅ s → .

  15. Exercise 1.20 Worked answer

    Consider the function mapping the plane to itself that takes a vector to its projection into the line y = x . These two each show that the map is linear, the first one in a way that is coordinate-bound (that is, it fixes a basis and then computes) and the second in a way that is more conceptual.

    1. Produce a matrix that describes the function’s action.

    2. Show that we can obtain this map by first rotating everything in the plane π / 4  radians clockwise, then projecting into the x -axis, and then rotating π / 4  radians counterclockwise.

    Back to Exercise 1.20

    Answer.

    1. Fixing

      s → = ( 1 1 )

      as the vector whose span is the line, the formula gives this action,

      ( x y ) ↦ ( x y ) ⋅ ( 1 1 ) ( 1 1 ) ⋅ ( 1 1 ) ⋅ ( 1 1 ) = x + y 2 ⋅ ( 1 1 ) = ( ( x + y ) / 2 ( x + y ) / 2 )

      which is the effect of this matrix.

      ( 1 / 2 1 / 2 1 / 2 1 / 2 )

    2. Rotating the entire plane π / 4  radians clockwise brings the y = x  line to lie on the x -axis. Now projecting and then rotating back has the desired effect.

  16. Exercise 1.21 Worked answer

    For a → , b → ∈ ℝ n let v → 1 be the projection of a → into the line spanned by b → , let v → 2 be the projection of v → 1 into the line spanned by a → , let v → 3 be the projection of v → 2 into the line spanned by b → , etc., back and forth between the spans of a → and b → . That is, v → i + 1 is the projection of v → i into the span of a → if i + 1 is even, and into the span of b → if i + 1 is odd. Must that sequence of vectors eventually settle down—must there be a sufficiently large  i such that v → i + 2 equals v → i and v → i + 3 equals v → i + 1 ? If so, what is the earliest such  i ?

    Back to Exercise 1.21

    Answer. The sequence need not settle down. With

    a → = ( 1 0 ) b → = ( 1 1 )

    the projections are these.

    v → 1 = ( 1 / 2 1 / 2 ) , v → 2 = ( 1 / 2 0 ) , v → 3 = ( 1 / 4 1 / 4 ) , …

    This sequence doesn’t repeat.

    Supplied-answer-only diagram: repeated perpendicular projections between the horizontal axis and a sloping line form progressively smaller right-triangle steps toward the origin.

Gram-Schmidt Orthogonalization

The prior subsection suggests that projecting v → into the line spanned by s → decomposes that vector into two parts

The vector v, its orthogonal projection proj_[s](v) onto the sloping line and the residual v minus proj_[s](v) form a right triangle. v → = proj [ s → ] ( v → ) + ( v → − proj [ s → ] ( v → ) )

that are orthogonal and so are “non-interacting.” We now develop that suggestion.

Definition 2.1 Vectors v → 1 , … , v → k ∈ ℝ n are mutually orthogonal when any two are orthogonal: if i ≠ j then the dot product v → i ⋅ v → j is zero.

Theorem 2.2 If the vectors in a set { v → 1 , … , v → k } ⊂ ℝ n are mutually orthogonal and nonzero then that set is linearly independent.

Proof Consider 0 → = c 1 v → 1 + c 2 v → 2 + ⋯ + c k v → k . For i ∈ { 1 , . . , k } , taking the dot product of v → i with both sides of the equation v → i ⋅ ( c 1 v → 1 + c 2 v → 2 + ⋯ + c k v → k ) = v → i ⋅ 0 → , which gives c i ⋅ ( v → i ⋅ v → i ) = 0 , shows that c i = 0 since v → i ≠ 0 → .

QED

Corollary 2.3 In a k  dimensional vector space, if the vectors in a size  k set are mutually orthogonal and nonzero then that set is a basis for the space.

Proof Any linearly independent size  k subset of a k  dimensional space is a basis.

QED

Of course, the converse of Corollary 2.3 does not hold—not every basis of every subspace of ℝ n has mutually orthogonal vectors. However, we can get the partial converse that for every subspace of ℝ n there is at least one basis consisting of mutually orthogonal vectors.

Example 2.4 The members β → 1 and β → 2 of this basis for ℝ 2 are not orthogonal.

B = ⟨ ( 4 2 ) , ( 1 3 ) ⟩

The nonorthogonal basis beta one=(4,2) and beta two=(1,3) is drawn on plane coordinate axes.

We will derive from B a new basis for the space ⟨ κ → 1 , κ → 2 ⟩ consisting of mutually orthogonal vectors. The first member of the new basis is just β → 1 .

κ → 1 = ( 4 2 )

For the second member of the new basis, we subtract from β → 2 the part in the direction of κ → 1 . This leaves the part of β → 2 that is orthogonal to κ → 1 .

κ → 2 = ( 1 3 ) − proj [ κ → 1 ] ( ( 1 3 ) ) = ( 1 3 ) − ( 2 1 ) = ( − 1 2 )

Gram-Schmidt keeps kappa one=(4,2) and replaces the second vector with kappa two=(minus 1,2), perpendicular to kappa one. The subtracted line component is shown head-to-tail.

By the corollary ⟨ κ → 1 , κ → 2 ⟩ is a basis for ℝ 2 .

Definition 2.5 An orthogonal basis for a vector space is a basis of mutually orthogonal vectors.

Example 2.6 To produce from this basis for ℝ 3

B = ⟨ ( 1 1 1 ) , ( 0 2 0 ) , ( 1 0 3 ) ⟩

an orthogonal basis, start by taking the first vector unchanged.

κ → 1 = ( 1 1 1 )

Get κ → 2 by subtracting from β → 2 its part in the direction of κ → 1 .

κ → 2 = ( 0 2 0 ) − proj [ κ → 1 ] ( ( 0 2 0 ) ) = ( 0 2 0 ) − ( 2 / 3 2 / 3 2 / 3 ) = ( − 2 / 3 4 / 3 − 2 / 3 )

Find κ → 3 by subtracting from β → 3 the part in the direction of κ → 1 and also the part in the direction of κ → 2 .

κ → 3 = ( 1 0 3 ) − proj [ κ → 1 ] ( ( 1 0 3 ) ) − proj [ κ → 2 ] ( ( 1 0 3 ) ) = ( − 1 0 1 )

As above, the corollary gives that the result is a basis for ℝ 3 .

⟨ ( 1 1 1 ) , ( − 2 / 3 4 / 3 − 2 / 3 ) , ( − 1 0 1 ) ⟩

Theorem 2.7 (Gram-Schmidt orthogonalization) If ⟨ β → 1 , … β → k ⟩ is a basis for a subspace of ℝ n then the vectors

κ → 1 = β → 1 κ → 2 = β → 2 − proj [ κ → 1 ] ( β → 2 ) κ → 3 = β → 3 − proj [ κ → 1 ] ( β → 3 ) − proj [ κ → 2 ] ( β → 3 ) ⋮ = κ → k = β → k − proj [ κ → 1 ] ( β → k ) − ⋯ − proj [ κ → k − 1 ] ( β → k )

form an orthogonal basis for the same subspace.

Remark 2.8 This is restricted to ℝ n only because we have not given a definition of orthogonality for other spaces.

Proof We will use induction to check that each κ → i is nonzero, is in the span of ⟨ β → 1 , … β → i ⟩ , and is orthogonal to all preceding vectors κ → 1 ⋅ κ → i = ⋯ = κ → i − 1 ⋅ κ → i = 0 . Then Corollary 2.3 gives that ⟨ κ → 1 , … κ → k ⟩ is a basis for the same space as is the starting basis.

We shall only cover the cases up to i = 3 , to give the sense of the argument. The full argument is Exercise 2.28.

The i = 1 case is trivial; taking κ → 1 to be β → 1 makes it a nonzero vector since β → 1 is a member of a basis, it is obviously in the span of ⟨ β → 1 ⟩ , and the ‘orthogonal to all preceding vectors’ condition is satisfied vacuously.

In the i = 2 case the expansion

κ → 2 = β → 2 − proj [ κ → 1 ] ( β → 2 ) = β → 2 − β → 2 ⋅ κ → 1 κ → 1 ⋅ κ → 1 ⋅ κ → 1 = β → 2 − β → 2 ⋅ κ → 1 κ → 1 ⋅ κ → 1 ⋅ β → 1

shows that κ → 2 ≠ 0 → or else this would be a non-trivial linear dependence among the β → ’s (it is nontrivial because the coefficient of β → 2 is 1 ). It also shows that κ → 2 is in the span of ⟨ β → 1 , β → 2 ⟩ . And, κ → 2 is orthogonal to the only preceding vector

κ → 1 ⋅ κ → 2 = κ → 1 ⋅ ( β → 2 − proj [ κ → 1 ] ( β → 2 ) ) = 0

because this projection is orthogonal.

The i = 3 case is the same as the i = 2 case except for one detail. As in the i = 2 case, expand the definition.

κ → 3 = β → 3 − β → 3 ⋅ κ → 1 κ → 1 ⋅ κ → 1 ⋅ κ → 1 − β → 3 ⋅ κ → 2 κ → 2 ⋅ κ → 2 ⋅ κ → 2 = β → 3 − β → 3 ⋅ κ → 1 κ → 1 ⋅ κ → 1 ⋅ β → 1 − β → 3 ⋅ κ → 2 κ → 2 ⋅ κ → 2 ⋅ ( β → 2 − β → 2 ⋅ κ → 1 κ → 1 ⋅ κ → 1 ⋅ β → 1 )

By the first line κ → 3 ≠ 0 → , since β → 3 isn’t in the span  [ β → 1 , β → 2 ] and therefore by the inductive hypothesis it isn’t in the span  [ κ → 1 , κ → 2 ] . By the second line κ → 3 is in the span of the first three β → ’s. Finally, the calculation below shows that κ → 3 is orthogonal to κ → 1 .

κ → 1 ⋅ κ → 3 = κ → 1 ⋅ ( β → 3 − proj [ κ → 1 ] ( β → 3 ) − proj [ κ → 2 ] ( β → 3 ) ) = κ → 1 ⋅ ( β → 3 − proj [ κ → 1 ] ( β → 3 ) ) − κ → 1 ⋅ proj [ κ → 2 ] ( β → 3 ) = 0

(Here is the difference with the i = 2 case: as happened for i = 2 the first term is  0 because this projection is orthogonal, but here the second term in the second line is  0 because κ → 1 is orthogonal to κ → 2 and so is orthogonal to any vector in the line spanned by κ → 2 .) A similar check shows that κ → 3 is also orthogonal to κ → 2 .

QED

In addition to having the vectors in the basis be orthogonal, we can also normalize each vector by dividing by its length, to end with an orthonormal basis. .

Example 2.9 From the orthogonal basis of Example 2.6, normalizing produces this orthonormal basis.

⟨ ( 1 / 3 1 / 3 1 / 3 ) , ( − 1 / 6 2 / 6 − 1 / 6 ) , ( − 1 / 2 0 1 / 2 ) ⟩

Besides its intuitive appeal, and its analogy with the standard basis ℰ n for ℝ n , an orthonormal basis also simplifies some computations. Exercise 2.22 is an example.

Exercises

  1. Exercise 2.10 Worked answer

    Normalize the lengths of these vectors.

    1. ( 1 2 )

    2. ( − 1 3 0 )

    3. ( 1 − 1 )

    Back to Exercise 2.10

    Answer.

    1. ( 1 / 5 2 / 5 ) = 1 5 ⋅ ( 1 2 )

    2. 1 10 ⋅ ( − 1 3 0 )

    3. 1 2 ⋅ ( 1 − 1 )

  2. Exercise 2.11 Worked answer

    Recommended. Perform Gram-Schmidt on this basis for  ℝ 2 .

    ⟨ ( 1 1 ) , ( − 1 2 ) ⟩

    Check that the resulting vectors are orthogonal.

    Back to Exercise 2.11

    Answer. Call the given basis B = ⟨ β → 1 , β → 2 ⟩ . First, κ → 1 = β → 1 . For the other, κ → 2 = β → 2 − proj [ κ → 1 ] ( β → 2 ) .

    κ → 2 = ( − 1 2 ) − ( − 1 2 ) ⋅ ( 1 1 ) ( 1 1 ) ⋅ ( 1 1 ) ⋅ ( 1 1 ) = ( − 3 / 2 3 / 2 )

    To check that they are orthogonal, just note that their dot product is zero.

  3. Exercise 2.12 Worked answer

    Recommended. Perform the Gram-Schmidt process on this basis for ℝ 3 .

    ⟨ ( 1 2 3 ) , ( 2 1 − 3 ) , ( 3 3 3 ) ⟩

    Back to Exercise 2.12

    Answer. Call the given basis B = ⟨ β → 1 , β → 2 , β → 3 ⟩ . First, κ → 1 = β → 1 .

    Next, κ → 2 = β → 2 − proj [ κ → 1 ] ( β → 2 ) .

    κ → 2 = ( 2 1 − 3 ) − ( 2 1 − 3 ) ⋅ ( 1 2 3 ) ( 1 2 3 ) ⋅ ( 1 2 3 ) ⋅ ( 1 2 3 ) = ( 33 / 14 24 / 14 − 27 / 14 )

    For the third, this is the formula

    κ → 3 = β → 3 − proj [ κ → 1 ] ( β → 3 ) − proj [ κ → 2 ] ( β → 3 )

    and here is the calculation.

    ( 3 3 3 ) − ( 3 3 3 ) ⋅ ( 1 2 3 ) ( 1 2 3 ) ⋅ ( 1 2 3 ) ⋅ ( 1 2 3 ) − ( 3 3 3 ) ⋅ ( 33 / 14 24 / 14 − 27 / 14 ) ( 33 / 14 24 / 14 − 27 / 14 ) ⋅ ( 33 / 14 24 / 14 − 27 / 14 ) ⋅ ( 33 / 14 24 / 14 − 27 / 14 ) = ( 9 / 19 − 9 / 19 3 / 19 )

  4. Exercise 2.13 Worked answer

    Recommended. Perform Gram-Schmidt on each of these bases for ℝ 2 .

    1. ⟨ ( 1 1 ) , ( 2 1 ) ⟩

    2. ⟨ ( 0 1 ) , ( − 1 3 ) ⟩

    3. ⟨ ( 0 1 ) , ( − 1 0 ) ⟩

    Then turn those orthogonal bases into orthonormal bases.

    Back to Exercise 2.13

    Answer.

    1. κ → 1 = ( 1 1 ) κ → 2 = ( 2 1 ) − proj [ κ → 1 ] ( ( 2 1 ) ) = ( 2 1 ) − ( 2 1 ) ⋅ ( 1 1 ) ( 1 1 ) ⋅ ( 1 1 ) ⋅ ( 1 1 ) = ( 2 1 ) − 3 2 ⋅ ( 1 1 ) = ( 1 / 2 − 1 / 2 )

      This is the corresponding orthonormal basis.

      ⟨ ( 1 / 2 1 / 2 ) , ( 2 / 2 − 2 / 2 ) ⟩

    2. κ → 1 = ( 0 1 ) κ → 2 = ( − 1 3 ) − proj [ κ → 1 ] ( ( − 1 3 ) ) = ( − 1 3 ) − ( − 1 3 ) ⋅ ( 0 1 ) ( 0 1 ) ⋅ ( 0 1 ) ⋅ ( 0 1 ) = ( − 1 3 ) − 3 1 ⋅ ( 0 1 ) = ( − 1 0 )

      Here is the orthonormal basis.

      ⟨ ( 0 1 ) , ( − 1 0 ) ⟩

    3. κ → 1 = ( 0 1 ) κ → 2 = ( − 1 0 ) − proj [ κ → 1 ] ( ( − 1 0 ) ) = ( − 1 0 ) − ( − 1 0 ) ⋅ ( 0 1 ) ( 0 1 ) ⋅ ( 0 1 ) ⋅ ( 0 1 ) = ( − 1 0 ) − 0 1 ⋅ ( 0 1 ) = ( − 1 0 )

    This is the associated orthonormal basis.

    ⟨ ( 0 1 ) , ( − 1 0 ) ⟩

  5. Exercise 2.14 Worked answer

    Perform the Gram-Schmidt process on each of these bases for ℝ 3 .

    1. ⟨ ( 2 2 2 ) , ( 1 0 − 1 ) , ( 0 3 1 ) ⟩

    2. ⟨ ( 1 − 1 0 ) , ( 0 1 0 ) , ( 2 3 1 ) ⟩

    Then turn those orthogonal bases into orthonormal bases.

    Back to Exercise 2.14

    Answer.

    1. The first basis vector is unchanged.

      κ → 1 = ( 2 2 2 )

      The second one comes from this calculation.

      κ → 2 = ( 1 0 − 1 ) − proj [ κ → 1 ] ( ( 1 0 − 1 ) ) = ( 1 0 − 1 ) − ( 1 0 − 1 ) ⋅ ( 2 2 2 ) ( 2 2 2 ) ⋅ ( 2 2 2 ) ⋅ ( 2 2 2 ) = ( 1 0 − 1 ) − 0 12 ⋅ ( 2 2 2 ) = ( 1 0 − 1 )

      For the third the arithmetic is uglier but it is a straightforward calculation.

      κ → 3 = ( 0 3 1 ) − proj [ κ → 1 ] ( ( 0 3 1 ) ) − proj [ κ → 2 ] ( ( 0 3 1 ) ) = ( 0 3 1 ) − ( 0 3 1 ) ⋅ ( 2 2 2 ) ( 2 2 2 ) ⋅ ( 2 2 2 ) ⋅ ( 2 2 2 ) − ( 0 3 1 ) ⋅ ( 1 0 − 1 ) ( 1 0 − 1 ) ⋅ ( 1 0 − 1 ) ⋅ ( 1 0 − 1 ) = ( 0 3 1 ) − 8 12 ⋅ ( 2 2 2 ) − − 1 2 ⋅ ( 1 0 − 1 ) = ( − 5 / 6 5 / 3 − 5 / 6 )

      This is the orthonormal basis.

      ⟨ ( 1 / 3 1 / 3 1 / 3 ) , ( 1 / 2 0 − 1 / 2 ) , ( − 1 / 6 2 / 6 − 1 / 6 ) ⟩

    2. The first basis vector is what was given.

      κ → 1 = ( 1 − 1 0 )

      The second is here.

      κ → 2 = ( 0 1 0 ) − proj [ κ → 1 ] ( ( 0 1 0 ) ) = ( 0 1 0 ) − ( 0 1 0 ) ⋅ ( 1 − 1 0 ) ( 1 − 1 0 ) ⋅ ( 1 − 1 0 ) ⋅ ( 1 − 1 0 ) = ( 0 1 0 ) − − 1 2 ⋅ ( 1 − 1 0 ) = ( 1 / 2 1 / 2 0 )

      Here is the third.

      κ → 3 = ( 2 3 1 ) − proj [ κ → 1 ] ( ( 2 3 1 ) ) − proj [ κ → 2 ] ( ( 2 3 1 ) ) = ( 2 3 1 ) − ( 2 3 1 ) ⋅ ( 1 − 1 0 ) ( 1 − 1 0 ) ⋅ ( 1 − 1 0 ) ⋅ ( 1 − 1 0 ) − ( 2 3 1 ) ⋅ ( 1 / 2 1 / 2 0 ) ( 1 / 2 1 / 2 0 ) ⋅ ( 1 / 2 1 / 2 0 ) ⋅ ( 1 / 2 1 / 2 0 ) = ( 2 3 1 ) − − 1 2 ⋅ ( 1 − 1 0 ) − 5 / 2 1 / 2 ⋅ ( 1 / 2 1 / 2 0 ) = ( 0 0 1 )

      Here is the associated orthonormal basis.

      ⟨ ( 1 / 2 − 1 / 2 0 ) , ( 1 / 2 1 / 2 0 ) ( 0 0 1 ) ⟩

  6. Exercise 2.15 Worked answer

    Recommended. Find an orthonormal basis for this subspace of ℝ 3 : the plane x − y + z = 0 .

    Back to Exercise 2.15

    Answer. We can parametrize the given space can in this way.

    { ( x y z ) ∣ x = y − z } = { ( 1 1 0 ) ⋅ y + ( − 1 0 1 ) ⋅ z ∣ y , z ∈ ℝ }

    So we take the basis

    ⟨ ( 1 1 0 ) , ( − 1 0 1 ) ⟩

    apply the Gram-Schmidt process to get this first basis vector

    κ → 1 = ( 1 1 0 )

    and this second one.

    κ → 2 = ( − 1 0 1 ) − proj [ κ → 1 ] ( ( − 1 0 1 ) ) = ( − 1 0 1 ) − ( − 1 0 1 ) ⋅ ( 1 1 0 ) ( 1 1 0 ) ⋅ ( 1 1 0 ) ⋅ ( 1 1 0 ) = ( − 1 0 1 ) − − 1 2 ⋅ ( 1 1 0 ) = ( − 1 / 2 1 / 2 1 )

    and then normalize.

    ⟨ ( 1 / 2 1 / 2 0 ) , ( − 1 / 6 1 / 6 2 / 6 ) ⟩

  7. Exercise 2.16 Worked answer

    Find an orthonormal basis for this subspace of ℝ 4 .

    { ( x y z w ) ∣ x − y − z + w = 0  and  x + z = 0 }

    Back to Exercise 2.16

    Answer. Reducing the linear system

    x − y − z + w = 0 x + z = 0 ⟶ − ρ 1 + ρ 2 ( x − y − z + w = 0 y + 2 z − w = 0

    and parametrizing gives this description of the subspace.

    { ( − 1 − 2 1 0 ) ⋅ z + ( 0 1 0 1 ) ⋅ w ∣ z , w ∈ ℝ }

    So we take the basis,

    ⟨ ( − 1 − 2 1 0 ) , ( 0 1 0 1 ) ⟩

    go through the Gram-Schmidt process with the first

    κ → 1 = ( − 1 − 2 1 0 )

    and second basis vectors

    κ → 2 = ( 0 1 0 1 ) − proj [ κ → 1 ] ( ( 0 1 0 1 ) ) = ( 0 1 0 1 ) − ( 0 1 0 1 ) ⋅ ( − 1 − 2 1 0 ) ( − 1 − 2 1 0 ) ⋅ ( − 1 − 2 1 0 ) ⋅ ( − 1 − 2 1 0 ) = ( 0 1 0 1 ) − − 2 6 ⋅ ( − 1 − 2 1 0 ) = ( − 1 / 3 1 / 3 1 / 3 1 )

    and finish by normalizing.

    ⟨ ( − 1 / 6 − 2 / 6 1 / 6 0 ) , ( − 3 / 6 3 / 6 3 / 6 3 / 2 ) ⟩

  8. Exercise 2.17 Worked answer

    Show that any linearly independent subset of ℝ n can be orthogonalized without changing its span.

    Back to Exercise 2.17

    Answer. A linearly independent subset of ℝ n is a basis for its own span. Apply Theorem 2.7.

    Remark. Here’s why the phrase ‘linearly independent’ is in the question. Dropping the phrase would require us to worry about two things. The first thing to worry about is that when we do the Gram-Schmidt process on a linearly dependent set then we get some zero vectors. For instance, with

    S = { ( 1 2 ) , ( 3 6 ) }

    we would get this.

    κ → 1 = ( 1 2 ) κ → 2 = ( 3 6 ) − proj [ κ → 1 ] ( ( 3 6 ) ) = ( 0 0 )

    This first thing is not so bad because the zero vector is by definition orthogonal to every other vector, so we could accept this situation as yielding an orthogonal set (although it of course can’t be normalized), or we just could modify the Gram-Schmidt procedure to throw out any zero vectors. The second thing to worry about if we drop the phrase ‘linearly independent’ from the question is that the set might be infinite. Of course, any subspace of the finite-dimensional ℝ n must also be finite-dimensional so only finitely many of its members are linearly independent, but nonetheless, a “process” that examines the vectors in an infinite set one at a time would at least require some more elaboration in this question. A linearly independent subset of ℝ n is automatically finite—in fact, of size  n or less—so the ‘linearly independent’ phrase obviates these concerns.

  9. Exercise 2.18 Worked answer

    What happens if we try to apply the Gram-Schmidt process to a finite set that is not a basis?

    Back to Exercise 2.18

    Answer. If that set is not linearly independent, then we get a zero vector. Otherwise (if our set is linearly independent but does not span the space), we are doing Gram-Schmidt on a set that is a basis for a subspace and so we get an orthogonal basis for a subspace.

  10. Exercise 2.19 Worked answer

    Recommended. What happens if we apply the Gram-Schmidt process to a basis that is already orthogonal?

    Back to Exercise 2.19

    Answer. The process leaves the basis unchanged.

  11. Exercise 2.20 Worked answer

    Let ⟨ κ → 1 , … , κ → k ⟩ be a set of mutually orthogonal vectors in ℝ n .

    1. Prove that for any v → in the space, the vector v → − ( proj [ κ → 1 ] ( v → ) + ⋯ + proj [ κ → k ] ( v → ) ) is orthogonal to each of κ → 1 , …, κ → k .

    2. Illustrate the prior item in ℝ 3 by using e → 1 as κ → 1 , using e → 2 as κ → 2 , and taking v → to have components 1 , 2 , and 3 .

    3. Show that proj [ κ → 1 ] ( v → ) + ⋯ + proj [ v → k ] ( v → ) is the vector in the span of the set of κ → ’s that is closest to v → . Hint. To the illustration done for the prior part, add a vector d 1 κ → 1 + d 2 κ → 2 and apply the Pythagorean Theorem to the resulting triangle.

    Back to Exercise 2.20

    Answer.

    1. The argument is as in the i = 3 case of the proof of Theorem 2.7. The dot product

      κ → i ⋅ ( v → − proj [ κ → 1 ] ( v → ) − ⋯ − proj [ v → k ] ( v → ) )

      can be written as the sum of terms of the form − κ → i ⋅ proj [ κ → j ] ( v → ) with j ≠ i , and the term κ → i ⋅ ( v → − proj [ κ → i ] ( v → ) ) . The first kind of term equals zero because the κ → ’s are mutually orthogonal. The other term is zero because this projection is orthogonal (that is, the projection definition makes it zero: κ → i ⋅ ( v → − proj [ κ → i ] ( v → ) ) = κ → i ⋅ v → − κ → i ⋅ ( ( v → ⋅ κ → i ) / ( κ → i ⋅ κ → i ) ) ⋅ κ → i equals, after all of the cancellation is done, zero).

    2. The vector v → is in black and the vector proj [ κ → 1 ] ( v → ) + proj [ v → 2 ] ( v → ) = 1 ⋅ e → 1 + 2 ⋅ e → 2 is in gray.

      Supplied-answer-only spatial diagram: a vector projects perpendicularly to a planar vector. A dashed drop and a right-angle mark distinguish the residual from the projection; light-gray lines show the planar components.

      The vector v → − ( proj [ κ → 1 ] ( v → ) + proj [ v → 2 ] ( v → ) ) lies on the dotted line connecting the black vector to the gray one, that is, it is orthogonal to the x y -plane.

    3. We get this diagram by following the hint.

      Supplied-answer-only spatial diagram: dashed construction lines compare the original vector with its planar projection and planar components. The solid black arrow is the original vector.

      The dashed triangle has a right angle where the gray vector 1 ⋅ e → 1 + 2 ⋅ e → 2 meets the vertical dashed line v → − ( 1 ⋅ e → 1 + 2 ⋅ e → 2 ) ; this is what first item of this question proved. The Pythagorean theorem then gives that the hypotenuse—the segment from v → to any other vector—is longer than the vertical dashed line.

      More formally, writing proj [ κ → 1 ] ( v → ) + ⋯ + proj [ v → k ] ( v → ) as c 1 ⋅ κ → 1 + ⋯ + c k ⋅ κ → k , consider any other vector in the span d 1 ⋅ κ → 1 + ⋯ + d k ⋅ κ → k . Note that

      v → − ( d 1 ⋅ κ → 1 + ⋯ + d k ⋅ κ → k ) = ( v → − ( c 1 ⋅ κ → 1 + ⋯ + c k ⋅ κ → k ) ) + ( ( c 1 ⋅ κ → 1 + ⋯ + c k ⋅ κ → k ) − ( d 1 ⋅ κ → 1 + ⋯ + d k ⋅ κ → k ) )

      and that ( v → − ( c 1 ⋅ κ → 1 + ⋯ + c k ⋅ κ → k ) ) ⋅ ( ( c 1 ⋅ κ → 1 + ⋯ + c k ⋅ κ → k ) − ( d 1 ⋅ κ → 1 + ⋯ + d k ⋅ κ → k ) ) = 0 (because the first item shows the v → − ( c 1 ⋅ κ → 1 + ⋯ + c k ⋅ κ → k ) is orthogonal to each κ → and so it is orthogonal to this linear combination of the κ → ’s). Now apply the Pythagorean Theorem (i.e., the Triangle Inequality).

  12. Exercise 2.21 Worked answer

    Find a nonzero vector in ℝ 3 that is orthogonal to both of these.

    ( 1 5 − 1 ) ( 2 2 0 )

    Back to Exercise 2.21

    Answer. One way to proceed is to find a third vector so that the three together make a basis for ℝ 3 , e.g.,

    β → 3 = ( 1 0 0 )

    (the second vector is not dependent on the third because it has a nonzero second component, and the first is not dependent on the second and third because of its nonzero third component), and then apply the Gram-Schmidt process. The first element of the new basis is this.

    κ → 1 = ( 1 5 − 1 )

    And this is the second element.

    κ → 2 = ( 2 2 0 ) − proj [ κ → 1 ] ( ( 2 2 0 ) ) = ( 2 2 0 ) − ( 2 2 0 ) ⋅ ( 1 5 − 1 ) ( 1 5 − 1 ) ⋅ ( 1 5 − 1 ) ⋅ ( 1 5 − 1 ) = ( 2 2 0 ) − 12 27 ⋅ ( 1 5 − 1 ) = ( 14 / 9 − 2 / 9 4 / 9 )

    Here is the final element.

    κ → 3 = ( 1 0 0 ) − proj [ κ → 1 ] ( ( 1 0 0 ) ) − proj [ κ → 2 ] ( ( 1 0 0 ) ) = ( 1 0 0 ) − ( 1 0 0 ) ⋅ ( 1 5 − 1 ) ( 1 5 − 1 ) ⋅ ( 1 5 − 1 ) ⋅ ( 1 5 − 1 ) − ( 1 0 0 ) ⋅ ( 14 / 9 − 2 / 9 4 / 9 ) ( 14 / 9 − 2 / 9 4 / 9 ) ⋅ ( 14 / 9 − 2 / 9 4 / 9 ) ⋅ ( 14 / 9 − 2 / 9 4 / 9 ) = ( 1 0 0 ) − 1 27 ⋅ ( 1 5 − 1 ) − 7 12 ⋅ ( 14 / 9 − 2 / 9 4 / 9 ) = ( 1 / 18 − 1 / 18 − 4 / 18 )

    The result κ → 3 is orthogonal to both κ → 1 and κ → 2 . It is therefore orthogonal to every vector in the span of the set { κ → 1 , κ → 2 } , including the two vectors given in the question.

  13. Exercise 2.22 Worked answer

    Recommended. One advantage of orthogonal bases is that they simplify finding the representation of a vector with respect to that basis.

    1. For this vector and this non-orthogonal basis for ℝ 2

      v → = ( 2 3 ) B = ⟨ ( 1 1 ) , ( 1 0 ) ⟩

      first represent the vector with respect to the basis. Then project the vector into the span of each basis vector [ β → 1 ] and [ β → 2 ] .

    2. With this orthogonal basis for ℝ 2

      K = ⟨ ( 1 1 ) , ( 1 − 1 ) ⟩

      represent the same vector v → with respect to the basis. Then project the vector into the span of each basis vector. Note that the coefficients in the representation and the projection are the same.

    3. Let K = ⟨ κ → 1 , … , κ → k ⟩ be an orthogonal basis for some subspace of ℝ n . Prove that for any v → in the subspace, the i -th component of the representation Rep K ( v → ) is the scalar coefficient ( v → ⋅ κ → i ) / ( κ → i ⋅ κ → i ) from proj [ κ → i ] ( v → ) .

    4. Prove that v → = proj [ κ → 1 ] ( v → ) + ⋯ + proj [ κ → k ] ( v → ) .

    Back to Exercise 2.22

    Answer.

    1. We can do the representation by eye.

      ( 2 3 ) = 3 ⋅ ( 1 1 ) + ( − 1 ) ⋅ ( 1 0 ) Rep B ( v → ) = ( 3 − 1 ) B

      The two projections are also easy.

      proj [ β → 1 ] ( ( 2 3 ) ) = ( 2 3 ) ⋅ ( 1 1 ) ( 1 1 ) ⋅ ( 1 1 ) ⋅ ( 1 1 ) = 5 2 ⋅ ( 1 1 ) proj [ β → 2 ] ( ( 2 3 ) ) = ( 2 3 ) ⋅ ( 1 0 ) ( 1 0 ) ⋅ ( 1 0 ) ⋅ ( 1 0 ) = 2 1 ⋅ ( 1 0 )

    2. As above, we can do the representation by eye

      ( 2 3 ) = ( 5 / 2 ) ⋅ ( 1 1 ) + ( − 1 / 2 ) ⋅ ( 1 − 1 )

      and the two projections are easy.

      proj [ β → 1 ] ( ( 2 3 ) ) = ( 2 3 ) ⋅ ( 1 1 ) ( 1 1 ) ⋅ ( 1 1 ) ⋅ ( 1 1 ) = 5 2 ⋅ ( 1 1 ) proj [ β → 2 ] ( ( 2 3 ) ) = ( 2 3 ) ⋅ ( 1 − 1 ) ( 1 − 1 ) ⋅ ( 1 − 1 ) ⋅ ( 1 − 1 ) = − 1 2 ⋅ ( 1 − 1 )

      Note the recurrence of the 5 / 2 and the − 1 / 2 .

    3. Represent v → with respect to the basis

      Rep K ( v → ) = ( r 1 ⋮ r k )

      so that v → = r 1 κ → 1 + ⋯ + r k κ → k . To determine r i , take the dot product of both sides with κ → i .

      v → ⋅ κ → i = ( r 1 κ → 1 + ⋯ + r k κ → k ) ⋅ κ → i = r 1 ⋅ 0 + ⋯ + r i ⋅ ( κ → i ⋅ κ → i ) + ⋯ + r k ⋅ 0

      Solving for r i yields the desired coefficient.

    4. This is a restatement of the prior item.

  14. Exercise 2.23 Worked answer

    Bessel’s Inequality. Consider these orthonormal sets

    B 1 = { e → 1 } B 2 = { e → 1 , e → 2 } B 3 = { e → 1 , e → 2 , e → 3 } B 4 = { e → 1 , e → 2 , e → 3 , e → 4 }

    along with the vector v → ∈ ℝ 4 whose components are 4 , 3 , 2 , and 1 .

    1. Find the coefficient c 1 for the projection of v → into the span of the vector in B 1 . Check that | v → | 2 ≥ | c 1 | 2 .

    2. Find the coefficients c 1 and c 2 for the projection of v → into the spans of the two vectors in B 2 . Check that | v → | 2 ≥ | c 1 | 2 + | c 2 | 2 .

    3. Find c 1 , c 2 , and c 3 associated with the vectors in B 3 , and c 1 , c 2 , c 3 , and c 4 for the vectors in B 4 . Check that | v → | 2 ≥ | c 1 | 2 + ⋯ + | c 3 | 2 and that | v → | 2 ≥ | c 1 | 2 + ⋯ + | c 4 | 2 .

    Show that this holds in general: where { κ → 1 , … , κ → k } is an orthonormal set and c i is coefficient of the projection of a vector v → from the space then | v → | 2 ≥ | c 1 | 2 + ⋯ + | c k | 2 . Hint. One way is to look at the inequality 0 ≤ | v → − ( c 1 κ → 1 + ⋯ + c k κ → k ) | 2 and expand the c ’s.

    Back to Exercise 2.23

    Answer. First, | v → | 2 = 4 2 + 3 2 + 2 2 + 1 2 = 50 .

    1. c 1 = 4

    2. c 1 = 4 , c 2 = 3

    3. c 1 = 4 , c 2 = 3 , c 3 = 2 , c 4 = 1

    For the proof, we will do only the k = 2 case because the completely general case is messier but no more enlightening. We follow the hint (recall that for any vector w → we have | w → | 2 = w → ⋅ w → ).

    0 ≤ ( v → − ( v → ⋅ κ → 1 κ → 1 ⋅ κ → 1 ⋅ κ → 1 + v → ⋅ κ → 2 κ → 2 ⋅ κ → 2 ⋅ κ → 2 ) ) ⋅ ( v → − ( v → ⋅ κ → 1 κ → 1 ⋅ κ → 1 ⋅ κ → 1 + v → ⋅ κ → 2 κ → 2 ⋅ κ → 2 ⋅ κ → 2 ) ) = v → ⋅ v → − 2 ⋅ v → ⋅ ( v → ⋅ κ → 1 κ → 1 ⋅ κ → 1 ⋅ κ → 1 + v → ⋅ κ → 2 κ → 2 ⋅ κ → 2 ⋅ κ → 2 ) + ( v → ⋅ κ → 1 κ → 1 ⋅ κ → 1 ⋅ κ → 1 + v → ⋅ κ → 2 κ → 2 ⋅ κ → 2 ⋅ κ → 2 ) ⋅ ( v → ⋅ κ → 1 κ → 1 ⋅ κ → 1 ⋅ κ → 1 + v → ⋅ κ → 2 κ → 2 ⋅ κ → 2 ⋅ κ → 2 ) = v → ⋅ v → − 2 ⋅ ( v → ⋅ κ → 1 κ → 1 ⋅ κ → 1 ⋅ ( v → ⋅ κ → 1 ) + v → ⋅ κ → 2 κ → 2 ⋅ κ → 2 ⋅ ( v → ⋅ κ → 2 ) ) + ( ( v → ⋅ κ → 1 κ → 1 ⋅ κ → 1 ) 2 ⋅ ( κ → 1 ⋅ κ → 1 ) + ( v → ⋅ κ → 2 κ → 2 ⋅ κ → 2 ) 2 ⋅ ( κ → 2 ⋅ κ → 2 ) )

    (The two mixed terms in the third part of the third line are zero because κ → 1 and κ → 2 are orthogonal.) The result now follows on gathering like terms and on recognizing that κ → 1 ⋅ κ → 1 = 1 and κ → 2 ⋅ κ → 2 = 1 because these vectors are members of an orthonormal set.

  15. Exercise 2.24 Worked answer

    Prove or disprove: every vector in ℝ n is in some orthogonal basis.

    Back to Exercise 2.24

    Answer. It is true, except for the zero vector. Every vector in ℝ n except the zero vector is in a basis, and that basis can be orthogonalized.

  16. Exercise 2.25 Worked answer

    Show that the columns of an n × n matrix form an orthonormal set if and only if the inverse of the matrix is its transpose. Produce such a matrix.

    Back to Exercise 2.25

    Answer. The 3 × 3 case gives the idea. The set

    { ( a d g ) , ( b e h ) , ( c f i ) }

    is orthonormal if and only if these nine conditions all hold

    ( a d g ) ⋅ ( a d g ) = 1 ( a d g ) ⋅ ( b e h ) = 0 ( a d g ) ⋅ ( c f i ) = 0 ( b e h ) ⋅ ( a d g ) = 0 ( b e h ) ⋅ ( b e h ) = 1 ( b e h ) ⋅ ( c f i ) = 0 ( c f i ) ⋅ ( a d g ) = 0 ( c f i ) ⋅ ( b e h ) = 0 ( c f i ) ⋅ ( c f i ) = 1

    (the three conditions in the lower left are redundant but nonetheless correct). Those, in turn, hold if and only if

    ( a d g b e h c f i ) ( a b c d e f g h i ) = ( 1 0 0 0 1 0 0 0 1 )

    as required.

    This is an example, the inverse of this matrix is its transpose.

    ( 1 / 2 1 / 2 0 − 1 / 2 1 / 2 0 0 0 1 )

  17. Exercise 2.26 Worked answer

    Does the proof of Theorem 2.2 fail to consider the possibility that the set of vectors is empty (i.e., that k = 0 )?

    Back to Exercise 2.26

    Answer. If the set is empty then the summation on the left side is the linear combination of the empty set of vectors, which by definition adds to the zero vector. In the second sentence, there is not such i , so the ‘if …then …’ implication is vacuously true.

  18. Exercise 2.27 Worked answer

    Theorem 2.7 describes a change of basis from any basis B = ⟨ β → 1 , … , β → k ⟩ to one that is orthogonal K = ⟨ κ → 1 , … , κ → k ⟩ . Consider the change of basis matrix Rep B , K ( id ) .

    1. Prove that the matrix Rep K , B ( id ) changing bases in the direction opposite to that of the theorem has an upper triangular shape—all of its entries below the main diagonal are zeros.

    2. Prove that the inverse of an upper triangular matrix is also upper triangular (if the matrix is invertible, that is). This shows that the matrix Rep B , K ( id ) changing bases in the direction described in the theorem is upper triangular.

    Back to Exercise 2.27

    Answer.

    1. Part of the induction argument proving Theorem 2.7 checks that κ → i is in the span of ⟨ β → 1 , … , β → i ⟩ . (The i = 3 case in the proof illustrates.) Thus, in the change of basis matrix Rep K , B ( id ) , the i -th column Rep B ( κ → i ) has components i + 1 through k that are zero.

    2. One way to see this is to recall the computational procedure that we use to find the inverse. We write the matrix, write the identity matrix next to it, and then we do Gauss-Jordan reduction. If the matrix starts out upper triangular then the Gauss-Jordan reduction involves only the Jordan half and these steps, when performed on the identity, will result in an upper triangular inverse matrix.

  19. Exercise 2.28 Worked answer

    Complete the induction argument in the proof of Theorem 2.7.

    Back to Exercise 2.28

    Answer. For the inductive step, we assume that for all j in  [ 1. . i ] , these three conditions are true of each κ → j : (i) each κ → j is nonzero, (ii) each κ → j is a linear combination of the vectors β → 1 , … , β → j , and (iii) each κ → j is orthogonal to all of the κ → m ’s prior to it (that is, with m < j ). With those inductive hypotheses, consider κ → i + 1 .

    κ → i + 1 = β → i + 1 − proj [ κ → 1 ] ( β i + 1 ) − proj [ κ → 2 ] ( β i + 1 ) − ⋯ − proj [ κ → i ] ( β i + 1 ) = β → i + 1 − β i + 1 ⋅ κ → 1 κ → 1 ⋅ κ → 1 ⋅ κ → 1 − β i + 1 ⋅ κ → 2 κ → 2 ⋅ κ → 2 ⋅ κ → 2 − ⋯ − β i + 1 ⋅ κ → i κ → i ⋅ κ → i ⋅ κ → i

    By the inductive assumption (ii) we can expand each κ → j into a linear combination of β → 1 , … , β → j

    = β → i + 1 − β → i + 1 ⋅ κ → 1 κ → 1 ⋅ κ → 1 ⋅ β → 1 − β → i + 1 ⋅ κ → 2 κ → 2 ⋅ κ → 2 ⋅ (  linear combination of  β → 1 , β → 2 ) − ⋯ − β → i + 1 ⋅ κ → i κ → i ⋅ κ → i ⋅ (  linear combination of  β → 1 , … , β → i )

    The fractions are scalars so this is a linear combination of linear combinations of β → 1 , … , β → i + 1 . It is therefore just a linear combination of β → 1 , … , β → i + 1 . Now, (i) it cannot sum to the zero vector because the equation would then describe a nontrivial linear relationship among the β → ’s that are given as members of a basis (the relationship is nontrivial because the coefficient of β → i + 1 is 1 ). Also, (ii) the equation gives κ → i + 1 as a combination of β → 1 , … , β → i + 1 . Finally, for (iii), consider κ → j ⋅ κ → i + 1 ; as in the i = 3 case, the dot product of κ → j with κ → i + 1 = β → i + 1 − proj [ κ → 1 ] ( β → i + 1 ) − ⋯ − proj [ κ → i ] ( β → i + 1 ) can be rewritten to give two kinds of terms, κ → j ⋅ ( β → i + 1 − proj [ κ → j ] ( β → i + 1 ) ) (which is zero because the projection is orthogonal) and κ → j ⋅ proj [ κ → m ] ( β → i + 1 ) with m ≠ j and m < i + 1 (which is zero because by the hypothesis (iii) the vectors κ → j and κ → m are orthogonal).

Projection Into a Subspace

This subsection uses material from the optional earlier subsection on Combining Subspaces.

The prior subsections project a vector into a line by decomposing it into two parts: the part in the line proj [ s → ] ( v → ) and the rest v → − proj [ s → ] ( v → ) . To generalize projection to arbitrary subspaces we will follow this decomposition idea.

Definition 3.1 Let a vector space be a direct sum V = M ⊕ N . Then for any v → ∈ V with v → = m → + n → where m → ∈ M , n → ∈ N , the projection of v → into M along N is proj M , N ( v → ) = m → .

This definition applies in spaces where we don’t have a ready definition of orthogonal. (Definitions of orthogonality for spaces other than the ℝ n are perfectly possible but we haven’t seen any in this book.)

Example 3.2 The space ℳ 2 × 2 of 2 × 2 matrices is the direct sum of these two.

M = { ( a b 0 0 ) ∣ a , b ∈ ℝ } N = { ( 0 0 c d ) ∣ c , d ∈ ℝ }

To project

A = ( 3 1 0 4 )

into M along N , we first fix bases for the two subspaces.

B M = ⟨ ( 1 0 0 0 ) , ( 0 1 0 0 ) ⟩ B N = ⟨ ( 0 0 1 0 ) , ( 0 0 0 1 ) ⟩

Their concatenation

B = B M ⌢ B N = ⟨ ( 1 0 0 0 ) , ( 0 1 0 0 ) , ( 0 0 1 0 ) , ( 0 0 0 1 ) ⟩

is a basis for the entire space because ℳ 2 × 2 is the direct sum. So we can use it to represent A .

( 3 1 0 4 ) = 3 ⋅ ( 1 0 0 0 ) + 1 ⋅ ( 0 1 0 0 ) + 0 ⋅ ( 0 0 1 0 ) + 4 ⋅ ( 0 0 0 1 )

The projection of A into M along N keeps the M part and drops the N part.

proj M , N ( ( 3 1 0 4 ) ) = 3 ⋅ ( 1 0 0 0 ) + 1 ⋅ ( 0 1 0 0 ) = ( 3 1 0 0 )

Example 3.3 Both subscripts on proj M , N ( v → ) are significant. The first subscript M matters because the result of the projection is a member of M . For an example showing that the second one matters, fix this plane subspace of ℝ 3 and its basis.

M = { ( x y z ) ∣ y − 2 z = 0 } B M = ⟨ ( 1 0 0 ) , ( 0 2 1 ) ⟩

We will compare the projections of this element of  ℝ 3

v → = ( 2 2 5 )

into M along these two subspaces (verification that ℝ 3 = M ⊕ N and ℝ 3 = M ⊕ N ^ is routine).

N = { k ( 0 0 1 ) ∣ k ∈ ℝ } N ^ = { k ( 0 1 − 2 ) ∣ k ∈ ℝ }

Here are natural bases for N and N ^ .

B N = ⟨ ( 0 0 1 ) ⟩ B N ^ = ⟨ ( 0 1 − 2 ) ⟩

To project into M along  N , represent v → with respect to the concatenation B M ⌢ B N

( 2 2 5 ) = 2 ⋅ ( 1 0 0 ) + 1 ⋅ ( 0 2 1 ) + 4 ⋅ ( 0 0 1 )

and drop the N term.

proj M , N ( v → ) = 2 ⋅ ( 1 0 0 ) + 1 ⋅ ( 0 2 1 ) = ( 2 2 1 )

To project into  M along  N ^ represent v → with respect to B M ⌢ B N ^

( 2 2 5 ) = 2 ⋅ ( 1 0 0 ) + ( 9 / 5 ) ⋅ ( 0 2 1 ) − ( 8 / 5 ) ⋅ ( 0 1 − 2 )

and omit the N ^ part.

proj M , N ^ ( v → ) = 2 ⋅ ( 1 0 0 ) + ( 9 / 5 ) ⋅ ( 0 2 1 ) = ( 2 18 / 5 9 / 5 )

So projecting along different subspaces can give different results.

These pictures compare the two maps. Both show that the projection is indeed ‘into’ the plane and ‘along’ the line.

Projection into a shaded plane M along a line N. A dotted segment parallel to N joins the original vector tip to its projection in M; this projection is not orthogonal.
Projection into the same shaded plane M along the different line N hat. The dotted segment is parallel to N hat and perpendicular to M, so this projection is orthogonal.

Notice that the projection along N is not orthogonal since there are members of the plane M that are not orthogonal to the dotted line. But the projection along N ^ is orthogonal.

We have seen two projection operations, orthogonal projection into a line as well as this subsections’s projection into an  M and along an  N , and we naturally ask whether they are related. The right-hand picture above suggests the answer—orthogonal projection into a line is a special case of this subsection’s projection; it is projection along a subspace perpendicular to the line.

Two perpendicular lines M and N through the origin illustrate orthogonal projection into M along its perpendicular complement N. A dotted segment joins the original vector tip to its image on M.

Definition 3.4 The orthogonal complement of a subspace M of ℝ n is

M ⊥ = { v → ∈ ℝ n ∣ v →  is perpendicular to all vectors in  M }

(read “ M perp”). The orthogonal projection proj M ( v → ) of a vector is its projection into M along M ⊥ .

Example 3.5 In ℝ 3 , to find the orthogonal complement of the plane

P = { ( x y z ) ∣ 3 x + 2 y − z = 0 }

we start with a basis for P .

B = ⟨ ( 1 0 3 ) , ( 0 1 2 ) ⟩

Any v → perpendicular to every vector in B is perpendicular to every vector in the span of B (the proof of this is Exercise 3.22). Therefore, the subspace P ⊥ consists of the vectors that satisfy these two conditions.

( 1 0 3 ) ⋅ ( v 1 v 2 v 3 ) = 0 ( 0 1 2 ) ⋅ ( v 1 v 2 v 3 ) = 0

Those conditions give a linear system.

P ⊥ = { ( v 1 v 2 v 3 ) ∣ ( 1 0 3 0 1 2 ) ( v 1 v 2 v 3 ) = ( 0 0 ) }

We are thus left with finding the null space of the map represented by the matrix, that is, with calculating the solution set of the homogeneous linear system.

v 1 + 3 v 3 = 0 v 2 + 2 v 3 = 0 ⟹ P ⊥ = { k ( − 3 − 2 1 ) ∣ k ∈ ℝ }

Example 3.6 Where M is the x y -plane subspace of ℝ 3 , what is M ⊥ ? A common first reaction is that M ⊥ is the y z -plane but that’s not right because some vectors from the y z -plane are not perpendicular to every vector in the x y -plane.

( 1 1 0 ) ⊥̸ ( 0 3 2 )

The xy-plane vector (1,1,0) and the yz-plane vector (0,3,2) are drawn with dashed coordinate constructions. They are not perpendicular; the adjacent formula gives an angle of approximately 0.94 radians.

θ = arccos ⁡ ( 1 ⋅ 0 + 1 ⋅ 3 + 0 ⋅ 2 2 ⋅ 13 ) ≈ 0.94  rad

Instead M ⊥ is the z -axis, since proceeding as in the prior example and taking the natural basis for the x y -plane gives this.

M ⊥ = { ( x y z ) ∣ ( 1 0 0 0 1 0 ) ( x y z ) = ( 0 0 ) } = { ( x y z ) ∣ x = 0  and  y = 0 }

Lemma 3.7 If M is a subspace of ℝ n then its orthogonal complement M ⊥ is also a subspace. The space is the direct sum of the two ℝ n = M ⊕ M ⊥ . For any v → ∈ ℝ n the vector v → − proj M ( v → ) is perpendicular to every vector in M .

Proof First, the orthogonal complement M ⊥ is a subspace of ℝ n because it is a null space, namely the null space of the orthogonal projection map.

To show that the space  ℝ n is the direct sum of the two, start with any basis B M = ⟨ μ → 1 , … , μ → k ⟩ for M . Expand it to a basis for the entire space and then apply the Gram-Schmidt process to get an orthogonal basis K = ⟨ κ → 1 , … , κ → n ⟩ for ℝ n . This K is the concatenation of two bases: ⟨ κ → 1 , … , κ → k ⟩ with the same number of members,  k , as B M , and D = ⟨ κ → k + 1 , … , κ → n ⟩ . The first is a basis for M so if we show that the second is a basis for M ⊥ then we will have that the entire space is the direct sum.

Exercise 2.22 from the prior subsection proves this about any orthogonal basis: each vector v → in the space is the sum of its orthogonal projections into the lines spanned by the basis vectors.

v → = proj [ κ → 1 ] ( v → ) + ⋯ + proj [ κ → n ] ( v → ) ( ∗ )

To check this, represent the vector as v → = r 1 κ → 1 + ⋯ + r n κ → n , apply κ → i to both sides v → ⋅ κ → i = ( r 1 κ → 1 + ⋯ + r n κ → n ) ⋅ κ → i = r 1 ⋅ 0 + ⋯ + r i ⋅ ( κ → i ⋅ κ → i ) + ⋯ + r n ⋅ 0 , and solve to get r i = ( v → ⋅ κ → i ) / ( κ → i ⋅ κ → i ) , as desired.

Any member of the span of D is orthogonal to any vector in M so the span of D is a subset of M ⊥ . To show that D is a basis for M ⊥ we need only show the other containment, that any w → ∈ M ⊥ is an element of the span of  D . The prior paragraph works for this. Any w → ∈ M ⊥ gives this on projections into basis vectors from M : proj [ κ → 1 ] ( w → ) = 0 → , … , proj [ κ → k ] ( w → ) = 0 → . Therefore equation ( ∗ ) gives that w → is a linear combination of κ → k + 1 , … , κ → n . Thus D is a basis for M ⊥ and ℝ n is the direct sum of the two.

The final sentence of the lemma is proved in much the same way. Write v → = proj [ κ → 1 ] ( v → ) + ⋯ + proj [ κ → n ] ( v → ) . Then proj M ( v → ) keeps only the M part and drops the M ⊥ part: proj M ( v → ) = proj [ κ → 1 ] ( v → ) + ⋯ + proj [ κ → k ] ( v → ) . Therefore v → − proj M ( v → ) consists of a linear combination of elements of M ⊥ and so is perpendicular to every vector in M .

QED

Given a subspace, we could compute the orthogonal projection into that subspace by following the steps of that proof: finding a basis, expanding it to a basis for the entire space, applying Gram-Schmidt to get an orthogonal basis, and projecting into each linear subspace. However we will instead use a convenient formula.

Theorem 3.8 Let M be a subspace of ℝ n with basis ⟨ β → 1 , … , β → k ⟩ and let A be the matrix whose columns are the β → ’s. Then for any v → ∈ ℝ n the orthogonal projection is proj M ( v → ) = c 1 β → 1 + ⋯ + c k β → k , where the coefficients c i are the entries of the vector ( A 𝖳 A ) − 1 A 𝖳 ⋅ v → . That is, proj M ( v → ) = A ( A 𝖳 A ) − 1 A 𝖳 ⋅ v → .

Proof The vector proj M ( v → ) is a member of M and so is a linear combination of basis vectors c 1 ⋅ β → 1 + ⋯ + c k ⋅ β → k . Since A ’s columns are the β → ’s, there is a c → ∈ ℝ k such that proj M ( v → ) = A c → . To find c → note that the vector v → − proj M ( v → ) is perpendicular to each member of the basis so

0 → = A 𝖳 ( v → − A c → ) = A 𝖳 v → − A 𝖳 A c →

and solving gives this (showing that A 𝖳 A is invertible is an exercise).

c → = ( A 𝖳 A ) − 1 A 𝖳 ⋅ v →

Therefore proj M ( v → ) = A ⋅ c → = A ( A 𝖳 A ) − 1 A 𝖳 ⋅ v → , as required.

QED

Example 3.9 To orthogonally project this vector into this subspace

v → = ( 1 − 1 1 ) P = { ( x y z ) ∣ x + z = 0 }

first make a matrix whose columns are a basis for the subspace

A = ( 0 1 1 0 0 − 1 )

and then compute.

A ( A 𝖳 A ) − 1 A 𝖳 = ( 0 1 1 0 0 − 1 ) ( 1 0 0 1 / 2 ) ( 0 1 0 1 0 − 1 ) = ( 1 / 2 0 − 1 / 2 0 1 0 − 1 / 2 0 1 / 2 )

With the matrix, calculating the orthogonal projection of any vector into P is easy.

proj P ( v → ) = ( 1 / 2 0 − 1 / 2 0 1 0 − 1 / 2 0 1 / 2 ) ( 1 − 1 1 ) = ( 0 − 1 0 )

Note, as a check, that this result is indeed in P .

Exercises

  1. Exercise 3.10 Worked answer

    Recommended. Project the vectors into M along N .

    1. ( 3 − 2 ) , M = { ( x y ) ∣ x + y = 0 } , N = { ( x y ) ∣ − x − 2 y = 0 }

    2. ( 1 2 ) , M = { ( x y ) ∣ x − y = 0 } , N = { ( x y ) ∣ 2 x + y = 0 }

    3. ( 3 0 1 ) , M = { ( x y z ) ∣ x + y = 0 } , N = { c ⋅ ( 1 0 1 ) ∣ c ∈ ℝ }

    Back to Exercise 3.10

    Answer.

    1. When bases for the subspaces

      B M = ⟨ ( 1 − 1 ) ⟩ B N = ⟨ ( 2 − 1 ) ⟩

      are concatenated

      B = B M ⌢ B N = ⟨ ( 1 − 1 ) , ( 2 − 1 ) ⟩

      and the given vector is represented

      ( 3 − 2 ) = 1 ⋅ ( 1 − 1 ) + 1 ⋅ ( 2 − 1 )

      then the answer comes from retaining the M  part and dropping the N  part.

      proj M , N ( ( 3 − 2 ) ) = ( 1 − 1 )

    2. When the bases

      B M = ⟨ ( 1 1 ) ⟩ B N ⟨ ( 1 − 2 ) ⟩

      are concatenated, and the vector is represented,

      ( 1 2 ) = ( 4 / 3 ) ⋅ ( 1 1 ) − ( 1 / 3 ) ⋅ ( 1 − 2 )

      then retaining only the M  part gives this answer.

      proj M , N ( ( 1 2 ) ) = ( 4 / 3 4 / 3 )

    3. With these bases

      B M = ⟨ ( 1 − 1 0 ) , ( 0 0 1 ) ⟩ B N = ⟨ ( 1 0 1 ) ⟩

      the representation with respect to the concatenation is this.

      ( 3 0 1 ) = 0 ⋅ ( 1 − 1 0 ) − 2 ⋅ ( 0 0 1 ) + 3 ⋅ ( 1 0 1 )

      and so the projection is this.

      proj M , N ( ( 3 0 1 ) ) = ( 0 0 − 2 )

  2. Exercise 3.11 Worked answer

    Recommended. Find M ⊥ .

    1. M = { ( x y ) ∣ x + y = 0 }

    2. M = { ( x y ) ∣ − 2 x + 3 y = 0 }

    3. M = { ( x y ) ∣ x − y = 0 }

    4. M = { 0 → }

    5. M = { ( x y ) ∣ x = 0 }

    6. M = { ( x y z ) ∣ − x + 3 y + z = 0 }

    7. M = { ( x y z ) ∣ x = 0  and  y + z = 0 }

    Back to Exercise 3.11

    Answer. As in Example 3.5, we can simplify the calculation by just finding the space of vectors perpendicular to all the the vectors in M ’s basis.

    1. Parametrizing to get

      M = { c ⋅ ( − 1 1 ) ∣ c ∈ ℝ }

      gives that

      M ⊥ { ( u v ) ∣ 0 = ( u v ) ⋅ ( − 1 1 ) } = { ( u v ) ∣ 0 = − u + v }

      Parametrizing the one-equation linear system gives this description.

      M ⊥ = { k ⋅ ( 1 1 ) ∣ k ∈ ℝ }

    2. As in the answer to the prior part, we can describe M as a span

      M = { c ⋅ ( 3 / 2 1 ) ∣ c ∈ ℝ } B M = ⟨ ( 3 / 2 1 ) ⟩

      and then M ⊥ is the set of vectors perpendicular to the one vector in this basis.

      M ⊥ = { ( u v ) ∣ ( 3 / 2 ) ⋅ u + 1 ⋅ v = 0 } = { k ⋅ ( − 2 / 3 1 ) ∣ k ∈ ℝ }

    3. Parametrizing the linear requirement in the description of M gives this basis.

      M = { c ⋅ ( 1 1 ) ∣ c ∈ ℝ } B M = ⟨ ( 1 1 ) ⟩

      Now, M ⊥ is the set of vectors perpendicular to (the one vector in) B M .

      M ⊥ = { ( u v ) ∣ u + v = 0 } = { k ⋅ ( − 1 1 ) ∣ k ∈ ℝ }

      (By the way, this answer checks with the first item in this question.)

    4. Every vector in the space is perpendicular to the zero vector so M ⊥ = ℝ n .

    5. The appropriate description and basis for M are routine.

      M = { y ⋅ ( 0 1 ) ∣ y ∈ ℝ } B M = ⟨ ( 0 1 ) ⟩

      Then

      M ⊥ = { ( u v ) ∣ 0 ⋅ u + 1 ⋅ v = 0 } = { k ⋅ ( 1 0 ) ∣ k ∈ ℝ }

      and so ( y -axis ) ⊥ = x -axis .

    6. The description of M is easy to find by parametrizing.

      M = { c ⋅ ( 3 1 0 ) + d ⋅ ( 1 0 1 ) ∣ c , d ∈ ℝ } B M = ⟨ ( 3 1 0 ) , ( 1 0 1 ) ⟩

      Finding M ⊥ here just requires solving a linear system with two equations

      3 u + v = 0 u + w = 0 ⟶ − ( 1 / 3 ) ρ 1 + ρ 2 ( 3 u + v = 0 − ( 1 / 3 ) v + w = 0

      and parametrizing.

      M ⊥ = { k ⋅ ( − 1 3 1 ) ∣ k ∈ ℝ }

    7. Here, M is one-dimensional

      M = { c ⋅ ( 0 − 1 1 ) ∣ c ∈ ℝ } B M = ⟨ ( 0 − 1 1 ) ⟩

      and as a result, M ⊥ is two-dimensional.

      M ⊥ = { ( u v w ) ∣ 0 ⋅ u − 1 ⋅ v + 1 ⋅ w = 0 } = { j ⋅ ( 1 0 0 ) + k ⋅ ( 0 1 1 ) ∣ j , k ∈ ℝ }

  3. Exercise 3.12 Worked answer

    Recommended. Find the orthogonal projection of the vector into the subspace.

    ( 1 2 0 ) S = [ { ( 0 2 0 ) , ( 1 − 1 1 ) } ]

    Back to Exercise 3.12

    Answer. Where

    A = ( 0 1 2 − 1 0 1 )

    straightforward, although tedious, calculation gives this.

    A ( A 𝖳 A ) − 1 A 𝖳 = ( 1 / 2 0 1 / 2 0 1 0 1 / 2 0 1 / 2 )

    When applied to the vector we get the projection.

    ( 1 / 2 0 1 / 2 0 1 0 1 / 2 0 1 / 2 ) ( 1 2 0 ) = ( 1 / 2 2 1 / 2 )

  4. Exercise 3.13 Worked answer

    Recommended. With the same subspace as in the prior problem, find the orthogonal projection of this vector.

    ( 1 2 − 1 )

    Back to Exercise 3.13

    Answer. Using the matrix calculation from the prior answer, this

    ( 1 / 2 0 − 1 / 2 0 1 0 − 1 / 2 0 1 / 2 ) ( 1 2 − 1 ) = ( 1 2 − 1 )

    shows the vector is in the subspace  S .

  5. Exercise 3.14 Worked answer

    Recommended. Let p → be the orthogonal projection of v → ∈ ℝ n onto a subspace  S . Show that p → is the point in the subspace closest to  v → .

    Back to Exercise 3.14

    Answer. Suppose w → ∈ S . Because this is orthogonal projection, the two vectors v → − p → and p → − w → are at a right angle. The Triangle Inequality applies and the hypotenuse v → − w → is therefore at least as long as v → − p → .

  6. Exercise 3.15 Worked answer

    This subsection shows how to project orthogonally in two ways, the method of Example 3.2 and 3.3, and the method of Theorem 3.8. To compare them, consider the plane P specified by 3 x + 2 y − z = 0 in ℝ 3 .

    1. Find a basis for P .

    2. Find P ⊥ and a basis for P ⊥ .

    3. Represent this vector with respect to the concatenation of the two bases from the prior item.

      v → = ( 1 1 2 )

    4. Find the orthogonal projection of v → into P by keeping only the P part from the prior item.

    5. Check that against the result from applying Theorem 3.8.

    Back to Exercise 3.15

    Answer.

    1. Parametrizing the equation leads to this basis for P .

      B P = ⟨ ( 1 0 3 ) , ( 0 1 2 ) ⟩

    2. Because ℝ 3 is three-dimensional and P is two-dimensional, the complement P ⊥ must be a line. Anyway, the calculation as in Example 3.5

      P ⊥ = { ( x y z ) ∣ ( 1 0 3 0 1 2 ) ( x y z ) = ( 0 0 ) }

      gives this basis for P ⊥ .

      B P ⊥ = ⟨ ( 3 2 − 1 ) ⟩

    3. ( 1 1 2 ) = ( 5 / 14 ) ⋅ ( 1 0 3 ) + ( 8 / 14 ) ⋅ ( 0 1 2 ) + ( 3 / 14 ) ⋅ ( 3 2 − 1 )

    4. proj P ( ( 1 1 2 ) ) = ( 5 / 14 8 / 14 31 / 14 )

    5. The matrix of the projection

      ( 1 0 0 1 3 2 ) ( ( 1 0 3 0 1 2 ) ( 1 0 0 1 3 2 ) ) − 1 ( 1 0 3 0 1 2 ) = ( 1 0 0 1 3 2 ) ( 10 6 6 5 ) − 1 ( 1 0 3 0 1 2 ) = 1 14 ( 5 − 6 3 − 6 10 2 3 2 13 )

      when applied to the vector, yields the expected result.

      1 14 ( 5 − 6 3 − 6 10 2 3 2 13 ) ( 1 1 2 ) = ( 5 / 14 8 / 14 31 / 14 )

  7. Exercise 3.16 Worked answer

    We have three ways to find the orthogonal projection of a vector into a line, the Definition 1.1 way from the first subsection of this section, the Example 3.2 and 3.3 way of representing the vector with respect to a basis for the space and then keeping the M  part, and the way of Theorem 3.8. For these cases, do all three ways.

    1. v → = ( 1 − 3 ) , M = { ( x y ) ∣ x + y = 0 }

    2. v → = ( 0 1 2 ) , M = { ( x y z ) ∣ x + z = 0  and  y = 0 }

    Back to Exercise 3.16

    Answer.

    1. Parametrizing gives this.

      M = { c ⋅ ( − 1 1 ) ∣ c ∈ ℝ }

      For the first way, we take the vector spanning the line M to be

      s → = ( − 1 1 )

      and the Definition 1.1 formula gives this.

      proj [ s → ] ( ( 1 − 3 ) ) = ( 1 − 3 ) ⋅ ( − 1 1 ) ( − 1 1 ) ⋅ ( − 1 1 ) ⋅ ( − 1 1 ) = − 4 2 ⋅ ( − 1 1 ) = ( 2 − 2 )

      For the second way, we fix

      B M = ⟨ ( − 1 1 ) ⟩

      and so (as in Example 3.5 and 3.6, we can just find the vectors perpendicular to all of the members of the basis)

      M ⊥ = { ( u v ) ∣ − 1 ⋅ u + 1 ⋅ v = 0 } = { k ⋅ ( 1 1 ) ∣ k ∈ ℝ } B M ⊥ = ⟨ ( 1 1 ) ⟩

      and representing the vector with respect to the concatenation gives this.

      ( 1 − 3 ) = − 2 ⋅ ( − 1 1 ) − 1 ⋅ ( 1 1 )

      Keeping the M  part yields the answer.

      proj M , M ⊥ ( ( 1 − 3 ) ) = ( 2 − 2 )

      The third part is also a simple calculation (there is a 1 × 1 matrix in the middle, and the inverse of it is also 1 × 1 )

      A ( A 𝖳 A ) − 1 A 𝖳 = ( − 1 1 ) ( ( − 1 1 ) ( − 1 1 ) ) − 1 ( − 1 1 ) = ( − 1 1 ) ( 2 ) − 1 ( − 1 1 ) = ( − 1 1 ) ( 1 / 2 ) ( − 1 1 ) = ( − 1 1 ) ( − 1 / 2 1 / 2 ) = ( 1 / 2 − 1 / 2 − 1 / 2 1 / 2 )

      which of course gives the same answer.

      proj M ( ( 1 − 3 ) ) = ( 1 / 2 − 1 / 2 − 1 / 2 1 / 2 ) ( 1 − 3 ) = ( 2 − 2 )

    2. Parametrization gives this.

      M = { c ⋅ ( − 1 0 1 ) ∣ c ∈ ℝ }

      With that, the formula for the first way gives this.

      ( 0 1 2 ) ⋅ ( − 1 0 1 ) ( − 1 0 1 ) ⋅ ( − 1 0 1 ) ⋅ ( − 1 0 1 ) = 2 2 ⋅ ( − 1 0 1 ) = ( − 1 0 1 )

      To proceed by the second method we find M ⊥ ,

      M ⊥ = { ( u v w ) ∣ − u + w = 0 } = { j ⋅ ( 1 0 1 ) + k ⋅ ( 0 1 0 ) ∣ j , k ∈ ℝ }

      find the representation of the given vector with respect to the concatenation of the bases B M and B M ⊥

      ( 0 1 2 ) = 1 ⋅ ( − 1 0 1 ) + 1 ⋅ ( 1 0 1 ) + 1 ⋅ ( 0 1 0 )

      and retain only the M  part.

      proj M ( ( 0 1 2 ) ) = 1 ⋅ ( − 1 0 1 ) = ( − 1 0 1 )

      Finally, for the third method, the matrix calculation

      A ( A 𝖳 A ) − 1 A 𝖳 = ( − 1 0 1 ) ( ( − 1 0 1 ) ( − 1 0 1 ) ) − 1 ( − 1 0 1 ) = ( − 1 0 1 ) ( 2 ) − 1 ( − 1 0 1 ) = ( − 1 0 1 ) ( 1 / 2 ) ( − 1 0 1 ) = ( − 1 0 1 ) ( − 1 / 2 0 1 / 2 ) = ( 1 / 2 0 − 1 / 2 0 0 0 − 1 / 2 0 1 / 2 )

      followed by matrix-vector multiplication

      proj M ( ( 0 1 2 ) ) ( 1 / 2 0 − 1 / 2 0 0 0 − 1 / 2 0 1 / 2 ) ( 0 1 2 ) = ( − 1 0 1 )

      gives the answer.

  8. Exercise 3.17 Worked answer

    Check that the operation of Definition 3.1 is well-defined. That is, in Example 3.2 and 3.3, doesn’t the answer depend on the choice of bases?

    Back to Exercise 3.17

    Answer. No, a decomposition of vectors v → = m → + n → into m → ∈ M and n → ∈ N does not depend on the bases chosen for the subspaces, as we showed in the Direct Sum subsection.

  9. Exercise 3.18 Worked answer

    What is the orthogonal projection into the trivial subspace?

    Back to Exercise 3.18

    Answer. The orthogonal projection of a vector into a subspace is a member of that subspace. Since a trivial subspace has only one member, 0 → , the projection of any vector must equal 0 → .

  10. Exercise 3.19 Worked answer

    What is the projection of v → into M along N if v → ∈ M ?

    Back to Exercise 3.19

    Answer. The projection into M along N of a v → ∈ M is v → . Decomposing v → = m → + n → gives m → = v → and n → = 0 → , and dropping the N  part but retaining the M  part results in a projection of m → = v → .

  11. Exercise 3.20 Worked answer

    Show that if M ⊆ ℝ n is a subspace with orthonormal basis ⟨ κ → 1 , … , κ → n ⟩ then the orthogonal projection of v → into M is this.

    ( v → ⋅ κ → 1 ) ⋅ κ → 1 + ⋯ + ( v → ⋅ κ → n ) ⋅ κ → n

    Back to Exercise 3.20

    Answer. The proof of Lemma 3.7 shows that each vector v → ∈ ℝ n is the sum of its orthogonal projections into the lines spanned by the basis vectors.

    v → = proj [ κ → 1 ] ( v → ) + ⋯ + proj [ κ → n ] ( v → ) = v → ⋅ κ → 1 κ → 1 ⋅ κ → 1 ⋅ κ → 1 + ⋯ + v → ⋅ κ → n κ → n ⋅ κ → n ⋅ κ → n

    Since the basis is orthonormal, the bottom of each fraction has κ → i ⋅ κ → i = 1 .

  12. Exercise 3.21 Worked answer

    Recommended. Prove that the map p : V → V is the projection into M along N if and only if the map id − p is the projection into N along M . (Recall the definition of the difference of two maps:  ( id − p ) ( v → ) = id ( v → ) − p ( v → ) = v → − p ( v → ) .)

    Back to Exercise 3.21

    Answer. If V = M ⊕ N then every vector decomposes uniquely as v → = m → + n → . For all v → the map p gives p ( v → ) = m → if and only if v → − p ( v → ) = n → , as required.

  13. Exercise 3.22 Worked answer

    Show that if a vector is perpendicular to every vector in a set then it is perpendicular to every vector in the span of that set.

    Back to Exercise 3.22

    Answer. Let v → be perpendicular to every w → ∈ S . Then v → ⋅ ( c 1 w → 1 + ⋯ + c n w → n ) = v → ⋅ ( c 1 w → 1 ) + ⋯ + v → ⋅ ( c n ⋅ w → n ) = c 1 ( v → ⋅ w → 1 ) + ⋯ + c n ( v → ⋅ w → n ) = c 1 ⋅ 0 + ⋯ + c n ⋅ 0 = 0 .

  14. Exercise 3.23 Worked answer

    True or false: the intersection of a subspace and its orthogonal complement is trivial.

    Back to Exercise 3.23

    Answer. True; the only vector orthogonal to itself is the zero vector.

  15. Exercise 3.24 Worked answer

    Show that the dimensions of orthogonal complements add to the dimension of the entire space.

    Back to Exercise 3.24

    Answer. This is immediate from the statement in Lemma 3.7 that the space is the direct sum of the two.

  16. Exercise 3.25 Worked answer

    Suppose that v → 1 , v → 2 ∈ ℝ n are such that for all complements M , N ⊆ ℝ n , the projections of v → 1 and v → 2 into M along N are equal. Must v → 1 equal v → 2 ? (If so, what if we relax the condition to: all orthogonal projections of the two are equal?)

    Back to Exercise 3.25

    Answer. The two must be equal, even only under the seemingly weaker condition that they yield the same result on all orthogonal projections. Consider the subspace M spanned by the set { v → 1 , v → 2 } . Since each is in M , the orthogonal projection of v → 1 into M is v → 1 and the orthogonal projection of v → 2 into M is v → 2 . For their projections into M to be equal, they must be equal.

  17. Exercise 3.26 Worked answer

    Recommended. Let M , N be subspaces of ℝ n . The perp operator acts on subspaces; we can ask how it interacts with other such operations.

    1. Show that two perps cancel: ( M ⊥ ) ⊥ = M .

    2. Prove that M ⊆ N implies that N ⊥ ⊆ M ⊥ .

    3. Show that ( M + N ) ⊥ = M ⊥ ∩ N ⊥ .

    Back to Exercise 3.26

    Answer.

    1. We will show that the sets are mutually inclusive, M ⊆ ( M ⊥ ) ⊥ and ( M ⊥ ) ⊥ ⊆ M . For the first, if m → ∈ M then by the definition of the perp operation, m → is perpendicular to every v → ∈ M ⊥ , and therefore (again by the definition of the perp operation) m → ∈ ( M ⊥ ) ⊥ . For the other direction, consider v → ∈ ( M ⊥ ) ⊥ . Lemma 3.7’s proof shows that ℝ n = M ⊕ M ⊥ and that we can give an orthogonal basis for the space ⟨ κ → 1 , … , κ → k , κ → k + 1 , … , κ → n ⟩ such that the first half ⟨ κ → 1 , … , κ → k ⟩ is a basis for M and the second half is a basis for M ⊥ . The proof also checks that each vector in the space is the sum of its orthogonal projections into the lines spanned by these basis vectors.

      v → = proj [ κ → 1 ] ( v → ) + ⋯ + proj [ κ → n ] ( v → )

      Because v → ∈ ( M ⊥ ) ⊥ , it is perpendicular to every vector in M ⊥ , and so the projections in the second half are all zero. Thus v → = proj [ κ → 1 ] ( v → ) + ⋯ + proj [ κ → k ] ( v → ) , which is a linear combination of vectors from M , and so v → ∈ M . (Remark. Here is a slicker way to do the second half: write the space both as M ⊕ M ⊥ and as M ⊥ ⊕ ( M ⊥ ) ⊥ . Because the first half showed that M ⊆ ( M ⊥ ) ⊥ and the prior sentence shows that the dimension of the two subspaces M and ( M ⊥ ) ⊥ are equal, we can conclude that M equals ( M ⊥ ) ⊥ .)

    2. Because M ⊆ N , any v → that is perpendicular to every vector in N is also perpendicular to every vector in M . But that sentence simply says that N ⊥ ⊆ M ⊥ .

    3. We will again show that the sets are equal by mutual inclusion. The first direction is easy; any v → perpendicular to every vector in M + N = { m → + n → ∣ m → ∈ M , n → ∈ N } is perpendicular to every vector of the form m → + 0 → (that is, every vector in M ) and every vector of the form 0 → + n → (every vector in N ), and so ( M + N ) ⊥ ⊆ M ⊥ ∩ N ⊥ . The second direction is also routine; any vector v → ∈ M ⊥ ∩ N ⊥ is perpendicular to any vector of the form c m → + d n → because v → ⋅ ( c m → + d n → ) = c ⋅ ( v → ⋅ m → ) + d ⋅ ( v → ⋅ n → ) = c ⋅ 0 + d ⋅ 0 = 0 .

  18. Exercise 3.27 Worked answer

    Recommended. The material in this subsection allows us to express a geometric relationship that we have not yet seen between the range space and the null space of a linear map.

    1. Represent f : ℝ 3 → ℝ given by

      ( v 1 v 2 v 3 ) ↦ 1 v 1 + 2 v 2 + 3 v 3

      with respect to the standard bases and show that

      ( 1 2 3 )

      is a member of the perp of the null space. Prove that 𝒩 ( f ) ⊥ is equal to the span of this vector.

    2. Generalize that to apply to any f : ℝ n → ℝ .

    3. Represent f : ℝ 3 → ℝ 2

      ( v 1 v 2 v 3 ) ↦ ( 1 v 1 + 2 v 2 + 3 v 3 4 v 1 + 5 v 2 + 6 v 3 )

      with respect to the standard bases and show that

      ( 1 2 3 ) , ( 4 5 6 )

      are both members of the perp of the null space. Prove that 𝒩 ( f ) ⊥ is the span of these two. (Hint. See the third item of Exercise 3.26.)

    4. Generalize that to apply to any f : ℝ n → ℝ m .

    In [Strang 93] this is called the Fundamental Theorem of Linear Algebra

    Back to Exercise 3.27

    Answer.

    1. The representation of

      ( v 1 v 2 v 3 ) ⟼ f 1 v 1 + 2 v 2 + 3 v 3

      is this.

      Rep ℰ 3 , ℰ 1 ( f ) = ( 1 2 3 )

      By the definition of f

      𝒩 ( f ) = { ( v 1 v 2 v 3 ) ∣ 1 v 1 + 2 v 2 + 3 v 3 = 0 } = { ( v 1 v 2 v 3 ) ∣ ( 1 2 3 ) ⋅ ( v 1 v 2 v 3 ) = 0 }

      and this second description exactly says this.

      𝒩 ( f ) ⊥ = [ { ( 1 2 3 ) } ]

    2. The generalization is that for any f : ℝ n → ℝ there is a vector h → so that

      ( v 1 ⋮ v n ) ⟼ f h 1 v 1 + ⋯ + h n v n

      and h → ∈ 𝒩 ( f ) ⊥ . We can prove this by, as in the prior item, representing f with respect to the standard bases and taking h → to be the column vector gotten by transposing the one row of that matrix representation.

    3. Of course,

      Rep ℰ 3 , ℰ 2 ( f ) = ( 1 2 3 4 5 6 )

      and so the null space is this set.

      𝒩 ( f ) { ( v 1 v 2 v 3 ) ∣ ( 1 2 3 4 5 6 ) ( v 1 v 2 v 3 ) = ( 0 0 ) }

      That description makes clear that

      ( 1 2 3 ) , ( 4 5 6 ) ∈ 𝒩 ( f ) ⊥

      and since 𝒩 ( f ) ⊥ is a subspace of ℝ n , the span of the two vectors is a subspace of the perp of the null space. To see that this containment is an equality, take

      M = [ { ( 1 2 3 ) } ] N = [ { ( 4 5 6 ) } ]

      in the third item of Exercise 3.26, as suggested in the hint.

    4. As above, generalizing from the specific case is easy: for any f : ℝ n → ℝ m the matrix H representing the map with respect to the standard bases describes the action

      ( v 1 ⋮ v n ) ⟼ f ( h 1 , 1 v 1 + h 1 , 2 v 2 + ⋯ + h 1 , n v n ⋮ h m , 1 v 1 + h m , 2 v 2 + ⋯ + h m , n v n )

      and the description of the null space gives that on transposing the m  rows of H

      h → 1 = ( h 1 , 1 h 1 , 2 ⋮ h 1 , n ) , … h → m = ( h m , 1 h m , 2 ⋮ h m , n )

      we have 𝒩 ( f ) ⊥ = [ { h → 1 , … , h → m } ] . ([Strang 93] describes this space as the transpose of the row space of H .)

  19. Exercise 3.28 Worked answer

    Define a projection to be a linear transformation t : V → V with the property that repeating the projection does nothing more than does the projection alone:  ( t ∘ t ) ( v → ) = t ( v → ) for all v → ∈ V .

    1. Show that orthogonal projection into a line has that property.

    2. Show that projection along a subspace has that property.

    3. Show that for any such t there is a basis B = ⟨ β → 1 , … , β → n ⟩ for V such that

      t ( β → i ) = { β → i i = 1 , 2 , … , r 0 → i = r + 1 , r + 2 , … , n

      where r is the rank of t .

    4. Conclude that every projection is a projection along a subspace.

    5. Also conclude that every projection has a representation

      Rep B , B ( t ) = ( I Z Z Z )

      in block partial-identity form.

    Back to Exercise 3.28

    Answer.

    1. First note that if a vector v → is already in the line then the orthogonal projection gives v → itself. One way to verify this is to apply the formula for projection into the line spanned by a vector s → , namely ( v → ⋅ s → / s → ⋅ s → ) ⋅ s → . Taking the line as { k ⋅ v → ∣ k ∈ ℝ } (the v → = 0 → case is separate but easy) gives ( v → ⋅ v → / v → ⋅ v → ) ⋅ v → , which simplifies to v → , as required.

      Now, that answers the question because after once projecting into the line, the result proj ℓ ( v → ) is in that line. The prior paragraph says that projecting into the same line again will have no effect.

    2. The argument here is similar to the one in the prior item. With V = M ⊕ N , the projection of v → = m → + n → is proj M , N ( v → ) = m → . Now repeating the projection will give proj M , N ( m → ) = m → , as required, because the decomposition of a member of M into the sum of a member of M and a member of N is m → = m → + 0 → . Thus, projecting twice into M along N has the same effect as projecting once.

    3. As suggested by the prior items, the condition gives that t leaves vectors in the range space unchanged, and hints that we should take β → 1 , …, β → r to be basis vectors for the range, that is, that we should take the range space of t for M (so that dim ⁡ ( M ) = r ). As for the complement, we write N for the null space of t and we will show that V = M ⊕ N .

      To show this, we can show that their intersection is trivial M ∩ N = { 0 → } and that they sum to the entire space M + N = V . For the first, if a vector m → is in the range space then there is a v → ∈ V with t ( v → ) = m → , and the condition on t gives that t ( m → ) = ( t ∘ t ) ( v → ) = t ( v → ) = m → , while if that same vector is also in the null space then t ( m → ) = 0 → and so the intersection of the range space and null space is trivial. For the second, to write an arbitrary v → as the sum of a vector from the range space and a vector from the null space, the fact that the condition t ( v → ) = t ( t ( v → ) ) can be rewritten as t ( v → − t ( v → ) ) = 0 → suggests taking v → = t ( v → ) + ( v → − t ( v → ) ) .

      To finish we taking a basis B = ⟨ β → 1 , … , β → n ⟩ for V where ⟨ β → 1 , … , β → r ⟩ is a basis for the range space M and ⟨ β → r + 1 , … , β → n ⟩ is a basis for the null space N .

    4. Every projection (as defined in this exercise) is a projection into its range space and along its null space.

    5. This also follows immediately from the third item.

  20. Exercise 3.29 Worked answer

    A square matrix is symmetric if each i , j entry equals the j , i entry (i.e., if the matrix equals its transpose). Show that the projection matrix A ( A 𝖳 A ) − 1 A 𝖳 is symmetric. [Strang 80] Hint. Find properties of transposes by looking in the index under ‘transpose’.

    Back to Exercise 3.29

    Answer. For any matrix M we have that ( M − 1 ) 𝖳 = ( M 𝖳 ) − 1 , and for any two matrices M , N we have that M N 𝖳 = N 𝖳 M 𝖳 (provided, of course, that the inverse and product are defined). Applying these two gives that the matrix equals its transpose.

    ( A ( A 𝖳 A ) − 1 A 𝖳 ) 𝖳 = ( A 𝖳 𝖳 ) ( ( ( A 𝖳 A ) − 1 ) 𝖳 ) ( A 𝖳 ) = ( A 𝖳 𝖳 ) ( ( ( A 𝖳 A ) 𝖳 ) − 1 ) ( A 𝖳 ) = A ( A 𝖳 A 𝖳 𝖳 ) − 1 A 𝖳 = A ( A 𝖳 A ) − 1 A 𝖳

References cited in this section

Strang 93

Gilbert Strang The Fundamental Theorem of Linear Algebra, American Mathematical Monthly, Nov. 1993, p. 848–855.

Strang 80

Gilbert Strang, Linear Algebra and its Applications, second edition, Harcourt Brace Jovanovich, 1980.