Original English by Jim Hefferon — 34 validated sections. The original mathematics and supplied answers below are preserved. This is a partial-book reading edition, not the complete book or an Everyday-English rewrite.

Source revision df2262e089a02651c127f1dd12649c4622ee1383; CC BY-SA 2.5 option, with original component credits retained. This is not an Everyday-English rewrite. The complete source section is included; whole-book include and conditional closure remains separate.

Source-preserving rebuild, navigation, source packaging and current deterministic checks: OpenAI Codex — GPT-6 Astra, Ultra effort. Jim Hefferon remains the author of the mathematics. Earlier intermediate-conversion runtime identity is not established by its retained receipts and is not reassigned to this rebuild. No human review or exhaustive proof certification is claimed.

Linear Geometry

If you have seen the elements of vectors then this section is an optional review. However, later work will refer to this material so if this is not a review then it is not optional.

In the first section we had to do a bit of work to show that there are only three types of solution sets—singleton, empty, and infinite. But this is easy to see geometrically in the case of systems with two equations and two unknowns. Draw each two-unknowns equation as a line in the plane and then the two lines could have a unique intersection, be parallel, or be the same line.

Unique solution

Two lines intersect at one point, illustrating the unique solution of the displayed two-equation system.
3 x + 2 y = 7 x − y = − 1

No solutions

Two distinct parallel lines do not intersect, illustrating a system with no solution.
3 x + 2 y = 7 3 x + 2 y = 4

Infinitely many
solutions

The two equations describe the same line, giving infinitely many common solutions.
3 x + 2 y = 7 6 x + 4 y = 14

These pictures aren’t a short way to prove the results from the prior section, because those results apply to linear systems with any number of variables. But they do provide a visual insight, another way of seeing those results.

This section develops what we need to express our results geometrically. In particular, while the two-dimensional case is familiar, to extend to systems with more than two unknowns we shall need some higher-dimensional geometry.

Vectors in Space

“Higher-dimensional geometry” sounds exotic. It is exotic—interesting and eye-opening. But it isn’t distant or unreachable.

We begin by defining one-dimensional space to be ℝ . To see that the definition is reasonable, picture a one-dimensional space

A line represents one-dimensional space before an origin and a unit direction are selected.

and pick a point to label 0 and another to label  1 .

A number line with points labelled 0 and 1 fixes an origin, scale and direction.

Now, with a scale and a direction, we have a correspondence with ℝ . For instance, to find the point matching + 2.17 , start at 0 and head in the direction of 1 , and go 2.17 times as far.

The basic idea here, combining magnitude with direction, is the key to extending to higher dimensions.

An object in an  ℝ n that is comprised of a magnitude and a direction is a vector (we use the same word as in the prior section because we shall show below how to describe such an object with a column vector). We can draw a vector as having some length and pointing in some direction.

An arrow represents a vector with a magnitude and direction.

There is a subtlety involved in the definition of a vector as consisting of a magnitude and a direction—these

Two arrows in different positions have the same length and direction and represent equal free vectors.

are equal, even though they start in different places They are equal because they have equal lengths and equal directions. Again: those vectors are not just alike, they are equal.

How can things that are in different places be equal? Think of a vector as representing a displacement (the word ‘vector’ is Latin for “carrier” or “traveler”). These two squares undergo displacements that are equal despite that they start in different places.

Two equal arrows displace two squares by the same amount even though their starting positions differ.

When we want to emphasize this property vectors have of not being anchored we refer to them as free vectors. Thus, these free vectors are equal, as each is a displacement of one over and two up.

Several equal arrows at different starting points all displace one unit right and two units up.

More generally, vectors in the plane are the same if and only if they have the same change in first components and the same change in second components: the vector extending from ( a 1 , a 2 ) to ( b 1 , b 2 ) equals the vector from ( c 1 , c 2 ) to ( d 1 , d 2 ) if and only if b 1 − a 1 = d 1 − c 1 and b 2 − a 2 = d 2 − c 2 .

Saying ‘the vector that, were it to start at ( a 1 , a 2 ) , would extend to ( b 1 , b 2 ) ’ would be unwieldy. We instead describe that vector as

( b 1 − a 1 b 2 − a 2 )

so that we represent the ‘one over and two up’ arrows shown above in this way.

( 1 2 )

We often draw the arrow as starting at the origin, and we then say it is in the canonical position (or natural position or standard position). When

v → = ( v 1 v 2 )

is in canonical position then it extends from the origin to the endpoint ( v 1 , v 2 ) .

We will typically say “the point

( 1 2 ) ”

rather than “the endpoint of the canonical position of” that vector. Thus, we will call each of these ℝ 2 .

{ ( x 1 , x 2 ) ∣ x 1 , x 2 ∈ ℝ } { ( x 1 x 2 ) ∣ x 1 , x 2 ∈ ℝ }

In the prior section we defined vectors and vector operations with an algebraic motivation;

r ⋅ ( v 1 v 2 ) = ( r v 1 r v 2 ) ( v 1 v 2 ) + ( w 1 w 2 ) = ( v 1 + w 1 v 2 + w 2 )

we can now understand those operations geometrically. For instance, if v → represents a displacement then 3 v → represents a displacement in the same direction but three times as far and − 1 v → represents a displacement of the same distance as v → but in the opposite direction.

Vectors illustrate scalar multiplication: the original vector, its negative, and three times that vector.

And, where v → and w → represent displacements, v → + w → represents those displacements combined.

Head-to-tail addition of two vectors gives a vector from the initial point to the final point.

The long arrow is the combined displacement in this sense: imagine that you are walking on a ship’s deck. Suppose that in one minute the ship’s motion gives it a displacement relative to the sea of v → , and in the same minute your walking gives you a displacement relative to the ship’s deck of w → . Then v → + w → is your displacement relative to the sea.

Another way to understand the vector sum is with the parallelogram rule. Draw the parallelogram formed by the vectors v → and w → . Then the sum v → + w → extends along the diagonal to the far corner.

The parallelogram rule shows the diagonal as the sum of the two side vectors.

The above drawings show how vectors and vector operations behave in ℝ 2 . We can extend to ℝ 3 , or to even higher-dimensional spaces where we have no pictures, with the obvious generalization: the free vector that, if it starts at ( a 1 , … , a n ) , ends at ( b 1 , … , b n ) , is represented by this column.

( b 1 − a 1 ⋮ b 1 − a 1 b n − a n )

Vectors are equal if they have the same representation. We aren’t too careful about distinguishing between a point and the vector whose canonical representation ends at that point.

ℝ n = { ( v 1 ⋮ v 1 v n ) ∣ v 1 , … , v n ∈ ℝ }

And, we do addition and scalar multiplication component-wise.

Having considered points, we next turn to lines. In ℝ 2 , the line through ( 1 , 2 ) and ( 3 , 1 ) is comprised of (the endpoints of) the vectors in this set.

A line is described by a fixed starting point plus all scalar multiples of one direction vector.

In the description the vector that is associated with the parameter  t

( 2 − 1 ) = ( 3 1 ) − ( 1 2 )

is the one shown in the picture as having its whole body in the line—it is a direction vector for the line. Note that points on the line to the left of x = 1 are described using negative values of t .

In ℝ 3 , the line through ( 1 , 2 , 1 ) and ( 0 , 3 , 2 ) is the set of (endpoints of) vectors of this form

A line in three-dimensional space passes through the points (1,2,1) and (0,3,2).

and lines in even higher-dimensional spaces work in the same way.

In ℝ 3 , a line uses one parameter so that a particle on that line would be free to move back and forth in one dimension. A plane involves two parameters. For example, the plane through the points ( 1 , 0 , 5 ) , ( 2 , 1 , − 3 ) , and ( − 2 , 4 , 0.5 ) consists of (endpoints of) the vectors in this set.

{ ( 1 0 5 ) + t ( 1 1 − 8 ) + s ( − 3 4 − 4.5 ) ∣ t , s ∈ ℝ }

The column vectors associated with the parameters come from these calculations.

( 1 1 − 8 ) = ( 2 1 − 3 ) − ( 1 0 5 ) ( − 3 4 − 4.5 ) = ( − 2 4 0.5 ) − ( 1 0 5 )

As with the line, note that we describe some points in this plane with negative t ’s or negative s ’s or both.

Calculus books often describe a plane by using a single linear equation.

The plane 2x+y+z=4 is shown in three-dimensional space together with its coordinate axes.

To translate from this to the vector description, think of this as a one-equation linear system and parametrize: x = 2 − y / 2 − z / 2 .

The plane 2x+y+z=4 is described by a base point and two independent direction vectors.

Shown in grey are the vectors associated with y and  z , offset from the origin by  2 units along the x -axis, so that their entire body lies in the plane. Thus the vector sum of the two, shown in black, has its entire body in the plane along with the rest of the parallelogram.

Generalizing, a set of the form { p → + t 1 v → 1 + t 2 v → 2 + ⋯ + t k v → k ∣ t 1 , … , t k ∈ ℝ } where v → 1 , … , v → k ∈ ℝ n and k ≤ n is a k -dimensional linear surface (or k -flat). For example, in ℝ 4

{ ( 2 π 3 − 0.5 ) + t ( 1 0 0 0 ) ∣ t ∈ ℝ }

is a line,

{ ( 0 0 0 0 ) + t ( 1 1 0 − 1 ) + s ( 2 0 1 0 ) ∣ t , s ∈ ℝ }

is a plane, and

{ ( 3 1 − 2 0.5 ) + r ( 0 0 0 − 1 ) + s ( 1 0 1 0 ) + t ( 2 0 1 0 ) ∣ r , s , t ∈ ℝ }

is a three-dimensional linear surface. Again, the intuition is that a line permits motion in one direction, a plane permits motion in combinations of two directions, etc. When the dimension of the linear surface is one less than the dimension of the space, that is, when in ℝ n we have an ( n − 1 ) -flat, the surface is called a hyperplane.

A description of a linear surface can be misleading about the dimension. For example, this

L = { ( 1 0 − 1 − 2 ) + t ( 1 1 0 − 1 ) + s ( 2 2 0 − 2 ) ∣ t , s ∈ ℝ }

is a degenerate plane because it is actually a line, since the vectors are multiples of each other and we can omit one.

L = { ( 1 0 − 1 − 2 ) + r ( 1 1 0 − 1 ) ∣ r ∈ ℝ }

We shall see in the Linear Independence section of Chapter Two what relationships among vectors causes the linear surface they generate to be degenerate.

We now can restate in geometric terms our conclusions from earlier. First, the solution set of a linear system with n unknowns is a linear surface in ℝ n . Specifically, it is a k -dimensional linear surface, where k is the number of free variables in an echelon form version of the system. For instance, in the single equation case the solution set is an n − 1 -dimensional hyperplane in ℝ n , where n ≥ 1 . Second, the solution set of a homogeneous linear system is a linear surface passing through the origin. Finally, we can view the general solution set of any linear system as being the solution set of its associated homogeneous system offset from the origin by a vector, namely by any particular solution.

Exercises

  1. Exercise 1.1 Worked answer

    Recommended. Find the canonical name for each vector.

    1. the vector from ( 2 , 1 ) to ( 4 , 2 ) in ℝ 2

    2. the vector from ( 3 , 3 ) to ( 2 , 5 ) in ℝ 2

    3. the vector from ( 1 , 0 , 6 ) to ( 5 , 0 , 3 ) in ℝ 3

    4. the vector from ( 6 , 8 , 8 ) to ( 6 , 8 , 8 ) in ℝ 3

    Back to Exercise 1.1

    Answer.

    1. ( 2 1 )

    2. ( − 1 2 )

    3. ( 4 0 − 3 )

    4. ( 0 0 0 )

  2. Exercise 1.2 Worked answer

    Recommended. Decide if the two vectors are equal.

    1. the vector from ( 5 , 3 ) to ( 6 , 2 ) and the vector from ( 1 , − 2 ) to ( 1 , 1 )

    2. the vector from ( 2 , 1 , 1 ) to ( 3 , 0 , 4 ) and the vector from ( 5 , 1 , 4 ) to ( 6 , 0 , 7 )

    Back to Exercise 1.2

    Answer.

    1. No, their canonical positions are different.

      ( 1 − 1 ) ( 0 3 )

    2. Yes, their canonical positions are the same.

      ( 1 − 1 3 )

  3. Exercise 1.3 Worked answer

    Recommended. Does ( 1 , 0 , 2 , 1 ) lie on the line through ( − 2 , 1 , 1 , 0 ) and ( 5 , 10 , − 1 , 4 ) ?

    Back to Exercise 1.3

    Answer. That line is this set.

    { ( − 2 1 1 0 ) + ( 7 9 − 2 4 ) t ∣ t ∈ ℝ }

    Note that this system

    − 2 + 7 t = 1 1 + 9 t = 0 1 − 2 t = 2 0 + 4 t = 1

    has no solution. Thus the given point is not in the line.

  4. Exercise 1.4 Worked answer

    Recommended.

    1. Describe the plane through ( 1 , 1 , 5 , − 1 ) , ( 2 , 2 , 2 , 0 ) , and ( 3 , 1 , 0 , 4 ) .

    2. Is the origin in that plane?

    Back to Exercise 1.4

    Answer.

    1. Note that

      ( 2 2 2 0 ) − ( 1 1 5 − 1 ) = ( 1 1 − 3 1 ) ( 3 1 0 4 ) − ( 1 1 5 − 1 ) = ( 2 0 − 5 5 )

      and so the plane is this set.

      { ( 1 1 5 − 1 ) + ( 1 1 − 3 1 ) t + ( 2 0 − 5 5 ) s ∣ t , s ∈ ℝ }

    2. No; this system

      1 + 1 t + 2 s = 0 1 + 1 t = 0 5 − 3 t − 5 s = 0 − 1 + 1 t + 5 s = 0

      has no solution.

  5. Exercise 1.5 Worked answer

    Give a vector description of each.

    1. the plane subset of  ℝ 3 with equation x − 2 y + z = 4

    2. the plane in  ℝ 3 with equation 2 x + y + 4 z = − 1

    3. the hyperplane subset of  ℝ 4 with equation x + y + z + w = 10

    Back to Exercise 1.5

    Answer.

    1. Think of x − 2 y + z = 4 as a one-equation linear system and parametrize with the variables y and  z to get x = 4 + 2 y − z . That gives this vector description of the plane.

      { ( x y z ) = ( 4 0 0 ) + ( 2 1 0 ) ⋅ y + ( − 1 0 1 ) ⋅ z ∣ y , z ∈ ℝ }

    2. Parametrizing gives x = − ( 1 / 2 ) − ( 1 / 2 ) y − 2 z , so this is the vector description.

      ( x y z ) = ( − 1 / 2 0 0 ) + ( − 1 / 2 1 0 ) ⋅ y + ( − 2 0 1 ) ⋅ z

    3. Here x = 10 − y − z − w and so we get a vector description with three parameters.

      ( x y z w ) = ( 10 0 0 0 ) + ( − 1 1 0 0 ) ⋅ y + ( − 1 0 1 0 ) ⋅ z + ( − 1 0 0 1 ) ⋅ w

  6. Exercise 1.6 Worked answer

    Describe the plane that contains this point and line.

    ( 2 0 3 ) { ( − 1 0 − 4 ) + ( 1 1 2 ) t ∣ t ∈ ℝ }

    Back to Exercise 1.6

    Answer. The vector

    ( 2 0 3 )

    is not in the line. Because

    ( 2 0 3 ) − ( − 1 0 − 4 ) = ( 3 0 7 )

    we can describe that plane in this way.

    { ( − 1 0 − 4 ) + m ( 1 1 2 ) + n ( 3 0 7 ) ∣ m , n ∈ ℝ }

  7. Exercise 1.7 Worked answer

    Recommended. Intersect these planes.

    { ( 1 1 1 ) t + ( 0 1 3 ) s ∣ t , s ∈ ℝ } { ( 1 1 0 ) + ( 0 3 0 ) k + ( 2 0 4 ) m ∣ k , m ∈ ℝ }

    Back to Exercise 1.7

    Answer. The points of coincidence are solutions of this system.

    t = 1 + 2 m t + s = 1 + 3 k t + 3 s = 4 m

    Gauss’s Method

    ( 1 0 0 − 2 1 1 1 − 3 0 1 1 3 0 − 4 0 ) ⟶ − ρ 1 + ρ 3 − ρ 1 + ρ 2 ( ( 1 0 0 − 2 1 0 1 − 3 2 0 0 3 0 − 2 − 1 ) ⟶ − 3 ρ 2 + ρ 3 ( ( 1 0 0 − 2 1 0 1 − 3 2 0 0 0 9 − 8 − 1 )

    gives k = − ( 1 / 9 ) + ( 8 / 9 ) m , so s = − ( 1 / 3 ) + ( 2 / 3 ) m and t = 1 + 2 m . The intersection is this.

    { ( 1 1 0 ) + ( 0 3 0 ) ( − 1 9 + 8 9 m ) + ( 2 0 4 ) m ∣ m ∈ ℝ } = { ( 1 2 / 3 0 ) + ( 2 8 / 3 4 ) m ∣ m ∈ ℝ }

  8. Exercise 1.8 Worked answer

    Recommended. Intersect each pair, if possible.

    1. { ( 1 1 2 ) + t ( 0 1 1 ) ∣ t ∈ ℝ } ,  { ( 1 3 − 2 ) + s ( 0 1 2 ) ∣ s ∈ ℝ }

    2. { ( 2 0 1 ) + t ( 1 1 − 1 ) ∣ t ∈ ℝ } ,  { s ( 0 1 2 ) + w ( 0 4 1 ) ∣ s , w ∈ ℝ }

    Back to Exercise 1.8

    Answer.

    1. The system

      1 = 1 1 + t = 3 + s 2 + t = − 2 + 2 s

      gives s = 6 and t = 8 , so this is the solution set.

      { ( 1 9 10 ) }

    2. This system

      2 + t = 0 t = s + 4 w 1 − t = 2 s + w

      gives t = − 2 , w = − 1 , and s = 2 so their intersection is this point.

      ( 0 − 2 3 )

  9. Exercise 1.9 Worked answer

    How should we define ℝ 0 ?

    Back to Exercise 1.9

    Answer. We shall later define it to be a set with one element—an “origin”.

  10. Exercise 1.10 Worked answer

    Puzzle. [Math. Mag., Jan. 1957] A person traveling eastward at a rate of 3 miles per hour finds that the wind appears to blow directly from the north. On doubling his speed it appears to come from the north east. What was the wind’s velocity?

    Back to Exercise 1.10

    Answer. This is how the answer was given in the cited source. The vector triangle is as follows, so w → = 3 2 from the north west.

    A wind triangle with a horizontal resultant and two diagonal vectors; a vertical segment marks the geometry of the construction.

  11. Exercise 1.11 Worked answer

    Euclid describes a plane as “a surface which lies evenly with the straight lines on itself”. Commentators such as Heron have interpreted this to mean, “(A plane surface is) such that, if a straight line pass through two points on it, the line coincides wholly with it at every spot, all ways”. (Translations from [Heath], pp. 171-172.) Do planes, as described in this section, have that property? Does this description adequately define planes?

    Back to Exercise 1.11

    Answer. Euclid no doubt is picturing a plane inside of ℝ 3 . Observe, however, that both ℝ 1 and ℝ 2 also satisfy that definition.

Length and Angle Measures

We’ve translated the first section’s results about solution sets into geometric terms, to better understand those sets. But we must be careful not to be misled by our own terms— labeling subsets of ℝ k of the forms { p → + t v → ∣ t ∈ ℝ } and { p → + t v → + s w → ∣ t , s ∈ ℝ } as ‘lines’ and ‘planes’ doesn’t make them act like the lines and planes of our past experience. Rather, we must ensure that the names suit the sets. While we can’t prove that the sets satisfy our intuition—we can’t prove anything about intuition—in this subsection we’ll observe that a result familiar from ℝ 2 and ℝ 3 , when generalized to arbitrary ℝ n , supports the idea that a line is straight and a plane is flat. Specifically, we’ll see how to do Euclidean geometry in a ‘plane’ by giving a definition of the angle between two ℝ n vectors, in the plane that they generate.

Definition 2.1 The length of a vector v → ∈ ℝ n is the square root of the sum of the squares of its components.

| v → | = v 1 2 + ⋯ + v n 2

Remark 2.2 This is a natural generalization of the Pythagorean Theorem. A classic motivating discussion is in [Polya].

For any nonzero v → , the vector v → / | v → | has length one. We say that the second normalizes v → to length one.

We can use that to get a formula for the angle between two vectors. Consider two vectors in ℝ 3 where neither is a multiple of the other

Two vectors in three-dimensional space are shown before translating them to a common origin.

(the special case of multiples will turn out below not to be an exception). They determine a two-dimensional plane— for instance, put them in canonical position and take the plane formed by the origin and the endpoints. In that plane consider the triangle with sides u → , v → , and u → − v → .

The two vectors are translated to a common origin and lie in one plane; their endpoints form a triangle.

Apply the Law of Cosines: | u → − v → | 2 = | u → | 2 + | v → | 2 − 2 | u → | | v → | cos ⁡ θ where θ is the angle between the vectors. The left side gives

( u 1 − v 1 ) 2 + ( u 2 − v 2 ) 2 + ( u 3 − v 3 ) 2 = ( u 1 2 − 2 u 1 v 1 + v 1 2 ) + ( u 2 2 − 2 u 2 v 2 + v 2 2 ) + ( u 3 2 − 2 u 3 v 3 + v 3 2 )

while the right side gives this.

( u 1 2 + u 2 2 + u 3 2 ) + ( v 1 2 + v 2 2 + v 3 2 ) − 2 | u → | | v → | cos ⁡ θ

Canceling squares u 1 2 , …, v 3 2 and dividing by 2 gives a formula for the angle.

θ = arccos ⁡ ( u 1 v 1 + u 2 v 2 + u 3 v 3 | u → | | v → | )

In higher dimensions we cannot draw pictures as above but we can instead make the argument analytically. First, the form of the numerator is clear; it comes from the middle terms of ( u i − v i ) 2 .

Definition 2.3 The dot product (or inner product or scalar product) of two n -component real vectors is the linear combination of their components.

u → ⋅ v → = u 1 v 1 + u 2 v 2 + ⋯ + u n v n

Note that the dot product of two vectors is a real number, not a vector, and that the dot product is only defined if the two vectors have the same number of components. Note also that dot product is related to length: u → ⋅ u → = u 1 u 1 + ⋯ + u n u n = | u → | 2 .

Remark 2.4 Some authors require that the first vector be a row vector and that the second vector be a column vector. We shall not be that strict and will allow the dot product operation between two column vectors.

Still reasoning analytically but guided by the pictures, we use the next theorem to argue that the triangle formed by the line segments making the bodies of u → , v → , and u → + v → in ℝ n lies in the planar subset of ℝ n generated by u → and v → (see the figure below).

Theorem 2.5 (Triangle Inequality) For any u → , v → ∈ ℝ n ,

| u → + v → | ≤ | u → | + | v → |

with equality if and only if one of the vectors is a nonnegative scalar multiple of the other one.

This is the source of the familiar saying, “The shortest distance between two points is in a straight line.”

Triangle-inequality diagram: consecutive vectors u and v connect start to finish; u+v takes the direct path.

Proof (We’ll use some algebraic properties of dot product that we have not yet checked, for instance that u → ⋅ ( a → + b → ) = u → ⋅ a → + u → ⋅ b → and that u → ⋅ v → = v → ⋅ u → . See Exercise 2.18.) Since all the numbers are positive, the inequality holds if and only if its square holds.

| u → + v → | 2 ≤ ( | u → | + | v → | ) 2 ( u → + v → ) ⋅ ( u → + v → ) ≤ | u → | 2 + 2 | u → | | v → | + | v → | 2 u → ⋅ u → + u → ⋅ v → + v → ⋅ u → + v → ⋅ v → ≤ u → ⋅ u → + 2 | u → | | v → | + v → ⋅ v → 2 u → ⋅ v → ≤ 2 | u → | | v → |

That, in turn, holds if and only if the relationship obtained by multiplying both sides by the nonnegative numbers | u → | and | v → |

2 ( | v → | u → ) ⋅ ( | u → | v → ) ≤ 2 | u → | 2 | v → | 2

and rewriting

0 ≤ | u → | 2 | v → | 2 − 2 ( | v → | u → ) ⋅ ( | u → | v → ) + | u → | 2 | v → | 2

is true. But factoring shows that it is true

0 ≤ ( | u → | v → − | v → | u → ) ⋅ ( | u → | v → − | v → | u → )

since it only says that the square of the length of the vector | u → | v → − | v → | u → is not negative. As for equality, it holds when, and only when, | u → | v → − | v → | u → is 0 → . The check that | u → | v → = | v → | u → if and only if one vector is a nonnegative real scalar multiple of the other is easy.

QED

This result supports the intuition that even in higher-dimensional spaces, lines are straight and planes are flat. We can easily check from the definition that linear surfaces have the property that for any two points in that surface, the line segment between them is contained in that surface. But if the linear surface were not flat then that would allow for a shortcut.

A bent surface contains a curved path from P to Q; a straight shortcut shown in grey leaves that surface.

Because the Triangle Inequality says that in any ℝ n the shortest cut between two endpoints is simply the line segment connecting them, linear surfaces have no bends.

Back to the definition of angle measure. The heart of the Triangle Inequality’s proof is the u → ⋅ v → ≤ | u → | | v → | line. We might wonder if some pairs of vectors satisfy the inequality in this way: while u → ⋅ v → is a large number, with absolute value bigger than the right-hand side, it is a negative large number. The next result says that does not happen.

Corollary 2.6 (Cauchy-Schwarz Inequality) For any u → , v → ∈ ℝ n ,

| u → ⋅ v → | ≤ | u → | | v → |

with equality if and only if one vector is a scalar multiple of the other.

Proof The Triangle Inequality’s proof shows that u → ⋅ v → ≤ | u → | | v → | so if u → ⋅ v → is positive or zero then we are done. If u → ⋅ v → is negative then this holds.

| u → ⋅ v → | = − ( u → ⋅ v → ) = ( − u → ) ⋅ v → ≤ | − u → | | v → | = | u → | | v → |

The equality condition is Exercise 2.19.

QED

The Cauchy-Schwarz inequality assures us that the next definition makes sense because the fraction has absolute value less than or equal to one.

Definition 2.7 The angle between two nonzero vectors u → , v → ∈ ℝ n is

θ = arccos ⁡ ( u → ⋅ v → | u → | | v → | )

(if either is the zero vector then we take the angle to be a right angle).

Corollary 2.8 Vectors from ℝ n are orthogonal, that is, perpendicular, if and only if their dot product is zero. They are parallel if and only if their dot product equals the product of their lengths.

Example 2.9 These vectors are orthogonal.

Two perpendicular arrows in a plane illustrate orthogonal vectors.

( 1 − 1 ) ⋅ ( 1 1 ) = 0

We’ve drawn the arrows away from canonical position but nevertheless the vectors are orthogonal.

Example 2.10 The ℝ 3 angle formula given at the start of this subsection is a special case of the definition. Between these two

Two arrows in three-dimensional space illustrate vectors that are not orthogonal.

the angle is

arccos ⁡ ( ( 1 ) ( 0 ) + ( 1 ) ( 3 ) + ( 0 ) ( 2 ) 1 2 + 1 2 + 0 2 0 2 + 3 2 + 2 2 ) = arccos ⁡ ( 3 2 13 )

approximately 0.94   radians . Notice that these vectors are not orthogonal. Although the y z -plane may appear to be perpendicular to the x y -plane, in fact the two planes are that way only in the weak sense that there are vectors in each orthogonal to all vectors in the other. Not every vector in each is orthogonal to all vectors in the other.

Exercises

  1. Exercise 2.11 Worked answer

    Recommended. Find the length of each vector.

    1. ( 3 1 )

    2. ( − 1 2 )

    3. ( 4 1 1 )

    4. ( 0 0 0 )

    5. ( 1 − 1 1 0 )

    Back to Exercise 2.11

    Answer.

    1. 3 2 + 1 2 = 10

    2. 5

    3. 18

    4. 0

    5. 3

  2. Exercise 2.12 Worked answer

    Recommended. Find the angle between each two, if it is defined.

    1. ( 1 2 ) , ( 1 4 )

    2. ( 1 2 0 ) , ( 0 4 1 )

    3. ( 1 2 ) , ( 1 4 − 1 )

    Back to Exercise 2.12

    Answer.

    1. arccos ⁡ ( 9 / 85 ) ≈ 0.22  radians

    2. arccos ⁡ ( 8 / 85 ) ≈ 0.52  radians

    3. Not defined.

  3. Exercise 2.13 Worked answer

    Recommended. [Ohanian] During maneuvers preceding the Battle of Jutland, the British battle cruiser Lion moved as follows (in nautical miles):  1.2 miles north, 6.1 miles 38 degrees east of south, 4.0 miles at 89 degrees east of north, and 6.5 miles at 31 degrees east of north. Find the distance between starting and ending positions. (Ignore the earth’s curvature.)

    Back to Exercise 2.13

    Answer. We express each displacement as a vector, rounded to one decimal place because that’s the accuracy of the problem’s statement, and add to find the total displacement (ignoring the curvature of the earth).

    ( 0.0 1.2 ) + ( 3.8 − 4.8 ) + ( 4.0 0.1 ) + ( 3.3 5.6 ) = ( 11.1 2.1 )

    The distance is 11.1 2 + 2.1 2 ≈ 11.3 .

  4. Exercise 2.14 Worked answer

    Find k so that these two vectors are perpendicular.

    ( k 1 ) ( 4 3 )

    Back to Exercise 2.14

    Answer. Solve ( k ) ( 4 ) + ( 1 ) ( 3 ) = 0 to get k = − 3 / 4 .

  5. Exercise 2.15 Worked answer

    Describe the set of vectors in ℝ 3 orthogonal to the one with entries 1 , 3 , and  − 1 .

    Back to Exercise 2.15

    Answer. We could describe the set

    { ( x y z ) ∣ 1 x + 3 y − 1 z = 0 }

    with parameters in this way.

    { ( − 3 1 0 ) y + ( 1 0 1 ) z ∣ y , z ∈ ℝ }

  6. Exercise 2.16 Worked answer

    Recommended.

    1. Find the angle between the diagonal of the unit square in ℝ 2 and any one of the axes.

    2. Find the angle between the diagonal of the unit cube in ℝ 3 and one of the axes.

    3. Find the angle between the diagonal of the unit cube in ℝ n and one of the axes.

    4. What is the limit, as n goes to ∞ , of the angle between the diagonal of the unit cube in ℝ n and any one of the axes?

    Back to Exercise 2.16

    Answer.

    1. We can use the x -axis.

      arccos ⁡ ( ( 1 ) ( 1 ) + ( 0 ) ( 1 ) 1 2 ) ≈ 0.79   radians

    2. Again, use the x -axis.

      arccos ⁡ ( ( 1 ) ( 1 ) + ( 0 ) ( 1 ) + ( 0 ) ( 1 ) 1 3 ) ≈ 0.96   radians

    3. The x -axis worked before and it will work again.

      arccos ⁡ ( ( 1 ) ( 1 ) + ⋯ + ( 0 ) ( 1 ) 1 n ) = arccos ⁡ ( 1 n )

    4. Using the formula from the prior item, lim n → ∞ arccos ⁡ ( 1 / n ) = π / 2  radians .

  7. Exercise 2.17 Worked answer

    Is any vector perpendicular to itself?

    Back to Exercise 2.17

    Answer. Clearly u 1 u 1 + ⋯ + u n u n is zero if and only if each u i is zero. So only 0 → ∈ ℝ n is perpendicular to itself.

  8. Exercise 2.18 Worked answer

    Describe the algebraic properties of dot product.

    1. Is it right-distributive over addition: ( u → + v → ) ⋅ w → = u → ⋅ w → + v → ⋅ w → ?

    2. Is it left-distributive (over addition)?

    3. Does it commute?

    4. Associate?

    5. How does it interact with scalar multiplication?

    As always, you must back any assertion with a suitable argument.

    Back to Exercise 2.18

    Answer. In each item below, assume that the vectors u → , v → , w → ∈ ℝ n have components u 1 , … , u n , v 1 , … , w n .

    1. Dot product is right-distributive.

      ( u → + v → ) ⋅ w → = [ ( u 1 ⋮ u 1 u n ) + ( v 1 ⋮ v 1 v n ) ] ⋅ ( w 1 ⋮ w 1 w n ) = ( u 1 + v 1 ⋮ u 1 + v 1 u n + v n ) ⋅ ( w 1 ⋮ w 1 w n ) = ( u 1 + v 1 ) w 1 + ⋯ + ( u n + v n ) w n = ( u 1 w 1 + ⋯ + u n w n ) + ( v 1 w 1 + ⋯ + v n w n ) = u → ⋅ w → + v → ⋅ w →

    2. Dot product is also left distributive: w → ⋅ ( u → + v → ) = w → ⋅ u → + w → ⋅ v → . The proof is just like the prior one.

    3. Dot product commutes.

      ( u 1 ⋮ u 1 u n ) ⋅ ( v 1 ⋮ v 1 v n ) = u 1 v 1 + ⋯ + u n v n = v 1 u 1 + ⋯ + v n u n = ( v 1 ⋮ v 1 v n ) ⋅ ( u 1 ⋮ u 1 u n )

    4. Because u → ⋅ v → is a scalar, not a vector, the expression ( u → ⋅ v → ) ⋅ w → makes no sense; the dot product of a scalar and a vector is not defined.

    5. This is a vague question so it has many answers. Some are (1)  k ( u → ⋅ v → ) = ( k u → ) ⋅ v → and k ( u → ⋅ v → ) = u → ⋅ ( k v → ) , (2)  k ( u → ⋅ v → ) ≠ ( k u → ) ⋅ ( k v → ) (in general; an example is easy to produce), and (3)  | k v → | = | k | | v → | (the connection between length and dot product is that the square of the length is the dot product of a vector with itself).

  9. Exercise 2.19 Worked answer

    Verify the equality condition in Corollary 2.6, the Cauchy-Schwarz Inequality.

    1. Show that if u → is a negative scalar multiple of v → then u → ⋅ v → and v → ⋅ u → are less than or equal to zero.

    2. Show that | u → ⋅ v → | = | u → | | v → | if and only if one vector is a scalar multiple of the other.

    Back to Exercise 2.19

    Answer.

    1. Verifying that ( k x → ) ⋅ y → = k ( x → ⋅ y → ) = x → ⋅ ( k y → ) for k ∈ ℝ and x → , y → ∈ ℝ n is easy. Now, for k ∈ ℝ and v → , w → ∈ ℝ n , if u → = k v → then u → ⋅ v → = ( k v → ) ⋅ v → = k ( v → ⋅ v → ) , which is k times a nonnegative real.

      The v → = k u → half is similar (actually, taking the k in this paragraph to be the reciprocal of the k above gives that we need only worry about the k = 0 case).

    2. We first consider the u → ⋅ v → ≥ 0 case. From the Triangle Inequality we know that u → ⋅ v → = | u → | | v → | if and only if one vector is a nonnegative scalar multiple of the other. But that’s all we need because the first part of this exercise shows that, in a context where the dot product of the two vectors is positive, the two statements ‘one vector is a scalar multiple of the other’ and ‘one vector is a nonnegative scalar multiple of the other’, are equivalent.

      We finish by considering the u → ⋅ v → < 0 case. Because 0 < | u → ⋅ v → | = − ( u → ⋅ v → ) = ( − u → ) ⋅ v → and | u → | | v → | = | − u → | | v → | , we have that 0 < ( − u → ) ⋅ v → = | − u → | | v → | . Now the prior paragraph applies to give that one of the two vectors − u → and v → is a scalar multiple of the other. But that’s equivalent to the assertion that one of the two vectors u → and v → is a scalar multiple of the other, as desired.

  10. Exercise 2.20 Worked answer

    Suppose that u → ⋅ v → = u → ⋅ w → and u → ≠ 0 → . Must v → = w → ?

    Back to Exercise 2.20

    Answer. No. These give an example.

    u → = ( 1 0 ) v → = ( 1 0 ) w → = ( 1 1 )

  11. Exercise 2.21 Worked answer

    Recommended. Does any vector have length zero except a zero vector? (If “yes”, produce an example. If “no”, prove it.)

    Back to Exercise 2.21

    Answer. We prove that a vector has length zero if and only if all its components are zero.

    Let u → ∈ ℝ n have components u 1 , … , u n . Recall that the square of any real number is greater than or equal to zero, with equality only when that real is zero. Thus | u → | 2 = u 1 2 + ⋯ + u n 2 is a sum of numbers greater than or equal to zero, and so is itself greater than or equal to zero, with equality if and only if each u i is zero. Hence | u → | = 0 if and only if all the components of u → are zero.

  12. Exercise 2.22 Worked answer

    Recommended. Find the midpoint of the line segment connecting ( x 1 , y 1 ) with ( x 2 , y 2 ) in ℝ 2 . Generalize to ℝ n .

    Back to Exercise 2.22

    Answer. We can easily check that

    ( x 1 + x 2 2 , y 1 + y 2 2 )

    is on the line connecting the two, and is equidistant from both. The generalization is obvious.

  13. Exercise 2.23 Worked answer

    Show that if v → ≠ 0 → then v → / | v → | has length one. What if v → = 0 → ?

    Back to Exercise 2.23

    Answer. Assume that v → ∈ ℝ n has components v 1 , … , v n . If v → ≠ 0 → then we have this.

    ( v 1 v 1 2 + ⋯ + v n 2 ) 2 + ⋯ + ( v n v 1 2 + ⋯ + v n 2 ) 2 = ( v 1 2 v 1 2 + ⋯ + v n 2 ) + ⋯ + ( v n 2 v 1 2 + ⋯ + v n 2 ) = 1

    If v → = 0 → then v → / | v → | is not defined.

  14. Exercise 2.24 Worked answer

    Show that if r ≥ 0 then r v → is r times as long as v → . What if r < 0 ?

    Back to Exercise 2.24

    Answer. For the first question, assume that v → ∈ ℝ n and r ≥ 0 , take the root, and factor.

    | r v → | = ( r v 1 ) 2 + ⋯ + ( r v n ) 2 = r 2 ( v 1 2 + ⋯ + v n 2 = r | v → |

    For the second question, the result is r times as long, but it points in the opposite direction in that r v → + ( − r ) v → = 0 → .

  15. Exercise 2.25 Worked answer

    Recommended. A vector v → ∈ ℝ n of length one is a unit vector. Show that the dot product of two unit vectors has absolute value less than or equal to one. Can ‘less than’ happen? Can ‘equal to’?

    Back to Exercise 2.25

    Answer. Assume that u → , v → ∈ ℝ n both have length 1 . Apply Cauchy-Schwarz: | u → ⋅ v → | ≤ | u → | | v → | = 1 .

    To see that ‘less than’ can happen, in ℝ 2 take

    u → = ( 1 0 ) v → = ( 0 1 )

    and note that u → ⋅ v → = 0 . For ‘equal to’, note that u → ⋅ u → = 1 .

  16. Exercise 2.26 Worked answer

    When a plane does not pass through the origin, performing operations on vectors whose bodies lie in it is more complicated than when the plane does pass through the origin. Consider the picture in this subsection of the plane

    { ( 2 0 0 ) + ( − 0.5 1 0 ) y + ( − 0.5 0 1 ) z ∣ y , z ∈ ℝ }

    and the three vectors with endpoints ( 2 , 0 , 0 ) , ( 1.5 , 1 , 0 ) , and ( 1.5 , 0 , 1 ) .

    1. Redraw the picture, including the vector starting at ( 2 , 0 , 0 ) whose body is in the plane, and that is twice as long as the vector shown in the plane whose endpoint is ( 1.5 , 1 , 0 ) . The endpoint of this vector is not ( 3 , 2 , 0 ) ; what is it?

    2. Redraw the picture, including the parallelogram in the plane that shows the sum of the vectors ending at ( 1.5 , 0 , 1 ) and ( 1.5 , 1 , 0 ) . The endpoint of the sum, on the diagonal, is not ( 3 , 1 , 1 ) ; what is it?

    Back to Exercise 2.26

    Answer.

    1. The vector shown

      A vector is drawn within a plane in three-dimensional space, with the coordinate axes shown for reference.

      is not the result of doubling

      ( 2 0 0 ) + ( − 0.5 1 0 ) ⋅ 1 = ( 1.5 1 0 )

      instead it is the result of doubling the parameter.

      ( 2 0 0 ) + ( − 0.5 1 0 ) ⋅ 2 = ( 1 2 0 )

      This compares the lengths.

      1 2 + 2 2 + 0 2 = ( 2 ⋅ 0.5 ) 2 + ( 2 ⋅ 1 ) 2 + ( 2 ⋅ 0 ) 2 = 2 2 ⋅ ( 0.5 2 + 1 2 + 0 2 ) = 2 ⋅ 0.5 2 + 1 2 + 0 2

    2. The vector

      The plane 2x+y+z=4 is described by a base point and two independent direction vectors.

      is not the result of adding

      ( ( 2 0 0 ) + ( − 0.5 1 0 ) ⋅ 1 ) + ( ( 2 0 0 ) + ( − 0.5 0 1 ) ⋅ 1 )

      instead it is

      ( 2 0 0 ) + ( − 0.5 1 0 ) ⋅ 1 + ( − 0.5 0 1 ) ⋅ 1 = ( 1 1 1 )

      which adds the parameters.

  17. Exercise 2.27 Worked answer

    Show that the line segments ( a 1 , a 2 ) ( b 1 , b 2 ) ― and ( c 1 , c 2 ) ( d 1 , d 2 ) ― have the same lengths and slopes if b 1 − a 1 = d 1 − c 1 and b 2 − a 2 = d 2 − c 2 . Is that only if?

    Back to Exercise 2.27

    Answer. The “if” half is straightforward. If b 1 − a 1 = d 1 − c 1 and b 2 − a 2 = d 2 − c 2 then

    ( b 1 − a 1 ) 2 + ( b 2 − a 2 ) 2 = ( d 1 − c 1 ) 2 + ( d 2 − c 2 ) 2

    so they have the same lengths, and the slopes are just as easy:

    b 2 − a 2 b 1 − a 1 = d 2 − c 2 d 1 − a 1

    (if the denominators are 0 they both have undefined slopes).

    For “only if”, assume that the two segments have the same length and slope (the case of undefined slopes is easy; we will do the case where both segments have a slope m ). Also assume, without loss of generality, that a 1 < b 1 and that c 1 < d 1 . The first segment is ( a 1 , a 2 ) ( b 1 , b 2 ) ― = { ( x , y ) ∣ y = m x + n 1 , x ∈ [ a 1 . . b 1 ] } (for some intercept n 1 ) and the second segment is ( c 1 , c 2 ) ( d 1 , d 2 ) ― = { ( x , y ) ∣ y = m x + n 2 , x ∈ [ c 1 . . d 1 ] } (for some n 2 ). Then the lengths of those segments are

    ( b 1 − a 1 ) 2 + ( ( m b 1 + n 1 ) − ( m a 1 + n 1 ) ) 2 = ( 1 + m 2 ) ( b 1 − a 1 ) 2

    and, similarly, ( 1 + m 2 ) ( d 1 − c 1 ) 2 . Therefore, | b 1 − a 1 | = | d 1 − c 1 | . Thus, as we assumed that a 1 < b 1 and c 1 < d 1 , we have that b 1 − a 1 = d 1 − c 1 .

    The other equality is similar.

  18. Exercise 2.28 Worked answer

    Is | u → 1 + ⋯ + u → n | ≤ | u → 1 | + ⋯ + | u → n | ? If it is true then it would generalize the Triangle Inequality.

    Back to Exercise 2.28

    Answer. Yes; we can prove this by induction.

    Assume that the vectors are in some ℝ k . Clearly the statement applies to one vector. The Triangle Inequality is this statement applied to two vectors. For an inductive step assume the statement is true for n or fewer vectors. Then this

    | u → 1 + ⋯ + u → n + u → n + 1 | ≤ | u → 1 + ⋯ + u → n | + | u → n + 1 |

    follows by the Triangle Inequality for two vectors. Now the inductive hypothesis, applied to the first summand on the right, gives that as less than or equal to | u → 1 | + ⋯ + | u → n | + | u → n + 1 | .

  19. Exercise 2.29 Worked answer

    What is the ratio between the sides in the Cauchy-Schwarz inequality?

    Back to Exercise 2.29

    Answer. By definition

    u → ⋅ v → | u → | | v → | = cos ⁡ θ

    where θ is the angle between the vectors. Thus the ratio is | cos ⁡ θ | .

  20. Exercise 2.30 Worked answer

    Why is the zero vector defined to be perpendicular to every vector?

    Back to Exercise 2.30

    Answer. So that the statement ‘vectors are orthogonal iff their dot product is zero’ has no exceptions.

  21. Exercise 2.31 Worked answer

    Describe the angle between two vectors in ℝ 1 .

    Back to Exercise 2.31

    Answer. We can find the angle between ( a ) and ( b ) (for a , b ≠ 0 ) with

    arccos ⁡ ( a b a 2 b 2 ) .

    If a or b is zero then the angle is π / 2 radians. Otherwise, if a and b are of opposite signs then the angle is π radians, else the angle is zero radians.

  22. Exercise 2.32 Worked answer

    Give a simple necessary and sufficient condition to determine whether the angle between two vectors is acute, right, or obtuse.

    Back to Exercise 2.32

    Answer. The angle between u → and v → is acute if u → ⋅ v → > 0 , is right if u → ⋅ v → = 0 , and is obtuse if u → ⋅ v → < 0 . That’s because, in the formula for the angle, the denominator is never negative.

  23. Exercise 2.33 Worked answer

    Generalize to ℝ n the converse of the Pythagorean Theorem, that if u → and v → are perpendicular then | u → + v → | 2 = | u → | 2 + | v → | 2 .

    Back to Exercise 2.33

    Answer. Suppose that u → , v → ∈ ℝ n . If u → and v → are perpendicular then

    | u → + v → | 2 = ( u → + v → ) ⋅ ( u → + v → ) = u → ⋅ u → + 2 u → ⋅ v → + v → ⋅ v → = u → ⋅ u → + v → ⋅ v → = | u → | 2 + | v → | 2

    (the third equality holds because u → ⋅ v → = 0 ).

  24. Exercise 2.34 Worked answer

    Show that | u → | = | v → | if and only if u → + v → and u → − v → are perpendicular. Give an example in ℝ 2 .

    Back to Exercise 2.34

    Answer. Where u → , v → ∈ ℝ n , the vectors u → + v → and u → − v → are perpendicular if and only if 0 = ( u → + v → ) ⋅ ( u → − v → ) = u → ⋅ u → − v → ⋅ v → , which shows that those two are perpendicular if and only if u → ⋅ u → = v → ⋅ v → . That holds if and only if | u → | = | v → | .

  25. Exercise 2.35 Worked answer

    Show that if a vector is perpendicular to each of two others then it is perpendicular to each vector in the plane they generate. (Remark. They could generate a degenerate plane—a line or a point—but the statement remains true.)

    Back to Exercise 2.35

    Answer. Suppose u → ∈ ℝ n is perpendicular to both v → ∈ ℝ n and w → ∈ ℝ n . Then, for any k , m ∈ ℝ we have this.

    u → ⋅ ( k v → + m w → ) = k ( u → ⋅ v → ) + m ( u → ⋅ w → ) = k ( 0 ) + m ( 0 ) = 0

  26. Exercise 2.36 Worked answer

    Prove that, where u → , v → ∈ ℝ n are nonzero vectors, the vector

    u → | u → | + v → | v → |

    bisects the angle between them. Illustrate in ℝ 2 .

    Back to Exercise 2.36

    Answer. We will show something more general: if | z → 1 | = | z → 2 | for z → 1 , z → 2 ∈ ℝ n , then z → 1 + z → 2 bisects the angle between z → 1 and z → 2

    Two equal-length vectors form a parallelogram. Its diagonal bisects their angle; matching marks identify equal segments in the construction.

    (we ignore the case where z → 1 and z → 2 are the zero vector).

    The z → 1 + z → 2 = 0 → case is easy. For the rest, by the definition of angle, we will be finished if we show this.

    z → 1 ⋅ ( z → 1 + z → 2 ) | z → 1 | | z → 1 + z → 2 | = z → 2 ⋅ ( z → 1 + z → 2 ) | z → 2 | | z → 1 + z → 2 |

    But distributing inside each expression gives

    z → 1 ⋅ z → 1 + z → 1 ⋅ z → 2 | z → 1 | | z → 1 + z → 2 | z → 2 ⋅ z → 1 + z → 2 ⋅ z → 2 | z → 2 | | z → 1 + z → 2 |

    and z → 1 ⋅ z → 1 = | z → 1 | 2 = | z → 2 | 2 = z → 2 ⋅ z → 2 , so the two are equal.

  27. Exercise 2.37 Worked answer

    Verify that the definition of angle is dimensionally correct: (1) if k > 0 then the cosine of the angle between k u → and v → equals the cosine of the angle between u → and v → , and (2) if k < 0 then the cosine of the angle between k u → and v → is the negative of the cosine of the angle between u → and v → .

    Back to Exercise 2.37

    Answer. We can show the two statements together. Let u → , v → ∈ ℝ n , write

    u → = ( u 1 ⋮ u 1 u n ) v → = ( v 1 ⋮ v 1 v n )

    and calculate.

    cos ⁡ θ = k u 1 v 1 + ⋯ + k u n v n ( k u 1 ) 2 + ⋯ + ( k u n ) 2 b 1 2 + ⋯ + b n 2 = k | k | u → ⋅ v → | u → | | v → | = ± u → ⋅ v → | u → | | v → |

  28. Exercise 2.38 Worked answer

    Recommended. Show that the inner product operation is linear: for u → , v → , w → ∈ ℝ n and k , m ∈ ℝ , u → ⋅ ( k v → + m w → ) = k ( u → ⋅ v → ) + m ( u → ⋅ w → ) .

    Back to Exercise 2.38

    Answer. Let

    u → = ( u 1 ⋮ u 1 u n ) , v → = ( v 1 ⋮ v 1 v n ) w → = ( w 1 ⋮ w 1 w n )

    and then

    u → ⋅ ( k v → + m w → ) = ( u 1 ⋮ u 1 u n ) ⋅ ( ( k v 1 ⋮ k v 1 k v n ) + ( m w 1 ⋮ m w 1 m w n ) ) = ( u 1 ⋮ u 1 u n ) ⋅ ( k v 1 + m w 1 ⋮ k v 1 + m w 1 k v n + m w n ) = u 1 ( k v 1 + m w 1 ) + ⋯ + u n ( k v n + m w n ) = k u 1 v 1 + m u 1 w 1 + ⋯ + k u n v n + m u n w n = ( k u 1 v 1 + ⋯ + k u n v n ) + ( m u 1 w 1 + ⋯ + m u n w n ) = k ( u → ⋅ v → ) + m ( u → ⋅ w → )

    as required.

  29. Exercise 2.39 Worked answer

    Puzzle. [Cleary] Astrologers claim to be able to recognize trends in personality and fortune that depend on an individual’s birthday by incorporating where the stars were 2000  years ago. Suppose that instead of star-gazers coming up with stuff, math teachers who like linear algebra (we’ll call them vectologers) had come up with a similar system as follows: Consider your birthday as a row vector ( month day ) . For instance, I was born on July  12 so my vector would be ( 7 12 ) . Vectologers have made the rule that how well individuals get along with each other depends on the angle between vectors. The smaller the angle, the more harmonious the relationship.

    1. Find the angle between your vector and mine, in radians.

    2. Would you get along better with me, or with a professor born on September  19 ?

    3. For maximum harmony in a relationship, when should the other person be born?

    4. Is there a person with whom you have a “worst case” relationship, i.e., your vector and theirs are orthogonal? If so, what are the birthdate(s) for such people? If not, explain why not.

    Back to Exercise 2.39

    Answer.

    1. For instance, a birthday of October  12 gives this.

      θ = arccos ⁡ ( ( 7 12 ) ⋅ ( 10 12 ) | ( 7 12 ) | ⋅ | ( 10 12 ) | ) = arccos ⁡ ( 214 244 193 ) ≈ 0.17  rad

    2. Applying the same equation to ( 9 19 ) gives about 0.09  radians.

    3. The angle will measure 0  radians if the other person is born on the same day. It will also measure 0 if one birthday is a scalar multiple of the other. For instance, a person born on Mar  6 would be harmonious with a person born on Feb  4 .

      Given a birthday, we can get Sage to plot the angle for other dates. This example shows the relationship of all dates with July 12.

        sage: plot3d(lambda x, y: math.acos((x*7+y*12)/(math.sqrt(7**2+12**2)*math.sqrt(x**2+y**2))),
                                                                       (1,12),(1,31))

      A three-dimensional plot shows the angle between the July 12 birthday vector and other month-day vectors.

    4. We want to maximize this.

      θ = arccos ⁡ ( ( 7 12 ) ⋅ ( m d ) | ( 7 12 ) | ⋅ | ( m d ) | )

      Of course, we cannot take m or d negative and so we cannot get a vector orthogonal to the given one. This Python script finds the largest angle by brute force.

        import math
        days={1:31,  # Jan
              2:29, 3:31, 4:30, 5:31, 6:30, 7:31, 8:31, 9:30, 10:31, 11:30, 12:31}
        BDAY=(7,12)
        max_res=0
        max_res_date=(-1,-1)
        for month in range(1,13):
            for day in range(1,days[month]+1):
                num=BDAY[0]*month+BDAY[1]*day
                denom=math.sqrt(BDAY[0]**2+BDAY[1]**2)*math.sqrt(month**2+day**2)
                if denom>0:
                    res=math.acos(min(num*1.0/denom,1))
                    print "day:",str(month),str(day)," angle:",str(res)
                    if res>max_res:
                        max_res=res
                        max_res_date=(month,day)
        print "For ",str(BDAY),"worst case",str(max_res),"rads, date",str(max_res_date)
        print "  That is ",180*max_res/math.pi,"degrees"

      The result is

        For  (7, 12) worst case 0.95958064648 rads, date (12, 1)
          That is  54.9799211457 degrees

      A more conceptual approach is to consider the relation of all points ( month , day ) to the point ( 7 , 12 ) . The picture below makes clear that the answer is either Dec  1 or Jan  31 , depending on which is further from the birthdate. The dashed line bisects the angle between the line from the origin to Dec  1 , and the line from the origin to Jan  31 . Birthdays above the line are furthest from Dec  1 and birthdays below the line are furthest from Jan  31 .

      The month-day plane is split by the angle bisector between January 31 and December 1; its sides identify the farther birthday direction.

  30. Exercise 2.40 Worked answer

    Puzzle. [Am. Math. Mon., Feb. 1933] A ship is sailing with speed and direction v → 1 ; the wind blows apparently (judging by the vane on the mast) in the direction of a vector a → ; on changing the direction and speed of the ship from v → 1 to v → 2 the apparent wind is in the direction of a vector b → .

    Find the vector velocity of the wind.

    Back to Exercise 2.40

    Answer. This is how the answer was given in the cited source. The actual velocity v → of the wind is the sum of the ship’s velocity and the apparent velocity of the wind. Without loss of generality we may assume a → and b → to be unit vectors, and may write

    v → = v → 1 + s a → = v → 2 + t b →

    where s and t are undetermined scalars. Take the dot product first by a → and then by b → to obtain

    s − t a → ⋅ b → = a → ⋅ ( v → 2 − v → 1 ) s a → ⋅ b → − t = b → ⋅ ( v → 2 − v → 1 )

    Multiply the second by a → ⋅ b → , subtract the result from the first, and find

    s = [ a → − ( a → ⋅ b → ) b → ] ⋅ ( v → 2 − v → 1 ) 1 − ( a → ⋅ b → ) 2 .

    Substituting in the original displayed equation, we get

    v → = v → 1 + [ a → − ( a → ⋅ b → ) b → ] ⋅ ( v → 2 − v → 1 ) a → 1 − ( a → ⋅ b → ) 2 .

  31. Exercise 2.41 Worked answer

    Verify the Cauchy-Schwarz inequality by first proving Lagrange’s identity:

    ( ∑ 1 ≤ j ≤ n a j b j ) 2 = ( ∑ 1 ≤ j ≤ n a j 2 ) ( ∑ 1 ≤ j ≤ n b j 2 ) − ∑ 1 ≤ k < j ≤ n ( a k b j − a j b k ) 2

    and then noting that the final term is positive. This result is an improvement over Cauchy-Schwarz because it gives a formula for the difference between the two sides. Interpret that difference in ℝ 2 .

    Back to Exercise 2.41

    Answer. We use induction on n .

    In the n = 1 base case the identity reduces to

    ( a 1 b 1 ) 2 = ( a 1 2 ) ( b 1 2 ) − 0

    and clearly holds.

    For the inductive step assume that the formula holds for the 0 , …, n cases. We will show that it then holds in the n + 1 case. Start with the right-hand side

    ( ∑ 1 ≤ j ≤ n + 1 a j 2 ) ( ∑ 1 ≤ j ≤ n + 1 b j 2 ) − ∑ 1 ≤ k < j ≤ n + 1 ( a k b j − a j b k ) 2 = [ ( ∑ 1 ≤ j ≤ n a j 2 ) + a n + 1 2 ] [ ( ∑ 1 ≤ j ≤ n b j 2 ) + b n + 1 2 ] − [ ∑ 1 ≤ k < j ≤ n ( a k b j − a j b k ) 2 + ∑ 1 ≤ k ≤ n ( a k b n + 1 − a n + 1 b k ) 2 ] = ( ∑ 1 ≤ j ≤ n a j 2 ) ( ∑ 1 ≤ j ≤ n b j 2 ) + ∑ 1 ≤ j ≤ n b j 2 a n + 1 2 + ∑ 1 ≤ j ≤ n a j 2 b n + 1 2 + a n + 1 2 b n + 1 2 − [ ∑ 1 ≤ k < j ≤ n ( a k b j − a j b k ) 2 + ∑ 1 ≤ k ≤ n ( a k b n + 1 − a n + 1 b k ) 2 ] = ( ∑ 1 ≤ j ≤ n a j 2 ) ( ∑ 1 ≤ j ≤ n b j 2 ) − ∑ 1 ≤ k < j ≤ n ( a k b j − a j b k ) 2 + ∑ 1 ≤ j ≤ n b j 2 a n + 1 2 + ∑ 1 ≤ j ≤ n a j 2 b n + 1 2 + a n + 1 2 b n + 1 2 − ∑ 1 ≤ k ≤ n ( a k b n + 1 − a n + 1 b k ) 2

    and apply the inductive hypothesis.

    = ( ∑ 1 ≤ j ≤ n a j b j ) 2 + ∑ 1 ≤ j ≤ n b j 2 a n + 1 2 + ∑ 1 ≤ j ≤ n a j 2 b n + 1 2 + a n + 1 2 b n + 1 2 − [ ∑ 1 ≤ k ≤ n a k 2 b n + 1 2 − 2 ∑ 1 ≤ k ≤ n a k b n + 1 a n + 1 b k + ∑ 1 ≤ k ≤ n a n + 1 2 b k 2 ] = ( ∑ 1 ≤ j ≤ n a j b j ) 2 + 2 ( ∑ 1 ≤ k ≤ n a k b n + 1 a n + 1 b k ) + a n + 1 2 b n + 1 2 = [ ( ∑ 1 ≤ j ≤ n a j b j ) + a n + 1 b n + 1 ] 2

    to derive the left-hand side.

References cited in this section

Math. Mag., Jan. 1957

M. S. Klamkin (proposer), Trickie T-27, Mathematics Magazine, volume 30 number 3 (Jan-Feb. 1957), p. 173.

Heath

T. Heath, Euclid’s Elements, volume 1, Dover, 1956.

Polya

G. Polya, Mathematics and Plausible Reasoning, Princeton University Press, 1954.

Ohanian

Hans O’Hanian, Physics, volume one, W. W. Norton, 1985.

Cleary

R. Cleary, private communication, Nov. 2011.

Am. Math. Mon., Feb. 1933

V. F. Ivanoff (proposer), T. C. Esty (solver), problem 3529, American Mathematical Monthly, vol. 39 no. 2 (Feb. 1933), p. 118.