Original English by Jim Hefferon — 34 validated sections. The original mathematics and supplied answers below are preserved. This is a partial-book reading edition, not the complete book or an Everyday-English rewrite.

Source, reuse and conversion details

Source revision df2262e089a02651c127f1dd12649c4622ee1383; CC BY-SA 2.5 option, with original component credits retained. This is not an Everyday-English rewrite. The complete active source topic is included. Commented alternatives remain in the editable source only. This topic requires the formulas from Orthogonal Projection Into a Line and Projection Into a Subspace; retain the preceding Projection reader for offline use.

AI-assisted conversion, not human-reviewed. The original historical computer transcripts are preserved as text; they have not been executed or certified for current software.

Prerequisite: Projection — original English local reader

Notes about the original source and supplied answers

These nine bounded findings are separate from the unchanged original text, formulas and diagrams. Opening them can reveal answers. They are not an exhaustive correctness audit or human review. Historical classroom computations are not operational safety guidance.

  1. Source note 1: The least-squares fit minimizes the sum of squared vertical residuals, not their total absolute length. In the printed coin data, m=79/140 gives squared cost 13/14 and absolute cost 9/7; m=17/30 gives squared cost 1 but the smaller absolute cost 1. The original wording and diagram are preserved.
  2. Source note 2: The original sentence repeats “the” before “coefficients”. The unchanged original text is retained.
  3. Source note 3: In supplied answer 7(a), the exact intercept 4259/1398 is correct but its printed decimal 3.239628 is not: it is approximately 3.046495. The seven-flight fit predicts 6317/2796, approximately 2.259299, failures at 31 degrees F (not 2.45); its four-failure threshold is -2666/71, approximately -37.549296 degrees F (not -29.94). These are mathematical checks of the historical classroom data, not safety guidance. Part (b) gives the correct all-flight coefficients; 2.79 is correctly rounded and the printed threshold 11 is a coarse rounding of 810/73, approximately 11.095890. Original formulas and answers remain unchanged.
  4. Source note 4: The final supplied answer states orthogonality but cites the Triangle Inequality for a conclusion that follows from Pythagoras: squared distance to w equals squared distance to the projection plus squared distance from the projection to w. The Triangle Inequality alone does not supply that lower bound. The original “hypoteneuse” spelling and reasoning are preserved.
  5. Source note 5: The source describes the asteroid belt as the remains of a planet that broke apart. NASA describes remnants from solar-system formation whose growth into planetary bodies was interrupted by Jupiter. Retain the exercise as a historical distance-fitting example, not as a current account of asteroid-belt formation; its original text is unchanged. Asteroid Facts — NASA.
  6. Source note 6: The source says this distance-fitting method helped discover Neptune. NASA credits the prediction to calculations of perturbations in Uranus’s orbit and the subsequent telescope search using Le Verrier’s position. Those dynamical calculations should not be conflated with this ordinal-distance least-squares exercise. The original historical assertion remains unchanged. 175 Years Ago: Astronomers Discover Neptune, the Eighth Planet — NASA.
  7. Source note 7: Question 8(d) requests Neptune’s predicted location, 8(e) requests Pluto’s, and 8(f) asks about accuracy for both. Supplied answer (d) gives the fitting coefficients only; answer (e) evaluates the Neptune index x=9 and compares its distance, and answer (f) evaluates Pluto at x=10 and compares its distance. The computations are present collectively but the last three answer labels do not match the corresponding tasks. Preserve the original labels and disclose this crosswalk separately.
  8. Source note 8: The original two-column matrix in supplied answer 3 has a middle row with two successive vdots commands but no column separator. Other rows contain two cells. The original math is preserved rather than silently inserting an ampersand; this is a source-level layout defect, not a conversion-introduced value change.
  9. Source note 9: The bridge-data description says “if a crossings” rather than the singular “if a crossing”. Its units and original wording are preserved.

Wide formulas, tables and historical transcripts scroll horizontally. Focus a region and use the arrow keys. Some original exercises require Internet data; no offline completion of those tasks is claimed.

Source-preserving rebuild, navigation, source packaging and current deterministic checks: OpenAI Codex — GPT-6 Astra, Ultra effort. Jim Hefferon remains the author of the mathematics. Earlier intermediate-conversion runtime identity is not established by its retained receipts and is not reassigned to this rebuild. No human review or exhaustive proof certification is claimed.

Line of Best Fit

This Topic requires the formulas from the subsections on Orthogonal Projection Into a Line and Projection Into a Subspace.

Scientists are often presented with a system that has no solution and they must find an answer anyway. More precisely, they must find a best answer. For instance, this is the result of flipping a penny, including some intermediate numbers.

number of flips 30 60 90
number of heads 16 34 51

Because of the randomness in this experiment we expect that the ratio of heads to flips will fluctuate around a penny’s long-term ratio of 50-50. So the system for such an experiment likely has no solution, and that’s what happened here.

30 m = 16 60 m = 34 90 m = 51

That is, the vector of data that we collected is not in the subspace where ideally it would be.

( 16 34 51 ) ∉ { m ( 30 60 90 ) ∣ m ∈ ℝ }

However, we have to do something so we look for the  m that most nearly works. An orthogonal projection of the data vector into the line subspace gives a best guess, the vector in the subspace closest to the data vector.

( 16 34 51 ) ⋅ ( 30 60 90 ) ( 30 60 90 ) ⋅ ( 30 60 90 ) ⋅ ( 30 60 90 ) = 7110 12600 ⋅ ( 30 60 90 )

The estimate ( m = 7110 / 12600 ≈ 0.56 ) is a bit more than one half, but not much more than half, so probably the penny is fair enough.

The line with the slope m ≈ 0.56 is the line of best fit for this data.

Three coin-flip data points at (30,16), (60,34) and (90,51), with labelled flips and heads axes. A dashed line through the origin has slope 79/140, the least-squares fit constrained to pass through the origin.

Minimizing the distance between the given vector and the vector used as the right-hand side minimizes the total of these vertical lengths, and consequently we say that the line comes from fitting by least-squares.

The same coin-fit line, with the original residual illustration exaggerated vertically by a factor of ten. Paired horizontal guide segments indicate the point and fitted-value levels; no vertical connecting bars are drawn.

This diagram exaggerates the vertical scale by a factor of ten to make the lengths more visible.

In the above equation the line must pass through ( 0 , 0 ) , because we take it to be the line whose slope is this coin’s true proportion of heads to flips. We can also handle cases where the line need not pass through the origin.

Here is the progression of world record times for the men’s mile race [Oakley & Baker]. In the early 1900’s many people wondered when, or if, this record would fall below the four minute mark. Here are the times that were in force on January first of each decade through the first half of that century.

year 1870 1880 1890 1900 1910 1920 1930 1940 1950
secs 268.8 264.5 258.4 255.6 255.6 252.6 250.4 246.4 241.4

We can use this to give a circa 1950 prediction of the date for 240 seconds, and then compare that to the actual date. As with the penny data, these numbers do not lie in a perfect line. That is, this system does not have an exact solution for the slope and intercept.

b + 1870 m = 268.8 b + 1880 m = 264.5 ⋮ b + 1950 m = 241.4

We find a best approximation by using orthogonal projection.

(Comments on the data. Restricting to the times at the start of each decade reduces the data entry burden, smooths the data to some extent, and gives much the same result as entering all of the dates and records. There are different sequences of times from competing standards bodies but the ones here are from [Wikipedia, Mens Mile]. We’ve started the plot at 1870 because at one point there were two classes of records, called ‘professional’ and ‘amateur’, and after a while the first class stopped being active so we’ve followed the second class.)

Write the linear system’s matrix of coefficients and also its vector of constants, the world record times.

A = ( 1 1870 1 1880 ⋮ ⋮ 1 1950 ) v → = ( 268.8 264.5 ⋮ 241.4 )

The ending result in the subsection on Projection into a Subspace gives the formula for the the coefficients b and m that make the linear combination of A ’s columns as close as possible to v → . Those coefficients are the entries of the vector ( A 𝖳 A ) − 1 A 𝖳 ⋅ v → .

Sage can do the computation for us.

sage: year = [1870, 1880, 1890, 1900, 1910, 1920, 1930, 1940, 1950]
sage: secs = [268.8, 264.5, 258.4, 255.6, 255.6, 252.6, 250.4, 246.4, 241.4]
sage: var('a, b, t')
(a, b, t)
sage: model(t) = a*t+b
sage: data = zip(year, secs)
sage: fit = find_fit(data, model, solution_dict=True)
sage: model.subs(fit)
t |--> -0.3048333333333295*t + 837.0872222222147
sage: g=points(data)+plot(model.subs(fit),(t,1860,1960),color='red',
....:                     figsize=3,fontsize=7,typeset='latex')
sage: g.save("four_minute_mile.pdf")
sage: g

Historical men’s mile-record data: year is horizontal, seconds vertical. Nine blue points and a descending red best-fit line appear. The original table gives the data values.

The progression makes a surprisingly good line. From the slope and intercept we predict 1958.73 ; the actual date of Roger Bannister’s record was 1954-May-06.

The final example compares team salaries from US major league baseball against the number of wins the team had, for the year 2002. In this year the Oakland Athletics used mathematical techniques to optimize the players that they fielded for the money that they could spend, as told in the film Moneyball. (Salaries are in millions of dollars and the number of wins is out of 162 games).

To do the computations we again use Sage.

sage: sal = [40, 40, 39, 42, 45, 42, 62, 34, 41, 57, 58, 63, 47, 75, 57, 78, 80, 50, 60, 93,
....:        77, 55, 95, 103, 79, 76, 108, 126, 95, 106]
sage: wins = [103, 94, 83, 79, 78, 72, 99, 55, 66, 81, 80, 84, 62, 97, 73, 95, 93, 56, 67,
....:        101, 78, 55, 92, 98, 74, 67, 93, 103, 75, 72]
sage: var('a, b, t')
(a, b, t)
sage: model(t) = a*t+b
sage: data = zip(sal,wins)
sage: fit = find_fit(data, model, solution_dict=True)
sage: model.subs(fit)
t |--> 0.2634981251436269*t + 63.06477642781477
sage: p = points(data,size=25)+plot(model.subs(fit),(t,30,130),color='red',typeset='latex')
sage: p.save('moneyball.pdf')

The graph is below. The team in the upper left, who paid little for many wins, is the Oakland A’s.

Historical baseball salary-versus-win scatter: salaries are horizontal and wins vertical. Blue points and an ascending red fitted line appear, including a low-salary, high-win outlier.

Judging this line by eye would be error-prone. So the equations give us a certainty about the ‘best’ in best fit. In addition, the model’s equation tells us roughly that by spending an additional million dollars a team owner can expect to buy 1 / 4 of a win (and that expectation is not very sure, thank goodness).

Exercises

The calculations here are best done on a computer. Some of the problems require data from the Internet.

  1. Exercise 1 Supplied answer

    Use least-squares to judge if the coin in this experiment is fair.

    flips 8 16 24 32 40
    heads 4 9 13 17 20

    Back to Exercise 1

    Answer. As with the first example discussed above, we are trying to find a best m to “solve” this system.

    8 m = 4 16 m = 9 24 m = 13 32 m = 17 40 m = 20

    Projecting into the linear subspace gives this

    ( 4 9 13 17 20 ) ⋅ ( 8 16 24 32 40 ) ( 8 16 24 32 40 ) ⋅ ( 8 16 24 32 40 ) ⋅ ( 8 16 24 32 40 ) = 1832 3520 ⋅ ( 8 16 24 32 40 )

    so the slope of the line of best fit is approximately 0.52 .

    Supplied-answer-only coin-data scatter. Positive points are (8,4), (16,9), (24,13), (32,17), and (40,20). Zero-based axes retain their original ticks.

  2. Exercise 2 Supplied answer

    For the men’s mile record, rather than give each of the many records and its exact date, we’ve “smoothed” the data somewhat by taking a periodic sample. Do the longer calculation and compare the conclusions.

    Back to Exercise 2

    Answer. With this input

    A = ( 1 1852.71 1 1858.88 ⋮ ⋮ 1 1985.54 1 1993.71 ) b = ( 292.0 285.0 ⋮ 226.32 224.39 )

    (the dates have been rounded to months, e.g., for a September record, the decimal .71 ≈ ( 8.5 / 12 ) was used), Maple responded with an intercept of b = 994.8276974 and a slope of m = − 0.3871993827 .

    Supplied-answer-only historical men’s mile-record scatter: years are horizontal and seconds vertical. The points descend from nearly 300 seconds in the nineteenth century toward roughly 224 seconds late in the twentieth century.

  3. Exercise 3 Supplied answer

    Find the line of best fit for the men’s 1500  meter run. How does the slope compare with that for the men’s mile? (The distances are close; a mile is about 1609  meters.)

    Back to Exercise 3

    Answer. With this input (the years are zeroed at 1900 )

    A := ( 1 .38 1 .54 ⋮ ⋮ 1 92.71 1 95.54 ) b = ( 249.0 246.2 ⋮ 208.86 207.37 )

    (the dates have been rounded to months, e.g., for a September record, the decimal .71 ≈ ( 8.5 / 12 ) was used), Maple gives an intercept of b = 243.1590327 and a slope of m = − 0.401647703 . The slope given in the body of this Topic for the men’s mile is quite close to this.

    Supplied-answer-only historical men’s 1500-metre record scatter: years are horizontal and seconds vertical; the points descend from about 250 toward about 206 seconds.

  4. Exercise 4 Supplied answer

    Find the line of best fit for the records for women’s mile.

    Back to Exercise 4

    Answer. With this input (the years are zeroed at 1900 )

    A = ( 1 21.46 1 32.63 ⋮ ⋮ 1 89.54 1 96.63 ) b = ( 373.2 327.5 ⋮ 255.61 252.56 )

    (the dates have been rounded to months, e.g., for a September record, the decimal .71 ≈ ( 8.5 / 12 ) was used), MAPLE gave an intercept of b = 378.7114894 and a slope of m = − 1.445753225 .

    Supplied-answer-only historical women’s mile-record scatter: years are horizontal and seconds vertical; the plotted records descend from about 374 toward about 254 seconds.

  5. Exercise 5 Supplied answer

    Do the lines of best fit for the men’s and women’s miles cross?

    Back to Exercise 5

    Answer. These are the equations of the lines for men’s and women’s mile (the vertical intercept term of the equation for the women’s mile has been adjusted from the answer above, to zero it at the year 0 , because that’s how the men’s mile equation was done).

    y = 994.8276974 − 0.3871993827 x y = 3125.6426 − 1.445753225 x

    Obviously the lines cross. A computer program is the easiest way to do the arithmetic: MuPAD gives x = 2012.949004 and y = 215.4150856 ( 215  seconds is 3  minutes and 35  seconds). Remark. Of course all of this projection is highly dubious —for one thing, the equation for the women is influenced by the quite slow early times —but it is nonetheless fun.

    Supplied-answer-only comparison of historical men’s and women’s mile-record traces. Men use longer dashes, women shorter dashes. The original traces do not show an extrapolated intersection.

  6. Exercise 6 Supplied answer

    (This illustrates that there are data sets for which a linear model is not right, and that the line of best fit doesn’t in that case have any predictive value.) In a highway restaurant a trucker told me that his boss often sends him by a roundabout route, using more gas but paying lower bridge tolls. He said that New York State calibrates the toll for each bridge across the Hudson, playing off the extra gas to get there from New York City against a lower crossing cost, to encourage people to go upstate. This table, from [Cost Of Tolls] and [Google Maps], lists for each toll crossing of the Hudson River, the distance to drive from Times Square in miles and the cost in US dollars for a passenger car (if a crossings has a one-way toll then it shows half that number).

    Crossing Distance Toll
    Lincoln Tunnel
    Holland Tunnel
    George Washington Bridge
    Verrazano-Narrows Bridge
    Tappan Zee Bridge
    Bear Mountain Bridge
    Newburgh-Beacon Bridge
    Mid-Hudson Bridge
    Kingston-Rhinecliff Bridge
    Rip Van Winkle Bridge
    2
    7
    8
    16
    27
    47
    67
    82
    102
    120
    6.00
    6.00
    6.00
    6.50
    2.50
    1.00
    1.00
    1.00
    1.00
    1.00

    Find the line of best fit and graph the data to show that the driver was practicing on my credulity.

    Back to Exercise 6

    Answer. Sage gives the line of best fit as toll = − 0.05 ⋅ dist + 5.63 .

    sage: dist = [2, 7, 8, 16, 27, 47, 67, 82, 102, 120]
    sage: toll = [6, 6, 6, 6.5, 2.5, 1, 1, 1, 1, 1]
    sage: var('a,b,t')
    (a, b, t)
    sage: model(t) = a*t+b
    sage: data = zip(dist,toll)
    sage: fit = find_fit(data, model, solution_dict=True)
    sage: model.subs(fit)
    t |--> -0.0508568169130319*t + 5.630955848442933
    sage: p = plot(model.subs(fit), (t,0,120))+points(data,size=25,color='red')
    sage: p.save('bridges.pdf')        

    But the graph shows that the equation has little predictive value.

    Supplied-answer-only distance-versus-bridge-toll plot: distance north of New York City is horizontal, dollar toll vertical. Ten red points and a descending blue fitted line appear. The source table lists all crossings and values.

    Apparently a better model is that (with only one intermediate exception) crossings in the city cost roughly the same as each other, and crossings upstate cost the same as each other.

  7. Exercise 7 Supplied answer

    When the space shuttle Challenger exploded in 1986, one of the criticisms made of NASA’s decision to launch was in the way they did the analysis of number of O-ring failures versus temperature (O-ring failure caused the explosion). Four O-ring failures would be fatal. NASA had data from 24 previous flights.

    temp  ∘ F 53 75 57 58 63 70 70 66 67 67 67
    failures 3 2 1 1 1 1 1 0 0 0 0


    68 69 70 70 72 73 75 76 76 78 79 80 81
    0 0 0 0 0 0 0 0 0 0 0 0 0

    The temperature that day was forecast to be 31 ∘ F .

    1. NASA based the decision to launch partially on a chart showing only the flights that had at least one O-ring failure. Find the line that best fits these seven flights. On the basis of this data, predict the number of O-ring failures when the temperature is 31 , and when the number of failures will exceed four.

    2. Find the line that best fits all 24 flights. On the basis of this extra data, predict the number of O-ring failures when the temperature is 31 , and when the number of failures will exceed four.

    Which do you think is the more accurate method of predicting? (An excellent discussion is in [Dalal, et. al.].)

    Back to Exercise 7

    Answer.

    1. A computer algebra system like MAPLE or MuPAD will give an intercept of b = 4259 / 1398 ≈ 3.239628 and a slope of m = − 71 / 2796 ≈ − 0.025393419 Plugging x = 31 into the equation yields a predicted number of O-ring failures of y = 2.45 (rounded to two places). Plugging in y = 4 and solving gives a temperature of x = − 29.94 ∘ F.

    2. On the basis of this information

      A = ( 1 53 1 75 ⋮ 1 80 1 81 ) b = ( 3 2 ⋮ 0 0 )

      MAPLE gives the intercept b = 187 / 40 = 4.675 and the slope m = − 73 / 1200 ≈ − 0.060833 . Here, plugging x = 31 into the equation predicts y = 2.79 O-ring failures (rounded to two places). Plugging in y = 4  failures gives a temperature of x = 11 ∘ F.

      Supplied-answer-only plot of the historical classroom O-ring dataset. Horizontal values are launch temperature in degrees Fahrenheit; vertical values are number of failures. The zero-failure flights remain on the baseline alongside all positive-failure flights.

  8. Exercise 8 Supplied answer

    This table lists the average distance from the sun to each of the first seven planets, using Earth’s average as a unit.

    Mercury Venus Earth Mars Jupiter Saturn Uranus
    0.39 0.72 1.00 1.52 5.20 9.54 19.2

    1. Plot the number of the planet (Mercury is 1 , etc.) versus the distance. Note that it does not look like a line, and so finding the line of best fit is not fruitful.

    2. It does, however look like an exponential curve. Therefore, plot the number of the planet versus the logarithm of the distance. Does this look like a line?

    3. The asteroid belt between Mars and Jupiter is what is left of a planet that broke apart. Renumber so that Jupiter is 6 , Saturn is 7 , and Uranus is 8 , and plot against the log again. Does this look better?

    4. Use least squares on that data to predict the location of Neptune.

    5. Repeat to predict where Pluto is.

    6. Is the formula accurate for Neptune and Pluto?

    This method was used to help discover Neptune (although the second item is misleading about the history; actually, the discovery of Neptune in position  9 prompted people to look for the “missing planet” in position  5 ). See [Gardner, 1970]

    Back to Exercise 8

    Answer.

    1. The plot is nonlinear.

      Supplied-answer-only plot of seven historical planet-index and solar-distance values. Planet index is horizontal; the distance values rise rapidly on the vertical axis.

    2. Here is the plot.

      Supplied-answer-only plot of the base-ten logarithms of the same seven solar-distance values, retaining the negative logarithms for the nearest planets and the zero value.

      There is perhaps a jog up between planet  4 and planet  5 .

    3. This plot seems even more linear.

      Supplied-answer-only logarithmic solar-distance plot with the original eighth data position included; the historical planet numbering and distances remain unchanged.

    4. With this input

      A = ( 1 1 1 2 1 3 1 4 1 6 1 7 1 8 ) b = ( − 0.40893539 − 0.1426675 0 0.18184359 0.71600334 0.97954837 1.2833012 )

      MuPAD gives that the intercept is b = − 0.6780677466 and the slope is m = 0.2372763818 .

    5. Plugging x = 9 into the equation y = − 0.6780677466 + 0.2372763818 x from the prior item gives that the log of the distance is 1.4574197 , so the expected distance is 28.669472 . The actual distance is about 30.003 .

    6. Plugging x = 10 into the same equation gives that the log of the distance is 1.6946961 , so the expected distance is 49.510362 . The actual distance is about 39.503 .

  9. Exercise 9 Supplied answer

    Suppose that W is a subspace of ℝ n for some  n and suppose that v → is not an element of  W . Let the orthogonal projection of v → into W be the vector proj W ( v → ) = p → . Show that p → is the element of  W that is closest to  v → .

    Back to Exercise 9

    Answer. For any w → ∈ W , the vectors v → − p → and p → − w → are orthogonal. So the Triangle Inequality applies to the triangle with those vectors as sides, and with v → − w → as hypoteneuse. Therefore v → − w → is at least as long as v → − p → .

References cited in this section

Oakley & Baker

Cletus O. Oakley, Justine C. Baker, Least Squares and the 3 : 40 Mile, Mathematics Teacher, Apr. 1977.

Wikipedia, Mens Mile

Mile run world record progression, http://en.wikipedia.org/wiki/Mile_run_world_record_progression, 2011-Apr-09.

Cost Of Tolls

Cost of Tolls, http://costoftolls.com/Tolls_in_New_York.html, 2012-Jan-07.

Google Maps

Directions—Google Maps, http://maps.google.com/help/maps/directions/, 2012-Jan-07.

Dalal, et. al.

Siddhartha R. Dalal, Edward B. Fowlkes, & Bruce Hoadley, Lesson Learned from Challenger: A Statistical Perspective, Stats: the Magazine for Students of Statistics, Fall 1989, p. 3.

Gardner, 1970

Martin Gardner, Mathematical Games, Some mathematical curiosities embedded in the solar system, Scientific American, April 1970, p. 108–112.