Line of Best Fit
This Topic requires the formulas from the subsections on Orthogonal Projection Into a Line and Projection Into a Subspace.
Scientists are often presented with a system that has no solution and they must find an answer anyway. More precisely, they must find a best answer. For instance, this is the result of flipping a penny, including some intermediate numbers.
| number of flips | 30 | 60 | 90 |
|---|---|---|---|
| number of heads | 16 | 34 | 51 |
Because of the randomness in this experiment we expect that the ratio of heads to flips will fluctuate around a penny’s long-term ratio of 50-50. So the system for such an experiment likely has no solution, and that’s what happened here.
That is, the vector of data that we collected is not in the subspace where ideally it would be.
However, we have to do something so we look for the that most nearly works. An orthogonal projection of the data vector into the line subspace gives a best guess, the vector in the subspace closest to the data vector.
The estimate () is a bit more than one half, but not much more than half, so probably the penny is fair enough.
The line with the slope is the line of best fit for this data.
Minimizing the distance between the given vector and the vector used as the right-hand side minimizes the total of these vertical lengths, and consequently we say that the line comes from fitting by least-squares.
This diagram exaggerates the vertical scale by a factor of ten to make the lengths more visible.
In the above equation the line must pass through , because we take it to be the line whose slope is this coin’s true proportion of heads to flips. We can also handle cases where the line need not pass through the origin.
Here is the progression of world record times for the men’s mile race [Oakley & Baker]. In the early 1900’s many people wondered when, or if, this record would fall below the four minute mark. Here are the times that were in force on January first of each decade through the first half of that century.
| year | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| secs |
We can use this to give a circa 1950 prediction of the date for seconds, and then compare that to the actual date. As with the penny data, these numbers do not lie in a perfect line. That is, this system does not have an exact solution for the slope and intercept.
We find a best approximation by using orthogonal projection.
(Comments on the data. Restricting to the times at the start of each decade reduces the data entry burden, smooths the data to some extent, and gives much the same result as entering all of the dates and records. There are different sequences of times from competing standards bodies but the ones here are from [Wikipedia, Mens Mile]. We’ve started the plot at 1870 because at one point there were two classes of records, called ‘professional’ and ‘amateur’, and after a while the first class stopped being active so we’ve followed the second class.)
Write the linear system’s matrix of coefficients and also its vector of constants, the world record times.
The ending result in the subsection on Projection into a Subspace gives the formula for the the coefficients and that make the linear combination of ’s columns as close as possible to . Those coefficients are the entries of the vector .
Sage can do the computation for us.
sage: year = [1870, 1880, 1890, 1900, 1910, 1920, 1930, 1940, 1950]
sage: secs = [268.8, 264.5, 258.4, 255.6, 255.6, 252.6, 250.4, 246.4, 241.4]
sage: var('a, b, t')
(a, b, t)
sage: model(t) = a*t+b
sage: data = zip(year, secs)
sage: fit = find_fit(data, model, solution_dict=True)
sage: model.subs(fit)
t |--> -0.3048333333333295*t + 837.0872222222147
sage: g=points(data)+plot(model.subs(fit),(t,1860,1960),color='red',
....: figsize=3,fontsize=7,typeset='latex')
sage: g.save("four_minute_mile.pdf")
sage: g
The progression makes a surprisingly good line. From the slope and intercept we predict ; the actual date of Roger Bannister’s record was 1954-May-06.
The final example compares team salaries from US major league baseball against the number of wins the team had, for the year 2002. In this year the Oakland Athletics used mathematical techniques to optimize the players that they fielded for the money that they could spend, as told in the film Moneyball. (Salaries are in millions of dollars and the number of wins is out of 162 games).
To do the computations we again use Sage.
sage: sal = [40, 40, 39, 42, 45, 42, 62, 34, 41, 57, 58, 63, 47, 75, 57, 78, 80, 50, 60, 93,
....: 77, 55, 95, 103, 79, 76, 108, 126, 95, 106]
sage: wins = [103, 94, 83, 79, 78, 72, 99, 55, 66, 81, 80, 84, 62, 97, 73, 95, 93, 56, 67,
....: 101, 78, 55, 92, 98, 74, 67, 93, 103, 75, 72]
sage: var('a, b, t')
(a, b, t)
sage: model(t) = a*t+b
sage: data = zip(sal,wins)
sage: fit = find_fit(data, model, solution_dict=True)
sage: model.subs(fit)
t |--> 0.2634981251436269*t + 63.06477642781477
sage: p = points(data,size=25)+plot(model.subs(fit),(t,30,130),color='red',typeset='latex')
sage: p.save('moneyball.pdf')
The graph is below. The team in the upper left, who paid little for many wins, is the Oakland A’s.
Judging this line by eye would be error-prone. So the equations give us a certainty about the ‘best’ in best fit. In addition, the model’s equation tells us roughly that by spending an additional million dollars a team owner can expect to buy of a win (and that expectation is not very sure, thank goodness).
Exercises
The calculations here are best done on a computer. Some of the problems require data from the Internet.
Exercise 1 Supplied answer
Use least-squares to judge if the coin in this experiment is fair.
flips 8 16 24 32 40 heads 4 9 13 17 20 Answer. As with the first example discussed above, we are trying to find a best to “solve” this system.
Projecting into the linear subspace gives this
so the slope of the line of best fit is approximately .
Exercise 2 Supplied answer
For the men’s mile record, rather than give each of the many records and its exact date, we’ve “smoothed” the data somewhat by taking a periodic sample. Do the longer calculation and compare the conclusions.
Answer. With this input
(the dates have been rounded to months, e.g., for a September record, the decimal was used), Maple responded with an intercept of and a slope of .
Exercise 3 Supplied answer
Find the line of best fit for the men’s meter run. How does the slope compare with that for the men’s mile? (The distances are close; a mile is about meters.)
Answer. With this input (the years are zeroed at )
(the dates have been rounded to months, e.g., for a September record, the decimal was used), Maple gives an intercept of and a slope of . The slope given in the body of this Topic for the men’s mile is quite close to this.
Exercise 4 Supplied answer
Find the line of best fit for the records for women’s mile.
Answer. With this input (the years are zeroed at )
(the dates have been rounded to months, e.g., for a September record, the decimal was used), MAPLE gave an intercept of and a slope of .
Exercise 5 Supplied answer
Do the lines of best fit for the men’s and women’s miles cross?
Answer. These are the equations of the lines for men’s and women’s mile (the vertical intercept term of the equation for the women’s mile has been adjusted from the answer above, to zero it at the year , because that’s how the men’s mile equation was done).
Obviously the lines cross. A computer program is the easiest way to do the arithmetic: MuPAD gives and ( seconds is minutes and seconds). Remark. Of course all of this projection is highly dubious —for one thing, the equation for the women is influenced by the quite slow early times —but it is nonetheless fun.
Exercise 6 Supplied answer
(This illustrates that there are data sets for which a linear model is not right, and that the line of best fit doesn’t in that case have any predictive value.) In a highway restaurant a trucker told me that his boss often sends him by a roundabout route, using more gas but paying lower bridge tolls. He said that New York State calibrates the toll for each bridge across the Hudson, playing off the extra gas to get there from New York City against a lower crossing cost, to encourage people to go upstate. This table, from [Cost Of Tolls] and [Google Maps], lists for each toll crossing of the Hudson River, the distance to drive from Times Square in miles and the cost in US dollars for a passenger car (if a crossings has a one-way toll then it shows half that number).
Crossing Distance Toll Lincoln Tunnel Holland Tunnel George Washington Bridge Verrazano-Narrows Bridge Tappan Zee Bridge Bear Mountain Bridge Newburgh-Beacon Bridge Mid-Hudson Bridge Kingston-Rhinecliff Bridge Rip Van Winkle Bridge Find the line of best fit and graph the data to show that the driver was practicing on my credulity.
Answer. Sage gives the line of best fit as .
sage: dist = [2, 7, 8, 16, 27, 47, 67, 82, 102, 120] sage: toll = [6, 6, 6, 6.5, 2.5, 1, 1, 1, 1, 1] sage: var('a,b,t') (a, b, t) sage: model(t) = a*t+b sage: data = zip(dist,toll) sage: fit = find_fit(data, model, solution_dict=True) sage: model.subs(fit) t |--> -0.0508568169130319*t + 5.630955848442933 sage: p = plot(model.subs(fit), (t,0,120))+points(data,size=25,color='red') sage: p.save('bridges.pdf')But the graph shows that the equation has little predictive value.
Apparently a better model is that (with only one intermediate exception) crossings in the city cost roughly the same as each other, and crossings upstate cost the same as each other.
Exercise 7 Supplied answer
When the space shuttle Challenger exploded in 1986, one of the criticisms made of NASA’s decision to launch was in the way they did the analysis of number of O-ring failures versus temperature (O-ring failure caused the explosion). Four O-ring failures would be fatal. NASA had data from previous flights.
temp F 53 75 57 58 63 70 70 66 67 67 67 failures 3 2 1 1 1 1 1 0 0 0 0
68 69 70 70 72 73 75 76 76 78 79 80 81 0 0 0 0 0 0 0 0 0 0 0 0 0 The temperature that day was forecast to be .
NASA based the decision to launch partially on a chart showing only the flights that had at least one O-ring failure. Find the line that best fits these seven flights. On the basis of this data, predict the number of O-ring failures when the temperature is , and when the number of failures will exceed four.
Find the line that best fits all 24 flights. On the basis of this extra data, predict the number of O-ring failures when the temperature is , and when the number of failures will exceed four.
Which do you think is the more accurate method of predicting? (An excellent discussion is in [Dalal, et. al.].)
Answer.
A computer algebra system like MAPLE or MuPAD will give an intercept of and a slope of Plugging into the equation yields a predicted number of O-ring failures of (rounded to two places). Plugging in and solving gives a temperature of F.
On the basis of this information
MAPLE gives the intercept and the slope . Here, plugging into the equation predicts O-ring failures (rounded to two places). Plugging in failures gives a temperature of F.
Exercise 8 Supplied answer
This table lists the average distance from the sun to each of the first seven planets, using Earth’s average as a unit.
Mercury Venus Earth Mars Jupiter Saturn Uranus Plot the number of the planet (Mercury is , etc.) versus the distance. Note that it does not look like a line, and so finding the line of best fit is not fruitful.
It does, however look like an exponential curve. Therefore, plot the number of the planet versus the logarithm of the distance. Does this look like a line?
The asteroid belt between Mars and Jupiter is what is left of a planet that broke apart. Renumber so that Jupiter is , Saturn is , and Uranus is , and plot against the log again. Does this look better?
Use least squares on that data to predict the location of Neptune.
Repeat to predict where Pluto is.
Is the formula accurate for Neptune and Pluto?
This method was used to help discover Neptune (although the second item is misleading about the history; actually, the discovery of Neptune in position prompted people to look for the “missing planet” in position ). See [Gardner, 1970]
Answer.
The plot is nonlinear.
Here is the plot.
There is perhaps a jog up between planet and planet .
This plot seems even more linear.
With this input
MuPAD gives that the intercept is and the slope is .
Plugging into the equation from the prior item gives that the log of the distance is , so the expected distance is . The actual distance is about .
Plugging into the same equation gives that the log of the distance is , so the expected distance is . The actual distance is about .
Exercise 9 Supplied answer
Suppose that is a subspace of for some and suppose that is not an element of . Let the orthogonal projection of into be the vector . Show that is the element of that is closest to .
Answer. For any , the vectors and are orthogonal. So the Triangle Inequality applies to the triangle with those vectors as sides, and with as hypoteneuse. Therefore is at least as long as .
References cited in this section
Oakley & Baker
Cletus O. Oakley, Justine C. Baker, Least Squares and the Mile, Mathematics Teacher, Apr. 1977.
Wikipedia, Mens Mile
Mile run world record progression, http://en.wikipedia.org/wiki/Mile_run_world_record_progression, 2011-Apr-09.
Cost Of Tolls
Cost of Tolls, http://costoftolls.com/Tolls_in_New_York.html, 2012-Jan-07.
Google Maps
Directions—Google Maps, http://maps.google.com/help/maps/directions/, 2012-Jan-07.
Dalal, et. al.
Siddhartha R. Dalal, Edward B. Fowlkes, & Bruce Hoadley, Lesson Learned from Challenger: A Statistical Perspective, Stats: the Magazine for Students of Statistics, Fall 1989, p. 3.
Gardner, 1970
Martin Gardner, Mathematical Games, Some mathematical curiosities embedded in the solar system, Scientific American, April 1970, p. 108–112.