26  Simple Linear Regression

Chapter 21 ended on an identity. For the forty stores, the correlation between marketing spend and revenue was 0.620, and the slope of the line through that cloud was

\[b_1 = r \, \frac{s_y}{s_x} = 4.27\]

which was offered as a curiosity and is in fact the whole of this chapter in one line. Correlation says how tightly two variables track a line. Regression takes the line seriously, and once you do, four things become available that a correlation cannot give you: a number with units attached, a statement of how much \(y\) changes per unit of \(x\), a prediction, and an honest interval around that prediction.

The same forty stores run through this chapter, so every figure can be checked against Chapter 21.

26.1 The Line, and What “Best” Means

The model says each observation is a point on a straight line plus an error:

\[y_i = \beta_0 + \beta_1 x_i + \varepsilon_i\]

\(\beta_0\) and \(\beta_1\) are population quantities nobody observes. What the data gives is estimates \(b_0\) and \(b_1\), and from them a fitted value for every store,

\[\hat{y}_i = b_0 + b_1 x_i\]

and a residual, the part the line failed to account for:

\[e_i = y_i - \hat{y}_i\]

Infinitely many lines can be drawn through a cloud of points. “Best” has to be defined before it can be found.

Least squares defines it: choose \(b_0\) and \(b_1\) to make

\[SSE = \sum_{i=1}^{n} e_i^{2} = \sum \left(y_i - b_0 - b_1 x_i\right)^{2}\]

as small as possible. Two choices inside that definition are worth questioning, because both are conventions rather than necessities.

Why vertical distances? Because the question is asymmetric. Regression predicts \(y\) from \(x\), so what matters is the error in \(y\). Measuring perpendicular distances instead gives a different line and a different method, and regressing \(x\) on \(y\) gives a third line again. Correlation was symmetric; regression is not, and choosing which variable goes on the left is a real decision.

Why squared? Squaring makes the problem have exactly one answer, reachable in closed form. It also punishes a residual of 20 four times as hard as two residuals of 10, which pulls the line hard toward distant points. That sensitivity is the price, and it is the same sensitivity that let one outlier wreck the correlation in Chapter 21.

Minimising that sum gives the two estimates directly:

\[b_1 = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sum (x_i - \bar{x})^{2}} = \frac{S_{xy}}{S_{xx}} \qquad\qquad b_0 = \bar{y} - b_1 \bar{x}\]

The second says something useful on its own: the fitted line always passes through \((\bar{x}, \bar{y})\). Whatever else it does, it goes through the middle of the data.

Divide the top and bottom of \(b_1\) by \(n-1\) and it becomes \(\operatorname{cov}(x,y)/s_x^2\), which is \(r \, s_y / s_x\). The identity from Chapter 21 is not a coincidence; it is the same formula written differently.

The first tab of the interactive above makes the definition literal. Move the intercept and slope sliders and the orange segments, which are the residuals, grow and shrink while the sum of their squares is reported. The curve underneath is that sum plotted against the slope. It has one minimum, it is a parabola, and the bottom of it is the fitted line. Every other line you can draw sits higher.

26.2 A Worked Example

Example

The forty stores of Chapter 21, spend and revenue both in thousands. Least squares returns

\[\widehat{\text{revenue}} = 137.11 + 4.27 \times \text{spend}\]

with \(SSE = 23{,}071\), and no other straight line gets that sum below 23,071.

The slope is the sentence that matters. A store spending one thousand more on marketing is associated with about 4,270 more in revenue. Note the units: revenue per unit of spend, thousands per thousand. Note also “associated with”, not “produces”, for every reason Chapter 21 gave.

The intercept needs more care. Taken literally it says a store spending nothing on marketing would earn 137,110. Whether that is a claim worth making depends entirely on whether zero spend is inside the range of the data. Here the lowest observed spend is 8.01 thousand, so the intercept sits outside the data and is best read as a number that positions the line rather than a fact about unmarketed stores. An intercept is only interpretable when \(x = 0\) is a real and observed possibility.

26.3 How Much of the Scatter the Line Accounts For

The total variation in revenue splits in two, exactly as it did in Chapter 19:

\[\underbrace{\sum (y_i - \bar{y})^2}_{SST} = \underbrace{\sum (\hat{y}_i - \bar{y})^2}_{SSR} + \underbrace{\sum (y_i - \hat{y}_i)^2}_{SSE}\]

the part the line accounts for plus the part it does not. For the forty stores:

\[37{,}480.93 = 14{,}409.55 + 23{,}071.38\]

and the ratio of the first piece to the whole is the coefficient of determination:

\[R^{2} = \frac{SSR}{SST} = \frac{14{,}409.55}{37{,}480.93} = 0.3845\]

which is exactly the \(r^2\) of Chapter 21. In simple regression with one predictor, \(R^2\) and \(r^2\) are the same number, and \(R\) is \(|r|\).

\(R^2\) is unitless and therefore easy to quote and easy to over-read. The number that says how wrong the line typically is, in the units anyone cares about, is the residual standard error:

\[s = \sqrt{\frac{SSE}{n-2}} = \sqrt{\frac{23{,}071.38}{38}} = 24.64\]

Revenue predictions from this line are typically out by about 24,640. Set that against an average revenue of 214,000 and the model looks a good deal less impressive than “\(R^2 = 0.38\)” makes it sound. Report both.

The divisor is \(n - 2\), not \(n - 1\), because two parameters were estimated from the data before the residuals were computed. Each estimate costs one degree of freedom, exactly as the single mean cost one in Chapter 15.

26.4 Inference on the Slope

The fitted slope is an estimate from one sample and carries a standard error:

\[SE(b_1) = \frac{s}{\sqrt{\sum (x_i - \bar{x})^2}} = \frac{s}{\sqrt{S_{xx}}}\]

Read that denominator. The slope is estimated more precisely when \(x\) is spread out, which is the same fact that appeared in Chapter 21 as restriction of range: a narrow band of \(x\) pins a line down badly.

The test of whether there is any relationship at all is a \(t\) test on \(n-2\) degrees of freedom:

\[H_0: \beta_1 = 0 \qquad t = \frac{b_1}{SE(b_1)}\]

For the forty stores, \(SE(b_1) = 0.8768\) and

\[t = \frac{4.2715}{0.8768} = 4.872 \quad \text{on 38 df}, \qquad p = 0.00002\]

That \(t\) of 4.872 should look familiar. It is exactly the statistic Chapter 21 computed to test whether the correlation was zero, to every decimal place. Testing \(\rho = 0\) and testing \(\beta_1 = 0\) are not similar tests; they are the same test, because a correlation of zero and a slope of zero are the same claim.

The same identity appears once more. R’s regression output also reports \(F = 23.733\) on 1 and 38 degrees of freedom, and \(4.872^2 = 23.733\). This is the \(F = t^2\) result from Chapter 19 arriving from a third direction, because a regression is an analysis of variance on a continuous predictor.

The interval is more use than the p-value, since it is in units:

\[b_1 \pm t_{0.975,\,38} \times SE(b_1) = 4.2715 \pm 2.024 \times 0.8768 = (2.50,\ 6.05)\]

So the data is consistent with an extra thousand of marketing spend being worth anywhere from about 2,500 to about 6,050 in revenue. That is a wide range, and it is the honest one. A report that gives only “\(p < 0.001\)” has concealed the fact that the effect could be less than half, or more than twice, what the point estimate says.

26.5 The Residuals Are the Diagnostic

Four conditions carry the inference above, and they are assumptions about the errors, not about \(x\) or \(y\) separately.

Linearity. The relationship really is a straight line. If it curves, every coefficient is describing something that is not there.

Independence. The errors are independent of each other. Time-ordered data breaks this most often, because today’s error resembles yesterday’s.

Constant variance. The spread of the errors is the same all along the line. When it fans out, the term is heteroscedasticity, and the standard errors it produces are wrong.

Approximate Normality of the errors. Needed for the \(t\) intervals, and the least demanding of the four once \(n\) passes a few dozen, for the Central Limit Theorem reasons of Chapter 14.

Nothing in that list is checked by looking at \(R^2\), and nothing in it is checked by looking at the p-value. All four are checked by looking at the residuals.

Two properties of the residuals are worth knowing precisely because they are useless as diagnostics. Least squares forces both:

\[\sum e_i = 0 \qquad\text{and}\qquad \sum x_i e_i = 0\]

The residuals always average exactly zero, and they are always exactly uncorrelated with \(x\). That holds for a perfect fit and for a catastrophic one. Anyone reporting either as evidence the model is sound has reported nothing at all.

What carries information is the pattern.

The plot to draw is residuals against fitted values, or against \(x\), which is the same plot in simple regression. A healthy one is a shapeless horizontal band.

What you see What it means
Shapeless band around zero The line is doing its job
An arch or a valley The relationship is curved; fit a curve or transform
A funnel widening to the right Non-constant variance; transform \(y\) or use robust errors
A drift with time order Errors are correlated; the standard errors are too small
One point far from the rest An outlier worth investigating before it is removed

Example

The second tab of the interactive above holds two datasets. Switch between them and watch the top plot, then the bottom one.

The forty stores give an unremarkable scatter and an unremarkable residual plot. The second dataset gives a scatter that also looks perfectly ordinary, a rising cloud with scatter around it, and a fit that reports \(R^2 = 0.30\) with a slope significant at \(p = 0.0003\). On the numbers alone it looks like a working model.

Its residual plot is an arch: strongly negative at both ends, positive through the middle. The relationship is curved, and the straight line is wrong. Adding a squared term takes \(R^2\) from 0.30 to 0.80, with the new term significant at \(p \approx 10^{-11}\).

A significant slope and a respectable \(R^2\) told you nothing about whether a line was the right shape. The residual plot told you immediately.

26.6 Two Intervals, and the One People Confuse

Ask the model about a store spending 20 thousand and it returns 222.54. That single number can carry two completely different intervals, answering two completely different questions.

A confidence interval for the mean response answers: where is the average revenue of all stores spending 20? Its standard error is

\[SE_{\text{mean}} = s\sqrt{\frac{1}{n} + \frac{(x_0 - \bar{x})^{2}}{S_{xx}}}\]

A prediction interval for one new observation answers: where will the revenue of one particular new store spending 20 fall? Its standard error carries an extra term:

\[SE_{\text{pred}} = s\sqrt{1 + \frac{1}{n} + \frac{(x_0 - \bar{x})^{2}}{S_{xx}}}\]

That leading 1 is the scatter of individual stores about the line. It does not shrink as \(n\) grows, because collecting more data tells you nothing about how unusual the next store will be.

Example

At a spend of 20, with \(s = 24.64\), \(n = 40\), \(\bar{x} = 18.00\):

Standard error 95% interval Width
Mean response 4.27 (213.9, 231.2) 17.3
One new store 25.01 (171.9, 273.2) 101.3

The prediction interval is 5.85 times wider. Both are correct; they answer different questions. Quoting the first when someone asked the second is the most common error in applied regression reporting, and it understates the real uncertainty by a factor of six.

Both formulas contain \((x_0 - \bar{x})^2\), so both bands are narrowest at the mean of \(x\) and widen either side, giving the hourglass shape in the third tab of the interactive above. A fitted line is pinned most firmly through the middle of its data and pivots at the ends.

The effect is much more visible on the confidence band than the prediction band. Drag the slider: the mean-response band roughly triples in width going from the middle of the data to its edge, while the prediction band barely moves, because the prediction band is dominated by that leading 1 and the line’s own wobble is a small part of it.

26.7 Extrapolation

The observed spend runs from 8.01 to 26.74 thousand. The fitted line will happily return a value for a spend of 50, or 200, or zero, and every one of those numbers is arithmetic rather than evidence.

The data contains no information about whether the relationship stays straight outside the range it covers, and most real relationships do not. Marketing spend surely shows diminishing returns eventually; nothing in a sample that never exceeded 27 thousand can tell you where that begins. The shaded regions in the third tab mark the territory the data never visited.

The standard errors above do widen as \(x_0\) moves away from \(\bar{x}\), which is honest as far as it goes, but they widen only for the uncertainty in a line that is assumed straight. They contain no allowance whatsoever for the line being the wrong shape out there, which is the actual risk.

Two habits follow. State the range of \(x\) the model was fitted on whenever the model is passed to anyone else, because the fitted equation does not carry that information and looks perfectly authoritative without it. And treat a prediction beyond the data as a hypothesis to be tested rather than a result to be acted on.

Three things in that output are worth carrying away. The slope least squares returns is the same 4.2715 that Chapter 21 produced from the correlation alone, and the t statistic testing it is the same 4.8717 that chapter used to test the correlation, because they are one test. The sums of squares add up, and their ratio is the \(r^2\) from the previous chapter under a new name. And the interval for one new store is five and a half times wider than the interval for the average store, from the same fit, at the same spend.

Recap

Chapters 21 and 22 are one relationship examined twice. Correlation reports how tightly two measured variables track a straight line, on a scale with no units, and refuses every other question. Regression draws the line, which buys a slope in real units, a statement of how much \(y\) moves per unit of \(x\), a prediction, and an interval around it. The arithmetic connecting them is not loose: the fitted slope is \(r\) rescaled by the two standard deviations, \(R^2\) is \(r^2\), and testing the slope against zero is literally the same test as testing the correlation against zero. What both chapters keep saying, in different ways, is that the summary number is the least interesting part. Chapter 21 put four wildly different datasets behind one correlation; Chapter 22 put a curved relationship behind a significant slope and a respectable \(R^2\), where only the residual plot gave it away. The instruction is the same both times, and it is to look at the picture. What this chapter has not done is handle more than one predictor. Real questions rarely have one: revenue depends on spend and store size and region, and those predictors are tangled with each other, so the effect of any one of them can only be stated once the others are accounted for. That is multiple regression, and with it comes a new set of problems, starting with what “holding the others constant” actually means when the predictors move together.


Summary

Concept Description
The Line
The Regression Model Each observation is a point on a straight line plus an independent error
Fitted Value The height of the fitted line at that observation's x
Residual Observed minus fitted, the part of y the line did not account for
Least Squares Choose the line that makes the sum of squared residuals as small as possible
Why Vertical Distances The question is to predict y from x, so the error that matters is the error in y
Why Squared It makes the answer unique and closed form, at the cost of sensitivity to outliers
The Slope Formula The cross product of deviations divided by the sum of squared deviations in x
The Intercept Formula The mean of y minus the slope times the mean of x
The Line Through the Means The fitted line always passes through the point of the two means
Reading It
The Link to Correlation The slope is r times the ratio of the standard deviation of y to that of x
Reading the Slope The change in y associated with a one unit change in x, in real units
Reading the Intercept Only interpretable when x equals zero is inside the range the data covers
Regression Is Not Symmetric Regressing y on x and x on y give different lines, unlike correlation
How Well It Fits
Partition of the Variance Total sum of squares equals the regression part plus the error part
R Squared The share of the variation in y the line accounts for
R Squared Equals r Squared In simple regression with one predictor they are the same number
Residual Standard Error The typical size of a residual, in the units of y, and the honest companion to R squared
Why n Minus Two Two parameters were estimated before the residuals were formed
Inference
Standard Error of the Slope The residual standard error divided by the square root of the spread in x
Spread of x Helps A wider range of x pins the slope down better, which is restriction of range again
Testing the Slope A t test on n minus two degrees of freedom, against a slope of zero
The Same Test as the Correlation Testing the slope against zero and the correlation against zero are one test
F Equals t Squared The regression F is the square of the slope t, as in Chapter 19
Interval for the Slope More use than the p-value, because it says how large the effect could be in units
The Residuals
Assumptions Linearity, independence, constant variance, and approximate Normality of the errors
Residuals Against Fitted Values The one plot that checks all four, and the only one that finds a wrong shape
Two Useless Residual Facts Residuals always sum to zero and are always uncorrelated with x, by construction
An Arch in the Residuals The relationship is curved and the straight line is the wrong model
A Funnel in the Residuals Non-constant variance, so the reported standard errors are wrong
Drift with Time Order The errors are correlated and the standard errors are too small
Prediction
Confidence Interval for the Mean Where the average y sits for all cases at that x
Prediction Interval for One Case Where one particular new case will fall, which is a different question
Why the Second Is Wider It carries the scatter of individual cases as well as the uncertainty in the line
The Hourglass Both bands are narrowest at the mean of x and widen either side as the line pivots
Extrapolation The fitted line returns numbers beyond the data that the data does not support
What the Standard Errors Do Not Cover They allow for uncertainty in a line assumed straight, not for the shape being wrong