27  Multiple Regression

Chapter 21 gave a warning and then walked away from it:

In business data the third variable is usually size. Large stores spend more on marketing and earn more revenue, so spend and revenue correlate across stores even if the marketing does nothing at all.

Chapter 22 then fitted a line to exactly that data and reported that every extra thousand of marketing spend came with 4,270 more revenue, significant at \(p = 0.00002\). Both chapters were careful to say “associated with” rather than “produces”, which is honest and also unsatisfying, because the question anyone actually has is whether the marketing is doing anything.

Multiple regression is the tool for that question. It holds other variables fixed while examining one, which is the closest thing to a controlled comparison that observational data allows. This chapter puts store size into the model and watches what happens to the marketing coefficient. The answer is not comfortable.

27.1 The Model With Several Predictors

The extension is the obvious one. With \(k\) predictors,

\[y_i = \beta_0 + \beta_1 x_{1i} + \beta_2 x_{2i} + \cdots + \beta_k x_{ki} + \varepsilon_i\]

and least squares does the same job it did in Chapter 22: choose the coefficients that make \(\sum e_i^2\) as small as possible. With one predictor that meant fitting a line through a scatter. With two it means fitting a plane through a cloud in three dimensions, and beyond that the picture gives out while the arithmetic does not.

Almost everything from Chapter 22 carries over unchanged. Each coefficient still gets a standard error and a \(t\) test on the residual degrees of freedom, now \(n - k - 1\) because \(k+1\) parameters have been estimated. \(SST\) still splits into \(SSR\) and \(SSE\). The residual standard error is still \(\sqrt{SSE/(n-k-1)}\) and still the honest measure of typical error.

What changes, and it changes completely, is what a coefficient means.

27.2 What “Holding the Others Constant” Means

In simple regression, \(b_1\) was the slope of a line: the change in \(y\) per unit of \(x\). In multiple regression, \(b_1\) is a partial coefficient:

the change in \(y\) associated with a one unit change in \(x_1\), among cases that are identical on every other predictor in the model.

That final clause is not decoration. It is the whole meaning, and it has three consequences that catch people out.

The coefficient depends on what else is in the model. The same predictor, the same data, a different set of companions, a different number. There is no such thing as “the effect of \(x_1\)” in isolation.

It is only as good as the list. Holding constant the variables you measured says nothing about the ones you did not. A confounder you never recorded is not controlled for by any amount of regression.

A coefficient can change sign. Not only shrink; reverse. This is Simpson’s paradox from Chapter 10 and Chapter 21, arriving as a regression coefficient.

Example

The forty stores of Chapters 21 and 22, now with each store’s floor area recorded, in hundreds of square feet. Size correlates 0.80 with marketing spend and 0.72 with revenue, which is exactly the pattern Chapter 21 described: big stores spend more, and big stores earn more.

Three models on the same data:

Model Spend coefficient Size coefficient \(R^2\)
revenue ~ spend 4.27 (\(p < 0.0001\)) 0.384
revenue ~ size 2.48 (\(p < 0.0001\)) 0.519
revenue ~ spend + size 0.84 (\(p = 0.52\)) 2.14 (\(p = 0.002\)) 0.524

The marketing coefficient falls from 4.27 to 0.84, a fifth of its former size, and its p-value goes from 0.00002 to 0.52. The size coefficient barely moves, from 2.48 to 2.14.

Read together, those three rows tell a single story. Almost all of the apparent marketing effect was store size wearing marketing’s clothes. Once floor area is in the model, spend has next to nothing left to explain, and the data can no longer distinguish its effect from zero.

Be careful about what that does and does not establish.

It does not prove marketing is worthless. The confidence interval for the adjusted coefficient runs from about \(-1.8\) to \(3.5\), which comfortably includes effects worth having. What collapsed is not the effect but the evidence for it: after size takes its share, too little independent variation in spend remains to measure anything precisely. Absence of evidence, as Chapter 16 insisted, is not evidence of absence.

It also does not prove size is the cause. Size is simply another observed variable, and something else again may drive both. What the model does is remove one specific alternative explanation, and there are always others.

The first tab of the interactive above runs this as a dial. Store size is held at a correlation of 0.72 with revenue while its tie to spend is varied. At zero, spend keeps its full 4.27, because a confounder unrelated to the predictor confounds nothing. As the tie tightens, the coefficient walks down, the interval swells, and somewhere around 0.7 it swallows zero.

That is a definition made visible. A confounder must be related to both variables, and how much damage it does depends on how tightly it is bound to the one you care about.

27.3 R Squared Can Only Rise

\(R^2\) has a property that makes it useless for choosing between models: adding a predictor can never lower it. Not “rarely”. Never. Least squares can always set the new coefficient to zero and do exactly as well as before, so it will only ever do better.

That includes predictors with no relationship to anything. Take the model with spend and size, \(R^2 = 0.524\), and add columns of pure random noise one at a time:

Noise columns added \(R^2\) Adjusted \(R^2\)
0 0.5242 0.4985
1 0.5243 0.4847
2 0.5266 0.4725
4 0.5290 0.4434
6 0.5418 0.4235

\(R^2\) climbs on nothing at all. A model comparison run on \(R^2\) will always pick the biggest model.

Adjusted \(R^2\) charges rent for each predictor:

\[R^{2}_{\text{adj}} = 1 - \frac{SSE / (n - k - 1)}{SST / (n - 1)} = 1 - (1 - R^{2}) \frac{n-1}{n-k-1}\]

Each extra predictor shrinks the denominator \(n - k - 1\), which pushes the adjusted figure down unless the predictor earns its place by reducing \(SSE\) enough to compensate. In the table above it falls the whole way, from 0.499 to 0.424, which is the correct verdict: the noise columns were worthless.

Two properties are worth knowing. Adjusted \(R^2\) can be negative, when the model is worse than simply using \(\bar{y}\). And it is not a proportion of anything, so the interpretation that \(R^2\) enjoys does not transfer. It is a model-comparison device, not a description.

The penalty is mild, as the second tab of the interactive above shows: with eight noise columns and forty cases, adjusted \(R^2\) falls but never becomes alarming. It is a check on casual over-fitting, not a defence against determined over-fitting. Chapter 24 takes the problem seriously, with out-of-sample testing rather than a penalty term.

27.4 Categorical Predictors

Regression needs numbers, and store format, region and payment method are not numbers. Dummy coding converts them: pick one level as the reference, and create an indicator column for each remaining level, holding 1 when the case is at that level and 0 otherwise.

With two levels, one dummy does it. Coding Compact as the reference,

\[\text{Flagship}_i = \begin{cases} 1 & \text{if store } i \text{ is a flagship} \\ 0 & \text{otherwise}\end{cases}\]

and the model \(\hat{y} = b_0 + b_1 x + b_2\,\text{Flagship}\) describes two parallel lines: intercept \(b_0\) for compact stores, intercept \(b_0 + b_2\) for flagships, the same slope \(b_1\) for both. So \(b_2\) is the vertical gap between them, which is the difference between the two formats at any given spend.

Example

A second set of forty stores, twenty of each format, where the categorical variable does real work.

Model Coefficients \(R^2\)
revenue ~ spend spend 7.65 0.822
revenue ~ spend + format spend 4.19, flagship +44.47 0.924

Ignoring format, each thousand of spend looks worth 7,650. Once format is in the model it is worth 4,190, and flagship stores earn about 44,470 more than compact ones at the same spend.

The pooled slope was too steep because it was doing two jobs at once: tracking the real within-format relationship and absorbing the fact that flagships both spend more and earn more. This is the same confounding as the size example, and the same phenomenon Chapter 21 drew as subgroups pointing one way inside a cloud sloping the other.

Three traps with dummy coding.

Use \(k-1\) dummies for \(k\) levels, never \(k\). Including one for every level makes them sum to a column of ones, which is the intercept, and the model becomes unsolvable. This is the dummy variable trap. Software that builds dummies for you avoids it automatically; software fed hand-made columns does not.

Every coefficient is relative to the reference level. Change the reference and every dummy coefficient changes, without the model changing at all. The fitted values, \(R^2\) and predictions are all identical. Report which level is the reference or the numbers cannot be read.

Do not code an unordered category as a number. Assigning 1, 2, 3 to three regions and entering it as one predictor claims that region 3 is to region 2 as region 2 is to region 1, and that the gaps are equal. That is a claim about a nominal scale that Chapter 1 warned against, and it is almost always false.

27.5 Interaction Terms

The dummy model above forced the two formats to share a slope. That is an assumption, and it can be tested by letting them differ. Multiply the two predictors together and add the product:

\[\hat{y} = b_0 + b_1 x + b_2 \text{Flagship} + b_3 (x \times \text{Flagship})\]

For compact stores the Flagship indicator is 0, so the fitted line is \(b_0 + b_1 x\). For flagships it is 1, so the line is \((b_0 + b_2) + (b_1 + b_3)x\). \(b_3\) is the difference in slopes, and testing it against zero tests whether the relationship between spend and revenue is the same in both formats.

This is the interaction of Chapter 20 in a new notation. There it was a term in an ANOVA table; here it is a product column. The meaning has not changed: the effect of one predictor depends on the level of another.

Example

Adding the interaction to the second dataset:

Term Estimate \(p\)
Intercept 149.22
spend 2.56 0.003
Flagship −14.39 0.50
spend × Flagship 3.27 0.007

which unpacks into two lines:

\[\text{Compact}: \ \widehat{\text{revenue}} = 149.22 + 2.56 \times \text{spend}\] \[\text{Flagship}: \ \widehat{\text{revenue}} = 134.83 + 5.83 \times \text{spend}\]

A thousand of marketing spend is worth about 2,560 in a compact store and 5,830 in a flagship, and the difference of 3,270 is unlikely to be chance at \(p = 0.007\). Comparing the two models directly gives \(F(1, 36) = 8.26\), the same \(p\). There is no single answer to “what is marketing worth here”, and the parallel-lines model was quietly imposing one.

Look at the Flagship row: \(-14.39\), with \(p = 0.50\). It is tempting to read that as “format does not matter”, and it would be badly wrong.

In a model containing an interaction, a main effect is the effect of that variable when the other is zero. So \(-14.39\) is the estimated gap between formats at a spend of zero, which is outside the data and of no interest to anybody. It is not the average difference between formats, and its p-value is not a test of whether format matters.

The rule follows Chapter 20’s exactly. Read the interaction first. If it is significant, the main effects are not the summary, and the right description is the set of separate lines, not the individual rows of the table. And do not drop a main effect while keeping its interaction: the model that results is almost never one anybody meant to fit.

27.6 Multicollinearity

When two predictors carry much the same information, the data cannot say which deserves the credit. Multicollinearity is that problem, and it is measured per predictor by the variance inflation factor:

\[VIF_j = \frac{1}{1 - R^{2}_{j}}\]

where \(R^2_j\) comes from regressing predictor \(j\) on all the others. A \(VIF\) of 1 means that predictor is unrelated to the rest; 4 means its coefficient’s variance is four times what it would otherwise be, so its standard error is doubled.

For spend and size in the main example, \(r = 0.80\) and \(VIF = 2.77\) for both, which is noticeable but ordinary. Conventional alarm levels are above 5, and above 10 almost universally.

Add to the model a second predictor that carries the same information as spend, and tighten the resemblance:

Correlation \(VIF\) \(SE\) of the spend slope \(t\)
0.00 1.00 0.89 4.81
0.80 2.78 1.48 2.76
0.90 5.26 2.04 1.95
0.95 10.26 2.84 1.35
0.99 50.25 6.30 0.52

The standard error grows sevenfold and a decisive \(t\) of 4.81 becomes an unremarkable 0.52, on data where nothing about the relationship changed.

What collinearity does not do is worth stating as plainly as what it does.

It does not bias the coefficients; they remain right on average. It does not harm prediction, because the model as a whole still fits, and a model built only to predict can carry collinear predictors happily. It does not affect \(R^2\), the residual standard error, or the overall \(F\) test.

What it destroys is the ability to attribute. The data contains almost no cases where one predictor moves and the other does not, so the question “what does this one do on its own?” has no evidence behind it, and the widened standard error is the model reporting that honestly.

The responses, in rough order of preference: decide that only one of the two is needed and drop the other; combine them into a single measure; collect data in which they come apart; or, if prediction is all you want, leave them and stop interpreting the individual coefficients. What is not a response is noticing that neither predictor is significant and concluding that neither matters. Together they may be doing a great deal, and the overall \(F\) test will usually say so.

Three things in that output are worth carrying away. The marketing coefficient that Chapter 22 reported with such confidence survives contact with one extra column of data as 0.84 with a p-value of 0.52, and nothing about the marketing data changed. Six columns of pure random noise raise \(R^2\) every single time, which is why no model was ever chosen well on that number. And in the interaction model, the row labelled with the group name is not the group effect, a misreading that has probably survived more peer reviews than any other line of regression output.

Looking Ahead

Multiple regression is the first tool in this book that produces a model rather than a verdict. Everything before it answered a question and stopped; this returns an object with parts, which can be examined, argued with, and used on data it has never seen. That is a large gain and it arrives with a large bill. A coefficient now means nothing without the list of what else was in the model. Adding predictors always flatters the fit, so goodness of fit has stopped being evidence. Predictors that overlap can make each other look worthless while jointly explaining a great deal. And underneath all of it sit the same four assumptions Chapter 22 set out, now harder to check because the scatterplot that made them obvious no longer exists once there are three or more dimensions. The next chapter is about not being fooled by your own model. It covers what the residuals look like when each assumption fails and how to see that with several predictors; which individual observations are quietly deciding the answer, through leverage, influence and Cook’s distance; what to do when the variance is not constant; and the problem that only appears once a model can be made arbitrarily complicated, which is that a model fitting its own data beautifully is weak evidence that it will fit anything else. That is overfitting, and the honest answer to it is to test the model on data it has not seen.


Summary

Concept Description
The Model
The Multiple Regression Model Several predictors, each with its own coefficient, and one error term
Fitting a Plane With two predictors least squares fits a plane; beyond that the picture gives out
Residual Degrees of Freedom n minus k minus one, since k plus one parameters were estimated
What a Coefficient Means
Partial Coefficient The change in y per unit of that predictor among cases identical on the others
Holding the Others Constant Not decoration but the entire meaning of every coefficient in the table
It Depends on the Model The same predictor with different companions is a different number
Only the Variables You Measured Variables you never recorded are not controlled for by any amount of regression
A Coefficient Can Reverse Adding a predictor can flip a sign, which is Simpson's paradox as a coefficient
Confounding
Confounding A variable related to both the predictor and the outcome, manufacturing association
The Collapse of an Effect Marketing spend falls from 4.27 to 0.84 once store floor area is in the model
What a Collapse Does Not Prove Not that the effect is zero; the interval still covers effects worth having
Choosing a Model
R Squared Never Falls Adding any predictor, however useless, can only raise it or leave it alone
Adjusted R Squared Charges each predictor against the degrees of freedom, so junk lowers it
Adjusted R Squared Can Be Negative When the model is worse than simply using the mean of y
A Mild Penalty Only It checks casual overfitting, not determined overfitting, which needs a test set
Categorical Predictors
Dummy Coding Turning a categorical predictor into indicator columns of zeros and ones
The Reference Level The omitted level; every dummy coefficient is a difference from it
Parallel Lines One dummy shifts the intercept and leaves the slope alone
The Dummy Variable Trap Including a dummy for every level makes them sum to the intercept and breaks the fit
Never Number an Unordered Category Coding three regions as 1, 2, 3 claims equal spacing on a nominal scale
Interactions
Interaction Term The product of two predictors, letting the effect of one depend on the other
Separate Slopes The interaction coefficient is the difference between the groups' slopes
The Main Effect in an Interaction Model The effect of that variable when the other is zero, not the average group effect
Read the Interaction First As in Chapter 20, a significant interaction demotes the main effects
Keep the Main Effects Dropping a main effect while keeping its interaction fits a model nobody meant
Collinearity
Multicollinearity Two or more predictors carrying much the same information
Variance Inflation Factor One over one minus the R squared from regressing that predictor on the others
What Collinearity Breaks Precision: standard errors inflate and individually significant effects vanish
What It Does Not Break Bias, prediction, R squared, the residual standard error and the overall F
Responses to Collinearity Drop one, combine them, collect better data, or stop interpreting them singly
The Overall F Test Often significant when no individual coefficient is, which is the tell for collinearity