25  Correlation

Every test from Chapter 16 to Chapter 20 needed groups. Two of them, three, or a two-way grid, but always groups, and always a measured outcome compared across them. A great many questions do not come in that shape.

Does marketing spend move with revenue? Does delivery time move with satisfaction? Does a store’s size move with its basket value? Both variables are measured, neither divides the data into categories, and cutting one into bands so that ANOVA can be run on it throws away most of what it knows.

This chapter takes the question directly. Correlation reduces the relationship between two measured variables to a single number between \(-1\) and \(+1\). That compression is the point of it and also the danger of it, because a great deal can be true of a scatterplot that no single number will report.

25.1 What r Measures

Start with covariance, which asks whether the two variables depart from their means together:

\[\operatorname{cov}(x, y) = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{n - 1}\]

When a point sits above the mean on both variables, or below on both, its contribution is positive. When it is above on one and below on the other, its contribution is negative. Add them up and the sign tells you which way the cloud leans.

Covariance is unusable on its own because it carries the units of both variables multiplied together. Rupees times minutes means nothing, and rescaling either variable changes the number without changing the relationship. Dividing by both standard deviations fixes this, and the result is Pearson’s correlation coefficient:

\[r = \frac{\operatorname{cov}(x, y)}{s_x s_y} = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum (x_i - \bar{x})^2 \sum (y_i - \bar{y})^2}}\]

Four properties follow, and all four matter.

It is bounded. \(-1 \le r \le +1\), with the extremes reached only when every point lies exactly on a straight line.

It is dimensionless. Multiply spend by a thousand, convert revenue to another currency, or add a constant to either, and \(r\) does not move. It measures shape, not scale.

It is symmetric. \(r(x, y) = r(y, x)\). Correlation has no notion of which variable explains which, which is the first hint that it cannot establish causation.

It measures straightness only. \(r\) is the degree to which the cloud approximates a straight line. A perfect but curved relationship can give an \(r\) of zero, and the third tab of the interactive above shows exactly that.

The conventional verbal labels, roughly 0.1 small, 0.3 moderate and 0.5 large in the social sciences, are worth no more than the effect size conventions of Chapter 17. What counts as a strong correlation is entirely a matter of field. An \(r\) of 0.3 between a cheap marketing signal and revenue may be worth a great deal of money; an \(r\) of 0.95 between two instruments that are supposed to measure the same physical quantity may be a failure.

25.2 r Squared, and Why It Is the Honest Number

Squaring the correlation gives the coefficient of determination:

\[r^{2} = \text{the proportion of the variance in } y \text{ that moves with } x\]

This is the quantity that behaves the way people expect \(r\) to behave, and the gap between them is wide enough to mislead.

\(r\) \(r^2\) Variance left unaccounted for
0.3 0.09 91%
0.5 0.25 75%
0.7 0.49 51%
0.9 0.81 19%

A correlation of 0.5 sounds like half of something. It is a quarter. A correlation of 0.7, which most people would call strong, still leaves more than half the variation in \(y\) unexplained by \(x\). Reporting \(r^2\) alongside \(r\) is the cheapest honesty available in this chapter.

25.3 A Worked Example

Example

Forty stores, monthly marketing spend and monthly revenue, both in thousands.

Mean Standard deviation
Spend 18.00 4.50
Revenue 214.00 31.00

The covariance is 86.50, in units of thousands multiplied by thousands, which is why nobody reports it. Dividing by the two standard deviations:

\[r = \frac{86.50}{4.50 \times 31.00} = 0.620\]

so \(r^2 = 0.38\). Marketing spend moves with about 38 per cent of the variance in revenue, and the other 62 per cent is doing something else.

One identity is worth carrying into the next chapter, because it is the bridge to it. The slope of the straight line fitted through this cloud is

\[b_1 = r \, \frac{s_y}{s_x} = 0.620 \times \frac{31.00}{4.50} = 4.27\]

meaning a store spending one thousand more on marketing is associated with about 4,270 more in revenue. Correlation and the regression line are the same fact expressed twice: \(r\) standardises both variables, the slope keeps their units. Chapter 22 takes the slope seriously.

25.4 Anscombe’s Quartet

In 1973 Francis Anscombe published four datasets of eleven points each. They share, to the decimals shown:

Set 1 Set 2 Set 3 Set 4
Mean of x 9.00 9.00 9.00 9.00
SD of x 3.317 3.317 3.317 3.317
Mean of y 7.50 7.50 7.50 7.50
SD of y 2.03 2.03 2.03 2.03
Pearson r 0.816 0.816 0.816 0.817
Fitted line \(3.00 + 0.50x\) \(3.00 + 0.50x\) \(3.00 + 0.50x\) \(3.00 + 0.50x\)

The four scatterplots have almost nothing in common. Set 1 is an ordinary linear relationship with scatter. Set 2 is a clean curve, with no straight-line component worth describing. Set 3 is a perfect straight line with a single point knocked well off it. Set 4 has ten points stacked at one value of \(x\) and one point far away, and that single point creates the entire correlation: delete it and \(r\) is undefined, because \(x\) has no variance left.

Open the second tab of the interactive above and look at them.

Spearman’s rank correlation, introduced below, is the one summary in that table that separates the four, running from 0.500 to 0.991. Even that would not tell you which picture you were looking at.

Anscombe’s point was not that summary statistics are useless. It was that they are answers to specific questions, and that the question “what does this relationship look like?” is not one of them. Matejka and Fitzmaurice made the same point again in 2017 with the Datasaurus Dozen, a set of datasets that share mean, standard deviation and correlation to two decimals while one of them is a drawing of a dinosaur.

The instruction that follows from both is short. Plot the data first.

25.5 Testing and Estimating a Correlation

The sample \(r\) estimates a population correlation \(\rho\). The null hypothesis is that there is none:

\[H_0: \rho = 0 \qquad H_1: \rho \neq 0\]

and the test statistic is a \(t\) on \(n - 2\) degrees of freedom:

\[t = \frac{r\sqrt{n-2}}{\sqrt{1 - r^{2}}}\]

For the forty stores, \(t = 0.620\sqrt{38}/\sqrt{1 - 0.385} = 4.872\) on 38 degrees of freedom, giving \(p = 0.00002\). The correlation is not zero.

A p-value against zero is a low bar, and an interval is more use. The sampling distribution of \(r\) is skewed whenever \(\rho\) is far from zero, because \(r\) is bounded at \(\pm 1\) and piles up against the boundary. Fisher’s z transformation straightens it out:

\[z = \tfrac{1}{2}\ln\!\left(\frac{1+r}{1-r}\right) = \operatorname{arctanh}(r), \qquad SE_z = \frac{1}{\sqrt{n-3}}\]

Build the interval on the \(z\) scale, where the distribution is approximately Normal, then transform the endpoints back with \(\tanh\). For the forty stores, \(z = 0.725\) and \(SE_z = 0.164\), giving a 95 per cent interval for \(\rho\) of

\[(0.382,\ 0.781)\]

Note that it is not symmetric about 0.620. That asymmetry is the whole reason for the transformation.

That interval is wider than most people expect from forty observations, and small samples are far worse. Four thousand samples drawn from a population where \(\rho\) is exactly 0.50 give:

\(n\) Middle 95% of sample correlations Full range seen
10 \(-0.17\) to \(0.86\) \(-0.71\) to \(0.96\)
20 \(0.11\) to \(0.78\) \(-0.48\) to \(0.89\)
50 \(0.25\) to \(0.69\) \(0.04\) to \(0.77\)
200 \(0.39\) to \(0.60\) \(0.29\) to \(0.67\)

With ten observations, a true correlation of 0.50 routinely produces sample correlations anywhere from slightly negative to nearly perfect. A reported \(r\) of 0.8 from \(n = 10\) is close to uninformative, and a reported \(r\) of 0.1 from the same \(n\) does not rule out a strong relationship.

Pull the sample size slider in the first tab of the interactive above and watch the interval breathe.

25.6 Rank Correlation

Pearson’s \(r\) assumes the relationship is straight and is badly disturbed by outliers, because it works with the values themselves. Replacing the values with their ranks fixes both problems at once.

Spearman’s \(\rho_s\) is Pearson’s \(r\) computed on the ranks. It measures monotonic association, whether \(y\) consistently rises or falls as \(x\) rises, without requiring the rise to be straight. It needs only ordinal data, which makes it the right choice for the Likert-type scales of Chapter 1.

Kendall’s \(\tau\) takes a different route, comparing every pair of observations and asking whether they are concordant, ordered the same way on both variables, or discordant:

\[\tau = \frac{(\text{concordant pairs}) - (\text{discordant pairs})}{\text{total pairs}}\]

Kendall’s \(\tau\) is smaller than Spearman’s \(\rho_s\) on the same data, typically by a third, so the two are not interchangeable as numbers even when they agree as verdicts. It handles ties better, has a more direct interpretation as a probability, and is preferred for small samples.

For the forty stores, Pearson gives 0.620, Spearman 0.567 and Kendall 0.405. All three say the same thing. When they disagree sharply, that disagreement is information: Pearson much larger than Spearman usually means an outlier is doing the work, and Spearman much larger than Pearson usually means the relationship is monotonic but curved.

Anscombe’s third set is the clearest case. Pearson reads 0.816 and Spearman reads 0.991, because in ranks the one displaced point is merely out of order rather than far away.

25.7 What Breaks r

Four failure modes account for most misread correlations. The third tab of the interactive above has one panel for each.

A single outlier. Twenty points of pure noise give \(r = -0.145\). Add one point out at \((10, 10)\) and \(r\) becomes \(+0.861\); move that same point to \((10, -10)\) and \(r\) becomes \(-0.897\). One observation in twenty-one swings the answer across nearly the whole scale, while Spearman stays at \(+0.027\) because ranks barely notice how far away a point is. Any correlation computed on fewer than a few dozen points should be plotted before it is believed.

Restricted range. If you can only observe part of the range of \(x\), \(r\) shrinks even though the relationship is unchanged. A correlation of 0.577 across the full range falls to 0.247 when only the middle half of \(x\) is observable, and to 0.095 on the middle quarter. This is why the correlation between an entrance test and later performance, computed among people who were admitted on that test, systematically understates how well the test works. The same applies to hired applicants, approved loans and surviving firms.

Curvature. Pearson measures the straight-line part of a relationship and nothing else. A tight parabolic relationship returns an \(r\) near zero, and so does Spearman if the curve is not monotonic. Neither number is wrong; they are answering a question the data does not fit.

Subgroups. Pooled across three groups the cloud can slope firmly upward while sloping downward inside every single group. This is Simpson’s paradox from Chapter 10, arriving as a correlation instead of a table. The pooled figure is not a compromise between the groups. It has the wrong sign.

A fifth deserves its own paragraph because it biases in a predictable direction. Measurement error attenuates correlation. If \(x\) and \(y\) are each measured with reliability \(\rho_{xx}\) and \(\rho_{yy}\), the correlation you observe is

\[r_{\text{observed}} = r_{\text{true}} \sqrt{\rho_{xx}\,\rho_{yy}}\]

Noise in the measurements can only push the observed correlation toward zero, never away from it. Two variables each measured with 80 per cent reliability will show a correlation 20 per cent smaller than the truth: a real \(\rho\) of 0.62 would be observed as about 0.50.

The formula can be inverted to correct for attenuation, but the correction is only as good as the reliability estimates it uses and can produce impossible values above 1 when they are poor. Its more useful role is as a reminder: a modest correlation between two noisily measured things is consistent with a strong relationship between the things themselves.

25.8 Correlation Is Not Causation

The slogan is the most repeated sentence in statistics and the least often acted on. It is worth setting out exactly what else could produce a correlation between \(x\) and \(y\).

A third variable causes both. Ice cream sales correlate with drowning deaths; summer causes both. In business data the third variable is usually size. Large stores spend more on marketing and earn more revenue, so spend and revenue correlate across stores even if the marketing does nothing at all.

The causation runs the other way. Firms that are already growing can afford larger marketing budgets. The correlation is identical whichever direction the arrow points, because \(r(x, y) = r(y, x)\).

Selection created it. The restricted-range panel above is one version. Another is conditioning on a common effect: among people already hired, ability and interview performance can correlate negatively simply because anyone weak on both was never hired.

Coincidence. With enough variables, some pairs correlate strongly by chance. This is the multiple-comparisons problem of Chapter 19 arriving through a correlation matrix, where 20 variables produce 190 correlations and roughly ten of them clear the 5 per cent bar with nothing behind them.

What would license a causal claim? Randomisation, as Chapter 13 described, because assigning \(x\) at random severs any path from a confounder into it. Failing that, a design that measures and adjusts for the plausible confounders, a clear temporal order, a mechanism that makes the claim sensible, and consistency across settings where the confounding structure differs.

None of that comes from \(r\). The correlation is where a causal investigation starts, not where it ends.

Three things in that output are worth carrying away. The four Anscombe sets agree on every summary statistic anyone would normally print, including the fitted line, and they are not remotely the same data. A single point added to twenty swings r from roughly zero to either end of the scale, and the rank correlation, which only knows the ordering, hardly registers it. And the slope computed as \(r \, s_y / s_x\) is exactly the slope least squares returns, which is the subject of the next chapter arriving early.

Looking Ahead

Correlation answers one question well and refuses every other question asked of it. It reports how tightly two measured variables track a straight line, on a scale that does not care about units, and that is genuinely useful. What it cannot do is say which variable moves which, predict a value of \(y\) from a value of \(x\), quantify by how much \(y\) changes per unit of \(x\), or hold a third variable constant while examining the first two. Every one of those is the job of regression, and the identity at the end of the worked example is the door into it: the fitted slope is nothing more than \(r\) rescaled from standard deviations back into the units the data came in. The next chapter takes that line seriously. It asks where the line comes from, what least squares is actually minimising, how much of the scatter the line accounts for, what the residuals reveal when the fit is wrong, and how to attach an interval to a prediction rather than a point.


Summary

Concept Description
What r Is
Covariance Whether two variables depart from their means together, in the units of both multiplied
Why Covariance Is Unusable Alone It carries the product of both units and changes when either variable is rescaled
Pearson's r Covariance divided by both standard deviations, which removes the units
Bounded It runs from minus one to plus one, reached only when every point lies on a line
Dimensionless Rescaling or shifting either variable leaves it unchanged
Symmetric r of x with y equals r of y with x, so it cannot say which causes which
Straight Lines Only It measures the straight-line part of an association and nothing else
Against Verbal Labels What counts as strong is a matter of field, not of a fixed table
r Squared and the Line
r Squared The share of the variance in y that moves with x, and the honest number to report
The Gap Between r and r Squared An r of 0.5 is a quarter of the variance, not a half
Slope From r The fitted slope is r times the ratio of the two standard deviations
One Number, Four Pictures
Anscombe's Quartet Four datasets sharing every summary statistic and looking nothing alike
The Datasaurus The same point made again in 2017, with a dataset shaped like a dinosaur
Plot the Data The one instruction that follows from both, and the cheapest safeguard available
Inference
Testing Whether rho Is Zero A t test on n minus two degrees of freedom, using r and the sample size
Fisher's z Transformation The arctanh of r, whose sampling distribution is approximately Normal
Why the Interval Is Asymmetric r is bounded, so its sampling distribution is skewed whenever rho is far from zero
Instability at Small n At n of ten, a true rho of 0.50 produces sample values from below zero to above 0.9
Rank Correlation
Spearman's Rank Correlation Pearson computed on the ranks; measures monotonic association, not straight-line
Kendall's Tau Concordant minus discordant pairs over all pairs; smaller than Spearman on the same data
When the Three Disagree Pearson above Spearman suggests an outlier; Spearman above Pearson suggests a curve
What Breaks r
A Single Outlier One point in twenty-one can move r from zero to either end of the scale
Restricted Range Observing only part of the range of x shrinks r without the relationship changing
Selection on the Predictor Correlations among those admitted, hired or approved understate the true association
Curvature A tight parabolic relationship returns an r near zero, and Spearman does no better
Subgroups Pooled the cloud can slope up while every group inside it slopes down
Attenuation by Measurement Error Measurement noise pushes the observed correlation toward zero, never away
Correcting for Attenuation Dividing by the square root of the two reliabilities, only as good as those estimates
Not Causation
A Third Variable A confounder causes both, and in business data that variable is usually size
Reverse Causation The arrow may point the other way, and r is identical either way
Selection Effects Conditioning on a common effect can manufacture a correlation from nothing
Coincidence in a Correlation Matrix Twenty variables give 190 correlations, about ten of which clear 5 per cent by chance
What Would License a Causal Claim Randomisation, or measured confounders, temporal order, mechanism and replication