21  Chi-Square Tests

Chapter 17 compared two groups on a measurement: average transaction value, average score. Plenty of outcomes are not measurements at all. Did the customer churn or not? Which of three payment methods did they use? Which region are they in?

For categorical outcomes there is no mean to compare, so the t-test has nothing to work with. What you have instead is a table of counts, and the question becomes whether the counts depart from what some hypothesis predicts. The chi-square tests answer exactly that, and they run on a single idea introduced back in Chapter 9: compare what you observed against what you would expect.

This chapter also settles a question Chapter 10 left hanging, and the answer is not the one that chapter implied.

21.1 Observed Against Expected

Every test in this chapter uses one statistic. For each cell of a table, take the difference between the observed count \(O\) and the expected count \(E\), square it so that shortfalls and excesses both count, divide by \(E\) so the comparison is proportionate, and add up across all cells:

\[\chi^{2} = \sum \frac{(O - E)^{2}}{E}\]

Dividing by \(E\) is what makes the statistic sensible: being 10 short of an expected 20 is a serious discrepancy, while being 10 short of an expected 5,000 is nothing. The squaring is why the test is always one-tailed. Departures in any direction make \(\chi^2\) larger, so only the upper tail is evidence against the null, and a small \(\chi^2\) means the data sits close to what was expected.

21.2 Goodness-of-Fit: One Categorical Variable

The goodness-of-fit test asks whether one categorical variable follows a claimed distribution.

\[H_0: \text{the proportions are as specified} \qquad H_1: \text{at least one differs}\]

Expected counts come straight from the claim: \(E_i = n \times p_i\). With \(k\) categories the degrees of freedom are

\[df = k - 1\]

losing one because the counts must sum to \(n\), so once you know \(k-1\) of them the last is fixed.

Example

A store believes its customers arrive evenly across the six working days. Over one week it records 240 visits: 30, 35, 40, 42, 45 and 48.

Under the null each day expects \(E = 240/6 = 40\):

\[\chi^{2} = \frac{(30-40)^2}{40} + \frac{(35-40)^2}{40} + \frac{(40-40)^2}{40} + \frac{(42-40)^2}{40} + \frac{(45-40)^2}{40} + \frac{(48-40)^2}{40}\]

\[= 2.50 + 0.625 + 0 + 0.10 + 0.625 + 1.60 = 5.45\]

With \(df = 5\) the critical value at 5% is 11.07 and \(p = 0.36\). Footfall this uneven is entirely ordinary under an even-arrival model, so the data gives no reason to abandon it.

21.3 Independence: Two Categorical Variables

The test of independence is the one most used in practice. It takes a contingency table of the kind built in Chapter 2 and asks whether the two classifications are related.

\[H_0: \text{the two variables are independent} \qquad H_1: \text{they are associated}\]

The expected counts follow from the definition of independence in Chapter 9. If row and column are independent, then \(P(\text{row } i \text{ and column } j) = P(\text{row } i) \times P(\text{column } j)\), and multiplying by \(n\) gives

\[E_{ij} = \frac{(\text{row } i \text{ total}) \times (\text{column } j \text{ total})}{n}\]

For an \(r \times c\) table:

\[df = (r-1)(c-1)\]

That degrees-of-freedom formula is worth a moment. Once the row and column totals are fixed, only \((r-1)(c-1)\) cells are free to vary; the rest are forced by arithmetic. In a 2 by 2 table that is a single cell, which is why \(df = 1\) no matter how large \(n\) is.

21.4 The Question Chapter 10 Left Open

Example

Chapters 9 and 10 used this table of 200 customers throughout:

Satisfied Not Satisfied Total
Online 70 30 100
In-Store 60 40 100
Total 130 70 200

Chapter 10 observed that \(P(\text{Satisfied} \mid \text{Online}) = 0.70\) while \(P(\text{Satisfied}) = 0.65\), and concluded the two were dependent. Let us now test that properly.

Expected counts under independence:

\[E_{\text{Online, Satisfied}} = \frac{100 \times 130}{200} = 65 \qquad E_{\text{Online, Not}} = \frac{100 \times 70}{200} = 35\]

and the same for the In-Store row. So:

\[\chi^{2} = \frac{(70-65)^2}{65} + \frac{(30-35)^2}{35} + \frac{(60-65)^2}{65} + \frac{(40-35)^2}{35} = 2.198\]

With \(df = (2-1)(2-1) = 1\), the critical value is 3.841 and \(p = 0.138\).

The test does not reject independence.

This deserves to be sat with, because it looks like a contradiction and is not.

Chapter 10 was doing descriptive work. Within those 200 customers, satisfaction really did differ by channel; that is an arithmetic fact about the sample and nothing here overturns it. This chapter is doing inferential work, and it asks a harder question: is the gap large enough to conclude that channel and satisfaction are associated in the population the sample came from? At \(p = 0.138\), the answer is that a gap this size turns up easily in samples drawn from a population where the two are entirely independent.

The general lesson is the one that separates Modules I to II from Module IV onwards. A pattern in a sample is not evidence of a pattern in the population until it has survived a test. Five percentage points across 200 customers is well inside what chance produces.

Note also what the non-significant result does not license. It does not show the two are independent, only that this sample does not establish otherwise. Scale the same table up in the interactive above: at four times the counts, with the proportions completely unchanged, \(\chi^2\) rises from 2.20 to 8.79 and \(p\) falls to 0.003. The association was always the same size. Only the evidence changed.

21.5 Where the Association Lives

A significant \(\chi^2\) tells you the table as a whole departs from independence. It does not tell you where. For that, look inside the sum at the contribution of each cell, scaled to a comparable size. The standardised residual is

\[r_{ij} = \frac{O_{ij} - E_{ij}}{\sqrt{E_{ij}}}\]

which is positive where a cell holds more than independence predicts and negative where it holds fewer. Squaring every residual and adding them back up returns \(\chi^2\) exactly, so the residuals are a decomposition of the statistic, not a separate calculation.

As a working guide, \(|r| > 2\) marks a cell worth a comment and \(|r| > 3\) one that is doing most of the work.

Example

Six hundred customers classified by age group and payment method, the last scenario in the interactive above:

Card Cash Wallet Total
Under 30 70 30 100 200
30 to 50 95 55 50 200
Over 50 85 95 20 200
Total 250 180 170 600

Here \(\chi^2 = 97.28\) on \(df = 4\), so independence is rejected emphatically. The residuals say what that means:

Card Cash Wallet
Under 30 −1.46 −3.87 +5.76
30 to 50 +1.28 −0.65 −0.89
Over 50 +0.18 +4.52 −4.87

Almost the whole statistic sits in two columns. Under-30s use wallets far more than independence predicts and cash far less; over-50s do the reverse. The card column is close to unremarkable, and the middle age group is unremarkable across the board. A report that stopped at “age and payment method are associated, \(p < 0.001\)” would have thrown away every one of those findings.

A technical caveat on the \(\pm 2\) rule. Under the null the Pearson residual above has variance somewhat less than one, so judging it against a standard Normal is conservative. The adjusted standardised residual fixes this by dividing by \(\sqrt{E_{ij}(1 - p_{i\cdot})(1 - p_{\cdot j})}\) instead, and is approximately standard Normal. R reports it as chisq.test(x)$stdres. For the table above the adjusted residuals are larger throughout, with the wallet cell reaching 8.33.

There is a small piece of arithmetic worth knowing here. In a 2 by 2 table all four adjusted residuals have the same magnitude, and that magnitude squared is \(\chi^2\). For the Chapter 9 customers each one is 1.482, and \(1.482^2 = 2.198\).

21.6 How Big Is the Association?

The statistic itself is useless as a measure of strength, because it grows with \(n\). Scale a table up and \(\chi^2\) scales with it, exactly as the interactive shows. The standard remedy divides the scaling back out. Cramér’s V is

\[V = \sqrt{\frac{\chi^{2}}{n \times \min(r - 1,\ c - 1)}}\]

which runs from 0 at perfect independence to 1 at perfect association, and is unchanged by multiplying every count by the same factor. In a 2 by 2 table the \(\min(\cdot)\) term is 1 and \(V\) reduces to the phi coefficient, \(\phi = \sqrt{\chi^2 / n}\), which is also the ordinary correlation between the two variables coded as 0 and 1.

Cohen’s conventions for \(V\) depend on the size of the table, because the same amount of association is spread over more cells as the table grows. Writing \(df^{*} = \min(r-1,\ c-1)\):

\(df^{*}\) Small Medium Large
1 0.10 0.30 0.50
2 0.07 0.21 0.35
3 0.06 0.17 0.29

These are guidance, not law, and the same warning attached to Cohen’s \(d\) in Chapter 17 applies here: a \(V\) of 0.10 on a variable worth millions may matter more than a \(V\) of 0.40 on one that does not.

The two worked examples sit on opposite sides of that table. Channel and satisfaction give \(V = 0.105\), barely into “small”, which is the honest summary of a five-point gap. Age and payment method give \(V = 0.285\) on \(df^{*} = 2\), comfortably past “medium” and close to “large”.

21.7 Assumptions, and Cochran’s Rule

Three conditions carry the chi-square tests, and the first is the one most often broken.

The table must contain counts. Not percentages, not proportions, not means, not currency. Running \(\chi^2\) on a table of percentages is equivalent to claiming \(n = 100\) and will produce a p-value that is simply wrong. If the data arrives as percentages, multiply back to counts before testing, and if the base is unknown, the test cannot be run at all.

Each observation appears in exactly one cell. The table must classify \(n\) independent units. Counting the same customer in two rows because they bought twice, or pooling repeated measurements from the same machine, breaks the test in the same way that clustered data broke the t-test in Chapter 17.

The expected counts must not be too small. The reference distribution is an approximation that improves as expected counts grow, and it degrades badly when they are tiny.

The rule for “too small” is more permissive than the one usually taught. The version most textbooks state, that every expected count must be at least 5, is stricter than Cochran’s (1954) rule, which is:

No more than 20 per cent of expected counts below 5, and no expected count below 1.

In a 2 by 2 table 20 per cent of four cells is less than one cell, so there the strict reading and Cochran’s rule coincide: all four expected counts should reach 5. In a 5 by 4 table, however, four of the twenty cells may fall below 5 without difficulty.

Two details are easy to get wrong. The rule constrains expected counts, not observed ones; an observed zero is perfectly acceptable as long as the cell expects enough. And when a large table fails the rule, the usual fix is to collapse sparse categories into a sensible “other” group, decided on substantive grounds and, ideally, before seeing the data.

Pull the sample size slider in the interactive down to the bottom. The proportions never change, so Cramér’s V stays where it was, but the expected counts fall through 5 and the widget flags Cochran’s rule. R flags it too, printing Chi-squared approximation may be incorrect whenever an expected count drops below 5. That warning is not decoration. Take it as an instruction to switch tests.

21.8 Yates’ Correction, and Why R Turns It On

The observed counts are whole numbers, while the chi-square distribution is continuous. Yates’ continuity correction patches the mismatch in 2 by 2 tables by shrinking every discrepancy by a half before squaring:

\[\chi^{2}_{\text{Yates}} = \sum \frac{\left(|O - E| - 0.5\right)^{2}}{E}\]

It always produces a smaller statistic and a larger p-value. On the Chapter 9 table it moves \(\chi^2\) from 2.198 to 1.780 and \(p\) from 0.138 to 0.182.

The correction is contested, and has been for decades. Its aim is to approximate Fisher’s exact test, which conditions on both sets of margins, and it does so well. The objection is that in most applications the margins are not fixed by the design, they are themselves random, and against that target the corrected test is markedly conservative: it rejects less often than its stated 5 per cent, which costs real power. Sokal and Rohlf recommended against routine use as early as 1981, and Campbell’s 2007 review of the evidence recommended an \(N-1\) variant of the ordinary test in its place for all but the smallest tables.

The practical trap is that R applies the correction by default. chisq.test(x) on a 2 by 2 table silently returns the corrected statistic, so a reader comparing R’s output to a hand calculation will find they disagree, and neither is arithmetically wrong. Pass correct = FALSE for the textbook statistic. Python’s scipy.stats.chi2_contingency has the same default, under the argument name correction.

The defensible positions are to use the uncorrected test when expected counts are comfortable, and Fisher’s exact test when they are not. Yates’ correction occupies an awkward middle that neither of those needs.

21.9 Fisher’s Exact Test

When the expected counts are too small for the approximation, stop approximating. Fisher’s exact test holds both sets of margins fixed and enumerates every table that could have produced them, computing the probability of each from the hypergeometric distribution. The p-value is the total probability of all tables at least as extreme as the one observed. No reference distribution is involved and no large-sample argument is needed, so the result is exact at any sample size.

Three things to know before reaching for it.

It is not only for small samples. Modern software runs Fisher’s test on a 2 by 2 table of any size, and on larger tables by network algorithms or Monte Carlo. Small samples are simply where it is the only defensible option.

Its conditioning is the same one that makes Yates conservative. Fisher’s test fixes both margins; when the design did not fix them, the test inherits the same conservatism. Where expected counts are healthy the ordinary uncorrected chi-square remains the better-calibrated choice.

Its agreement with Yates is not a coincidence. On the Chapter 9 table Fisher gives \(p = 0.1819\) against Yates’ 0.1821. The correction was built to approximate exactly this, and on comfortable data it succeeds.

21.10 What the Test Does Not Tell You

A rejected null in this chapter licenses one sentence: the two classifications are associated in the population. Four things it does not license.

Causation. An association between region and churn says nothing about which causes which, or whether a third variable drives both. Chapter 10’s warning about conditional probability applies here unchanged, and the whole of the discussion of confounding from Chapter 13 sits behind it.

Direction or pattern. Chi-square is a single number summarising departure in any direction. Whether satisfaction rises or falls with the channel is a question for the residuals and the percentages, not for the statistic.

Importance. With a large enough \(n\) any association at all becomes significant, including one small enough to be of no consequence to anybody. This is why Cramér’s V belongs in the report beside the p-value, not instead of it and not omitted.

Order. The test treats every category as unordered. Applied to a table whose rows run “strongly disagree” through “strongly agree”, it throws the ordering away and tests a weaker hypothesis than the one you almost certainly meant. For ordered categories the linear-by-linear association test, and for an ordered exposure against a binary outcome the Cochran-Armitage trend test, use the ordering and are considerably more powerful for detecting a trend.

Three things in that output are worth carrying away. R’s default for a 2 by 2 table is not the statistic this chapter derived, so correct = FALSE has to be typed deliberately. The residuals for the payment table localise an association that the single p-value only asserts. And the small table, holding exactly the proportions of the large one, returns a p-value that the software itself warns you not to trust.

Recap

Chapters 17 and 18 ask one question of two kinds of outcome. When the outcome is a measurement, comparison runs through means, and the t-test with Welch’s correction is the default tool. When the outcome is a category there is no mean to compare, so comparison runs through counts, and \(\chi^2\) measures how far the observed table sits from the one independence would produce. Both chapters end in the same place: a p-value is not a finding on its own. Cohen’s \(d\) and Cramér’s \(V\) say how large the effect is, the group means and the residuals say where and in which direction it lies, and the assumptions decide whether any of it can be believed at all. The next part lifts a restriction that has been in force since Chapter 16, that a comparison involves exactly two things. Three or more groups at once introduces a genuinely new problem, since testing every pair separately inflates the error rate far past the level anybody agreed to, and analysis of variance is the answer to it.


Summary

Concept Description
The Statistic
The Chi-Square Statistic The sum over all cells of the squared gap between observed and expected, divided by expected
Why the Expected Count Divides It makes a shortfall of ten serious against an expectation of twenty and trivial against five thousand
Why the Test Is One-Tailed Squaring makes departures in every direction enlarge the statistic, so only the upper tail is evidence
Goodness of Fit
Goodness-of-Fit Test Asks whether one categorical variable follows a claimed set of proportions
Expected Counts Under a Claim Each expected count is the sample size multiplied by the proportion the null claims for that category
Degrees of Freedom, One Variable Categories minus one, because the counts are forced to sum to n
Independence
Test of Independence Asks whether two classifications of the same units are related in the population
Expected Counts Under Independence Row total times column total divided by n, which is the multiplication rule of Chapter 9 in table form
Degrees of Freedom, Two Variables Rows minus one times columns minus one, which is one for any 2 by 2 table however large n is
Descriptive Against Inferential A gap within the sample is arithmetic; whether it holds in the population is what the test decides
A Non-Significant Result Failure to reject is not evidence of independence, only an absence of evidence against it
Where the Association Lives
Standardised Residuals The gap between observed and expected in a cell, divided by the square root of the expected count
Reading the Residuals Roughly, past two is worth a comment and past three is doing most of the work
Adjusted Standardised Residuals A rescaling that is approximately standard Normal under the null, reported by R as stdres
How Big It Is
Cramer's V The square root of chi-square over n times the smaller of rows minus one and columns minus one
The Phi Coefficient The 2 by 2 case of V, equal to the correlation between the two variables coded as zero and one
Effect Size Conventions for V Thresholds fall as the table grows, from 0.10, 0.30 and 0.50 at df of one to 0.06, 0.17 and 0.29 at three
Effect Size Is Independent of n Scaling every count leaves V unchanged while chi-square and the p-value move a long way
Assumptions
Counts, Not Percentages The test requires raw frequencies; run on percentages it silently assumes a sample size of one hundred
One Observation, One Cell Each unit must be classified once, or the table overstates the evidence exactly as clustering does
Cochran's Rule No more than a fifth of expected counts below five, and none below one, which is milder than the usual rule
Expected, Not Observed The rule constrains expected counts, so an observed zero is no problem where the cell expects enough
Collapsing Sparse Categories The usual repair for a sparse table, decided on substantive grounds and before seeing the data
Corrections and Exact Tests
Yates' Continuity Correction Shrinks every discrepancy in a 2 by 2 table by a half before squaring, always raising the p-value
The Case Against Yates It targets a test that conditions on margins the design did not fix, and is conservative as a result
What R Does by Default chisq.test applies Yates to 2 by 2 tables unless correct = FALSE is passed, as does SciPy
Fisher's Exact Test Enumerates every table with the same margins and sums the probabilities of those at least as extreme
When Fisher Is the Right Choice When expected counts are too small for the approximation, where it is exact at any sample size
What the Test Does Not Say
Association Is Not Causation A table shows that two classifications move together, never which one moves the other
Significance Is Not Importance Large n makes trivial associations significant, which is why V belongs beside the p-value
Ordered Categories Chi-square discards order; a linear-by-linear or Cochran-Armitage test uses it and finds trends sooner