22 One-Way ANOVA
Every test since Chapter 16 has compared at most two things. One mean against a claimed value, two means against each other, two classifications against independence. Business rarely stops at two. There are four store formats, not two. Five suppliers. Three training programmes, six regions, a dozen product lines.
The obvious extension is to run the two-group test of Chapter 17 on every pair and see which come back significant. That approach is not merely inelegant. It is wrong, and it is wrong in a way that gets worse the more groups you have.
This chapter explains why, and gives the test that replaces it: the one-way analysis of variance, which compares any number of group means with a single test, and asks its question by comparing two variances rather than any means at all.
22.1 Why Not Just Run All the t-Tests
With \(k\) groups the number of distinct pairs is
\[\binom{k}{2} = \frac{k(k-1)}{2}\]
so three groups need 3 tests, four need 6, and six need 15. Each is run at \(\alpha = 0.05\), meaning each has a 5 per cent chance of firing when nothing is there. The quantity that matters is not the error rate of any one test but the family-wise error rate: the chance that at least one of them fires by accident.
If the tests were independent, the arithmetic would be
\[\text{FWER} = 1 - (1 - \alpha)^{m}\]
for \(m\) comparisons, giving 14.3 per cent at three groups and 26.5 per cent at six comparisons.
The tests are not independent, because every pair is computed from the same groups and, if pooled, the same variance estimate. So that formula overstates the damage. It is worth knowing by how much, and the honest way to find out is to simulate it: generate \(k\) groups of 15 observations with no real difference anywhere, test every pair, and count how often at least one comes back significant. Twenty thousand repetitions give this:
| Groups | Pairwise tests | Formula | Actual | One F test |
|---|---|---|---|---|
| 3 | 3 | 14.3% | 13.0% | 5.0% |
| 4 | 6 | 26.5% | 22.2% | 5.1% |
| 6 | 15 | 53.7% | 40.0% | 5.0% |
The formula is optimistic about the correction and pessimistic about the damage, but the damage is real and it is large. With six groups and nothing whatsoever going on, the all-pairs approach finds something to report two times in five. The F test of this chapter holds at the 5 per cent that was actually asked for, whatever \(k\) is.
Run the simulation yourself in the second tab of the interactive above.
This is the same problem that Chapter 16 raised under the heading of multiple analyses, now arriving by a route that looks entirely innocent. Nobody running six t-tests feels they are fishing. They are simply answering the six questions that the data suggests. The error rate does not care about the analyst’s intentions; it only counts the tests.
22.2 What Analysis of Variance Actually Compares
The name is the single most confusing thing about the method, so it is worth settling immediately. ANOVA tests a hypothesis about means:
\[H_0: \mu_1 = \mu_2 = \cdots = \mu_k \qquad H_1: \text{at least one differs}\]
It does so by comparing variances, and the idea behind that is neat. Suppose the null is true and all groups come from one population with variance \(\sigma^2\). Then there are two independent ways to estimate \(\sigma^2\) from the data.
The first looks inside the groups. Pool the scatter of observations around their own group means. This estimates \(\sigma^2\) whether the null is true or not, because shifting a group’s mean does nothing to the spread around it.
The second looks between the groups. If all groups share one mean, the observed group means are just \(k\) sample means from one population, and their scatter estimates \(\sigma^2/n\), so multiplying by \(n\) estimates \(\sigma^2\). But if the groups have genuinely different means, that scatter contains the real differences as well, and the estimate comes out too large.
So: the within estimate is always honest, and the between estimate is honest only under the null and inflated otherwise. Their ratio is the test.
Formally, the total variation is partitioned. Writing \(\bar{x}\) for the grand mean and \(\bar{x}_j\) for the mean of group \(j\):
\[\underbrace{\sum_{j}\sum_{i} (x_{ij} - \bar{x})^{2}}_{SST} = \underbrace{\sum_{j} n_j (\bar{x}_j - \bar{x})^{2}}_{SSB} + \underbrace{\sum_{j}\sum_{i} (x_{ij} - \bar{x}_j)^{2}}_{SSW}\]
Every observation’s distance from the grand mean splits cleanly into how far its group sits from the grand mean, plus how far it sits from its own group. The degrees of freedom split the same way:
\[\underbrace{N - 1}_{\text{total}} = \underbrace{(k - 1)}_{\text{between}} + \underbrace{(N - k)}_{\text{within}}\]
Dividing each sum of squares by its degrees of freedom gives a mean square, which is just a variance, and the test statistic is their ratio:
\[F = \frac{MSB}{MSW} = \frac{SSB / (k-1)}{SSW / (N-k)}\]
Under the null this sits near 1. The larger it gets, the harder the null is to believe, and, as with \(\chi^2\), only the upper tail counts.
The first tab of the interactive above is this sentence made draggable. Pull the four formats apart and \(MSB\) climbs while \(MSW\) does not move at all. Increase the scatter inside each format and \(MSW\) climbs while \(MSB\) stays put. \(F\) is significant only when the gaps between the groups are large relative to the noise inside them, which is why neither number means anything on its own.
22.3 The ANOVA Table
Every piece of software reports the same five-column table, and reading it is most of the skill.
| Source | Sum of squares | df | Mean square | F |
|---|---|---|---|---|
| Between groups | \(SSB\) | \(k-1\) | \(MSB = SSB/(k-1)\) | \(MSB/MSW\) |
| Within groups | \(SSW\) | \(N-k\) | \(MSW = SSW/(N-k)\) | |
| Total | \(SST\) | \(N-1\) |
Two arithmetic checks are free and worth making every time. The sums of squares must add up, and so must the degrees of freedom. If they do not, something has been mis-entered.
22.4 A Worked Example
Example
A retailer operates four store formats and wants to know whether average basket value differs between them. Thirty transactions are sampled from each, so \(N = 120\).
| Format | \(n\) | Mean | Standard deviation |
|---|---|---|---|
| Compact | 30 | 4,790 | 620 |
| Standard | 30 | 4,920 | 620 |
| Flagship | 30 | 5,360 | 620 |
| Online | 30 | 5,180 | 620 |
The grand mean is \((4790 + 4920 + 5360 + 5180)/4 = 5{,}062.5\).
Between-group sum of squares. Each group has the same \(n\), so:
\[SSB = 30\left[(-272.5)^2 + (-142.5)^2 + (297.5)^2 + (117.5)^2\right] = 30 \times 196{,}875 = 5{,}906{,}250\]
Within-group sum of squares. Each group’s variance is \(620^2 = 384{,}400\), and each contributes \((n-1)\) times its variance:
\[SSW = 4 \times 29 \times 384{,}400 = 44{,}590{,}400\]
The table.
| Source | Sum of squares | df | Mean square | F |
|---|---|---|---|---|
| Between formats | 5,906,250 | 3 | 1,968,750 | 5.122 |
| Within formats | 44,590,400 | 116 | 384,400 | |
| Total | 50,496,650 | 119 |
The critical value at \(F_{3,\,116}\) and 5 per cent is 2.683, and \(p = 0.0023\). The four formats do not all have the same average basket value.
Read that conclusion carefully, because it is weaker than it sounds and weaker than most people report it as. A significant F says only that the group means are not all equal. It does not say which ones differ, it does not say how many do, and it does not say by how much. One format far from the other three produces a significant F; so does a slow gradient across all four. Distinguishing those cases is a separate job, taken up below.
When \(k = 2\), ANOVA and the pooled t-test of Chapter 17 are the same test wearing different clothes. The relationship is exact:
\[F = t^{2}\]
and the two p-values agree to every decimal place. The F-test on two groups is Student’s pooled test, not Welch’s, which is worth remembering given the argument Chapter 17 made for preferring Welch. The code below demonstrates the identity.
22.5 How Much Does the Factor Explain?
The p-value says the formats differ. It says nothing about how much of what happens to a basket is explained by which store it happened in. For that, take the share of total variation that the factor accounts for. Eta-squared is exactly the partition already computed:
\[\eta^{2} = \frac{SSB}{SST}\]
For the store formats, \(5{,}906{,}250 / 50{,}496{,}650 = 0.117\). Store format accounts for about 12 per cent of the variation in basket value, and the remaining 88 per cent is differences between one customer and the next inside the same format. That is the bar drawn under the plot in the first tab of the interactive.
Eta-squared has a flaw that matters in small studies: it is biased upward. Even when every population mean is identical, the observed group means will differ a little by chance, so \(SSB\) is never zero and \(\eta^2\) is never zero. On average \(SSB\) picks up \((k-1) \times MSW\) of pure noise. Omega-squared subtracts it:
\[\omega^{2} = \frac{SSB - (k-1)\,MSW}{SST + MSW}\]
For the store formats this gives \(\omega^2 = 0.093\) against \(\eta^2 = 0.117\). The gap is the noise that \(\eta^2\) was counting as signal, and it widens as \(k\) grows and as \(n\) shrinks. With three groups of five, \(\eta^2\) can read 0.20 on data with no effect at all. Report \(\omega^2\) when the sample is small, and watch the two converge in the interactive as you flatten the separation slider.
Cohen’s conventions for \(\eta^2\) are 0.01 small, 0.06 medium and 0.14 large. Some software instead reports Cohen’s \(f\), which is the same information rescaled:
\[f = \sqrt{\frac{\eta^{2}}{1 - \eta^{2}}}\]
with conventions 0.10, 0.25 and 0.40. The two sets agree: an \(f\) of 0.25 is an \(\eta^2\) of 0.059, and an \(f\) of 0.40 is an \(\eta^2\) of 0.138. The store formats give \(f = 0.364\), between medium and large.
One naming trap. Software usually reports partial eta-squared, which divides \(SSB\) by \(SSB + SSW\) rather than by \(SST\). In a one-way design those are the same quantity, since \(SST = SSB + SSW\). In the two-way designs of the next chapter they are not, and the two can differ substantially.
22.6 Which Pairs Differ
A significant F licenses exactly one further step: finding out which groups are responsible. Doing that with ordinary t-tests would reintroduce the error rate this chapter began by rejecting, so the comparison needs a threshold that accounts for the whole family at once.
Tukey’s Honest Significant Difference does this by referring the largest plausible gap to the studentised range distribution, which describes how far apart the extremes of \(k\) sample means fall when all the populations are identical. Two group means differ significantly when
\[|\bar{x}_i - \bar{x}_j| > q_{\alpha,\,k,\,N-k} \sqrt{\frac{MSW}{n}}\]
The whole set of comparisons then holds at \(\alpha\), not each one separately. The name is a fair description: it is the smallest difference Tukey was prepared to call real once the number of questions had been counted.
Example
For the store formats, \(q_{0.05,\,4,\,116} = 3.686\) and \(\sqrt{MSW/n} = \sqrt{384{,}400/30} = 113.2\), so any two formats must differ by more than \(3.686 \times 113.2 = 417\) to be called apart.
| Pair | Difference | Tukey \(p\) | Unadjusted \(p\) |
|---|---|---|---|
| Standard vs Compact | +130 | 0.849 | 0.418 |
| Flagship vs Compact | +570 | 0.003 | 0.001 |
| Online vs Compact | +390 | 0.076 | 0.016 |
| Flagship vs Standard | +440 | 0.035 | 0.007 |
| Online vs Standard | +260 | 0.369 | 0.107 |
| Online vs Flagship | −180 | 0.675 | 0.263 |
Of six pairs, two survive. Flagship beats Compact and Flagship beats Standard. Everything else, including the comparison between Flagship and Online that a glance at the means might have suggested, is inside what chance produces.
The third row is the one to sit with. Online runs 390 above Compact, an unadjusted t-test calls that significant at \(p = 0.016\), and Tukey does not, at \(p = 0.076\). Nothing about the data changed between those two numbers. What changed is the question. The unadjusted test answers “is this pair different, considered alone?” Tukey answers “is this pair different, given that I looked at all six?” The second is the question that was actually asked, because the table was inspected before the pair was chosen.
This is the entire content of the multiple-comparisons problem, in one row of one table.
Tukey is the default for all-pairs comparisons with roughly equal group sizes, but it is not the only option.
Bonferroni simply tests each pair at \(\alpha/m\). It is the easiest to explain and applies to any set of tests, not just pairwise ones, but it is conservative and loses power as \(m\) grows. On the store formats it gives 0.003, 0.042 and 0.098 for the three interesting pairs against Tukey’s 0.003, 0.035 and 0.076. Same verdicts here, slightly less power. Holm’s step-down variant is uniformly better than plain Bonferroni and should be preferred to it whenever Bonferroni is being considered at all.
Scheffé allows any contrast at all, including comparisons of averages of groups, and pays for that generality by being the most conservative of the family. Use it when the comparisons were not decided in advance.
Dunnett compares every group against one control and nothing else. Because it makes \(k-1\) comparisons rather than \(k(k-1)/2\), it is markedly more powerful than Tukey when a control group is genuinely what the design is about.
Fisher’s LSD applies no correction whatsoever. It is defensible only as the protected version, run after a significant F and only with \(k = 3\); beyond that it does not control the family-wise rate at all.
Games-Howell is the one to reach for when the group variances differ, since it does not assume a common \(MSW\). It stands in the same relation to Tukey as Welch’s test does to Student’s.
22.7 Assumptions, and Welch’s ANOVA
Three conditions carry the F test, and they are the Chapter 17 list with one group added.
Independence. Observations are independent within and between groups. This remains the assumption most often broken without anyone noticing, and no test finds it for you.
Approximate Normality of the residuals. What must be roughly Normal is the scatter within groups, not the raw outcome pooled across them, which will look multimodal whenever the group means genuinely differ. As with the t-test, the Central Limit Theorem makes this the least demanding of the three once group sizes reach a couple of dozen.
Equal variances across groups. This one bites, and it bites hardest exactly where analysts are least likely to look.
The classic F test pools all the within-group scatter into a single \(MSW\), which presumes every group has the same \(\sigma^2\). Welch’s ANOVA drops that presumption, weighting each group by \(n_j / s_j^2\) and adjusting the denominator degrees of freedom downward, in exactly the spirit of the Welch t-test of Chapter 17.
How much does it matter? Simulate four groups with no real difference at all, vary the spreads and the group sizes, and count how often each test rejects at 5 per cent. On twenty thousand replications:
| Configuration | Classic F | Welch |
|---|---|---|
| Equal sizes, equal spread | 4.8% | 5.0% |
| Small group has the large spread | 26.5% | 5.1% |
| Large group has the large spread | 1.5% | 5.2% |
When the smallest group is the noisiest, the classic F test fires more than five times as often as it claims to. When the largest group is the noisiest, it fires barely a third as often, which is not safety but power quietly thrown away. And when the assumption holds, Welch costs essentially nothing. The code below runs a shorter version of this in your browser, so the figures will wobble a little around these.
The instinct at this point is to test for equal variances first, with Levene’s or Bartlett’s test, and use the result to choose. Do not. The objection is the one Chapter 17 raised against preliminary variance tests, and it has not weakened: those tests have poor power at exactly the small sample sizes where unequal variances do the most damage, so they most often fail to detect the problem precisely when it matters. They also make the final error rate conditional on a first test, which is not the error rate anybody reports.
Welch’s ANOVA holds its stated rate across all of these configurations and loses almost nothing when variances really are equal. That makes it the better default, and it pairs naturally with Games-Howell for the follow-up comparisons.
Note that R’s aov() and summary(aov()) give the classic test. Welch’s version is oneway.test(y ~ g, var.equal = FALSE), which is in fact that function’s default, so the trap runs in the opposite direction from chisq.test(): here the function most people reach for is the one that assumes equal variances, and you have to know the other exists.
When Normality is genuinely doubtful and the groups are small, the distribution-free alternative is the Kruskal-Wallis test, the several-group extension of the Mann-Whitney U test from Chapter 17. It works on ranks and tests a slightly different hypothesis. On the store formats it gives \(p = 0.0026\) against the F test’s \(0.0023\), which is the usual outcome when the assumptions were fine all along.
Three things in that output are worth carrying away. The sums of squares and the degrees of freedom each add up, which is the free check to make on any ANOVA table you are handed. Omega-squared comes in below eta-squared every time, and the gap is the noise that eta-squared was counting as signal. And the classic F test, run on data containing no difference whatsoever, calls a quarter of the samples significant as soon as the smallest group happens to be the noisiest.
Looking Ahead
One-way ANOVA answers the question this module has been building towards: do these \(k\) groups differ on this measurement, asked once, at the error rate you actually chose. What it handles is a single grouping factor. Real designs rarely stop there either. A retailer wants to know about store format and region, and, more interestingly, whether the effect of one depends on the level of the other, which is a question neither variable can answer alone. A trainer wants to compare three methods measured on the same people before and after, where the observations within a person are not independent and the paired logic of Chapter 17 has to be extended rather than abandoned. The next chapter takes both: two-way ANOVA, where the interaction between two factors becomes the most interesting term in the table, and repeated measures ANOVA, which recovers the power that pairing offered when there are more than two occasions.
Summary
| Concept | Description |
|---|---|
| Why Not Many t-Tests | |
| The Multiple Comparisons Problem | Testing every pair separately inflates the error rate far past the level anyone chose |
| Number of Pairwise Tests | k groups give k times k minus one over two pairs, so six groups need fifteen tests |
| Family-Wise Error Rate | The chance that at least one test in the family fires when nothing is there |
| Why the Formula Overstates It | One minus 0.95 to the power m assumes independence, and pairwise tests share their data |
| What One F Test Buys | A single test of all k means at once, holding the stated rate whatever k is |
| What ANOVA Compares | |
| Null Hypothesis of ANOVA | That every population mean is equal, against at least one being different |
| Variances to Test Means | The hypothesis is about means, but the evidence is a ratio of two variance estimates |
| Within-Group Estimate | Scatter of observations around their own group mean, honest whether the null holds or not |
| Between-Group Estimate | Scatter of the group means around the grand mean, inflated when the null is false |
| The Arithmetic | |
| Partition of the Sum of Squares | Total equals between plus within, exactly, for every dataset |
| Partition of Degrees of Freedom | N minus one splits into k minus one and N minus k |
| Mean Square | A sum of squares divided by its degrees of freedom, which is simply a variance |
| The F Ratio | Between-group mean square over within-group mean square, near one under the null |
| Why the Test Is One-Tailed | Departures in any direction enlarge the numerator, so only the upper tail is evidence |
| Reading the Table | |
| The ANOVA Table | Source, sum of squares, degrees of freedom, mean square and F, in five columns |
| Two Free Arithmetic Checks | The sums of squares must add up and so must the degrees of freedom |
| What a Significant F Says | Only that the means are not all equal, never which ones differ or by how much |
| F Equals t Squared | With two groups, ANOVA is Student's pooled t-test and the p-values agree exactly |
| Effect Size | |
| Eta-Squared | Share of total variation explained by the factor, SSB divided by SST |
| Upward Bias of Eta-Squared | It is never zero even under the null, because SSB collects k minus one mean squares of noise |
| Omega-Squared | Eta-squared with that noise subtracted, always smaller, preferred in small samples |
| Cohen's f | The same effect size rescaled, with 0.10, 0.25 and 0.40 matching 0.01, 0.06 and 0.14 |
| Partial Eta-Squared | Identical to eta-squared in a one-way design, but not in the two-way designs ahead |
| Which Pairs Differ | |
| Post-Hoc Tests | Pairwise comparisons run after a significant F, with the family-wise rate controlled |
| Tukey's HSD | Compares each gap against the studentised range, holding all pairs together at alpha |
| The Studentised Range | The distribution of the spread between the largest and smallest of k sample means |
| Tukey Against an Unadjusted Test | The same data can be significant unadjusted and not significant under Tukey |
| Bonferroni and Holm | Test each pair at alpha over m; simple and general but conservative, and Holm is strictly better |
| Scheffe | Covers every contrast, not only pairs, and is the most conservative of the family |
| Dunnett | Compares each group against one control only, and is more powerful when that is the design |
| Fisher's LSD | No adjustment at all, defensible only after a significant F and only with three groups |
| Games-Howell | The unequal-variance post-hoc test, standing to Tukey as Welch stands to Student |
| Assumptions | |
| Assumptions of ANOVA | Independence, approximate Normality of the residuals, and equal variances across groups |
| Normality of Residuals | What must be roughly Normal is the scatter inside groups, not the pooled outcome |
| Equal Variances | The assumption that bites, because the classic test pools all scatter into one estimate |
| Welch's ANOVA | Weights each group by n over its own variance and adjusts the denominator degrees of freedom |
| Against Preliminary Variance Tests | Levene first then choose has poor power and makes the final error rate conditional |
| What R Does by Default | aov() is the classic test; oneway.test() is Welch and already defaults to it |
| Kruskal-Wallis | The rank-based alternative for several groups, extending Mann-Whitney from Chapter 17 |