23 Two-Way and Repeated Measures ANOVA
Chapter 19 compared several groups formed by one grouping variable. Two restrictions came with that, and both are worth lifting.
The first is that real designs usually have more than one factor. A retailer runs a promotion across several store formats and wants to know about the promotion, the format, and, most usefully, whether the promotion that works in one format is the same one that works in another. That last question belongs to neither factor on its own. It is the interaction, and it is the reason two-way ANOVA exists.
The second is that Chapter 19 assumed every observation came from a different unit. Chapter 17 showed what pairing buys when there are two measurements per unit. With three or more, the same logic gives repeated measures ANOVA, and the gain is often enormous.
23.1 Two Factors at Once
With two factors, A at \(a\) levels and B at \(b\) levels, the data splits into \(a \times b\) cells, and the total variation partitions into four pieces rather than two:
\[SST = SS_{A} + SS_{B} + SS_{AB} + SS_{E}\]
Each of the first three gets its own F test against the same error term, so a two-way ANOVA asks three questions at once:
\[H_0^{(A)}: \text{the A means are all equal}\] \[H_0^{(B)}: \text{the B means are all equal}\] \[H_0^{(AB)}: \text{the effect of A is the same at every level of B}\]
The first two are main effects, each averaging over the other factor. The third is the interaction, and it is the one that has no one-way counterpart.
The degrees of freedom split the same way. With \(n\) observations per cell and \(N = abn\) in total:
| Source | df |
|---|---|
| Factor A | \(a - 1\) |
| Factor B | \(b - 1\) |
| Interaction | \((a-1)(b-1)\) |
| Error | \(ab(n-1)\) |
| Total | \(N - 1\) |
The interaction degrees of freedom should look familiar. They are the degrees of freedom of a contingency table from Chapter 18, and for the same reason: once the row and column effects are fixed, only \((a-1)(b-1)\) cells are free to depart from what those effects predict.
23.2 The Interaction Is the Point
The cleanest way to understand an interaction is to ask what the cell means would look like without one. If the two factors act independently, each cell mean is the grand mean plus a row effect plus a column effect:
\[\bar{x}_{ij} = \bar{x} + \alpha_i + \beta_j\]
Plot that and the lines are parallel. Factor B shifts every level of A by the same amount, so the shape of one factor’s effect does not depend on the other. The interaction term is exactly what is left over when this additive model fails:
\[(\alpha\beta)_{ij} = \bar{x}_{ij} - \bar{x} - \alpha_i - \beta_j\]
and \(SS_{AB}\) is \(n\) times the sum of those squared leftovers. A significant interaction means the lines are not parallel, which means the effect of one factor changes depending on the level of the other.
Drag the interaction slider in the first tab of the interactive above. At zero the lines are exactly parallel and the interaction F is zero. As you pull it up the lines splay and cross, and the interaction F climbs with the square of the slider. Notice what does not move: both main effects stay pinned at the same F, because in a balanced design the three terms are orthogonal and each carries information the others do not.
23.3 A Worked Example
Example
A retailer tests three promotions in two store formats, sampling 20 transactions from each of the six combinations, so \(N = 120\). Average basket value by cell:
| None | Discount | Bundle | Format mean | |
|---|---|---|---|---|
| Compact | 4,600 | 5,100 | 4,750 | 4,816.7 |
| Flagship | 5,200 | 5,250 | 5,900 | 5,450.0 |
| Promotion mean | 4,900 | 5,175 | 5,325 | 5,133.3 |
Within-cell standard deviation is 450 throughout, giving \(MS_E = 202{,}500\) on 114 degrees of freedom.
| Source | Sum of squares | df | Mean square | F | p |
|---|---|---|---|---|---|
| Store format | 12,033,333 | 1 | 12,033,333 | 59.42 | < 0.001 |
| Promotion | 3,716,667 | 2 | 1,858,333 | 9.18 | 0.0002 |
| Format × Promotion | 5,016,667 | 2 | 2,508,333 | 12.39 | < 0.001 |
| Error | 23,085,000 | 114 | 202,500 | ||
| Total | 43,851,667 | 119 |
All three are significant. The order in which to read them, however, is not the order in which they are printed.
Read the interaction first, and read the main effects only in its light.
The main effect of promotion says that averaged across both formats, Bundle does best at 5,325, Discount next at 5,175, None worst at 4,900. A report that stopped there would recommend bundles.
That recommendation is wrong for half the estate. Look at the cells. In compact stores the best promotion is Discount, at 5,100 against 4,750 for Bundle. In flagship stores the best is Bundle, at 5,900 against 5,250 for Discount. The marginal mean that produced the recommendation is an average over two groups that want opposite things, and it describes neither.
This is why a significant interaction demotes the main effects. They are not wrong as arithmetic; they are answers to a question nobody should be asking once the interaction is known to be there.
23.4 When the Interaction Is Significant
The standard follow-up is a simple effects analysis: run the effect of one factor separately within each level of the other, rather than averaged over it. Here that means a one-way ANOVA of promotion inside each format, followed by Tukey as in Chapter 19.
| Format | Simple effect of promotion | Tukey, the pairs that separate |
|---|---|---|
| Compact | \(F(2, 57) = 6.50\), \(p = 0.003\) | Discount beats None by 500 (\(p = 0.003\)); Discount beats Bundle by 350 (\(p = 0.044\)) |
| Flagship | \(F(2, 57) = 15.06\), \(p < 0.001\) | Bundle beats None by 700 (\(p < 0.001\)); Bundle beats Discount by 650 (\(p < 0.001\)) |
Two sentences now replace the wrong one. In compact stores, discount; in flagship stores, bundle. In neither format does the other promotion beat doing nothing by a detectable margin.
A caution that applies to simple effects as much as to anything else in Chapter 19. Splitting a two-way design into simple effects multiplies the number of tests, so the family-wise argument has not gone away. Decide in advance which simple effects the question actually requires, or carry a correction across the set.
The opposite mistake is also common. When the interaction is not significant, do not go hunting through simple effects anyway. The main effects are then the right summary, and they are more powerful because they use all the data.
23.5 Effect Size, and Why Partial Now Matters
Chapter 19 noted that eta-squared and partial eta-squared coincide in a one-way design and part company in a two-way one. Here is where that bites. For a term \(X\):
\[\eta^{2}_{X} = \frac{SS_{X}}{SST} \qquad\qquad \eta^{2}_{p,X} = \frac{SS_{X}}{SS_{X} + SS_{E}}\]
Ordinary eta-squared divides by everything, so each term’s share is reduced by however much the other factors happen to explain. Partial eta-squared divides only by that term plus error, which asks a different question: of the variation this term could possibly have explained, how much did it?
| Term | \(\eta^2\) | Partial \(\eta^2\) |
|---|---|---|
| Store format | 0.274 | 0.343 |
| Promotion | 0.085 | 0.139 |
| Interaction | 0.114 | 0.179 |
The partial values are larger, always, and by more when the other terms are strong.
Two consequences follow, and both cause trouble in published work.
Partial eta-squares do not sum to anything. Ordinary eta-squares partition the total and add up to at most 1. Partial ones can easily sum past 1, so reading them as shares of a fixed pie is a mistake.
They are not comparable across studies with different designs. Adding a second factor that explains a lot of variance shrinks the error term, which raises the partial eta-squared of every other term without those terms having changed at all. Two studies reporting partial \(\eta^2 = 0.14\) for the same treatment may not have found the same size of effect.
SPSS reports partial eta-squared by default and labels it simply “Eta Squared” in some output, which is how a great many papers came to report the partial value while calling it the ordinary one. Say which you mean.
23.6 Unbalanced Designs, and the Types of Sums of Squares
Everything so far assumed balance: the same \(n\) in every cell. Balance is worth more than convenience. In a balanced design the factors are orthogonal, each sum of squares is independent of the others, and the partition is unambiguous.
When cells have different sizes, the factors become correlated and the sums of squares overlap. Some variation could be credited to either factor, and the answer depends on the order in which you ask. That ambiguity is what the types of sums of squares are about.
Type I is sequential. Each term is credited with what it explains after the terms before it. Its answers depend on the order of the terms in the model.
Type II gives each main effect what it explains after the other main effects, ignoring the interaction. It has the most power when the interaction is genuinely absent.
Type III gives each term what it explains after every other term including the interaction. It is order-independent, and it is what most software outside R means by an ANOVA table.
Example
Keep the six cell means exactly as they were, but make the cells uneven: 12 transactions in compact-none and 8 in flagship-bundle, 20 everywhere else. Now fit the same model twice, differing only in the order of the two factors:
| Terms in this order | Promotion sum of squares | Promotion p |
|---|---|---|
| format, then promotion | 1,588,269 | 0.023 |
| promotion, then format | 712,257 | 0.178 |
Same data, same model, same software. Promotion is significant in the first line and not significant in the second, and the only thing that changed is which of the two words came first in the formula. On the balanced data, both orders gave 3,716,667 and identical p-values.
The interaction sum of squares is 3,674,231 either way, because it is entered last in both orders and therefore gets whatever is left. That is the clue to what Type I is doing.
Practical guidance, since this is a genuine cross-software trap.
R’s anova() and summary(aov()) give Type I. This is a deliberate choice, not an oversight, but it means that on unbalanced data R’s default output is order-dependent and will not match SPSS, SAS or Stata, all of which report Type III by default. Use car::Anova(model, type = 3), and set orthogonal contrasts first with options(contrasts = c("contr.sum", "contr.poly")), or the Type III main effects are not meaningful.
The most reliable advice is the dullest: design for balance where you can, and where you cannot, say in the write-up which type you used. An ANOVA table on unbalanced data that does not state its type is not fully reported.
23.7 Repeated Measures: The Same Units, More Than Once
Chapter 17 made the case for the paired t-test. When each measurement in one group has a natural partner in the other, working with the differences removes all the variation between units and leaves only the variation in the change. The gain was often dramatic.
With three or more measurements per unit, the same idea becomes repeated measures ANOVA. The unit, whether a person, a store or a machine, is treated as a blocking factor, and its variation is pulled out of the error term:
\[SST = \underbrace{SS_{\text{subjects}}}_{\text{removed}} + SS_{\text{treatment}} + SS_{\text{error}}\]
The treatment F is then
\[F = \frac{SS_{\text{treatment}} / (k-1)}{SS_{\text{error}} / (n-1)(k-1)}\]
with \(n\) subjects and \(k\) occasions. Note the error degrees of freedom: \((n-1)(k-1)\) rather than the \(N - k\) of Chapter 19. Fewer degrees of freedom, but usually a very much smaller error term.
Example
Twelve stores, average basket value in three consecutive quarters. Quarterly means are 5,218, 5,408 and 5,548, so the picture is a gentle rise. The stores, however, range from about 4,350 to about 6,410 on average, and those differences dwarf the quarterly movement.
Analysed as if the three quarters were unrelated groups of stores:
\[F(2,\ 33) = 0.76, \qquad p = 0.475\]
Nothing. Analysed as repeated measures on the identical numbers:
\[F(2,\ 22) = 35.92, \qquad p < 0.001\]
The partition explains the reversal completely:
| Source | Sum of squares | df | Mean square |
|---|---|---|---|
| Between stores | 14,098,700 | 11 | 1,281,700 |
| Between quarters | 659,441 | 2 | 329,720 |
| Residual | 201,960 | 22 | 9,180 |
| Total | 14,960,100 | 35 |
Of the total variation, 94 per cent is differences between stores. The first analysis leaves all of it in the error term, giving a residual mean square of 433,353. The second removes it, leaving 9,180, a denominator forty-seven times smaller. Partial \(\eta^2\) is 0.766.
The lesson is the one Chapter 17 drew, now at a scale that is hard to ignore. The design is a fact about how the data was collected, not a modelling choice. These stores were measured three times each. An analysis that pretends otherwise is not being conservative; it is answering a different question badly, and it will report no effect where a large and consistent one exists.
Pull the between-stores slider in the second tab of the interactive above. The repeated measures F does not move at all, because widening the gaps between stores changes only the term it has already removed. The naive F collapses.
23.8 Sphericity
Repeated measures ANOVA buys its power with an assumption that has no counterpart in Chapter 19. Sphericity requires that the variances of the differences between every pair of occasions are equal:
\[\operatorname{Var}(Q_1 - Q_2) = \operatorname{Var}(Q_1 - Q_3) = \operatorname{Var}(Q_2 - Q_3)\]
With only two occasions there is a single difference and sphericity is automatic, which is why Chapter 17’s paired test never mentioned it. From three occasions on it is a real restriction, and it is often violated, because measurements taken close together in time usually correlate more strongly than measurements far apart.
When it fails, the F test is liberal: it rejects more often than its stated rate.
The standard repair does not change \(F\). It shrinks the degrees of freedom by a factor \(\varepsilon\) that measures how far from spherical the data is, and refers the same \(F\) to a stricter distribution:
\[F \sim F_{\,\varepsilon(k-1),\ \varepsilon(n-1)(k-1)}\]
\(\varepsilon\) runs from \(1/(k-1)\) at the worst possible violation to 1 when sphericity holds exactly. Greenhouse-Geisser is the conservative estimate; Huynh-Feldt corrects its downward bias and is less conservative. The common advice is to use Greenhouse-Geisser when its \(\varepsilon\) is below about 0.75 and Huynh-Feldt otherwise.
For the store data the three difference variances are 13,546, 12,281 and 29,253, so the assumption is visibly strained. Greenhouse-Geisser gives \(\varepsilon = 0.739\), taking the test from \(F(2, 22)\) to \(F(1.48, 16.25)\) and the p-value from 0.00000012 to 0.0000037. The conclusion is untouched, which is the usual outcome when the effect is large.
Mauchly’s test of sphericity exists, and the instinct is to run it and correct only if it is significant. Resist it, for the third time in three chapters.
On this data Mauchly gives \(W = 0.646\), \(p = 0.113\). It does not reject. Yet \(\varepsilon\) is 0.739, the difference variances differ by more than a factor of two, and the correction changes the p-value by a factor of thirty. With twelve stores the test simply has no power to detect a violation that is plainly visible in the numbers, and it is also sensitive to non-Normality in a way that makes a rejection hard to interpret.
Push the slider in the third tab of the interactive above and watch which notices first. Epsilon responds immediately; Mauchly’s p-value stays comfortable for a long while.
The defensible practice is to apply the correction routinely. When sphericity does hold, \(\varepsilon\) is near 1 and the correction costs almost nothing. When it does not, the correction is what keeps the stated error rate honest. Testing first gains nothing and makes the final error rate conditional on a test that was never reliable.
Two alternatives are worth knowing. A multivariate analysis of the same data, MANOVA on the set of difference contrasts, makes no sphericity assumption at all, at the cost of power when sphericity does hold and a requirement that \(n\) exceed \(k\). A linear mixed model with a random intercept for each unit handles the whole problem more generally, copes with missing occasions rather than dropping the whole unit, and is where most modern repeated-measures work ends up. The Friedman test is the rank-based option when Normality is the worry rather than sphericity.
Three things in that output are worth carrying away. The interaction row is the one that decides how the other two should be read, and here it says there is no single best promotion. The promotion p-value on the unbalanced data moves from 0.023 to 0.178 on nothing but the order of two words in a formula. And the twelve stores, analysed as twelve stores rather than thirty-six unrelated observations, turn a flat p of 0.475 into one below 0.0000002.
Recap
Chapters 19 and 20 take the comparison of groups as far as the analysis of variance carries it. Chapter 19 established the machinery: partition the total variation, compare between-group against within-group, and control the error rate across the whole family rather than one pair at a time. Chapter 20 added the two things that make the method fit real designs. A second factor brings the interaction, which is usually the most interesting term in the table and which demotes the main effects whenever it is significant. Repeated measurement of the same units brings blocking, which removes the variation between units from the error term and can turn a null result into a decisive one without touching a single number. Running through both chapters, and through Chapter 17 before them, is one recurring piece of advice: do not test an assumption in order to decide which test to run. Use Welch rather than screening with Levene, apply the Greenhouse-Geisser correction rather than screening with Mauchly. The screening tests have least power exactly where the problem does most damage. What ANOVA has never done is describe a relationship between two things that both vary continuously. Comparing groups requires groups, and cutting a continuous variable into bands to make some throws away most of what it knows. The next part asks the question directly: how do two measured variables move together, and can one be used to predict the other? That is correlation and regression.
Summary
| Concept | Description |
|---|---|
| Two Factors | |
| Two-Way ANOVA | Tests two grouping factors and their interaction on one measured outcome |
| Main Effect | The effect of one factor averaged across every level of the other |
| Interaction | Whether the effect of one factor changes with the level of the other |
| Four-Part Partition | Total splits into factor A, factor B, the interaction and error |
| Interaction Degrees of Freedom | Rows minus one times columns minus one, the same count as a contingency table |
| The Additive Model | Each cell mean is the grand mean plus a row effect plus a column effect |
| Parallel Lines | What an interaction plot shows when the two factors act independently |
| Reading the Result | |
| Reading the Table in the Right Order | Read the interaction first; the main effects mean less once it is significant |
| Why a Marginal Mean Can Mislead | It averages over groups that want opposite things and describes neither |
| Simple Effects | The effect of one factor computed separately within each level of the other |
| Multiplicity in Simple Effects | Splitting a design into simple effects multiplies the tests, so the family still counts |
| When There Is No Interaction | The main effects are then the right summary and use all the data |
| Effect Size | |
| Eta-Squared in a Two-Way Design | A term's sum of squares over the total, reduced by whatever the other terms explain |
| Partial Eta-Squared | A term's sum of squares over itself plus error, always the larger of the two |
| Partial Values Do Not Sum | They can exceed one in total, so they are not shares of a fixed whole |
| Not Comparable Across Designs | Adding a strong second factor raises every other term's partial value |
| What SPSS Reports | Partial eta-squared, sometimes labelled simply as eta squared, which has confused a literature |
| Balance and the Types | |
| Balanced Design | The same number of observations in every cell, which keeps the factors orthogonal |
| Why Balance Matters | With unequal cells the factors correlate and the sums of squares overlap |
| Type I Sums of Squares | Sequential: each term gets what it explains after the terms before it |
| Type II Sums of Squares | Each main effect after the other main effect, ignoring the interaction |
| Type III Sums of Squares | Each term after every other term, order-independent, the default outside R |
| What R Does by Default | anova() and summary(aov()) give Type I, so unbalanced output is order-dependent |
| The Interaction Is Unaffected | Entered last in every ordering, so its sum of squares does not move |
| Reporting the Type | An ANOVA table on unbalanced data is not fully reported until its type is stated |
| Repeated Measures | |
| Repeated Measures ANOVA | Extends the paired t-test of Chapter 17 to three or more occasions |
| Subject as a Blocking Factor | Its variation is pulled out of the error term rather than left in it |
| Error Degrees of Freedom | Subjects minus one times occasions minus one, fewer than a one-way design |
| The Cost of Ignoring the Design | Between-subject variation swamps the error term and hides a real effect |
| Design Is Not a Choice | How the data was collected is a fact, not a modelling option |
| Sphericity | |
| Sphericity | The variances of the differences between every pair of occasions are equal |
| Automatic With Two Occasions | With two occasions there is one difference, which is why the paired test never mentions it |
| Effect of Violating Sphericity | The F test becomes liberal and rejects more often than its stated rate |
| Epsilon | How far from spherical the data is, running from one over k minus one up to one |
| Greenhouse-Geisser and Huynh-Feldt | Two estimates of epsilon; the first is conservative, the second corrects its bias |
| Against Mauchly's Test | It has almost no power at realistic sample sizes, so correct routinely instead |
| Alternatives | |
| Multivariate Alternative | MANOVA on the difference contrasts, which assumes no sphericity at all |
| Linear Mixed Models | A random intercept per unit, which also copes with missing occasions |
| Friedman Test | The rank-based alternative when Normality rather than sphericity is the worry |