20 Comparing Two Groups
Chapter 16 tested a single mean against a claimed value. Useful, but it is rarely the question a business actually has. The real questions compare things: did the new checkout page outperform the old one, do the Chennai and Mumbai branches differ, did the training course change anything?
All of these compare two groups, and answering them takes the machinery of Chapter 16 pointed at a difference rather than at a single mean. The logic is unchanged. What changes is that you must first decide what kind of two groups you have, because that decision picks the test, and getting it wrong wastes real information.
This chapter also parts company with most textbooks on one point. The two-sample test they teach first is not the one you should use by default, and the reasons are measurable.
20.1 Two Designs, Two Tests
Before choosing a test, ask one question: is there a meaningful pairing between the two sets of measurements?
- Independent samples: the two groups contain different, unrelated units. Customers in Chennai and customers in Mumbai. Users shown version A and users shown version B. No row in one group belongs with any particular row in the other, and the two groups need not even be the same size.
- Paired samples: each measurement in one group has a natural partner in the other. The same employee before and after training. The same store this quarter and last. Left eye and right eye. The pairing is a fact about the data, not a choice.
Analysing paired data with an independent-samples test is one of the most costly mistakes in applied statistics, and it is costly in a specific and avoidable way. The independent test has to work around the variation between people, which is usually the largest variation present. The paired test never sees it: by looking only at each person’s change, the between-person differences cancel out entirely.
The third tab of the interactive above shows the damage. The same twelve people, the same real improvement of roughly four points, analysed both ways: the paired test returns \(p = 0.0025\) while the independent test returns \(p = 0.44\). One finds the effect, the other misses it completely. Nothing about the data changed, only the question asked of it.
Slide that same control down to zero, though, and the two tests converge. Pairing buys power only when the paired units genuinely resemble each other. Pairing on something irrelevant costs you degrees of freedom and gains nothing, so the pairing has to be real, not invented to look sophisticated.
20.2 The Independent-Samples t-Test
For two independent groups, the null hypothesis is that their population means are equal:
\[H_0: \mu_1 = \mu_2 \qquad\qquad H_1: \mu_1 \neq \mu_2\]
The test statistic has the same shape as every one in Chapter 16, an observed difference divided by its standard error:
\[t = \frac{\bar{x}_1 - \bar{x}_2}{SE(\bar{x}_1 - \bar{x}_2)}\]
The two versions of this test differ only in how that standard error is computed, and therefore in what they assume about the two populations.
Welch’s t-test estimates each group’s variability separately and makes no assumption that the two are equal:
\[SE = \sqrt{\frac{s_1^{2}}{n_1} + \frac{s_2^{2}}{n_2}}\]
Its degrees of freedom are not a whole number, since they are approximated from the data by the Welch-Satterthwaite formula:
\[df \approx \frac{\left(\dfrac{s_1^{2}}{n_1} + \dfrac{s_2^{2}}{n_2}\right)^{2}} {\dfrac{(s_1^{2}/n_1)^{2}}{n_1 - 1} + \dfrac{(s_2^{2}/n_2)^{2}}{n_2 - 1}}\]
Student’s t-test, the pooled-variance version, instead assumes the two populations share one variance and estimates it from both groups together:
\[s_p^{2} = \frac{(n_1-1)s_1^{2} + (n_2-1)s_2^{2}}{n_1 + n_2 - 2} \qquad SE = s_p\sqrt{\frac{1}{n_1} + \frac{1}{n_2}} \qquad df = n_1 + n_2 - 2\]
20.3 Why Welch Should Be Your Default
Most introductory textbooks present Student’s pooled test as the two-sample t-test and mention Welch as a special case for when variances differ. That ordering is backwards, and the cost of following it is measurable.
When the two groups have unequal variances and unequal sizes, Student’s test stops delivering the error rate it promises. Which way it fails depends on how the imbalance lines up:
| Situation | Student’s actual false-positive rate | Welch |
|---|---|---|
| Equal sizes, equal spreads | 5% | 5% |
| Larger spread in the larger group | far below 5% | 5% |
| Larger spread in the smaller group | far above 5% | 5% |
The second row sounds harmless and is not: a test that rejects too rarely is quietly throwing away power you paid for. The third row is the dangerous one, and the second tab of the interactive above measures it. Set the groups to 20 and 60 with a threefold difference in spread, run four thousand tests on data with no real difference at all, and Student’s test calls about one in five of them significant. Welch stays at 5% in every configuration.
A common response is to test for equal variances first, with Levene’s or an F-test, and pick the test based on the result. This two-step procedure is worse than simply using Welch, for three reasons.
The preliminary test has poor power: detecting even a 1.5-fold difference in standard deviations reliably takes in the region of 120 observations, so with typical sample sizes it usually fails to detect exactly the inequality that matters. The final error rate becomes conditional on that first test’s outcome, so neither stage delivers its stated \(\alpha\). And it is unnecessary, because Welch is nearly as powerful as Student even when the variances genuinely are equal, so there is almost nothing to gain by choosing.
Your software has already settled this. R’s t.test() defaults to var.equal = FALSE, which is Welch; getting Student’s test requires explicitly asking for it. So the default output of the most-used statistical language in the world is the test this chapter recommends, while the textbook convention teaches the other one first.
The practical rule is short: use Welch unless you have a specific reason not to, and do not run a preliminary variance test to decide.
Example
A retailer compares average transaction value at two store formats. The compact stores sample \(n_1 = 45\) with \(\bar{x}_1 = ₹4{,}820\) and \(s_1 = ₹910\); the flagship stores sample \(n_2 = 32\) with \(\bar{x}_2 = ₹5{,}260\) and \(s_2 = ₹1{,}340\).
The spreads clearly differ, and so do the sample sizes, which is precisely the situation where the pooled test misbehaves. Using Welch:
\[SE = \sqrt{\frac{910^{2}}{45} + \frac{1340^{2}}{32}} = \sqrt{18{,}402 + 56{,}113} = ₹273.0\]
\[t = \frac{4820 - 5260}{273.0} = -1.61 \qquad df \approx 50.8 \qquad p = 0.113\]
At the 5% level there is not sufficient evidence that the formats differ in average transaction value. Note what the conclusion does not say: it does not say the formats are the same. The 95% confidence interval for the difference runs from about ₹988 below to ₹108 above, which is wide enough to contain differences a retailer would certainly act on.
20.4 The Paired t-Test
When the data is paired, the trick is to stop treating it as two samples at all. Compute each pair’s difference, \(d_i = x_{i,\text{after}} - x_{i,\text{before}}\), and the problem collapses to a one-sample test on those differences against a null of zero, which is exactly Chapter 16’s test:
\[t = \frac{\bar{d}}{s_d / \sqrt{n}} \qquad\qquad df = n - 1\]
where \(n\) is the number of pairs, not the number of measurements.
Example
Twelve employees take a training course and are scored before and after, using the data in the code below. Eleven improve and one slips back by a point. The changes average \(\bar{d} = 4.33\) with \(s_d = 2.93\):
\[t = \frac{4.33}{2.93/\sqrt{12}} = \frac{4.33}{0.846} = 5.12 \qquad df = 11 \qquad p = 0.00034\]
Strong evidence that scores improved. Those same twenty-four numbers analysed as two independent groups give \(t = 0.96\) and \(p = 0.35\), nowhere near significance, because the employees differ from one another by far more than any of them changed.
This is why paired designs are worth engineering into a study when they are possible. Measuring the same units twice removes the between-unit variation from the comparison for free, which is usually the cheapest power you will ever buy. A before-and-after design on 20 stores can beat a two-group design on 200.
20.5 How Big Is the Difference?
Chapter 16 warned that a p-value confounds effect size with sample size. For a two-group comparison the standard remedy is Cohen’s d, which expresses the difference in standard deviations rather than in the original units:
\[d = \frac{\bar{x}_1 - \bar{x}_2}{s_p}\]
| \(|d|\) | Conventional label |
|---|---|
| 0.2 | Small |
| 0.5 | Medium |
| 0.8 | Large |
For small samples, \(d\) is biased slightly upward. Multiplying by a correction factor gives Hedges’ g, which is preferred when \(n\) is below about 20 per group:
\[g = d \times \left(1 - \frac{3}{4(n_1 + n_2) - 9}\right)\]
Cohen’s thresholds are conventions he himself offered reluctantly and hedged heavily, not laws. What counts as a large effect depends entirely on the field and the decision: a \(d\) of 0.1 in a national vaccination programme may matter enormously, while a \(d\) of 1.0 in a pilot study of six people may be noise. Report the number and interpret it in context rather than reading the label off the table.
The first tab of the interactive makes the independence of these two quantities concrete. Hold the true difference fixed and drag \(n\) from 30 to 200: the p-value collapses from 0.02 to below 0.0001, while Cohen’s d does not move at all. The p-value is answering “is this real?”; d is answering “is this big?”. Reporting only the first is what lets a trivial difference be presented as an important finding.
20.6 Assumptions, and What to Do When They Fail
Both tests assume the observations are a random sample, that observations are independent within each group, and that the sampling distribution of the mean is approximately Normal, which by Chapter 14 follows from the population being roughly Normal or \(n\) being adequate.
Note what is not on that list for Welch: equal variances. That assumption belongs to the pooled test alone, which is much of the argument for not using it.
Independence within groups is the assumption most often violated without anyone noticing, and no test detects it for you. Three customers from the same household, five readings from the same machine, or a class of students taught by one teacher are not independent observations. Treating clustered data as independent understates the standard error and inflates significance, which is the same failure mode as the intra-cluster correlation discussed in Chapter 13.
When the Normality condition is genuinely doubtful and \(n\) is small, the usual alternatives are the Mann-Whitney U test for independent groups and the Wilcoxon signed-rank test for paired data. Both work on ranks rather than values, so they make no Normality assumption, at the cost of testing a slightly different hypothesis and losing some power when Normality does hold. For a skewed but positive variable such as income, transforming with logarithms first is often the better move, since it keeps the t-test and its interpretable effect size.
Application
Two things in that output are worth carrying away. The paired and independent analyses of the identical twelve pairs reach opposite conclusions, which is the cost of ignoring a design feature that was free to use. And Student’s test, run four thousand times on data containing no difference whatsoever, calls a fifth of them significant. Neither of these is an exotic edge case; unequal group sizes with unequal spread is the normal condition of observational business data.
Looking Ahead
Two groups is where comparison starts, not where it ends. The next chapter handles the case where the outcome is a category rather than a measurement, so the question becomes whether two classifications are associated at all: the chi-square tests of goodness-of-fit and independence, which take the contingency tables of Chapter 2 and the expected-frequency logic of Chapter 9 and turn them into a formal test.
Summary
| Concept | Description |
|---|---|
| Two Designs | |
| Independent Samples | Two groups of different, unrelated units, which need not be the same size |
| Paired Samples | Each measurement in one group has a natural partner in the other, such as before and after |
| Choosing Between Them | Ask whether a meaningful pairing exists; the pairing is a fact about the data, not a choice |
| Cost of Ignoring Pairing | The independent test must work around between-unit variation, which the paired test removes entirely |
| When Pairing Does Not Help | Pairing buys power only when paired units genuinely resemble each other |
| The Independent Test | |
| Two-Sample Hypotheses | The null is that the two population means are equal, against the alternative that they differ |
| Welch's t-Test | Estimates each group's variance separately and assumes nothing about their equality |
| Welch-Satterthwaite Degrees of Freedom | A non-integer approximation of the degrees of freedom computed from the two sample variances |
| Student's Pooled t-Test | Assumes both populations share one variance, estimated from the two groups combined |
| Pooled Variance | A weighted average of the two sample variances, weighted by degrees of freedom |
| Why Welch | |
| Why Welch by Default | It holds the stated error rate across all configurations, and loses almost no power when variances are equal |
| Failure with Unequal Spread and Size | With unequal variances and sizes, the pooled test rejects far too often or far too rarely |
| Against Preliminary Variance Tests | Testing variances first has poor power, makes the final error rate conditional, and gains nothing |
| What R Does by Default | R's t.test uses var.equal = FALSE, so its default output is already Welch's test |
| The Paired Test | |
| Paired t-Test | Reduces paired data to the differences, then applies the one-sample test of Chapter 16 |
| Degrees of Freedom for Paired Data | Based on the number of pairs minus one, not the total number of measurements |
| Effect Size | |
| Cohen's d | The difference in means expressed in pooled standard deviations |
| Effect Size Conventions | Roughly 0.2 small, 0.5 medium and 0.8 large, offered as rough guidance rather than law |
| Hedges' g | A small-sample correction to Cohen's d, preferred below about twenty per group |
| Effect Size Is Independent of n | Raising n shrinks the p-value while leaving the effect size unchanged |
| Assumptions | |
| Assumptions of the t-Tests | Random sampling, independence within groups, and approximate Normality of the sampling distribution |
| Independence Within Groups | The assumption most often violated unnoticed, and no test detects it for you |
| Clustered Data | Households, machines or classrooms make observations dependent and inflate significance |
| Distribution-Free Alternatives | Mann-Whitney U for independent groups and Wilcoxon signed-rank for paired data |
| Transformation | Taking logarithms of a skewed positive variable often preserves the t-test and its effect size |