18  Hypothesis Testing

Chapter 15 asked what values of \(\mu\) are consistent with the data. This chapter asks the mirror question: someone has claimed a specific value, and you want to know whether your data contradicts it. Average transaction value used to be ₹5,000; has it moved? The defect rate is supposed to be 2%; is it worse?

Hypothesis testing is the formal machinery for answering that, and it runs on exactly the sampling distribution built in Chapter 14. It is also the most misused procedure in applied statistics, so this chapter spends as much effort on what a test does not tell you as on how to run one.

18.1 The Logic of a Test

A hypothesis test is an argument by contradiction. You provisionally assume the claim you doubt, work out what the data should look like under that assumption, and then ask how strange your actual data would be if the assumption held. If it would be very strange, you conclude the assumption is probably wrong.

The courtroom is the standard analogy and it is a good one. The defendant is presumed innocent, the prosecution must produce evidence strong enough to overcome that presumption, and a verdict of “not guilty” is not a finding of innocence, it is a finding that guilt was not established. Every one of those features has a direct counterpart in a statistical test.

  • Null hypothesis \(H_0\): the claim being tested, always a statement of no effect, no difference, or no change. It is the presumption of innocence, and it is what the test assumes true while computing.
  • Alternative hypothesis \(H_1\) or \(H_a\): what you conclude if the null is rejected. It carries the burden of proof.

The null always contains the equality, because it must specify a single distribution precisely enough to compute with:

\[H_0: \mu = 5000 \qquad\qquad H_1: \mu \neq 5000\]

The asymmetry between the two hypotheses is deliberate and it is the most commonly forgotten feature of the whole procedure. A test can reject \(H_0\), or it can fail to reject \(H_0\). It can never accept \(H_0\) or prove it true.

“We failed to reject the null” means the evidence was not strong enough, which may be because the null is true, or because the sample was too small to detect a real effect. Those two situations look identical in the output and are distinguished only by the power analysis later in this chapter. Reporting “we proved there is no difference” from a non-significant test is simply wrong, and it is wrong in a way that has distorted whole research literatures.

18.2 The Test Statistic and the p-Value

A test statistic measures how far the data sits from \(H_0\), in units of standard error. For a mean with \(\sigma\) unknown, it is the same quantity that produced the \(t\)-interval in Chapter 15:

\[t = \frac{\bar{x} - \mu_0}{s/\sqrt{n}}\]

The p-value converts that distance into a probability. Formally, it is the probability, computed assuming \(H_0\) is true, of obtaining a result at least as extreme as the one observed.

A small p-value means the observed data would be unusual if \(H_0\) were true, which counts as evidence against \(H_0\). A large p-value means the data is unremarkable under \(H_0\), which is not evidence for anything much.

The first tab of the interactive above is this definition drawn. The curve is the distribution of \(t\) assuming \(H_0\) is true, the blue line is the \(t\) your data produced, and the orange area is the p-value: the share of samples that would land at least that far out if the null held. Slide the sample mean and watch the orange area grow and shrink. Notice that the curve itself never moves, because it describes the null’s world, not yours.

18.3 Significance Level and the Decision Rule

The significance level \(\alpha\) is a threshold chosen before seeing the data, setting how much evidence you demand. The decision rule is then simply:

\[p \leq \alpha \;\Rightarrow\; \text{reject } H_0 \qquad\qquad p > \alpha \;\Rightarrow\; \text{fail to reject } H_0\]

Equivalently, and identically, compare the test statistic against the critical value that cuts off \(\alpha\) in the tail. The p-value approach and the critical-value approach always agree; they are the same comparison made in different units.

Conventional levels are \(\alpha = 0.05\), \(0.01\), and \(0.10\). There is nothing mathematically special about 0.05; it is a convention inherited from Fisher’s early tables, and treating it as a law of nature causes many of the problems discussed later in this chapter.

18.4 One-Tailed and Two-Tailed Tests

The alternative hypothesis decides which tail or tails count as extreme.

  • Two-tailed: \(H_1: \mu \neq \mu_0\). Departures in either direction are evidence. The significance level is split between both tails, \(\alpha/2\) in each.
  • One-tailed: \(H_1: \mu > \mu_0\) or \(H_1: \mu < \mu_0\). Only one direction counts, and the whole of \(\alpha\) sits in that tail.

For the same data, a one-tailed p-value is exactly half the two-tailed one, which makes a one-tailed test easier to pass.

That last fact is precisely why the choice must be made before looking at the data, and justified by the research question rather than by the result. Running a two-tailed test, finding \(p = 0.08\), and switching to one-tailed to obtain \(p = 0.04\) is not analysis, it is manufacturing a result. A one-tailed test is legitimate only when a departure in the opposite direction would genuinely lead to the same decision as no departure at all, which is rarer than its popularity suggests.

18.5 A Worked Test

Example

A retail chain’s average transaction value has historically been ₹5,000. After a pricing change, a random sample of \(n = 64\) transactions gives \(\bar{x} = ₹5{,}250\) with \(s = ₹800\). Has the mean changed, at \(\alpha = 0.05\)?

Step 1, state the hypotheses. \(H_0: \mu = 5000\) against \(H_1: \mu \neq 5000\), two-tailed, because a fall would matter as much as a rise.

Step 2, compute the standard error. \(s/\sqrt{n} = 800/8 = 100\).

Step 3, compute the test statistic.

\[t = \frac{5250 - 5000}{100} = 2.50 \qquad df = 63\]

Step 4, find the p-value. \(P(|T_{63}| \geq 2.50) = 0.015\).

Step 5, decide. Since \(0.015 \leq 0.05\), reject \(H_0\).

Step 6, state the conclusion in context. There is sufficient evidence at the 5% level to conclude that mean transaction value has changed from ₹5,000. The sample suggests an increase of about ₹250.

Step 6 is the step most often skipped, and it is the only one a decision-maker reads. “Reject \(H_0\), \(p = 0.015\)” is not a conclusion, it is an intermediate result. A conclusion names the variable, the direction, the size of the effect, and the population it applies to.

The connection to Chapter 15

The 95% confidence interval for this sample is \(5250 \pm 1.998 \times 100 = (5050,\; 5450)\). Notice that ₹5,000 falls outside it, and that the two-tailed test rejected \(H_0: \mu = 5000\) at the 5% level. This is not a coincidence.

A two-tailed test at level \(\alpha\) rejects \(H_0: \mu = \mu_0\) exactly when the \((1-\alpha)\) confidence interval excludes \(\mu_0\).

The interval and the test are the same statement in different clothes. The interval is usually the more useful of the two, because it reports the plausible range of effect sizes rather than a single reject-or-not verdict.

18.6 Two Ways to Be Wrong

A test produces a verdict about an unknown truth, so two distinct errors are possible.

\(H_0\) is actually true \(H_0\) is actually false
Reject \(H_0\) Type I error (probability \(\alpha\)) Correct decision (probability \(1-\beta\))
Fail to reject \(H_0\) Correct decision (probability \(1-\alpha\)) Type II error (probability \(\beta\))
  • A Type I error is a false positive: concluding there is an effect when there is none. Its probability is exactly \(\alpha\), the level you chose.
  • A Type II error is a false negative: missing an effect that is really there. Its probability is \(\beta\), which you do not choose directly.
  • Power is \(1 - \beta\), the probability of detecting an effect that genuinely exists.

At a fixed sample size, \(\alpha\) and \(\beta\) trade directly against one another. Tightening \(\alpha\) from 0.05 to 0.01 to guard against false positives necessarily raises \(\beta\) and lowers power, because the critical line moves further out and more real effects fall short of it. The second tab of the interactive shows this directly: drag \(\alpha\) down and watch the red region shrink while the orange one grows.

The only way to reduce both errors at once is to increase \(n\), which narrows both distributions and separates them.

Power depends on four things, and it is worth knowing which of them you control:

  • Effect size: bigger true effects are easier to detect. Not under your control.
  • Sample size \(n\): more data means more power. Under your control, at a cost.
  • Significance level \(\alpha\): a laxer threshold raises power but also false positives. Under your control.
  • Population variability \(\sigma\): noisier populations reduce power. Occasionally reducible through better measurement or a better design.

The conventional target is power of 0.80, meaning an 80% chance of detecting an effect of the size you care about, and it should be computed before the study is run.

An underpowered study is worse than no study, and this is not a rhetorical flourish. Set the interactive to an effect of 180 with \(n = 64\) and power comes out at about 44%: if the effect is real, this design misses it more often than it finds it. A non-significant result from such a study carries essentially no information, because the test was never capable of detecting the effect in the first place. Yet such results are routinely written up as evidence of no effect.

There is a second, subtler cost. Among underpowered studies, the ones that do reach significance must have landed on an unusually large sample estimate, so published effect sizes from underpowered research are systematically inflated.

Application

The power table is the part worth acting on. At \(n = 64\) this study has a 44% chance of detecting a ₹180 shift, which is a coin flip. Reaching the conventional 80% takes \(n = 156\), and the table brackets it between 100 and 200. That calculation costs nothing and takes a minute, and doing it before collecting data is the difference between a study that can answer its question and one that cannot.

18.7 What a p-Value Is Not

Misuse of p-values became serious enough that in 2016 the American Statistical Association issued a formal statement on them, the first time in its history it had taken a public position on a specific matter of statistical practice. Its definition is the one given earlier in this chapter: the probability, under a specified model, of a result at least as extreme as the one observed.

The statement sets out six principles, and they are worth reading as written.

  1. P-values can indicate how incompatible the data are with a specified statistical model.
  2. P-values do not measure the probability that the studied hypothesis is true, or the probability that the data were produced by random chance alone.
  3. Scientific conclusions and business or policy decisions should not be based only on whether a p-value passes a specific threshold.
  4. Proper inference requires full reporting and transparency.
  5. A p-value, or statistical significance, does not measure the size of an effect or the importance of a result.
  6. By itself, a p-value does not provide a good measure of evidence regarding a model or hypothesis.

Principle 2 is the one that trips up almost everyone. A p-value of 0.03 does not mean there is a 3% chance \(H_0\) is true.

The p-value is computed assuming \(H_0\) is true, so it cannot also be a statement about how likely that assumption is. It is \(P(\text{data this extreme} \mid H_0)\), not \(P(H_0 \mid \text{data})\). Chapter 10 already showed, via Bayes’ Theorem, that these two conditional probabilities can differ enormously, and the medical-testing example there is the same trap in different clothing: a test’s false-positive rate is not the probability that a positive result is wrong.

Getting from a p-value to \(P(H_0 \mid \text{data})\) requires a prior probability for \(H_0\), which the p-value does not contain and the test never asked for.

A short catalogue of the other common misreadings:

  • “\(p = 0.06\) means there is no effect.” No. It means this sample did not provide enough evidence at your chosen threshold. The difference between \(p = 0.049\) and \(p = 0.051\) is not the difference between a real effect and no effect, and treating 0.05 as a cliff edge is exactly what Principle 3 warns against.
  • “A smaller p-value means a bigger effect.” No. The p-value confounds effect size with sample size. A trivial effect measured on a million observations produces a minuscule p-value.
  • “\(p = 0.03\) means the result replicates 97% of the time.” No. Replication probability depends on the true effect and the power of the replication study, not on the original p-value.
  • “The result is significant, so it matters.” A statistical statement, not a business one. See the next section.

18.8 Statistical Versus Practical Significance

Statistical significance says an observed effect is unlikely to be pure sampling noise. Practical significance says the effect is large enough to matter. These are different questions, and the second one statistics cannot answer for you.

The link between them is sample size. Because the test statistic contains \(\sqrt{n}\), any effect that is not exactly zero becomes statistically significant once \(n\) is large enough.

Example

An e-commerce firm tests a checkout redesign on \(n = 500{,}000\) users. Conversion rises from 4.00% to 4.06%. With that sample the difference is statistically significant, \(p < 0.001\).

Is it worth deploying? That depends on whether 0.06 percentage points covers the engineering cost, the risk of the change, and the opportunity cost of not building something else. The test cannot say. It has established only that the 0.06 points is probably real, not that it is worth having.

Conversely, a promising effect in a study of 20 customers may fail to reach significance while still being large enough to justify a bigger trial. “Not significant” and “not worth pursuing” are different findings.

The practical remedy is simple and it is what good analysis does routinely: always report the effect size and its confidence interval alongside the test. “Conversion rose 0.06 points, 95% CI 0.03 to 0.09 points, \(p < 0.001\)” tells a decision-maker everything the p-value alone conceals: the direction, the magnitude, and the precision. This is why Chapter 15 came first.

18.9 How Tests Go Wrong in Practice

Two failure modes account for a large share of unreliable published findings, and both follow directly from the definition of \(\alpha\).

Multiple comparisons. If \(\alpha = 0.05\), then one test in twenty produces a false positive when the null is true. Run 20 independent tests on data with no real effects and you should expect one to come out significant. The probability of at least one false positive across \(m\) independent tests is \(1 - (1-\alpha)^m\), which reaches 40% at \(m = 10\) and 64% at \(m = 20\). Corrections such as Bonferroni, which tests each hypothesis at \(\alpha/m\), exist precisely to control this.

p-hacking. The same arithmetic applies to undisclosed flexibility in a single analysis. Trying several outcome variables, several subgroups, several ways of excluding outliers, or repeatedly adding data and re-testing until \(p\) drops below 0.05, all inflate the false-positive rate far above the stated \(\alpha\) while reporting only the test that worked. This is what ASA Principle 4 is about: the reported p-value is only meaningful if you also know how many analyses were run to obtain it.

The defences are procedural rather than mathematical. Decide the hypothesis, the test, and the sample size before collecting data. Report every analysis performed, not just the successful one. Correct for multiple comparisons when testing many hypotheses. Treat exploratory findings as hypotheses to be confirmed on fresh data rather than as results. None of this requires advanced statistics; it requires writing the plan down first.

Recap

Chapters 15 and 16 are two views of one calculation. A confidence interval reports the range of parameter values consistent with the data; a hypothesis test asks whether one specific value survives contact with it. They agree by construction, since a two-tailed test at level \(\alpha\) rejects exactly the values the \((1-\alpha)\) interval excludes. Between them they complete the inferential chain begun in Chapter 13: draw a sample properly, understand how its statistics vary, estimate the parameter with honest uncertainty, and test specific claims about it. What this module has deliberately not done is treat a p-value as a verdict. The effect size, its interval, the power of the design, and the number of analyses that produced the result all belong in the report alongside it.


Summary

Concept Description
The Logic of a Test
Hypothesis Test A formal argument by contradiction: assume a claim, then ask how strange the data would be if it held
Null Hypothesis The claim being tested, always stating no effect, no difference, or no change
Alternative Hypothesis What is concluded if the null is rejected; it carries the burden of proof
Why the Null Contains the Equality The null must pin down a single distribution precisely enough for probabilities to be computed
Failing to Reject Is Not Accepting A test can reject the null or fail to reject it, but can never prove it true
Machinery
Test Statistic A measure of how far the data lies from the null, expressed in standard errors
p-Value The probability, assuming the null is true, of a result at least as extreme as the one observed
Significance Level The threshold alpha, chosen before seeing the data, setting how much evidence is demanded
Critical Value The value of the test statistic that cuts off alpha in the tail of the null distribution
Decision Rule Reject the null when p is at most alpha; the p-value and critical-value approaches always agree
Tails
Two-Tailed Test Used when departures in either direction matter, splitting alpha between both tails
One-Tailed Test Used when only one direction matters, placing all of alpha in that tail
Choosing the Tails in Advance A one-tailed p-value is half the two-tailed one, so the choice must precede seeing the data
Running and Reporting
Steps of a Test State hypotheses, compute the standard error, compute the statistic, find p, decide, then conclude in context
Stating the Conclusion in Context A conclusion names the variable, the direction, the effect size, and the population, not just the verdict
Duality with the Confidence Interval A two-tailed test rejects a value exactly when the corresponding confidence interval excludes it
Errors and Power
Type I Error Rejecting a true null; a false positive, whose probability is exactly the chosen alpha
Type II Error Failing to reject a false null; a false negative, with probability beta
Power One minus beta, the probability of detecting an effect that genuinely exists
The Alpha-Beta Trade-off At fixed sample size, lowering alpha raises beta; only more data reduces both
What Power Depends On Effect size, sample size, significance level, and population variability
The 80 Percent Convention Power of 0.80 is the usual target, and should be computed before the study is run
Danger of Underpowered Studies A non-significant result from a low-power design carries almost no information
Inflated Published Effect Sizes Among underpowered studies, only unusually large estimates reach significance, so published effects are biased upward
What p-Values Are Not
The ASA Statement The American Statistical Association's 2016 statement setting out six principles on p-value use
p-Value Is Not P(H0 Given Data) The p-value assumes the null is true, so it cannot also be the probability that the null is true
p-Value Does Not Measure Effect Size The p-value confounds effect size with sample size and says nothing about importance
Bright-Line Thinking Treating 0.05 as a cliff edge is unsupported; 0.049 and 0.051 carry nearly identical evidence
Significance vs Importance
Statistical Significance An observed effect is unlikely to be explained by sampling noise alone
Practical Significance An effect is large enough to matter for the decision at hand, which statistics cannot determine
Large n Makes Everything Significant Because the statistic contains the square root of n, any non-zero effect becomes significant eventually
Report the Effect Size Always report the effect size and its confidence interval alongside the test result
How Tests Go Wrong
Multiple Comparisons Running many tests inflates the false-positive rate; the chance of at least one exceeds 60 percent at twenty tests
p-Hacking Undisclosed flexibility in analysis choices inflates false positives far above the stated alpha
Procedural Defences Pre-specify the hypothesis and sample size, report every analysis, correct for multiplicity, confirm on fresh data