36 What the Data Says on Its Own
Chapter 29 read the file and never answered the question. This chapter answers it, or tries to, using nothing but the data and the tools of Modules III, IV and V. No model yet. Just five months, one measurement, and the question of whether the summer has a shape.
It turns out to have two answers, and which one gets reported depends on a choice the analyst makes after seeing the data.
That is the chapter. The statistics are all standard, all correctly applied, and they do not agree.
36.1 Five Months, Five Intervals
The first honest picture is Chapter 15’s: each month’s mean with an interval around it, drawn on top of the readings that produced it.
Two things are visible at once. July and August sit near 60 parts per billion while May and September sit near 25, so the summer clearly has a middle. And June’s interval is enormous.
The width of an interval is not a statement about how variable a month was. It is mostly a statement about how many days went into it, and Chapter 29 already established why June has nine. The formula is \(\bar{x} \pm t\,s/\sqrt{n}\), and of the three moving parts, \(n\) is doing the work here.
Example
The five months, on the 116 days that have a reading.
| Month | n | Mean | SD | 95% interval | Width |
|---|---|---|---|---|---|
| May | 26 | 23.6 | 22.2 | 14.6 to 32.6 | 18.0 |
| June | 9 | 29.4 | 18.2 | 15.4 to 43.4 | 28.0 |
| July | 26 | 59.1 | 31.6 | 46.3 to 71.9 | 25.6 |
| August | 26 | 60.0 | 39.7 | 43.9 to 76.0 | 32.1 |
| September | 29 | 31.4 | 24.1 | 22.3 to 40.6 | 18.4 |
June has the smallest standard deviation of the five months and the second widest interval. Its readings are less spread out than May’s, 18.2 against 22.2, and its interval is 1.56 times as wide.
August’s interval is wider still, at 32.1, but August has the spread to justify it: a standard deviation of 39.7 on 26 days. June has neither.
June’s interval overlaps May’s and September’s. It does not quite overlap July’s, missing by 2.9 ppb. Hold on to that near miss.
36.2 The Test That Says Yes
Chapter 19’s question, applied without modification: are the five monthly means different?
\(F = 8.536\) on 4 and 111 degrees of freedom, \(p = 4.8 \times 10^{-6}\).
Overwhelming. If the analysis stopped here it would report that ozone varies significantly across the summer, and that sentence would be true.
But Chapter 19 attached two conditions to that \(F\), and Chapter 29 already gave reason to suspect one of them. Ozone has a skewness of \(+1.226\). A right skewed response does not usually produce five groups of equal variance and symmetric residuals.
Example
Both conditions fail on the raw scale.
| Check | Value | Verdict |
|---|---|---|
| Ratio of largest to smallest group SD | 2.18 | marginal |
| Bartlett’s test of equal variance | \(K^2 = 13.45\), \(p = 0.0093\) | fails |
| Skewness of the ANOVA residuals | \(+1.059\) | fails |
The residuals inherit the response’s skew almost exactly, which is what a right tail does when you subtract group means from it and nothing else.
On the log scale both repair themselves: Bartlett gives \(p = 0.355\) and the residual skewness falls to \(-0.609\). The quantile plots in the code below show it directly, a bowed line becoming a straight one.
So the assumption failed. The usual next move is to worry. Here is what actually happens when you take the worry seriously and run the question four different ways.
| Route | Statistic | \(p\) |
|---|---|---|
| Ordinary ANOVA, raw scale | \(F(4, 111) = 8.536\) | \(4.8 \times 10^{-6}\) |
| ANOVA on the log | \(F(4, 111) = 9.163\) | \(2.0 \times 10^{-6}\) |
| Welch, no pooling of variances | \(F(4, 42.7) = 8.027\) | \(6.4 \times 10^{-5}\) |
| Kruskal-Wallis, ranks only | \(H = 29.267\) | \(6.9 \times 10^{-6}\) |
Four routes, four p-values, the largest of them 0.000064.
The assumption that failed did not change the answer. This is worth sitting with, because it is the common case and textbooks rarely say so. An \(F\) test with unequal variances and skewed residuals is not thereby wrong; it is less trustworthy, and the way to find out how much less is to run the alternatives and see whether they move. Here they do not.
When a conclusion survives four methods, the method has stopped being the interesting part of the problem.
36.3 The Test Does Not Say Which
An \(F\) of 8.536 says the five months are not all the same. It does not say which ones differ, and the difference between those two statements is where most reported findings go wrong.
Chapter 19 dealt with this: ten pairs are available from five groups, and testing all ten at the usual threshold means accepting roughly a one in two chance of at least one false positive. Tukey’s method prices that in.
Four of the ten pairs come out significant: July and August each differ from May, and September differs from each of July and August. That is a coherent picture. The middle of the summer is higher than either end.
None of the four pairs involving June is significant. Not against May, not against September, and, the interesting ones, not against July or August either.
June, the month Chapter 29 spent its whole length on, turns out to be statistically invisible. Nine days can be different from everything and prove nothing.
Example
Take one pair, July against June, a difference of 29.7 ppb, and test it four defensible ways.
| Rule | What it does | \(p\) |
|---|---|---|
| Welch \(t\)-test | uses only these two months, pools nothing | 0.0022 |
| Pooled \(t\)-test | borrows the shared variance from all five | 0.0102 |
| Tukey | pays for the ten comparisons available | 0.0749 |
| Bonferroni | pays for them more bluntly | 0.1023 |
One dataset, four correct procedures, and 0.05 falls in the middle of them.
Nothing here is a mistake. Welch declines to assume equal variance, which is reasonable given Bartlett. The pooled test uses the shared estimate, which is reasonable given the log scale repairs it. Tukey and Bonferroni charge for the nine other comparisons the analyst could have run, which is reasonable because the analyst could have.
August against June does the same thing, 0.0043 and 0.0083 against 0.0623 and 0.0831. Those are the only two pairs of the ten where the rule decides the verdict, and both of them involve June.
The uncomfortable part is not that the four rules disagree. It is that the analyst chooses among them after seeing that they disagree.
Nothing in the output labels the choice. A paper reporting “July is significantly higher than June, \(p = 0.002\)” has run a real test and quoted it correctly. So has a paper reporting “July and June could not be distinguished, \(p = 0.07\).” A reader of either has no way to know that the other exists.
The defence is procedural rather than statistical: decide the rule before looking, and report the ones you did not use. The eight pairs where all four rules agree need no such protection, which is exactly why the two where they disagree deserve it.
36.4 And the Correction Makes It Worse
Chapter 29 ended with a correction: June’s nine measured days were 2.73 mph windier than its twenty-one unmeasured ones, and wind costs 5.55 ppb of ozone per mph, so June’s mean is biased low by about 10.6 ppb. Apply that correction to every month and June moves from 29.4 to 40.1.
Here is what it does to the evidence.
| As measured | Corrected | |
|---|---|---|
| \(F\) | 8.536 | 7.609 |
| \(p\) for \(F\) | \(4.8 \times 10^{-6}\) | \(1.9 \times 10^{-5}\) |
| July against June, Tukey | 0.0749 | 0.5815 |
| August against June, Tukey | 0.0623 | 0.4743 |
The correction makes the seasonal finding weaker, not stronger. June stops being a quiet month at the start of the summer and becomes an unremarkable one in the middle of it, and the two comparisons that were hovering at the threshold move a long way from it.
The overall conclusion survives, because it never depended on June: four pairs are still significant and they are the same four.
But note the incentive. An analyst hoping for a clean seasonal story has a reason not to apply a correction they have already shown is warranted, and no reader of the final table could tell whether it was applied. The correction is not in the numbers. It is in the analyst.
36.5 How Often Is It Bad?
A mean is not what anybody actually wants to know about air quality. Nobody breathes an average. The question a reader has is how often the air was bad, and that is Module III’s question rather than Module IV’s.
Take 100 ppb as the threshold. Seven of the 116 measured days cleared it, which is 6.0 per cent, with a binomial interval running from 2.5 to 12.0 per cent.
By month the counts are 1, 0, 2, 4 and 0. August did it four times in 26 days and September never did in 29.
June never did it either, in nine days, and that fact is worth precisely nothing. The exact interval for zero out of nine runs from 0 to 33.6 per cent. September’s zero out of 29 runs from 0 to 11.9 per cent. Two identical observations, “never happened”, carrying completely different amounts of information, and a table of percentages would print 0.0 for both.
Example
Chapter 10’s conditional probability, applied to the obvious suspect.
| Days | Exceeded 100 | Rate | |
|---|---|---|---|
| Temperature at least 85 F | 31 | 4 | 12.9% |
| Temperature below 85 F | 85 | 3 | 3.5% |
A hot day is 3.7 times more likely to be a bad air day, which is a large effect and almost certainly real on physical grounds.
Fisher’s exact test on the same table returns \(p = 0.0811\), with an odds ratio of 3.99 and an interval from 0.63 to 28.99.
Seven exceedances is not enough to prove a ratio of four. The interval spans a factor of forty-six. This is the Chapter 29 lesson in a new place: the effect is big, the evidence is thin, and the \(p\)-value is reporting the second of those rather than the first.
One last check, which is Chapter 12’s and takes ten seconds.
The mean is 42.1 and the standard deviation 33.0. Feed those to a normal distribution and ask for the chance of clearing 100 ppb: 0.0397, so 4.6 days out of 116. The data says 7. A lognormal fitted to the same readings says 9.9. Neither is right, and the normal is wrong in the more dangerous direction, understating a tail that matters.
Then ask the normal model something it was never asked:
\[P(\text{ozone} < 0) = 0.101\]
The normal distribution, fitted to these readings, predicts 11.7 days of negative ozone in a summer of 116.
That is not a subtle violation to be detected by a test. It is a model asserting something impossible about a tenth of the data, and it is visible in one line of arithmetic that nobody runs. A quantile plot would have shown it. So would drawing the fitted curve over the histogram, which the code below does.
The lesson is not that normality tests are needed. It is that a distribution has a domain, and a positive quantity with a mean near its own standard deviation cannot be normal, whatever a test says.
36.6 What Can Be Said
Three claims survive everything in this chapter, and they are worth separating from the ones that do not.
Ozone varies across the summer. \(F = 8.536\), and three other methods agree, including two that assume nothing the data violates. This is not in doubt.
The middle of the summer is higher than either end. July and August exceed May, and September falls below July and August. Four pairs, significant under every rule including the most conservative.
Nothing can be said about June. Not that it is low, not that it is like May, not that it differs from July. Its nine days were biased low by a known mechanism, correcting for that bias moves it to the middle of the pack, and no pairwise comparison involving it reaches significance under any adjusted rule.
What cannot be said, and would be easy to write: that ozone peaks in July and August and is lowest in May and June. The first half is supported. The second half quietly includes a month that has not earned a place in the sentence.
36.6.1 Five Months, and the Days Behind Them
Chapter 15’s interval, drawn over the readings that produced it, so the reason June’s is wide is on the same picture as the width.
36.6.2 The Test, and Whether It Was Allowed
Chapter 19’s \(F\), then the two conditions it came with, then the same question asked three other ways.
36.6.3 Which Pair, and by Whose Rule
All ten comparisons under four defensible procedures. Two of the ten change verdict depending on which is used, and both involve June.
36.6.4 How Often Was the Air Bad?
Module III on the same data: exceedance rates with exact intervals, one conditional probability, and what a normal distribution would have claimed.
Four outputs, and two different answers to one question. Five monthly means whose intervals are set more by how many days survived than by how variable the month was, with June’s the second widest on the smallest spread. An \(F\) of 8.536 that violates both of its stated conditions and is confirmed anyway by three methods that do not need them. Ten pairwise comparisons of which four are significant under every rule, four involve June and none of those reach significance under any adjusted rule, and two flip verdict depending on a choice made after seeing the data. And a normal distribution, fitted to a positive quantity, quietly predicting 11.7 days of negative ozone.
Recap
These two chapters were about everything that happens before a model, and how much of the eventual answer is settled there. Chapter 29 read the file as a file:
MonthandDaywere a date taken apart and stored as integers, 37 of the 153 ozone readings were missing, and three tests comparing measured against unmeasured days on solar radiation, wind and temperature all came back flat. Every one of those tests was correctly run and all three were misleading, because the thing that predicted a hole was not a measurement but the calendar, and the calendar disagreed at \(\chi^2 = 44.75\) on 4 degrees of freedom. Twenty-one of June’s thirty days were never measured, the nine that survived were the windy ones, and wind costs 5.55 ppb of ozone per mph, so June’s 29.4 is biased low by about 10.6. Chapter 30 then asked the question the data was collected for, and got a clean answer and a dirty one. Clean: ozone varies across the summer, \(F = 8.536\), confirmed by a log-scale ANOVA, by Welch and by Kruskal-Wallis, so the failure of both stated assumptions changed nothing. Dirty: the overall test says nothing about which months, and of the ten pairs, the two that involve June and a peak month flip across 0.05 depending on whether the analyst pays for multiplicity, while applying Chapter 29’s own correction moves them from 0.075 to 0.582. Underneath both chapters sits the same shape: a right skewed response whose skewness of \(+1.226\) falls to \(-0.555\) under a logarithm, which repairs the ANOVA’s assumptions, and whose raw-scale normal model predicts negative ozone on a tenth of the days. That shape is now a decision waiting to be made rather than a curiosity, and it is the first thing the next chapter has to settle.
Summary
| Concept | Description |
|---|---|
| Before Anything | |
| No Model Yet | Modules III to V on the data as Chapter 29 left it |
| Two Answers, One Dataset | Both correct, and the reported one depends on a later choice |
| Five Intervals | |
| The Interval Over the Readings | The reason an interval is wide belongs on the same picture |
| Width Is Mostly n | Of x-bar, s and root n, the sample size does the work here |
| June's Smallest SD, Second Widest Interval | 18.2 against May's 22.2, and 1.56 times the width |
| August Earns Its Width | A spread of 39.7 on 26 days justifies an interval of 32.1 |
| A Near Miss of 2.9 ppb | June's interval fails to reach July's, which matters later |
| The Test | |
| F = 8.536 | On 4 and 111 df, p = 4.8e-06, and not in doubt |
| Two Conditions | Equal variance and symmetric residuals, both stated by Chapter 19 |
| Bartlett Fails | K-squared 13.45, p = 0.0093, on an SD ratio of 2.18 |
| Residuals Inherit the Skew | Residual skewness +1.059 against the response's +1.226 |
| What the Log Repairs | Bartlett rises to 0.355 and residual skew falls to -0.609 |
| Robustness | |
| Four Routes, One Answer | Raw, log, Welch and Kruskal-Wallis, largest p 0.000064 |
| A Failed Assumption That Cost Nothing | Less trustworthy is not wrong; run the alternatives and look |
| Which Pair | |
| Overall Is Not Pairwise | An F says not all equal, and never says which |
| Ten Pairs From Five Groups | Testing all ten at 0.05 risks a false positive about half the time |
| Four Significant Pairs | July and August above May, September below both |
| None of Them Involve June | Not May, not September, not July, not August |
| Statistically Invisible | Nine days can differ from everything and prove nothing |
| Whose Rule | |
| Welch, Pooled, Tukey, Bonferroni | Four correct procedures for July against June |
| 0.05 in the Middle of Them | 0.0022, 0.0102, 0.0749, 0.1023, and the rule decides |
| The Rule Is Chosen After Looking | Nothing in the output labels the choice as a choice |
| Decide the Rule First | And report the rules you did not use |
| The Correction | |
| Applying the Correction | June moves from 29.4 to 40.1, into the middle of the summer |
| F Falls to 7.609 | From 8.536, so the evidence gets weaker, not stronger |
| 0.0749 Becomes 0.5815 | The pairs nearest the threshold move furthest from it |
| The Incentive Not to Correct | No reader of the final table can tell whether it was applied |
| How Often | |
| Nobody Breathes an Average | The reader wants to know how often the air was bad |
| Seven Days Over 100 ppb | 6.0 per cent, with an exact interval of 2.5 to 12.0 |
| Zero of Nine Versus Zero of 29 | Both print as 0.0 per cent; the intervals reach 33.6 and 11.9 |
| Hot Days Are 3.7 Times Worse | 12.9 per cent against 3.5 per cent, and physically plausible |
| And Fisher Says 0.0811 | Seven exceedances cannot prove a ratio of four |
| The Wrong Model | |
| The Normal Tail Understates | 4.6 days predicted against 7 observed, wrong in the worse direction |
| Negative Ozone on a Tenth of Days | P(ozone < 0) = 0.101, which is 11.7 days of the 116 |
| A Distribution Has a Domain | A positive quantity with sd near its mean cannot be normal |
| The Verdict | |
| What Survives | Ozone varies; the middle exceeds both ends; June says nothing |
| What Would Have Been Easy to Write | Peaks in July and August and is lowest in May and June |