35 The Question and the Data
Twenty-eight chapters have each taught a technique on data chosen to show that technique working. This module does the opposite. One dataset, one question, and every module in the book applied to it in the order a real analysis actually runs.
The data is the New York air quality series: daily readings of ground level ozone, solar radiation, wind speed and temperature, from May to September. One hundred and fifty-three days, six columns. It ships with R under the name airquality, which means anybody reading this can reproduce every number on the page.
The question is the obvious one. Does ozone vary across the summer, and what explains it?
That question will eventually need a model, and Chapter 31 will build one. This chapter never gets there, because the first thing a real analysis discovers is that the file is not what it appears to be, and the discovery changes the answer before a model is ever fitted.
35.1 What the File Says It Is
Six columns and a hundred and fifty-three rows, which is exactly what a summary function reports and exactly what a naive reader takes away.
Two of those columns are not measurements. Month and Day hold the integers 5 to 9 and 1 to 31, and a date has been taken apart and stored as two numbers. Nothing in the file objects. Ask for the mean of Month and R returns 6.99, a number that is arithmetically correct and refers to nothing. Chapter 1 called this the difference between a measurement and a label, and this is what it looks like when the distinction is not enforced by the storage.
It matters here for a practical reason rather than a pedantic one. Because the calendar is stored as ordinary numbers, anything that happened on a schedule will be invisible to a summary and visible only if somebody thinks to look along the calendar. Something did.
35.2 The Count Is Not the Audit
Thirty-seven of the 153 ozone readings are missing, which is 24.2 per cent. Seven solar readings are missing. Wind and temperature are complete.
Drop every row with anything missing and 111 remain, so the convenience of a complete-case analysis costs 42 rows, more than a quarter of the data, before a single estimate is computed. That is the cost every default reports.
What no count reports is which days went missing, and that is the only part that decides whether the drop was safe. Twenty-four per cent missing at random is an inconvenience. Twenty-four per cent missing for a reason is a different dataset from the one you think you have, and the two are indistinguishable from the count.
So the audit is not “how many” but “are the missing days different from the measured ones”, and that question can be asked directly, because wind and temperature were recorded on every single day including the days ozone was not.
Example
Compare the 116 days that have an ozone reading against the 37 that do not, on each of the other measured columns.
| Column | Measured days | Missing days | Gap | \(p\) |
|---|---|---|---|---|
| Solar.R | 184.8 | 189.5 | 4.71 | 0.785 |
| Wind | 9.9 | 10.3 | 0.39 | 0.545 |
| Temp | 77.9 | 77.9 | 0.05 | 0.979 |
Three comparisons, three flat results. The days without an ozone reading were no sunnier, no windier and no warmer than the days with one.
A reasonable analyst stops here, concludes the holes are harmless, drops the incomplete rows and gets on with it. That conclusion is wrong, and nothing in this table reveals it.
35.3 Where the Holes Are
Draw one square per day, five rows for five months, filled where there is a reading. The picture answers in one second a question three hypothesis tests could not answer at all.
June is almost empty. Twenty-one of its thirty days were never measured, against five missing in May, five in July, five in August and one in September.
The three tests were silent because they tested the wrong thing. Each asked whether some measurement differed between measured and unmeasured days. The reason the days went missing was not a measurement. It was the calendar, and the calendar was sitting in the file the whole time as two integer columns that no summary would ever compare against anything.
This is the chapter’s central point, and it survives being stated plainly: a missing value audit that only looks at columns will miss any pattern that lives in the rows.
Example
The same 153 days, tabulated by month.
| Month | Measured | Missing | Days |
|---|---|---|---|
| May | 26 | 5 | 31 |
| June | 9 | 21 | 30 |
| July | 26 | 5 | 31 |
| August | 26 | 5 | 31 |
| September | 29 | 1 | 30 |
Chapter 18’s test applies exactly as written: \(\chi^2 = 44.75\) on 4 degrees of freedom, \(p = 4.5 \times 10^{-9}\).
Whether a day was measured depends on the month, and it is not a close call. June contributes 30 of the days and 21 of the 37 holes.
The standard vocabulary for this has three terms, and it is worth having them because they decide what you are allowed to do next.
Missing completely at random (MCAR) means the holes are unrelated to anything, observed or not. Dropping the rows is safe and costs only precision.
Missing at random (MAR) is the badly named middle case: the holes depend on things you did observe. Here they depend on the month, and the month is in the file, so the dependence can be modelled and partly undone.
Missing not at random (MNAR) means the holes depend on the unobserved value itself, for example an instrument that fails precisely when ozone is high. Nothing in the data can rule this out, and no amount of cleverness recovers from it.
This dataset is at least MAR. The three flat tests would have supported MCAR, and MCAR is the assumption a complete-case analysis silently makes. The difference between those two assumptions is worth roughly ten parts per billion, as the next section shows.
35.4 What the Hole Costs
Knowing the holes are concentrated in June is not yet an answer. It matters only if the nine days that were measured in June are unrepresentative of the twenty-one that were not, and whether they are can be checked, because wind and temperature exist for all thirty.
Inside June, the nine measured days were 12.2 mph on average and the twenty-one unmeasured days 9.4 mph. The measured days were the windy ones.
Wind is not a neutral variable here. Across all 116 measured days, the correlation between wind and ozone is \(-0.602\), and a simple regression puts the slope at \(-5.55\) parts per billion per mph, with \(p = 9.3 \times 10^{-13}\). Wind disperses ozone. Chapter 21 would call that a strong negative relationship and Chapter 22 would fit exactly this line.
So the arithmetic is unavoidable. The nine measured June days were 2.73 mph windier than the twenty-one that were not, and at 5.55 ppb per mph that is 15.2 ppb of ozone. June’s nine surviving days are not a small sample of June. They are a biased one, biased low, by an amount the data itself can estimate.
Example
Two different corrections follow, and they answer two different questions. Keeping them apart matters.
| ppb | |
|---|---|
| June’s mean over the nine measured days | 29.4 |
| Estimate for the twenty-one unmeasured days | 44.6 |
| June as a whole, the thirty days recombined | 40.1 |
The measured mean of 29.4 makes June the second quietest month of the summer, below September’s 31.4. Corrected, June sits at roughly 40 and moves above September. A ranking changes, on a dataset where nothing was mistyped and no test was misapplied.
The correction is crude and should not be reported as a finding. It uses one predictor and no uncertainty. Its purpose is to establish the size and direction of what was lost, and both are large enough to matter.
There is one more thing in the June comparison worth extracting, because it is the single most misread number in applied work.
The wind gap inside June was tested and returned \(p = 0.0900\). Above the usual threshold. A report that stopped there would say the measured and unmeasured June days did not differ significantly in wind.
Now ask Chapter 16’s other question. The gap was 2.73 mph against a pooled standard deviation of 3.61, so the standardised effect is 0.756, which is large. With nine days in one group and twenty-one in the other, a gap that size is detected about 43 per cent of the time. Fewer than half. Reaching the conventional 80 per cent would take about twenty-nine days in each group, and June only has thirty days in total.
That \(p\) of 0.09 is a statement about how few days survived, not about the wind. The test was underpowered by exactly the same shortage of data that made the test necessary, which is the trap: the missingness both creates the question and destroys the ability to answer it.
35.5 The Shape Before the Model
One more thing has to be established before any model, and it is Module II’s business.
Across the 116 days with a reading, ozone has a mean of 42.1 and a median of 31.5. The mean sits 10.6 ppb above the median, the standard deviation is 33.0, and the skewness is \(+1.226\). The largest reading, 168 ppb, is 3.8 standard deviations above the mean.
Chapter 7 defined that coefficient and this is what it is for. A skewness of 1.2 is not a curiosity about the histogram; it is a warning about every technique that is about to be applied. The mean is not the typical day. Intervals built on a symmetric assumption will be wrong in a predictable direction, and a regression that assumes constant, symmetric error will have residuals that fan out.
Take logs and the skewness becomes \(-0.555\), near enough to symmetric to work with. That single fact is the reason Chapter 31 will end up modelling the logarithm rather than the raw value, and it is knowable now, before any model exists, from a histogram and one coefficient.
35.6 What This Chapter Bought
Nothing here was an analysis. No estimate has been reported, no hypothesis tested about the question that was actually asked. And yet three things are now known that change what the rest of the work is allowed to say.
The missing days are not missing at random, so any complete-case result carries a bias whose direction is known and whose size is roughly ten parts per billion in the month most affected.
June’s estimate rests on nine days, so whatever interval is reported for it will be the widest on the page, and honestly so.
Ozone is right skewed, so the response will need a transformation and the untransformed mean is not the summary anybody should quote.
None of the three would have been discovered by fitting the model first and checking the diagnostics afterwards. The file had to be read as a file before it could be read as data.
35.6.1 Reading the File for What It Is
The audit that every analysis should open with, and the map that the audit cannot produce on its own.
35.6.2 The Variable Nobody Tested
The calendar was in the file all along. Chapter 18’s test applied to it, and then the same comparisons from before, run inside June where they finally have something to find.
35.6.3 What the Hole Costs
Wind speaks for the days ozone cannot, and the price of the missing days is estimated in the units of the question.
35.6.4 The Shape Before the Model
Module II applied to the response, two chapters before a model needs it.
Four outputs, and the analysis has not started. A file that reports 153 rows, 6 columns and 37 missing ozone readings, with three tests agreeing that the missing days look like the measured ones. A calendar that disagrees at \(\chi^2 = 44.75\) on 4 degrees of freedom, because 21 of June’s 30 days were never measured. A correction that moves June from 29.4 to about 40 and past September in the ranking, resting on a wind gap whose own test returned \(p = 0.09\) with 43 per cent power. And a response with skewness \(+1.226\) that a logarithm brings to \(-0.555\), which decides the form of a model that does not yet exist.
Looking Ahead
This module runs one analysis from the file to the report, and the point of it is that the chapters stop being separable. Chapter 30 takes the data as this chapter leaves it and asks what it says on its own: monthly means with the intervals of Chapter 15, the comparison across months of Chapter 19, and the care required to state a finding when one of the five groups rests on nine days. Chapter 31 builds the model, ozone on temperature, wind and solar radiation, and then does the thing Chapter 24 insisted on, which is to check it rather than admire it, including whether the logarithm that Module II recommended actually earns its place. Chapter 32 reports the result to somebody who was not there, under the constraints Module VII established, and it is where the arc closes: the finding has to survive a single look from a reader who will not ask. What this chapter has fixed is what all three are allowed to claim. Any estimate that ignores June’s twenty-one missing days is biased low by a known amount in a known direction, and saying so is not a caveat, it is part of the result.
Summary
| Concept | Description |
|---|---|
| The Capstone | |
| One Dataset, Every Module | The techniques stop being separable once one question runs through them |
| The Question First | Does ozone vary across the summer, and what explains it |
| What the File Is | |
| A Date Stored as Two Integers | Month and Day are numbers, and nothing in the file objects |
| The Mean of Month Is 6.99 | Arithmetically correct and referring to nothing at all |
| Why That Matters Here | A pattern on the calendar is invisible to any column summary |
| The Count | |
| Twenty-four Per Cent Missing | 37 of 153 ozone readings, the number every default reports |
| The Cost of a Listwise Drop | 111 rows survive, so convenience costs 42 of them up front |
| How Many Is Not Which | The count cannot distinguish an inconvenience from a different dataset |
| Testing the Holes Against the Columns | Wind and temperature exist on every day, including the missing ones |
| Three Flat Results | Solar p 0.785, wind 0.545, temp 0.979, and all three are misleading |
| Where the Holes Are | |
| The Missingness Map | One square per day, five rows, filled where there is a reading |
| Twenty-one of Thirty | June's holes, against five in each of three other months |
| Chi-Square on the Calendar | 44.75 on 4 df, p = 4.5e-09, using Chapter 18 unchanged |
| Columns Versus Rows | A column-by-column audit misses any pattern living in the rows |
| Naming the Mechanism | |
| MCAR | Holes unrelated to anything, so dropping rows costs only precision |
| MAR | Holes depend on something observed, so they can be partly undone |
| MNAR | Holes depend on the unobserved value, and nothing recovers from it |
| What a Complete-Case Analysis Assumes | MCAR, silently, which is the assumption the three tests supported |
| What the Hole Costs | |
| A Complete Column Can Speak for a Missing One | Inside June, the nine measured days were the windy ones |
| Wind Disperses Ozone | r of -0.602 and -5.55 ppb per mph across 116 days |
| Biased, Not Merely Small | A 2.73 mph gap times 5.55 is 15.2 ppb of ozone |
| Two Different Corrections | 44.6 for the unmeasured days, 40.1 for June as a whole |
| A Ranking Changes | June moves from below September to above it |
| Why the Correction Is Not a Finding | One predictor, no uncertainty; it sizes the loss, it does not report it |
| Power, Not Evidence | |
| p = 0.09 with 43 Per Cent Power | A large effect, d of 0.756, caught fewer than half the times it occurs |
| Underpowered by the Same Shortage | The missing days create the question and destroy the power to answer it |
| Shape | |
| Skewness Before the Model | Mean 42.1, median 31.5, skewness +1.226 on the 116 measured days |
| The Mean Is Not the Typical Day | And intervals built on symmetry will be wrong in a predictable direction |
| What a Log Fixes | Skewness falls to -0.555, which decides the model two chapters early |
| What It Bought | |
| Three Things Now Known | Not random, resting on nine days, and right skewed |
| Read the File Before the Data | None of the three survives being left until the diagnostics |