34 Analysis From End to End
Every module so far has taught a technique on data chosen to show that technique working. The data arrived clean, the assumption held, the effect was there to be found, and the chapter ended when the number appeared. That is the right way to learn a method and a misleading picture of what using one is like.
This module runs a single analysis from the raw file to the finished report, and the difference shows up immediately. The chapters stop being separable. A decision made in Module II about the shape of a variable determines the form of a model three chapters later. An audit of missing values, which is not a statistical technique at all, changes the size of an estimate and the order of a ranking. A diagnostic plot sends the work back to a choice made before the model existed.
The dataset is the New York air quality series, 153 daily readings of ozone, solar radiation, wind and temperature from May to September. It is small enough to hold in your head, old enough to be free of any commercial sensitivity, and it ships with R, so every number in this module can be reproduced from a single line of code. It also has a flaw, discovered in Chapter 29, that no summary of it reports.
The question is the one anybody would ask: does ozone vary across the summer, and what explains it?
Chapter 29 reads the file as a file, before reading it as data, and finds that a quarter of the ozone readings are missing for a reason. Chapter 30 asks what the data says on its own, with the intervals of Module IV and the comparison of Module V. Chapter 31 builds a model and then, more importantly, checks it. Chapter 32 reports the result to somebody who was not in the room.
Nothing new is taught here. That is the point: everything used is something you already have, and the difficulty is entirely in the order, the judgement and the honesty.