29  Data Visualization

Every module so far has ended in a number. A mean, an interval, a p-value, a coefficient, a diagnostic. Those numbers are the analysis, and they are not the finding, because a finding is something a reader understands and acts on, and almost nobody understands a table.

This is not a complaint about presentation. A table asks the reader to redo the analysis. Forty residuals contain the fact that one store is running the regression, and no one has ever noticed it by reading down the column. Five regional totals contain the fact that the regions are indistinguishable, and that fact survives a table only if the reader does the subtraction themselves. The picture is not a courtesy extended after the work is finished. For a good many findings it is the only form in which the finding exists.

The trouble is that a chart is also the easiest place in the whole pipeline to mislead, and the only place where it can be done with every individual number correct.

This module treats visualization as a technical subject with right and wrong answers, rather than a matter of taste. Chapter 25 sets out the grammar: a chart is a mapping from variables onto visual channels, those channels are read with measurably different accuracy, and choosing a chart is choosing a task to give the reader’s eye. Chapter 26 works through the four questions data is usually asked, about distribution, comparison, relationship and change, and which form each one deserves.

Chapter 27 is about the ways a chart lies while remaining factually accurate, and Chapter 28 is about designing for a reader who will never get the chance to ask you what you meant.

Nothing here is software specific. The examples run in R and in Python, and the arguments would survive a change of either.