31 Choosing the Form the Data Deserves
Chapter 25 settled which visual channels are read accurately. That is half the problem. It does not tell you whether this particular column of numbers wants a histogram, a box plot, a strip of points or none of the three.
Before the four questions data is usually asked, a debt from Module II is worth collecting. Chapter 5 gave the mean, Chapter 6 the standard deviation, Chapter 7 Karl Pearson’s coefficient of skewness and Chapter 8 the moment-based kurtosis. Four numbers describing the shape of a distribution, computed carefully, over four chapters.
Not one of those chapters drew a distribution.
That was not an oversight so much as the ordinary way the subject is taught, and this chapter is where the bill comes due, because two datasets can agree on all four of those numbers to two decimal places and look nothing whatever alike.
31.1 The Four Questions
Almost every chart in practice is answering one of four questions, and knowing which one narrows the choice before any aesthetic judgement is needed.
| Question | Shape of the data | Usual forms |
|---|---|---|
| Distribution: what does this one variable look like? | One column | Histogram, density, box plot, strip, violin |
| Comparison: how do these categories differ? | One measure by one category | Dot plot, bar chart |
| Relationship: do these two move together? | Two columns | Scatter, binned scatter, smoother |
| Change: what happened over time? | A measure indexed by time | Line, slope chart, small multiples |
The list is short because the questions are few. What makes the choice hard is that the same data answers different questions depending on what you want to say, and that the default form for each of these four is the one that conceals the most.
For distribution the default is the histogram, and its bin width decides the finding. For comparison the default is the bar chart, which spends a great deal of ink to encode with length what position would encode better. For relationship it is the scatter, which stops working somewhere around a few thousand points. For change it is the line chart with every series on it, which stops working at about five.
31.2 Distribution, and the Bin Width Argument
A histogram has one parameter and it is not a display setting. The bin width decides how many groups the reader sees.
The trade is the same one that runs through all of statistics. Wide bins average over more observations, so the bars are stable, and they average away real structure. Narrow bins keep the structure and show every accident of sampling as though it were a feature. There is no width that is correct for all purposes, which is why the standard rules exist and why they disagree.
| Rule | Bins or width | Assumes |
|---|---|---|
| Sturges | \(\lceil \log_2 n \rceil + 1\) bins | The data is roughly Normal and \(n\) is small |
| Scott | \(h = 3.49\,s\,n^{-1/3}\) | The data is roughly Normal, minimising squared error |
| Freedman and Diaconis | \(h = 2\,\text{IQR}\,n^{-1/3}\) | Nothing much; uses the IQR, so outliers do not widen it |
Sturges is the oldest, is derived for small samples from a Normal distribution, and is the default in R’s hist(). At large \(n\) it gives far too few bins, because \(\log_2 n\) grows very slowly: ten thousand observations get fourteen bins and a million get twenty-one.
Freedman and Diaconis is the safe default for real data, precisely because it makes the fewest assumptions and is not dragged around by a handful of extreme values.
Example
Six thousand order values from a retailer. The same six thousand numbers, drawn six times.
| Bins | Bin width | What the reader sees |
|---|---|---|
| 6 | 11.7 | One group, centred near 50 |
| 10 | 7.0 | One group, centred near 50 |
| 18 | 3.9 | One group, slightly flat-topped |
| 40 | 1.8 | Two separate groups |
| 90 | 0.8 | Two groups, clearly |
| 300 | 0.2 | Two groups, and a lot of noise |
There are two customer segments in this business, a lower-spending one and a higher-spending one, with very few customers between them. At the first three widths that fact is invisible. At the last three it is the most obvious thing on the page.
Now apply the rules to those same numbers. Sturges asks for roughly fourteen bins. Scott and Freedman and Diaconis ask for well over a hundred. The default histogram in R hides the segments; the better-founded rules reveal them, and nobody had to be dishonest for that to happen.
Bin width is not the only decision. Bin placement matters too, and it is usually made silently by the software.
With the segments near 42 and 58, bins of width 10 starting at zero put both humps inside the bins running 40 to 50 and 50 to 60, and the dip between them lands in the middle of nothing. Shift every boundary by five and the same width resolves the two groups.
This is why the histogram is the one chart worth drawing more than once before believing. If the story survives three sensible bin widths it is probably in the data. If it appears at one width and vanishes at the next, it was in the bins.
31.3 What Four Numbers Could Not Tell You
Chapter 21 showed Anscombe’s quartet: four datasets with the same mean, variance, correlation and regression line, and four completely different scatter plots. The same demonstration can be built for the distribution of a single variable, and it lands harder, because the summaries it defeats are the ones Module II spent four chapters computing.
Below are two datasets of twenty thousand order values. They agree, to two decimal places, on the mean, the standard deviation, the coefficient of skewness and the excess kurtosis.
One of them is a single bell. The other is two clearly separated customer segments with a gap between them where almost nobody sits.
Example
| Statistic | Dataset A | Dataset B |
|---|---|---|
| Mean | 50.00 | 50.00 |
| Standard deviation | 10.00 | 10.00 |
| Skewness | 0.00 | -0.01 |
| Excess kurtosis | 0.03 | 0.03 |
| Median | 50.0 | 50.2 |
| Observations within 5 of the mean | 38.5% | 12.6% |
Every number Module II taught you to compute agrees. The last row is the one that gives it away, and it is not a summary anybody reports.
Chapter 7 offered this reading: “\(Sk \approx 0\): roughly symmetric; the mean is a trustworthy measure of the center.” Dataset B is perfectly symmetric and its skewness is essentially zero, so it passes that test. Its mean is 50. Its mean is also a value that almost none of its customers spends, because it falls in the valley between the two segments.
The coefficient was not wrong. The inference drawn from it was, and no additional coefficient would have caught it, because the four moments were matched on purpose. What catches it is the picture, in about a second.
The general point is worth stating plainly, because it applies well beyond this example.
A summary statistic is a projection, and every projection discards. The mean discards everything but the centre of mass. Adding the standard deviation recovers the spread and nothing else. Skewness and kurtosis add two more directions, and four numbers is still four numbers against a distribution that has infinitely many degrees of freedom.
This is not an argument against summaries. It is an argument about their order of use: plot first, summarise second, and quote the summary only once you have seen what it is standing in for.
31.4 The Box Plot, and the Five Numbers It Draws
The box plot, due to Tukey, draws a five-number summary: minimum, lower quartile, median, upper quartile, maximum, with the whiskers usually stopping at the most extreme point within 1.5 times the interquartile range of the box and anything beyond shown individually.
It is excellent at the job it was built for, which is comparing many distributions at once. Twenty box plots side by side are readable; twenty histograms are not. It is also compact, robust to outliers, and shows the median rather than the mean, which matters for skewed data.
What it cannot do is show shape, and this is not a limitation to be worked around but a definition. A box plot draws five numbers. Any two datasets agreeing on those five numbers produce identical box plots, however different they are.
Example
Four datasets of four thousand observations each, constructed to share a five-number summary exactly: minimum 0, lower quartile 20, median 50, upper quartile 80, maximum 100. Their means are all exactly 50.
| Dataset | Shape | Standard deviation |
|---|---|---|
| Even | Roughly flat across the range | 31.10 |
| Two groups | Dense near 20 and near 80, hollow between | 30.30 |
| One centre | Piled up around 50 | 26.33 |
| Both edges | Piled at 0 and 100, hollow in the middle | 36.31 |
The four box plots are identical, mark for mark. The four densities are not remotely alike, and one of them is the exact opposite of another: “One centre” has most of its mass where “Both edges” has almost none.
Note the last column. The standard deviations differ by 38 per cent between the extremes, and the box plot does not show that either, since the standard deviation is not one of its five numbers.
Three practical responses, in increasing order of effort.
Overlay the points. For up to a few hundred observations, draw the box and scatter the actual points over it with a little horizontal jitter. Nothing is lost and the shape returns.
Use a violin plot for larger \(n\), which is a density drawn symmetrically about the axis. It shows the shape the box plot hides, at the cost of a smoothing parameter that has the same arbitrariness as bin width.
Keep the box plot and check a histogram once. The most common situation is many groups to compare, where the box plot is genuinely the right display. Draw the histograms once during the analysis, satisfy yourself that no group is bimodal, then publish the box plots.
31.5 Comparison
Comparing a measure across categories is the commonest chart in business, and the bar chart is its default. Chapter 25 explains why the dot plot usually beats it: the bar encodes with length, which is the third channel and demands a zero baseline, while the dot encodes with position, which is the first and does not.
Bars earn their place in two situations. When the quantity is a count or an amount that genuinely starts at zero and the zero is meaningful to the comparison. And when there are few categories and the values differ enough that the ink is not misleading anybody.
Dots win when the variation is small relative to the level, which in business data is most of the time. Revenue by region, satisfaction by branch, delivery time by depot: all of these vary by a few per cent around a level far from zero, and all of them produce bar charts that are either uninformative from zero or dishonest when truncated.
Two decisions matter more than the choice between them.
Sort by value, not by name. Alphabetical order is an accident of the labels. Sorting by the measure turns the chart into a ranking that can be read at a glance, and makes the near ties visible as near ties.
Put the categories on the vertical axis when the labels are words. Horizontal bars and dots leave room for readable labels; vertical ones force rotated text, which is slower to read and looks like an apology.
31.6 Relationship, and Too Many Points
The scatter plot is the best chart in statistics and it has one failure mode: it stops working when the points start landing on each other.
Overplotting is not a cosmetic problem. Once a region is saturated, a cell holding fifty points is exactly the same colour as a cell holding one, so the densest and most important part of the picture is the part carrying no information at all. The eye then reads the outline of the cloud, which is set by the rare extreme points, and the shape of the bulk is invisible.
Three remedies, roughly by how many points they survive.
Transparency. Draw each point at low opacity so that overlap accumulates into shading. Cheap, keeps every observation, and stops working somewhere around a hundred thousand points, when even 1 per cent opacity saturates.
Bin the plane. A two-dimensional histogram, drawn as squares or hexagons shaded by count, scales to any number of rows because the work no longer depends on \(n\). The cost is that individual points, including the genuinely interesting outliers, disappear into their cells.
Sample. Draw a random few thousand. Honest and often sufficient, and worth saying in the caption, because a reader is entitled to know they are not looking at all of it.
The best version of the last two combined: bin the bulk, and draw individually only the points in sparsely populated cells, so the density is shaded and the outliers survive.
Example
Thirty thousand points with a correlation of 0.55, plotted at full opacity on a 300 by 300 grid of cells.
| Points drawn | 30,000 |
| Distinct cells they land in | 14,918 |
| Points per occupied cell | 2.0 |
| Busiest cell holds | 13 points |
| Cells holding exactly one point | 7,484 |
Half of the occupied cells hold a single point, and those cells contain only a quarter of the data. The other three quarters are stacked on top of each other, and at full opacity every one of those cells is rendered in exactly the same black as a cell holding one.
The correlation is 0.55, which is a strong and perfectly visible relationship, and the opaque version of this chart cannot show it. Turn the opacity down to 2 per cent and the whole density gradient appears from the same thirty thousand numbers.
31.7 Change Over Time
Time is the one variable that earns a line. A line between two points asserts that the quantity passed through the values in between, which is true for a measurement that exists continuously and false for most other things.
Connect points when the underlying quantity is continuous in time: a price, a temperature, a queue length, a running total. Do not connect them across categories, however tempting, because there is no “between” for region or product line to pass through.
The failure mode of the line chart is the spaghetti plot: every series on one pair of axes, distinguished by colour. It works up to about five series, which is roughly the number of hues a reader can hold apart without constantly returning to the legend, and past that it becomes a picture of the fact that there are many series.
The remedy is small multiples: one small panel per series, all drawn on identical axes, arranged in a grid. Chapter 25’s ranking explains why this works rather than merely looking tidy. Within a panel, comparison is position on a common scale, the first channel. Across panels, it is position on non-aligned scales, the second. The spaghetti plot uses the first channel within a series and colour, the last, to tell the series apart.
A useful middle course when one series matters most: draw all of them in grey, and the one that matters in a single strong colour, labelled directly. The context is preserved, the subject is unmistakable, and no legend is needed.
A log scale is honest in three situations and evasive everywhere else.
When the quantity grows multiplicatively, such as compound interest, adoption curves or anything spreading, a log scale turns exponential growth into a straight line, and a change of slope into a change of growth rate, which is the thing anybody actually wants to see.
When the range spans orders of magnitude, so that a linear axis devotes almost all its space to the largest values and crushes everything else to the baseline.
When ratios are the comparison, because equal distances on a log axis are equal ratios. A doubling occupies the same height wherever it starts, so a 5 to 10 rise looks exactly as large as 50 to 100, which is right if you care about growth and wrong if you care about amounts.
Everywhere else it flattens differences, and a chart that flattens a difference somebody would rather you did not notice is doing the same work as a truncated axis, only in the opposite direction. Label a log axis clearly and never use one without saying so, because a reader who misses it will misread every distance on the chart.
Three outputs, one argument. Twenty thousand numbers can match another twenty thousand on the mean, the standard deviation, the skewness and the kurtosis and still be a different distribution entirely. Four datasets can share a five-number summary exactly and share no shape at all. And thirty thousand points drawn at full opacity put three quarters of themselves underneath each other, so that the densest part of the picture is the part telling you nothing.
Recap
Chapters 25 and 26 are the two halves of choosing a chart. Chapter 25 established that a chart is a mapping onto visual channels, that those channels are read with measurably different accuracy, and that position beats length beats angle beats area beats colour. Chapter 26 asked the question that comes first in practice: given this data and this question, what form is honest? The four questions turn out to be few, and the default answer to each conceals the most. A histogram’s bin width is an argument, not a setting, and the same six thousand orders are one customer segment or two depending on where the edges fall. A box plot draws five numbers, so four datasets sharing those five numbers produce four identical pictures over four unrecognisably different shapes. A scatter plot saturates, and at thirty thousand points three quarters of the ink lies underneath other ink, so the densest region is the one carrying no information. And a line chart holds about five series before the legend starts doing work the eye cannot do, which is what small multiples exist to fix. Running under all of it is the debt this chapter collected from Module II. Two datasets agreed to two decimal places on the mean, the standard deviation, the skewness and the excess kurtosis, and one of them was a bell while the other was two segments with a hole in the middle where its own mean sat. The coefficients were right. Every inference anybody would have drawn from them was wrong. What none of this has addressed is the chart that is not badly chosen but deliberately built to mislead, where every number is accurate, every label correct, and the picture still argues for something the data does not support. The next chapter takes those apart one at a time.
Summary
| Concept | Description |
|---|---|
| Choosing a Form | |
| The Four Questions | Distribution, comparison, relationship, change over time |
| Defaults Conceal the Most | Histogram, bar chart, scatter and multi-series line each hide something |
| The Histogram | |
| Bin Width Is an Argument | One parameter that decides how many groups the reader sees |
| The Bias and Variance Trade | Wide bins are stable and hide structure; narrow bins show noise as feature |
| Sturges | Log base 2 of n plus one; derived for small Normal samples |
| Scott | 3.49 times s over the cube root of n; assumes roughly Normal |
| Freedman and Diaconis | 2 times the IQR over the cube root of n; assumes little, resists outliers |
| R's Default Histogram | Sturges, which at twenty thousand observations gives fourteen bins |
| Bin Placement | Where the edges fall matters as much as how wide they are |
| Draw It Three Ways | If a feature survives three sensible widths it is in the data |
| What Summaries Discard | |
| A Summary Is a Projection | Every summary discards, and four numbers cannot pin down a distribution |
| Matching Four Moments | Two datasets agreeing on mean, SD, skewness and kurtosis to two decimals |
| What Skewness Near Zero Does Not Mean | A perfectly symmetric distribution can still be two separate groups |
| The Mean in a Valley | Dataset B's mean is a value almost none of its customers spends |
| Plot First, Summarise Second | Quote a summary only once you have seen what it is standing in for |
| The Box Plot | |
| The Five-Number Summary | Minimum, lower quartile, median, upper quartile, maximum |
| What a Box Plot Is Good At | Comparing many distributions at once, compactly and robustly |
| What a Box Plot Cannot Show | Shape, by definition, since it draws five numbers and nothing else |
| Identical Boxes, Four Shapes | Even, two groups, one centre and both edges, all drawing the same box |
| The Standard Deviation Is Not in the Box | Their standard deviations differ by 38 per cent and the box hides it |
| Overlay the Points | For a few hundred observations, jitter the raw points over the box |
| Violin Plots | A density drawn symmetrically, with bin width's arbitrariness returning |
| Comparison | |
| Dot Plot Against Bar Chart | Position beats length, and needs no zero baseline |
| When Bars Earn Their Place | When the quantity genuinely starts at zero and the zero matters |
| Sort by Value | Alphabetical order is an accident of the labels, not of the data |
| Categories on the Vertical Axis | Horizontal layout leaves room for labels nobody has to rotate |
| Too Many Points | |
| Overplotting | Points landing on each other, which is not a cosmetic problem |
| Saturation Destroys Density | A cell with fifty points is the same colour as a cell with one |
| Transparency | Cheap, keeps every observation, fails around a hundred thousand points |
| Binning the Plane | A two-dimensional histogram, which scales to any number of rows |
| Sampling and Saying So | Honest, often sufficient, and the reader is entitled to be told |
| Change Over Time | |
| Connect Points Only Across Time | A line asserts the quantity passed through the values in between |
| The Spaghetti Plot | Every series on one pair of axes, readable up to about five |
| Small Multiples | One panel per series on identical axes, using the first two channels |
| Grey the Rest | Draw the context in grey and the subject in one strong colour, labelled |
| Log Scales | |
| When a Log Scale Is Honest | Multiplicative growth, ranges spanning orders of magnitude, ratio comparisons |
| Equal Distance, Equal Ratio | A doubling occupies the same height wherever it starts |
| Never Use One Silently | A reader who misses the log axis misreads every distance on the chart |