30  The Grammar of Graphics

Twenty-four chapters have ended in a table. Chapter 24 closed by observing that tables are where findings go to be ignored, which was meant as a complaint about presentation and is really something harder: a table asks the reader to do the analysis a second time. Forty rows of residuals contain the information that one observation is running the model, and nobody has ever seen it there.

The usual response is to say that a picture is worth a thousand words, add a chart, and move on. That skips the only interesting question. A chart is not a decoration applied to a finding; it is a translation of numbers into something the eye can do arithmetic on, and eyes are extremely good at some kinds of arithmetic and extremely bad at others.

This chapter is about which is which. It is not about how to call a plotting function, and it takes no position on which software you use. It is about the fact that when you choose a chart type you are choosing a task to give your reader’s visual system, that these tasks have measurably different accuracies, and that the chart types people reach for by habit tend to sit at the wrong end of that list.

30.1 A Chart Is a Mapping

Strip away the software and every statistical graphic is the same object: a set of rules assigning variables to visual properties.

Part What it is Example
Data The rows and columns Forty stores, each with spend and revenue
Mapping Which variable goes to which visual property Spend to horizontal position, revenue to vertical position
Geometry The mark that gets drawn One point per store
Scale How data values become visual values 8 to 27 on spend becomes 0 to 500 pixels
Guides The parts that let it be read back Axes, ticks, labels, legend

This decomposition is due to Leland Wilkinson, whose book The Grammar of Graphics gives it its name, and it is the idea underneath ggplot2 in R and several libraries elsewhere. You do not need any of that software to use the idea, and the idea is the useful part.

The reason it matters is that it separates two decisions that are usually run together. “Should this be a bar chart or a pie chart?” is not a question about chart types at all. It is a question about whether to map revenue to length or to angle, and once it is asked that way, it has an answer that does not depend on taste.

The visual properties available are limited, and there are roughly eight of them in two dimensions.

Position along a common scale, position along scales that are not aligned, length, angle or slope, area, volume or curvature, colour lightness or saturation, and colour hue.

Everything else is a combination. A bar chart maps a category to position and a quantity to length. A pie chart maps a category to nothing at all and a quantity to angle. A scatter plot maps two quantities to two positions. A bubble chart adds a third quantity to area. A heatmap maps a quantity to colour lightness.

That list is not in a random order. It is in order of how accurately people read it, which is the subject of the next section, and it is the single most useful piece of knowledge in this chapter.

30.2 The Channels, Ranked

In 1984 William Cleveland and Robert McGill published a study in the Journal of the American Statistical Association that did something graphical design had mostly not done: it measured people. Subjects were shown the same quantitative comparison encoded different ways and asked to judge what percentage the smaller value was of the larger. The errors were recorded.

The resulting ordering, from most accurately read to least, is:

Channel Typical use
1 Position along a common scale Dot plot, scatter plot, any shared axis
2 Position along non-aligned scales Small multiples with separate panels
3 Length Bar chart, stacked segment
4 Angle or slope Pie chart, line chart gradients
5 Area Bubble chart, treemap
6 Volume or curvature Three-dimensional anything
7 Colour saturation or lightness Heatmap, choropleth
8 Colour hue Categorical colouring

Two things follow immediately and neither is a matter of opinion.

The ordering is about accuracy, not about beauty or engagement. A treemap can be the right choice when the question is “roughly which of these is big” and there are two hundred of them. It is the wrong choice when the question is “which of these two is bigger”.

A chart type does not sit at one position on this list. It sits wherever its mapping puts it. A bar chart used to compare categories relies on length, position 3. The same bar chart read against a gridline relies on position, position 1, which is why gridlines are worth having.

Example

The forty stores of Chapters 21 to 24, grouped into five regions of eight, and their total revenue.

Region Revenue Share Angle in a pie
Central 1793.1 20.95% 75.4 degrees
North 1777.2 20.76% 74.7 degrees
South 1719.7 20.09% 72.3 degrees
West 1653.4 19.32% 69.5 degrees
East 1616.6 18.88% 68.0 degrees

Ask the obvious question: which region earned the most?

In a pie chart you are being asked to tell 75.4 degrees from 74.7 degrees. The difference between the leading region and the runner-up is seven tenths of one degree, which no reader will see, and the difference between the best and worst of the five spans 7.4 degrees out of 360. Every slice looks like a fifth of the circle because every slice is about a fifth of the circle.

Put the same five numbers as dots on a shared horizontal axis and the ranking reads instantly, the gaps are visible, and the near tie between Central and North is visible as a near tie, which is the honest answer and the one the pie cannot give.

Nothing about the data improved. The question moved from the fourth channel to the first.

30.3 Why the Ranking Has That Order

The order is not arbitrary, and the reason is worth knowing because it lets you reason about cases the list does not cover.

Position along a common scale wins because it turns a comparison into a subtraction. Two marks on one axis differ by a distance, and judging which of two distances is larger, when both are measured from the same baseline, is close to the easiest thing a visual system does. Nothing has to be estimated; the answer is the offset.

Length is next because it is position with a free endpoint. A bar is read as a distance too, but each bar’s distance starts somewhere, and if the starting points are not aligned the comparison degrades to position along non-aligned scales. This is exactly what happens in the middle of a stacked bar chart, where only the bottom segment sits on a common baseline and every segment above it floats.

Angle is harder because the eye does not measure angles, it recognises them. Judgements cluster near the familiar values of 45, 90 and 180 degrees and are poor between them, and the same angle looks different depending on its orientation, which is why rotating a pie chart changes which slice looks biggest without changing any number in it.

Area and volume are worse still for a reason that has nothing to do with charts. Perceived magnitude does not rise in proportion to physical magnitude. Stevens’ power law describes the relationship as perceived magnitude rising as physical magnitude raised to some exponent, and for the area of a plane figure that exponent is around 0.7 rather than 1. A shape of genuinely twice the area does not look twice as big; it looks about 1.6 times as big.

Colour is last because it is not ordered until you force it to be, and because the eye judges a colour against its neighbours rather than against a legend. The same grey square looks lighter on a dark background and darker on a light one, and no key in the corner fixes that.

30.4 Area, and the Two Mistakes

Area deserves its own section because it is the channel people misuse without noticing, and because the misuse has two separate causes that push in opposite directions.

The first mistake is geometric and it is simply an error. To show a quantity as the size of a circle, the quantity must be proportional to the area, and area goes as the square of the radius. Set the radius proportional to the value and a doubling of the value becomes a quadrupling of the ink. The fix is arithmetic, not judgement: make the radius proportional to the square root of the value.

\[r_i = r_{\max}\sqrt{\frac{v_i}{v_{\max}}}\]

The second is perceptual and it cannot be fixed by arithmetic. Even a correctly sized circle is underread, because of the Stevens exponent of roughly 0.7 mentioned above. A circle of exactly twice the area is seen as about 1.6 times the quantity.

Notice that the two errors run in opposite directions, which is why the misuse survives: the geometric mistake inflates, the perceptual effect deflates, and the result is merely wrong rather than obviously absurd.

Example

Four product lines with values 10, 20, 30 and 40.

Line Value Radius set to the value Area that produces Radius set to the square root Area that produces
A 10 1.00 3.1 1.00 3.1
B 20 2.00 12.6 1.41 6.3
C 30 3.00 28.3 1.73 9.4
D 40 4.00 50.3 2.00 12.6

D is four times A. Drawn the naive way it carries sixteen times the ink. Drawn correctly it carries four times, as intended.

And on top of that, here is what a correct area is actually read as, with the exponent at 0.7:

True ratio Read as Shortfall
1.5 times 1.33 -11%
2 times 1.62 -19%
4 times 2.64 -34%
10 times 5.01 -50%

Put the two together. A value that has doubled, drawn with radius proportional to the value, produces four times the area, which is read as \(4^{0.7} = 2.64\) times. The reader concludes the quantity rose by 164 per cent when it rose by 100. The mistake and the compensating misperception do not cancel; they leave a 32 per cent overstatement.

30.5 Colour

Colour is two channels wearing one name, and almost every colour problem in practice comes from confusing them.

Hue is the property that distinguishes red from blue from green. Hue has no natural order: nobody can say whether green is more than orange. That makes it excellent for categories and disqualifying for quantities.

Lightness and saturation are ordered. Pale to deep is a sequence everyone reads the same way, which makes them usable for a quantity, though as the ranking says, usable comes last.

From this, three rules follow without any appeal to taste.

A quantity gets a sequential ramp, varying in lightness, so that the order of the data is the order of the ink.

A category gets distinct hues, and not very many of them, because the number of hues a reader can hold apart without constant reference to the legend is about six.

A quantity with a meaningful midpoint, such as change against a target, or a residual against zero, gets a diverging scale: two hues meeting at a neutral centre, so that the sign is visible before the size.

The rainbow scale, which varies hue to show a quantity, breaks the first rule and is still the default in a good deal of scientific software. It invents boundaries where the data is smooth, because the eye sees the yellow band as an edge, and it collapses real differences elsewhere. Chapter 27 returns to it as a way of lying without intending to.

Whatever the palette, roughly one man in twelve and one woman in two hundred has some form of colour vision deficiency, most often in distinguishing red from green. A chart whose entire message is carried by a red against a green is unreadable to a substantial minority of any audience, and the reader will usually not mention it.

Two defences, and the second is the real one. Choose palettes that vary in lightness as well as hue, so the marks stay distinguishable in greyscale. Better, do not let colour carry the message alone: label the series directly, or vary the shape, so that colour is a convenience rather than the only key.

30.6 Matching the Channel to the Variable

Chapter 2 divided variables into nominal, ordinal, interval and ratio. That classification has been quietly waiting for this chapter, because it decides which channels are honest.

A channel has properties: it may carry identity, order, difference or ratio. A variable has properties too. The rule is that the channel must not claim more than the variable has.

Variable Honest channels What goes wrong otherwise
Nominal (region, product) Position, hue, shape Length says one category is more than another
Ordinal (small, medium, large) Position, lightness Length claims the gaps between levels are equal
Interval (temperature, dates) Position Length claims a ratio that an arbitrary zero cannot support
Ratio (revenue, spend, counts) Position, length Angle and area are honest but read poorly

The middle column is short on purpose. Position appears in every row, because position claims only order and difference, which every variable above the nominal level has. That is the deepest reason it sits at the top of Cleveland and McGill’s list, and it is why the humble dot plot is so hard to beat.

30.7 Where the Zero Argument Actually Comes From

“Bar charts must start at zero” is repeated as a rule and resented as a rule, because the people repeating it usually cannot say why, and because everyone has seen a case where starting at zero throws away the whole signal.

The grammar answers it in one line. Length encodes ratio, and a ratio is measured from zero. A bar’s length is its message, so where the bar begins is part of the encoding: move the origin and you have changed what the picture says, even though no number changed.

Position carries no such claim. A dot at 1793 and a dot at 1617 on an axis running from 1600 to 1800 are two positions, and the distance between them reads as a difference of 176.5, which is exactly true. Nothing is measured from the axis end, so nothing is distorted by where the axis ends.

So the rule is not about bars and it is not about zero. It is: truncate a position axis freely, and never truncate a length. If the interesting variation is small and sits far from zero, that is not an argument for a truncated bar chart, it is an argument for a dot plot.

Example

The same five regions, drawn as bars two ways.

Region Revenue Bar height from 0 Bar height from 1600
Central 1793.1 100% 100%
North 1777.2 99% 92%
South 1719.7 96% 62%
West 1653.4 92% 28%
East 1616.6 90% 9%

Central earns 1.11 times what East earns. Truncated at 1600, Central’s bar is 11.63 times the height of East’s. The truncation multiplies the apparent ratio by a factor of 10.5, and it does so without altering a single number, a single label or a single tick mark. Everything on the chart is accurate and the chart is a lie.

Draw the identical numbers as dots on an axis from 1600 to 1800 and nothing is exaggerated, because no mark’s size is carrying a claim. The axis is tight, the differences are visible, and the picture says only what is true.

The three outputs above are one argument in three pieces. Five numbers that are genuinely within two percentage points of each other are unreadable as angles and obvious as positions. A quantity encoded as circle size is wrong by a factor of four before perception has even been consulted, and then underread by a third once it has. And a baseline moved by 1600 turns a ratio of 1.11 into an apparent ratio of 11.63 while every label on the chart stays true.

Looking Ahead

This chapter has been about a single shift of question. Not “what chart should I use”, which has no general answer, but “what am I asking the reader’s eye to do”, which has a measured one. Once a chart is seen as a mapping from variables onto channels, several arguments that are normally settled by taste settle themselves instead. Pie charts are not forbidden because someone disapproves of them, they are weak because angle is the fourth channel and because five near-equal shares differ by less than a degree. Bubble charts are not wrong in principle, they are wrong when the radius was set to the value, and underread even when it was not. Rainbow scales are not merely ugly, they map an ordered quantity onto an unordered property. And bars start at zero not because of a convention but because length is a ratio, so its origin is part of what the picture claims, while position carries no such claim and may be cropped as tightly as the data deserves. What none of this has settled is the choice that comes first in practice. Knowing that position beats angle does not tell you whether this particular column of numbers wants a histogram, a box plot, a strip of points or none of the three, nor what to do when there are forty thousand rows and the points have covered the page. The next chapter takes the four questions data is usually asked, about distribution, comparison, relationship and change over time, and works out which form each one deserves, starting with the choice that quietly decides what a histogram says: the width of its bins.


Summary

Concept Description
A Chart as a Mapping
A Chart as a Mapping A set of rules assigning variables to properties the eye can read
Data, Mapping, Geometry, Scale, Guides The five parts every statistical graphic is built from
The Grammar of Graphics Wilkinson's decomposition, and the idea underneath ggplot2
The Real Question Not which chart type, but which visual task the reader is given
The Channels
Visual Channels The eight or so properties available: position, length, angle, area, colour
Position on a Common Scale The most accurately read channel, because comparison becomes subtraction
Length Accurate, but only from a shared starting point
Angle Judged poorly except near 45, 90 and 180 degrees
Area Underread, and easy to encode wrongly in the first place
Colour Lightness Ordered, and deliberately imprecise; suitable for impressions
Colour Hue Unordered, so it suits categories and ruins quantities
The Ranking
The Cleveland and McGill Ranking The 1984 experiment that measured how accurately each channel is read
Accuracy, Not Beauty The ranking says nothing about which chart is engaging or attractive
A Chart Type Has No Fixed Rank A chart sits wherever its mapping puts it, not where its name does
Gridlines Promote Length to Position A bar read against a gridline is being read as position, not length
Why That Order
Why Position Wins It claims only order and difference, which nearly every variable has
The Floating Baseline In a stacked bar only the bottom segment sits on a common baseline
Why Angle Is Hard The eye recognises angles rather than measuring them
Rotating a Pie Changes which slice looks largest without changing any number
Area
Stevens' Power Law Perceived magnitude rises as physical magnitude to a power
The Exponent for Area About 0.7, so twice the area is seen as about 1.6 times the quantity
The Radius Mistake Setting radius proportional to value makes a doubling quadruple the ink
Sizing a Circle Correctly Radius proportional to the square root of the value
The Two Errors Do Not Cancel The geometric inflation and the perceptual shortfall leave 32% overstatement
Colour
Hue Has No Order Nobody can say whether green is more than orange
Sequential Scales One hue varying in lightness, for an ordinary quantity
Diverging Scales Two hues meeting at a neutral centre, when there is a real midpoint
The Rainbow Problem Mapping an ordered quantity onto an unordered property, still a default
Six Hues Roughly the number of hues a reader holds apart without the legend
Colour Vision Deficiency About one man in twelve, most often red against green
Do Not Let Colour Carry It Alone Label directly or vary shape, so colour is a convenience not the key
Channel and Variable
Matching Channel to Variable The channel must not claim more than the variable has
Nominal and Length Length says one category is more than another, which is meaningless
Ordinal and Equal Spacing Length claims the gaps between small, medium and large are equal
Position Suits Everything Position appears in every row of the table, which is why it is hard to beat
The Zero Argument
Length Encodes Ratio So its origin is part of the encoding, and moving it changes the claim
Position Encodes Difference So it may be cropped as tightly as the data deserves
The Zero Baseline Rule, Derived Truncate a position axis freely, and never truncate a length
Truncation Without Falsehood 1.11 to 1 drawn as 11.63 to 1, with every label on the chart correct
The Dot Plot The answer whenever the variation is small and sits far from zero