16  Sampling Distributions and the Central Limit Theorem

Chapter 13 was about drawing a good sample. This chapter is about what happens next: you have computed \(\bar{x}\) from your one sample, and you need to know how close it is likely to be to the population mean \(\mu\) you actually care about. You cannot answer that by looking at your sample, because your sample is all you have.

The answer comes from a thought experiment. Imagine drawing not one sample but every possible sample of size \(n\), computing \(\bar{x}\) for each. Those values have a distribution of their own, the sampling distribution, and it turns out to be remarkably predictable: centred on \(\mu\), narrower by a factor of \(\sqrt{n}\), and, thanks to the Central Limit Theorem, approximately Normal almost regardless of what the population looks like.

That result is the hinge of this book. Every confidence interval and every hypothesis test in the chapters ahead is an application of it.

16.1 Three Distributions, Easily Confused

Almost every difficulty students have with inference comes from blurring three distinct distributions. They are worth separating carefully before going any further.

  • The population distribution: the values of every unit in the population. Its mean is \(\mu\) and its standard deviation is \(\sigma\). Its shape can be anything at all, and nothing you do changes it.
  • The sample distribution: the \(n\) values in the one sample you actually collected. It is an imperfect picture of the population distribution, and it looks more like the population as \(n\) grows.
  • The sampling distribution: the distribution of a statistic, such as \(\bar{x}\), across all possible samples of size \(n\). Nobody ever observes this distribution directly; it is a theoretical object, and it is the one that inference is built on.
Population distribution Sample distribution Sampling distribution
What each value is One unit One unit One whole sample’s \(\bar{x}\)
How many values \(N\) \(n\) Every possible sample
Centre \(\mu\) \(\bar{x}\) \(\mu\)
Spread \(\sigma\) \(s\) \(\sigma / \sqrt{n}\)
Shape Anything Roughly the population’s Approximately Normal for adequate \(n\)
Ever observed? Rarely Yes, it is your data No, it is theoretical

The interactive at the top of this chapter shows the first and third of these side by side. The upper chart is the population and never changes as you draw samples. The lower chart is the sampling distribution being built up one sample at a time, where every single bar represents not a unit but an entire sample reduced to its mean. Setting \(n = 1\) makes the two charts converge, which is the clearest way to see that they are different objects that happen to coincide in that one degenerate case.

16.2 The Sampling Distribution of the Mean

For a random sample of size \(n\) drawn from a population with mean \(\mu\) and standard deviation \(\sigma\), the sampling distribution of \(\bar{x}\) has two exactly known properties, whatever the shape of the population:

\[E(\bar{x}) = \mu \qquad\qquad \text{Var}(\bar{x}) = \frac{\sigma^{2}}{n}\]

The standard deviation of a sampling distribution has its own name, the standard error, to keep it clearly distinct from the standard deviation of the population:

\[\sigma_{\bar{x}} = \frac{\sigma}{\sqrt{n}}\]

The first result says the sample mean is an unbiased estimator of \(\mu\): it is not systematically too high or too low. The second says its precision improves as \(n\) grows, but only with the square root of \(n\).

Unbiasedness is a property of the procedure, not of your particular sample. It does not promise that your \(\bar{x}\) is close to \(\mu\); it promises that the sample means would average out to \(\mu\) across all possible samples. Any single \(\bar{x}\) can still be badly wrong, and the standard error is precisely the measure of how wrong it is likely to be.

Example

Monthly customer transactions at a retail chain have \(\mu = ₹5{,}000\) and \(\sigma = ₹800\), the same figures used in Chapter 12. For a random sample of \(n = 64\) transactions:

\[\sigma_{\bar{x}} = \frac{800}{\sqrt{64}} = \frac{800}{8} = ₹100\]

Individual transactions vary with a standard deviation of ₹800, but sample means of 64 transactions vary with a standard deviation of only ₹100. The averaging has removed seven eighths of the variability.

To halve the standard error to ₹50, you would need \(n = 256\), four times the sample, since \(\sqrt{256} = 16\) is twice \(\sqrt{64} = 8\).

The \(\sqrt{n}\) in the denominator is the single most consequential fact in applied statistics, and it is the same fact as the sample-size formula from Chapter 13 read backwards. Precision is expensive: four times the data buys only twice the precision. This is why survey samples of 1,000 to 2,000 are so common, and why pushing beyond them is rarely worth the cost. Going from 1,000 to 4,000 respondents quadruples the fieldwork budget to halve the margin of error, and does nothing at all about bias.

16.3 The Finite Population Correction

The formula \(\sigma_{\bar{x}} = \sigma/\sqrt{n}\) assumes sampling with replacement, or equivalently an infinite population. When sampling without replacement from a finite population, and the sample is a substantial share of it, the true standard error is smaller, because each unit drawn removes some of the remaining uncertainty. The adjustment is the finite population correction:

\[\sigma_{\bar{x}} = \frac{\sigma}{\sqrt{n}} \times \sqrt{\frac{N-n}{N-1}}\]

The conventional rule is to apply it when \(n\) exceeds 5% of \(N\), and to ignore it otherwise, since the factor is then very close to 1.

Example

Sampling \(n = 30\) employees from a firm of \(N = 300\) gives a sampling fraction of 10%, above the 5% threshold, so the correction applies:

\[\sqrt{\frac{300-30}{300-1}} = \sqrt{\frac{270}{299}} = \sqrt{0.9030} \approx 0.950\]

The standard error is 5% smaller than the uncorrected formula suggests. Taking the correction to its logical end, if \(n = N\) the factor becomes 0: a census has no sampling error at all, because there is nothing left to be uncertain about.

This is the formal reason the sample-size formula in Chapter 13 does not contain \(N\). For the large populations of most business and social research, \(n/N\) is a tiny fraction, the correction factor is essentially 1, and population size genuinely does not matter. A sample of 1,000 tells you almost exactly as much about a country of 1.4 billion as about a city of 1 million.

16.4 The Central Limit Theorem

Everything so far holds for any population shape, but says nothing about the shape of the sampling distribution. The Central Limit Theorem supplies it, and it is the reason the Normal distribution of Chapter 12 governs statistical inference even for data that is nothing like Normal.

For independent, identically distributed observations from a population with mean \(\mu\) and finite standard deviation \(\sigma\), as \(n\) grows:

\[\bar{x} \;\dot\sim\; N\!\left(\mu,\; \frac{\sigma^{2}}{n}\right) \qquad\text{equivalently}\qquad Z = \frac{\bar{x} - \mu}{\sigma/\sqrt{n}} \;\dot\sim\; N(0,\,1)\]

The population may be skewed, bimodal, or discrete. Provided \(\sigma\) is finite and \(n\) is adequate, the distribution of the sample mean approaches Normal regardless.

The bimodal setting in the interactive above is the most persuasive demonstration of this. Choose it, and the population has two separate humps with hardly any units near the overall mean at all. Yet at \(n = 25\) the sample means pile up in a single clean Normal bell centred exactly on \(\mu\), a value that is itself atypical of the population. Averaging does not preserve the shape of what you averaged, and that is the entire point.

The Central Limit Theorem is routinely misquoted. Three things it does not say:

  • It does not say your data becomes Normal. The population distribution never changes. Only the distribution of the sample mean becomes Normal, and collecting more data does not normalise skewed data.
  • It does not license Normal-based methods on any large sample. It concerns the sampling distribution of a statistic, not the individual observations, so a prediction interval for one new customer’s spend still needs the actual skewed distribution.
  • It does not repair a biased sample. It assumes the observations are a random sample. Applied to a convenience sample, it delivers a beautifully Normal sampling distribution centred on the wrong value.

The Central Limit Theorem is often confused with the Law of Large Numbers, which the coin-flip demonstration in Chapter 9 illustrated. They answer different questions.

  • Law of Large Numbers: as \(n\) grows, \(\bar{x}\) converges to \(\mu\). Formally, for any \(a > 0\), \(P(|\bar{x}_n - \mu| < a) \to 1\). It tells you the sample mean ends up in the right place.
  • Central Limit Theorem: as \(n\) grows, the distribution of \(\bar{x}\) around \(\mu\) approaches \(N(\mu, \sigma^2/n)\). It tells you the shape and spread of the error around that place.

The Law of Large Numbers alone would leave you knowing \(\bar{x}\) is probably close to \(\mu\) without any way to say how close. The Central Limit Theorem is what turns “probably close” into a number, and therefore what makes confidence intervals possible at all.

16.5 How Large Is “Large Enough”?

Textbooks commonly give \(n \geq 30\) as the threshold at which the Central Limit Theorem can be relied upon. This is a rule of thumb, not a theorem, and how well it works depends entirely on how skewed the population is.

Population shape Roughly adequate \(n\)
Already Normal Any \(n\), including 1
Symmetric, non-Normal (uniform, bimodal) 5 to 15
Moderately skewed 30, the familiar rule
Heavily skewed (incomes, insurance claims, waiting times) 100 or more
Extremely skewed, or heavy-tailed with outliers Several hundred, and check directly

The \(n \geq 30\) rule is comfortably wrong for heavily skewed data, which is exactly the kind that dominates business analytics: incomes, transaction values, claim sizes, and waiting times are all strongly right-skewed. Simulation studies of heavily skewed populations find that samples in the region of \(n = 100\) still fail a normality check a noticeable share of the time, with reliability only becoming dependable somewhere around \(n = 200\).

You can watch this yourself in the interactive above. Select Heavily skewed and set \(n = 10\): the sampling distribution is visibly lopsided and the widget tells you so. Raise \(n\) to 100 and it settles under the theoretical curve. The rule is not the theorem, and the honest approach is to check the sampling distribution for your own data rather than trust a number from a textbook.

When \(n\) is genuinely too small for the Central Limit Theorem and the population is not Normal, the usual remedies are to transform the variable, for example by taking logarithms of a right-skewed income measure, to use a distribution-free method that makes no normality assumption, or to bootstrap the sampling distribution directly by resampling your own data. Each of these appears later in this book; the point here is simply that “the sample is too small” has answers other than giving up.

Application

Read down the skew column. The population itself is strongly right-skewed, and at \(n = 2\) the sampling distribution inherits almost all of that skew. It falls steadily as \(n\) rises and is close to zero only well past \(n = 30\). Meanwhile the mean(xbar) column sits on the population mean at every sample size, including \(n = 2\), because unbiasedness holds immediately and has nothing to do with the Central Limit Theorem. The two properties arrive at different speeds, and conflating them is a common source of error.

16.6 The Sampling Distribution of a Proportion

A great deal of business data is not a measurement but a yes-or-no outcome: a customer churned or did not, an item was defective or was not, a voter supports a party or does not. The statistic of interest is then the sample proportion \(\hat{p}\), and it has a sampling distribution of exactly the same character.

\[E(\hat{p}) = p \qquad\qquad \sigma_{\hat{p}} = \sqrt{\frac{p(1-p)}{n}}\]

For adequate \(n\) the distribution of \(\hat{p}\) is approximately Normal, so the same standardisation applies:

\[Z = \frac{\hat{p} - p}{\sqrt{p(1-p)/n}} \;\dot\sim\; N(0,\,1)\]

The usual conditions for this approximation are the success-failure conditions, that the expected number of successes and of failures are both at least 10:

\[np \geq 10 \qquad \text{and} \qquad n(1-p) \geq 10\]

This is the Binomial distribution of Chapter 11 seen from a different angle. A count of successes is Binomial with mean \(np\) and variance \(np(1-p)\); dividing through by \(n\) to turn the count into a proportion gives mean \(p\) and variance \(p(1-p)/n\). The Central Limit Theorem then supplies the Normal shape, which is precisely the Normal approximation to the Binomial that Chapter 12 mentioned in passing.

Example

A bank’s historical churn rate is \(p = 0.20\). In a random sample of \(n = 400\) customers, what is the probability that more than 24% churn?

First check the conditions: \(np = 400 \times 0.20 = 80\) and \(n(1-p) = 320\), both comfortably above 10, so the Normal approximation is safe.

\[\sigma_{\hat{p}} = \sqrt{\frac{0.20 \times 0.80}{400}} = \sqrt{0.0004} = 0.02\]

\[Z = \frac{0.24 - 0.20}{0.02} = 2.0 \qquad\Rightarrow\qquad P(\hat{p} > 0.24) = P(Z > 2) \approx 0.0228\]

Only about 2.3% of samples of 400 would show a churn rate above 24% if the true rate were still 20%. Observing 24% is therefore reasonably strong evidence that churn has genuinely risen, which is exactly the logic formalised as a hypothesis test in Chapter 16.

Note that \(\sigma_{\hat{p}}\) depends on \(p\), the very quantity being estimated. In practice \(\hat{p}\) is substituted for it, which is acceptable for large \(n\) but circular for small \(n\). Note also that \(p(1-p)\) is largest at \(p = 0.5\) and shrinks towards the extremes, so proportions near 50% are the hardest to pin down and rare events near 1% are, in relative terms, harder still.

16.7 Why This Chapter Is the Hinge of the Book

Everything up to Chapter 12 described data you already had. Chapter 13 explained how to obtain a sample worth trusting. This chapter supplies the missing link that turns one into the other, because it answers the question inference actually rests on: how far from \(\mu\) is a single \(\bar{x}\) likely to be?

The answer, \(\sigma/\sqrt{n}\) with an approximately Normal shape, is what makes the next two chapters possible:

  • Estimation (Chapter 15) inverts the sampling distribution. If \(\bar{x}\) lies within about \(1.96\) standard errors of \(\mu\) in 95% of samples, then an interval of \(\bar{x} \pm 1.96\,\sigma_{\bar{x}}\) captures \(\mu\) in 95% of samples. That interval is a confidence interval, and it is nothing more than this chapter’s result rearranged.
  • Hypothesis testing (Chapter 16) uses the sampling distribution as a yardstick. Assume a value for \(\mu\), ask how surprising the observed \(\bar{x}\) would be if that assumption were true, and answer using the sampling distribution. The churn example above already performed exactly this calculation.

One practical caveat before moving on. Every formula in this chapter uses \(\sigma\), the population standard deviation, which in real work is almost never known. Substituting the sample standard deviation \(s\) introduces extra uncertainty, and for small \(n\) the standardised statistic then follows a \(t\)-distribution rather than a Normal one. That refinement is taken up in Chapter 15; for now, note that the logic is unchanged and only the reference distribution differs.

Recap

Chapters 13 and 14 together build the bridge from data you have to conclusions about a population you do not: how to draw a sample worth trusting, and then how far a statistic computed from it is likely to sit from the parameter it estimates. The sample mean is unbiased, its standard error is \(\sigma/\sqrt{n}\), and for adequate \(n\) its distribution is approximately Normal whatever the population looks like. The next part inverts that result into the question a working analyst actually asks: given the one \(\bar{x}\) I have, what range of values for \(\mu\) is consistent with it? That inversion produces the confidence interval, and with it the \(t\)-distribution for the realistic case where \(\sigma\) must be estimated from the sample too.


Summary

Concept Description
Three Distributions
Population Distribution The values of every unit in the population, with mean mu and standard deviation sigma, of any shape
Sample Distribution The n values in the one sample actually collected, an imperfect picture of the population
Sampling Distribution The distribution of a statistic across all possible samples of size n, theoretical and never observed
Statistic as a Random Variable A statistic varies from sample to sample, so it is itself a random variable with its own distribution
The Sampling Distribution of the Mean
Expected Value of the Sample Mean The expected value of the sample mean equals the population mean, whatever the population's shape
Unbiased Estimator An estimator whose expected value equals the parameter, so it is not systematically too high or low
Variance of the Sample Mean The variance of the sample mean is the population variance divided by n
Standard Error of the Mean The standard deviation of the sampling distribution, equal to sigma divided by the square root of n
The Square Root Law Precision improves only with the square root of n, so four times the data buys twice the precision
Finite Populations
Finite Population Correction A factor applied when sampling without replacement from a substantial share of a finite population
The Five Percent Rule Apply the finite population correction when the sample exceeds five percent of the population
Why Population Size Rarely Matters Because the correction is negligible for large N, sample size rather than population size drives precision
The Central Limit Theorem
Central Limit Theorem For adequate n the sample mean is approximately Normal regardless of the population's shape
Standardised Form of the CLT Subtracting mu and dividing by the standard error turns the sample mean into a standard Normal variable
What the CLT Does Not Say It does not normalise your data, does not license Normal methods on individual values, and does not fix bias
Law of Large Numbers As n grows the sample mean converges to the population mean, so it ends up in the right place
LLN Versus CLT The law of large numbers gives the location of the sample mean; the CLT gives the shape and spread of its error
How Large Is Large Enough
The n >= 30 Rule of Thumb The conventional threshold at which the CLT is assumed adequate, useful but not a theorem
When 30 Is Not Enough Heavily skewed populations may need n of 100 or more, and around 200 before normality is dependable
Remedies for Small n Transform the variable, use a distribution-free method, or bootstrap the sampling distribution directly
Proportions
Sample Proportion The proportion of successes in a sample, whose expected value is the population proportion p
Standard Error of a Proportion The square root of p times one minus p divided by n
Success-Failure Conditions The Normal approximation for a proportion requires np and n(1-p) to both be at least 10
Link to the Binomial A proportion is a Binomial count divided by n, which is why its mean is p and its variance p(1-p)/n
Where Proportions Are Hardest to Estimate The standard error is largest at p = 0.5 and shrinks towards the extremes
Why It Matters
Foundation for Estimation A confidence interval is the sampling distribution inverted to give a range of plausible values for mu
Foundation for Hypothesis Testing A hypothesis test uses the sampling distribution as a yardstick for how surprising an observed statistic is
When Sigma Is Unknown Substituting the sample standard deviation for sigma leads to the t-distribution rather than the Normal