17  Estimation and Confidence Intervals

Chapter 14 worked forwards: given a population with known \(\mu\) and \(\sigma\), it described how \(\bar{x}\) would behave across samples. Real analysis runs the other way. You have one sample, you do not know \(\mu\), and you need to say something defensible about it.

This chapter performs that inversion. A single number estimate, \(\bar{x} = ₹5{,}240\), is almost useless on its own because it carries no indication of how much it could be wrong by. An interval estimate, ₹5,240 ± ₹180, carries its own uncertainty with it, and the machinery for building one comes directly from the sampling distribution of the previous chapter.

Along the way this chapter introduces the \(t\)-distribution, needed because the realistic case has \(\sigma\) unknown, and confronts what is probably the most widely misunderstood sentence in applied statistics: what “95% confident” actually means.

17.1 Point and Interval Estimation

An estimator is a rule for computing an estimate of a parameter from sample data; an estimate is the number that rule produces on your particular sample.

  • A point estimate is a single value offered as the best guess at a parameter. The sample mean \(\bar{x}\) estimates \(\mu\); the sample proportion \(\hat{p}\) estimates \(p\); the sample standard deviation \(s\) estimates \(\sigma\).
  • An interval estimate is a range of values, together with a stated level of confidence that the procedure producing it captures the parameter.

A point estimate on its own is almost always insufficient, and reporting one without a measure of its uncertainty is among the most common failures in applied analytics. “Average customer spend is ₹5,240” invites the reader to treat a number computed from 40 customers as if it were a fact about all of them. The honest version states the interval, and the interval is frequently wide enough to change the decision.

17.2 What Makes a Good Estimator?

Several different rules could estimate the same parameter. The sample mean, the sample median, and the midpoint of the range all estimate the centre of a symmetric population. Statisticians judge between them on four properties.

  • Unbiasedness: The estimator’s expected value equals the parameter, \(E(\hat{\theta}) = \theta\), so it is not systematically too high or too low. Chapter 14 established this for \(\bar{x}\).
  • Efficiency: Among unbiased estimators, the one with the smallest variance is preferred, because it lands closer to the target more often. For a Normal population the sample mean is more efficient than the sample median.
  • Consistency: The estimator converges on the parameter as \(n\) grows, which is the Law of Large Numbers applied to an estimator.
  • Sufficiency: The estimator uses all the relevant information in the sample, wasting none of it.

This is where the \(n-1\) in the sample variance formula from Chapter 6 finally earns its explanation. Dividing the sum of squared deviations by \(n\) produces a biased estimator that systematically underestimates \(\sigma^2\), because the deviations are measured from \(\bar{x}\), which is itself pulled towards the data. Dividing by \(n-1\) corrects exactly for that, and is the reason every statistical package defaults to it.

\[s^{2} = \frac{\sum (x_i - \bar{x})^{2}}{n-1}\]

Unbiasedness and efficiency can pull in different directions, and unbiasedness is not automatically the more important of the two. An estimator that is slightly biased but much less variable can be closer to the truth on average than an unbiased but wildly variable one. This trade-off, usually framed as bias against variance, reappears as a central theme in predictive modelling later in this book.

17.3 Confidence Interval for a Mean, \(\sigma\) Known

Chapter 14 established that \(\bar{x}\) is approximately \(N(\mu,\, \sigma^2/n)\), so 95% of sample means fall within \(1.96\) standard errors of \(\mu\). Turning that around: for 95% of samples, an interval of \(1.96\) standard errors either side of \(\bar{x}\) contains \(\mu\). That interval is the confidence interval.

\[\bar{x} \pm z_{\alpha/2}\,\frac{\sigma}{\sqrt{n}}\]

The quantity added and subtracted is the margin of error. The confidence level is \((1-\alpha)\), and \(z_{\alpha/2}\) is the standard normal value cutting off \(\alpha/2\) in each tail.

Confidence level \(\alpha\) \(z_{\alpha/2}\)
90% 0.10 1.645
95% 0.05 1.960
99% 0.01 2.576

Example

A retail chain knows from long experience that transaction amounts have \(\sigma = ₹800\). A random sample of \(n = 64\) transactions gives \(\bar{x} = ₹5{,}240\). A 95% confidence interval for the mean transaction value:

\[₹5{,}240 \pm 1.96 \times \frac{800}{\sqrt{64}} = ₹5{,}240 \pm 1.96 \times 100 = ₹5{,}240 \pm ₹196\]

giving \((₹5{,}044,\; ₹5{,}436)\). The margin of error is ₹196, and the interval is ₹392 wide.

17.4 Interpreting a Confidence Interval

The correct interpretation is a statement about the procedure, not about the particular interval in front of you.

If we repeated this sampling procedure many times, 95% of the intervals it produces would contain the true population mean.

The interactive at the top of this chapter is that sentence made literal. Each line is one sample’s interval, the red line is the fixed true \(\mu\), and green intervals are the ones that caught it. Draw a few hundred and the proportion of green settles near the stated confidence level.

The near-universal misstatement is: “there is a 95% probability that \(\mu\) lies between ₹5,044 and ₹5,436.”

This is wrong, and the interactive shows why. In the frequentist framework \(\mu\) is a fixed constant, not a random variable. It does not move. Once your interval is computed, both endpoints are fixed numbers too, so \(\mu\) either is inside that interval or it is not. There is no probability left to speak of; the answer is already determined, you simply do not know which it is. The 95% describes how often the method succeeds, and the randomness lives entirely in which sample you happened to draw.

Three further misreadings worth naming:

  • “95% of the data lies in the interval.” No. The interval estimates the location of the mean, and is typically far narrower than the data. A confidence interval and a prediction interval are different objects.
  • “95% of future sample means will fall in this interval.” No. That is a different and wider interval again.
  • “A wider interval means the data is more variable.” Not necessarily. Width also depends on \(n\) and on the confidence level, so a wide interval may simply reflect a small sample or a demanding confidence level.

These are not beginner errors that professionals have outgrown. Surveys of researchers and statistics students repeatedly find that a majority endorse at least one of the false statements above, and the misreading is common in published work. Watching the intervals move while \(\mu\) stays fixed is the most reliable cure, which is exactly what the interactive is for.

17.5 The \(t\)-Distribution

The interval above assumed \(\sigma\) was known, which it almost never is. Replacing \(\sigma\) with the sample standard deviation \(s\) introduces a second source of uncertainty: \(s\) is itself an estimate and varies from sample to sample. The standardised quantity then no longer follows a Normal distribution:

\[t = \frac{\bar{x} - \mu}{s/\sqrt{n}}\]

This statistic follows Student’s \(t\)-distribution with \(n-1\) degrees of freedom. It is symmetric and bell-shaped like the Normal, but with heavier tails, reflecting the extra uncertainty from estimating \(\sigma\). As \(n\) grows, \(s\) becomes a reliable estimate of \(\sigma\) and the \(t\)-distribution converges on the standard Normal.

The distribution owes its odd name to its origin in industry rather than academia. William Sealy Gosset, a chemist at the Guinness brewery in Dublin, needed to draw conclusions about brewing quality from the very small samples that testing allowed. He found that for small \(n\) the distribution of the standardised mean departed from the Normal, and worked out the correct distribution, publishing it in Biometrika in 1908. Guinness forbade employees from publishing under their own names, so the paper appeared under the pseudonym “Student”, and the name stuck to one of the most-used distributions in statistics.

Degrees of freedom is the number of independent pieces of information left after estimating what the procedure needed to estimate. Having used the sample to compute \(\bar{x}\), only \(n-1\) of the deviations from it are free to vary, since they must sum to zero. That is why \(df = n-1\) here.

\(df\) \(t_{0.025}\) vs \(z = 1.96\)
1 12.706 far heavier tails
5 2.571 noticeably wider
10 2.228 wider
30 2.042 close
100 1.984 almost identical
\(\infty\) 1.960 identical

17.6 Confidence Interval for a Mean, \(\sigma\) Unknown

With \(\sigma\) estimated by \(s\), the confidence interval uses the \(t\) critical value in place of \(z\):

\[\bar{x} \pm t_{\alpha/2,\,n-1}\,\frac{s}{\sqrt{n}}\]

This is the interval used in practice in nearly all real work. The conditions are that the sample is random, and that either the population is approximately Normal or \(n\) is large enough for the Central Limit Theorem of Chapter 14 to apply.

Example

The sample of \(n = 25\) delivery times used in the code below gives \(\bar{x} = 43.44\) minutes and \(s = 4.12\) minutes, so the standard error is \(4.12/\sqrt{25} = 0.824\). For a 95% interval, \(df = 24\) and \(t_{0.025,\,24} = 2.064\):

\[43.44 \pm 2.064 \times 0.824 = 43.44 \pm 1.70\]

giving \((41.74,\; 45.14)\) minutes. Had \(z = 1.96\) been used instead, the margin would have been \(1.96 \times 0.824 = 1.62\), producing a narrower interval that claims more precision than the data supports.

Using \(z\) when \(\sigma\) has been estimated from a small sample does not merely make the interval slightly too narrow, it makes the stated confidence level a fiction. The interactive at the top of this chapter measures the damage directly: switch it to Use z with \(n = 5\), draw a few hundred intervals, and observed coverage lands near 86% rather than the 95% claimed. Roughly one interval in seven misses the mean while advertising one in twenty. Switch back to \(t\) and coverage returns to 95%.

Application

The t.test() line in R is how this is done in practice, and it is worth noticing that it returns the interval rather than a single number: the language’s own default output treats the estimate as a range. The final comparison shows the cost of the shortcut. With \(n = 25\) the \(z\) interval is only a little narrower, but the gap widens sharply as \(n\) falls, which is exactly the regime Gosset was working in.

17.7 Confidence Interval for a Proportion

For a yes-or-no outcome, Chapter 14 gave the sampling distribution of \(\hat{p}\). Inverting it in the same way produces the textbook interval, known as the Wald interval:

\[\hat{p} \pm z_{\alpha/2}\,\sqrt{\frac{\hat{p}(1-\hat{p})}{n}}\]

with the usual success-failure conditions, \(n\hat{p} \geq 10\) and \(n(1-\hat{p}) \geq 10\).

Example

Of \(n = 400\) sampled customers, 88 churned, so \(\hat{p} = 88/400 = 0.22\). Conditions: \(n\hat{p} = 88\) and \(n(1-\hat{p}) = 312\), both well above 10.

\[0.22 \pm 1.96\sqrt{\frac{0.22 \times 0.78}{400}} = 0.22 \pm 1.96 \times 0.0207 = 0.22 \pm 0.0406\]

giving \((0.179,\; 0.261)\), or churn somewhere between about 18% and 26%.

The Wald interval is the one every introductory textbook teaches, and it is known to perform badly. Its actual coverage oscillates erratically with both \(n\) and \(p\) rather than settling at the stated level, and the problem does not reliably disappear at sample sizes usually considered comfortable. A well-known review of the topic described its coverage behaviour as persistently chaotic and unacceptably poor.

The failure is worst exactly where it matters most: when \(\hat{p}\) is near 0 or 1, the interval can run outside \([0, 1]\), and when \(\hat{p} = 0\) the formula returns an interval of zero width, confidently asserting the true proportion is exactly zero on the basis of having seen no successes.

Two better intervals are easy to use and worth knowing.

The Wilson score interval is the standard recommendation for general reporting, holding its stated coverage across nearly all combinations of \(n\) and \(p\):

\[\frac{\hat{p} + \dfrac{z^{2}}{2n} \;\pm\; z\sqrt{\dfrac{\hat{p}(1-\hat{p})}{n} + \dfrac{z^{2}}{4n^{2}}}}{1 + \dfrac{z^{2}}{n}}\]

The Agresti-Coull interval is the same idea in a form you can do by hand. For 95% confidence, simply add two successes and two failures to the observed counts, then apply the ordinary Wald formula to the adjusted numbers. This “plus-four” rule is remarkably effective, and it is a good default whenever a proportion has to be estimated quickly.

17.8 Margin of Error and Sample Size

Every interval in this chapter has the same shape, an estimate plus or minus a margin of error, and the margin is governed by three quantities:

\[E = z_{\alpha/2}\,\frac{\sigma}{\sqrt{n}}\]

  • Confidence level: higher confidence widens the interval. Certainty and precision trade off directly against each other; a 100% confidence interval is the whole real line and says nothing.
  • Variability: a more variable population widens the interval, and this is not under the analyst’s control.
  • Sample size: more data narrows the interval, but only as \(\sqrt{n}\).

Rearranging for \(n\) recovers the sample-size formula from Chapter 13, which is now revealed to be the confidence interval solved backwards:

\[n = \left(\frac{z_{\alpha/2}\,\sigma}{E}\right)^{2} \qquad\qquad n = \left(\frac{z_{\alpha/2}}{E}\right)^{2} p(1-p)\]

For a proportion when no prior estimate of \(p\) exists, use \(p = 0.5\). Since \(p(1-p)\) is maximised there, this gives the largest \(n\) any true value could require, and therefore a conservative sample size that is guaranteed to be sufficient. It is also why so many opinion polls report a margin of error near 3%: at \(p = 0.5\) and 95% confidence, \(n = 1{,}068\) gives exactly that.

Application

The three intervals agree closely at \(\hat{p} = 0.22\) with \(n = 400\), which is the comfortable case. The zero-successes example is where they part company completely: Wald returns an interval of zero width, asserting with total confidence that the true rate is exactly 0, while Agresti-Coull returns a sensible range. Observing no defects in 20 items is genuinely weak evidence that the defect rate is zero, and only one of these methods says so.

Looking Ahead

A confidence interval answers “what values of \(\mu\) are consistent with my data?” The next chapter asks the mirror question: “is a specific claimed value of \(\mu\) consistent with my data?” That is hypothesis testing, and it turns out to be the same sampling-distribution logic pointed in a different direction, with the confidence interval and the test arriving at matching conclusions by construction.


Summary

Concept Description
Estimation Basics
Estimator and Estimate A rule for computing an estimate from sample data, and the number that rule produces
Point Estimate A single value offered as the best guess at a population parameter
Interval Estimate A range of values with a stated confidence that the procedure captures the parameter
Properties of Estimators
Unbiasedness The estimator's expected value equals the parameter, so it is not systematically off
Efficiency Among unbiased estimators, the one with the smallest variance is preferred
Consistency The estimator converges on the parameter as the sample size grows
Sufficiency The estimator uses all the relevant information the sample contains
Why n Minus One Dividing by n minus one corrects the downward bias from measuring deviations about the sample mean
Bias-Variance Trade-off A slightly biased but much less variable estimator can beat an unbiased but erratic one
Confidence Intervals
Confidence Interval An interval built from a sample that captures the parameter a stated proportion of the time
Margin of Error The quantity added and subtracted around the point estimate
Confidence Level The long-run proportion of intervals from this procedure that contain the parameter
Z Interval for a Mean The sample mean plus or minus z times sigma over the square root of n, used when sigma is known
Interpreting Them
Correct Interpretation A statement about the procedure: repeated sampling would produce intervals capturing mu that often
The Probability Misstatement Saying there is a 95 percent probability mu lies in this interval is wrong, since mu is fixed and so are the endpoints
Other Common Misreadings It does not contain 95 percent of the data, nor 95 percent of future sample means
The t-Distribution
Student's t-Distribution The distribution of the standardised mean when sigma is estimated by s, with heavier tails than the Normal
Degrees of Freedom The independent information remaining after estimation; n minus one for a one-sample interval
Origin of the Name Published in 1908 by W. S. Gosset of the Guinness brewery under the pseudonym Student
T Interval for a Mean The sample mean plus or minus t times s over the square root of n, the interval used in practice
Cost of Using z Wrongly Using z when sigma is estimated makes the interval too narrow and the stated confidence a fiction
Proportions
Wald Interval for a Proportion The textbook proportion interval, p-hat plus or minus z times the standard error
Why Wald Fails Its coverage oscillates erratically with n and p, and at zero successes it returns zero width
Wilson Score Interval A better interval that holds its stated coverage across nearly all values of n and p
Agresti-Coull and the Plus-Four Rule Adding two successes and two failures before applying the Wald formula, an easy and effective fix
Precision and Sample Size
What Drives Interval Width Confidence level and population variability widen it; larger n narrows it, but only as the square root
Sample Size from Margin of Error Rearranging the margin of error formula for n recovers the sample size formula of Chapter 13
Conservative p of One Half Using p equal to one half maximises p times one minus p and so gives a guaranteed sufficient sample size