15 Sampling and Sampling Methods
Sampling is the process of selecting a subset of units from a population to study. It lets us make inferences about populations when examining every unit is impractical or impossible. The quality of those inferences depends critically on how the sample is selected, since the sampling method decides whether your conclusions will be reliable or misleading.
This chapter introduces the vocabulary and techniques of sampling. You will learn the difference between probability and non-probability methods, work through four probability sampling designs and three non-probability ones, understand bias and representativeness, and see why random sampling is what unlocks the statistical inference methods that form the backbone of every chapter that follows.
15.1 Why Sample?
In most real research it is impractical or impossible to study an entire population, the complete set of all units of interest. Instead we select a sample, a subset of the population, and use statistics computed from the sample to make inferences about population parameters.
- Population: The complete set of all units (individuals, objects, measurements) of interest. Its size is denoted \(N\).
- Sample: A subset of the population selected for study. Its size is denoted \(n\).
- Parameter: A numerical characteristic of the population, such as the population mean \(\mu\) or population proportion \(\pi\). Almost always unknown.
- Statistic: A numerical characteristic of a sample, such as the sample mean \(\bar{x}\) or sample proportion \(\hat{p}\). Computed from data, and used to estimate the parameter.
The distinction between a parameter and a statistic is the hinge on which the rest of this book turns. A parameter is a fixed but unknown number describing the population. A statistic is a number you can actually compute, which varies from sample to sample. Everything in Module IV is about how much a statistic can be trusted as a stand-in for the parameter it estimates.
A well-designed sample can deliver reliable insight at a fraction of the cost and time a full census would need. There are four standard reasons to sample rather than count everything:
- Cost: Surveying 1,000 households costs a tiny fraction of surveying 300 million.
- Time: A sample can be collected and analysed in weeks; a census takes years.
- Practicality: Some populations cannot be fully enumerated at all, such as all potential future customers.
- Destructive testing: Testing the breaking strength of every component would leave nothing to sell.
Poor sampling introduces bias, and bias is not fixed by collecting more data. A larger sample drawn the wrong way gives you a wrong answer with narrower error bars, which is worse than a small honest sample, because it looks more convincing. This point returns at the end of the chapter and again in every chapter after it.
15.2 The Sampling Frame
Before selecting a sample we must define a sampling frame, the actual list of units from which the sample will be drawn. Ideally the frame matches the population exactly. When it does not, frame error occurs, and the sample can only ever represent the frame, never the population the frame was supposed to stand for.
Example
- Population: All registered voters in a state.
- Sampling frame: The voter registry maintained by the election commission.
- Frame error: If the registry is outdated, newly registered voters are excluded from the frame and therefore have zero probability of selection, no matter how carefully the sample is drawn from it.
Frame error is often unavoidable, but it must always be documented and discussed when presenting results. A telephone survey frames the population as “people with telephones,” an online panel frames it as “people who joined an online panel,” and neither is the general public. The most damaging sampling errors in practice come not from the selection method but from a frame nobody examined.
15.3 Probability vs. Non-Probability Sampling
Sampling methods fall into two broad families, and the difference between them decides whether formal statistical inference is available to you at all.
Probability sampling: Every unit in the population has a known, non-zero probability of being selected. That randomisation eliminates selection bias in expectation, allows margins of error and confidence intervals to be computed, and enables valid inference to the population. It requires a complete, accessible sampling frame.
Non-probability sampling: Units are selected by judgment, convenience, or other non-random criteria. The resulting bias cannot be quantified, so confidence intervals and hypothesis tests are not available. It remains useful in exploratory and qualitative research, but should be avoided whenever the goal is inference about a population.
| Aspect | Probability Sampling | Non-Probability Sampling |
|---|---|---|
| Selection | Random, known probabilities | Judgment, convenience, quota |
| Sampling frame | Required | Not required |
| Bias | Controlled in expectation | Present and unquantifiable |
| Margin of error | Computable | Not computable |
| Confidence intervals | Valid | Not valid |
| Cost and speed | Higher cost, slower | Cheaper, faster |
| Appropriate for | Inference about a population | Exploration, pilots, qualitative work |
15.4 Four Probability Sampling Methods
15.4.1 1. Simple Random Sampling
Every possible sample of size \(n\) has an equal probability of being selected. This is the gold standard of unbiased sampling, and the design that every formula in the rest of this book quietly assumes.
Advantages: Conceptually simple and easy to explain; unbiased, since every unit has an equal chance of selection; and the sampling distributions of its statistics are exactly the well-understood ones covered in the next chapter.
Disadvantages: It may fail to represent rare or dispersed subgroups purely by chance; it can be inefficient for large, heterogeneous populations; and it requires a complete sampling frame.
Example
A university wants to estimate the average GPA of its 5,000 students. It generates 200 random student ID numbers from the registry and computes the mean GPA of those 200 students.
15.4.2 2. Stratified Sampling
Divide the population into strata, subgroups that are internally homogeneous, then sample independently within each stratum. This guarantees that every important group appears in the sample rather than leaving its presence to chance.
When to stratify: You know a variable that splits the population into meaningful groups, those groups differ on the variable of interest, and you want guaranteed representation from each.
Allocation methods:
- Proportional allocation: Each stratum contributes in proportion to its share of the population.
- Optimal (Neyman) allocation: Strata that are larger or more variable contribute more, which minimises the variance of the overall estimate for a fixed total sample size.
Example
A retailer with 1,000 stores across three regions (East 300, West 400, South 300) wants to sample 100 stores. Under proportional allocation the sampling fraction is \(f = 100/1000 = 0.10\), so the sample takes East = 30, West = 40, South = 30.
Stratification reduces the variance of an estimate when the strata are internally homogeneous but differ from each other, which is exactly when simple random sampling is most at risk of drawing an unrepresentative mix. The gain is real and measurable, not theoretical, and the code later in this chapter measures it directly on a simulated population.
15.4.3 3. Systematic Sampling
Select every \(k\)-th unit after a random start, where \(k\) is the sampling interval:
\[k = \frac{N}{n}\]
The procedure has three steps: compute \(k = N/n\); choose a random start between 1 and \(k\); then select every \(k\)-th unit from there onwards.
Example
An inspector samples 50 units from the 10,000 produced daily. The interval is \(k = 10{,}000 / 50 = 200\). With a random start of 47, the sampled units are 47, 247, 447, and so on up to 9,847.
Systematic sampling is badly biased whenever the population has a cyclical pattern whose period matches the interval \(k\). Sampling every 7th day always lands on the same weekday; sampling every 12th item on a production line that cycles through 12 moulds always inspects the same mould. Before using this method, always ask whether the frame has a hidden periodicity, and if it might, shuffle the frame first or use simple random sampling instead.
15.4.4 4. Cluster Sampling
Divide the population into clusters, naturally occurring groups such as schools, villages, or clinics, then randomly select whole clusters and survey every unit inside the ones chosen.
When to use it: The population is naturally grouped; a complete unit-level sampling frame is unavailable or prohibitively costly; or travel and contact costs are high.
Example
A health agency estimates vaccination coverage by randomly selecting 50 clinics from a national list, then surveying every patient registered at those 50 clinics.
Cluster sampling is much cheaper than reaching randomly scattered individuals, but it is also less precise for the same sample size, because units inside a cluster tend to resemble one another. This is called intra-cluster correlation, and it means 500 patients drawn from 50 clinics carry less information than 500 patients drawn at random from the whole country.
Note the contrast with stratified sampling, which is easy to confuse with it. Stratified sampling divides the population into groups and samples from every group, seeking homogeneity within strata. Cluster sampling divides the population into groups and samples some groups entirely, and works best when each cluster is itself a miniature of the population.
| Method | Frame needed | Cost | Precision for a given \(n\) | Main risk |
|---|---|---|---|---|
| Simple random | Complete unit list | High | Baseline | Misses small subgroups by chance |
| Stratified | List plus stratum labels | High | Best when strata differ | Needs a good stratifying variable |
| Systematic | Ordered list | Low | Similar to simple random | Hidden periodicity in the frame |
| Cluster | Cluster list only | Lowest | Worst, from intra-cluster correlation | Clusters not miniature populations |
Application
Look at how far the cluster estimate lands from the true mean. Taking one whole region as a cluster is cheap, but because the regions genuinely differ in sales, the estimate inherits that region’s level rather than the country’s. Cluster sampling only works when each cluster is a miniature of the population, and here it plainly is not. That single number is the whole trade-off in this chapter made concrete.
15.5 Non-Probability Methods (to Avoid for Inference)
- Convenience sampling: Select the units that are easiest to reach. Fast and cheap, heavily biased, and not generalisable. Surveying shoppers outside one store on a weekday afternoon samples people who shop there on weekday afternoons.
- Judgment (purposive) sampling: The researcher selects units believed to be “representative.” The quality depends entirely on the expert’s judgment, which cannot be verified from the data, and no formal inference is possible.
- Quota sampling: Set target counts for subgroups, then fill each quota by non-random selection. It superficially resembles stratified sampling and is often mistaken for it, but the selection within each quota is not random, so the bias remains and cannot be measured.
Quota sampling deserves particular caution, because it produces a sample whose demographic composition matches the population exactly and therefore looks representative in a summary table. Matching on the variables you checked says nothing about the variables you did not. A quota sample can be perfectly balanced on age and gender and still be badly wrong on the outcome being measured, which is precisely what makes it dangerous.
15.6 Sampling Bias and Representativeness
Bias occurs when the expected value of a sample statistic differs from the population parameter it is meant to estimate. It is a systematic error, not a random one, which is why it does not average out over repeated sampling.
Three sources account for most of it in practice:
- Selection bias: The selection procedure systematically excludes or under-represents part of the population.
- Non-response bias: Those who respond differ systematically from those who do not, so the achieved sample is not the sample that was drawn.
- Measurement bias: The instrument itself distorts the values recorded, through leading questions, faulty calibration, or social desirability effects.
A representative sample reflects the population’s relevant characteristics. Representativeness is achieved through random probability selection, not through judgment about which units look typical.
Non-response bias is the one that most often survives an otherwise careful design. A perfect probability sample with a 20% response rate is, in practice, a self-selected sample of the 20% willing to answer. Always report the response rate alongside the sample size, and compare respondents against known population characteristics wherever the frame allows it.
15.7 Sample Size Determination
The sample size needed depends on four things: the precision you require, since a smaller margin of error demands a larger \(n\); the confidence level, since 99% confidence demands more than 95%; the variability of the population, since heterogeneous populations demand more; and the budget and time actually available.
For estimating a population mean:
\[n = \left( \frac{z_{\alpha/2}\,\sigma}{E} \right)^{2}\]
where \(z_{\alpha/2}\) is the critical value (1.96 for 95% confidence), \(\sigma\) is the population standard deviation or an estimate of it, and \(E\) is the desired margin of error.
Example
Estimate average household income to within \(E = ₹50{,}000\) at 95% confidence, where prior data suggests \(\sigma \approx ₹200{,}000\):
\[n = \left( \frac{1.96 \times 200{,}000}{50{,}000} \right)^{2} = (7.84)^{2} \approx 61.5\]
Rounding up, roughly 62 households are needed. Sample size is always rounded up, never down, since rounding down would leave the margin of error wider than specified.
Notice that \(n\) depends on \(E^{2}\) in the denominator. Halving the margin of error therefore quadruples the required sample size, and the population size \(N\) does not appear in the formula at all. This surprises most people: estimating the mean income of a city of one million and a country of one billion to the same precision needs essentially the same sample. What matters is the variability of the population, not its size.
Application
This is the claim of Section 2 put to the test rather than asserted. Both designs draw exactly 100 stores, and both are unbiased, so neither estimate is systematically wrong. What differs is the spread of the estimates across repeated samples, and stratification cuts it substantially, because it removes region-to-region variation from the sampling error entirely. Precision is bought here by design rather than by paying for a bigger sample.
15.8 Sampling Distributions
Every simulation above rests on one idea. A statistic computed from a random sample is itself a random variable: draw a different sample and you get a different value. The distribution of those values across all possible samples is called the sampling distribution of the statistic, and its standard deviation is the standard error.
For the sample mean drawn from a population with standard deviation \(\sigma\):
\[\sigma_{\bar{x}} = \frac{\sigma}{\sqrt{n}}\]
The key result, and the reason random sampling is so powerful, is that even when the population is heavily skewed, the distribution of the sample mean becomes approximately Normal as \(n\) grows. This is the Central Limit Theorem, previewed back in Chapter 12 and taken up properly in the next chapter. It is what makes the Normal distribution the workhorse of inference even for data that is nothing like Normal.
Note also the \(\sqrt{n}\) in the denominator. It is the same fact as the sample-size formula seen from the other side: precision improves with the square root of sample size, so four times the data buys only twice the precision.
15.9 Best Practices
- Use probability sampling whenever inference to a population is the goal.
- Document the method fully, including the frame, the selection procedure, and the response rate, so the work can be reproduced and criticised.
- Check for bias by comparing the achieved sample against known population characteristics.
- Stratify when subgroup representation matters or when a known variable divides the population sharply.
- Address non-response through follow-up contacts and, where necessary, weighting.
- Pilot the procedure on a small scale before committing the full budget.
15.10 Key Formulas
| Concept | Formula |
|---|---|
| Sampling interval (systematic) | \(k = N / n\) |
| Sampling fraction | \(f = n / N\) |
| Proportional allocation | \(n_h = n \times (N_h / N)\) |
| Sample size for a mean | \(n = \left( z_{\alpha/2}\,\sigma / E \right)^{2}\) |
| Standard error of the mean | \(\sigma_{\bar{x}} = \sigma / \sqrt{n}\) |
15.11 Exercises
Stratified design: A university has 3,000 undergraduates and 1,000 graduate students. You want a stratified random sample of 400 students using proportional allocation. How many undergraduates and how many graduates should you select?
-
Choosing a method: For each scenario, identify the most appropriate probability method and justify the choice.
- Estimating average income across a nation of 200 million people.
- Quality inspection of identical items coming off an assembly line.
- A health survey across urban and rural regions where both must be represented.
- A customer satisfaction survey where store locations are clustered in malls.
Bias recognition: A bank surveys customers at branch locations during business hours. Identify the likely sources of sampling bias and suggest improvements.
Sample size: You want to estimate the average spending of online shoppers with 95% confidence and a margin of error of ₹1,000. Prior data suggests \(\sigma \approx ₹8{,}000\). What sample size do you need? What happens to that figure if you tighten the margin to ₹500?
Sampling vs. census: When is a census preferable to sampling? Discuss the trade-offs of cost, time, and precision, and give one situation where a census is actually less accurate than a good sample.
Looking Ahead
This chapter has been about getting the sample right, because everything that follows depends on it. The next chapter turns to what happens after the sample is drawn: how a statistic such as \(\bar{x}\) varies from sample to sample, why that variation is predictable, and how the Central Limit Theorem makes the Normal distribution of Chapter 12 the foundation for the confidence intervals and hypothesis tests of Chapters 15 and 16.
Summary
| Concept | Description |
|---|---|
| Core Vocabulary | |
| Sampling | Selecting a subset of units from a population in order to draw conclusions about the whole |
| Population | The complete set of all units of interest, of size N, whose characteristics are usually unknown |
| Sample | A subset of the population actually observed, of size n |
| Parameter | A numerical characteristic of the population, such as the population mean, fixed but unknown |
| Statistic | A numerical characteristic of a sample, computed from data and used to estimate a parameter |
| Reasons to Sample | Cost, time, practicality, and destructive testing, any of which can rule out a full census |
| The Sampling Frame | |
| Sampling Frame | The actual list of units from which the sample is drawn, ideally matching the population exactly |
| Frame Error | The mismatch between the sampling frame and the population, which no selection method can repair |
| Two Families of Method | |
| Probability Sampling | Every unit has a known, non-zero chance of selection, which is what permits formal inference |
| Non-Probability Sampling | Units chosen by judgment or convenience, leaving bias present and unquantifiable |
| Probability Methods | |
| Simple Random Sampling | Every possible sample of size n is equally likely, the unbiased benchmark all formulas assume |
| Stratified Sampling | Dividing the population into homogeneous strata and sampling within every one of them |
| Proportional Allocation | Allocating the sample across strata in proportion to each stratum's share of the population |
| Optimal (Neyman) Allocation | Allocating more of the sample to strata that are larger or more variable, minimising overall variance |
| Systematic Sampling | Selecting every k-th unit after a random start |
| Sampling Interval | The step size k equals N divided by n |
| Cluster Sampling | Randomly selecting whole clusters and surveying every unit within the chosen ones |
| Intra-Cluster Correlation | The tendency of units within a cluster to resemble each other, which reduces effective sample size |
| Stratified vs. Cluster | Stratified samples from every group seeking homogeneity within; cluster samples some groups entirely |
| Non-Probability Methods | |
| Convenience Sampling | Selecting the easiest units to reach, producing heavily biased and non-generalisable results |
| Judgment Sampling | Selecting units the researcher believes are representative, with bias that cannot be verified |
| Quota Sampling | Filling subgroup quotas by non-random selection, which looks representative but is not |
| Bias and Representativeness | |
| Bias | A systematic difference between a statistic's expected value and the parameter, which more data cannot fix |
| Selection Bias | Systematic exclusion or under-representation of part of the population by the selection procedure |
| Non-Response Bias | Systematic differences between those who respond and those who do not, turning a probability sample self-selected |
| Measurement Bias | Distortion introduced by the measuring instrument itself, through wording, calibration, or social desirability |
| Representative Sample | A sample reflecting the population's relevant characteristics, achieved by randomisation rather than judgment |
| Sample Size | |
| Sample Size Formula | n equals the square of z times sigma divided by the margin of error E, always rounded up |
| Precision and Sample Size | Because E is squared in the formula, halving the margin of error quadruples the sample size needed |
| Towards Inference | |
| Sampling Distribution | The distribution of a statistic across all possible samples, since the statistic is itself a random variable |
| Standard Error | The standard deviation of a sampling distribution; for the mean it is sigma divided by the square root of n |
| Central Limit Theorem (Preview) | Sample means become approximately Normal as n grows, whatever the shape of the population |