Sampling error is the difference between a result calculated from a sample and the corresponding value in the population or data generating process (DGP). Sampling error occurs because different random samples contain different observations.
Even when a sample is collected perfectly, sampling error is expected because a sample is only a subset of the larger population or DGP.
Imagine taking two different random samples from the same population or DGP.
Because the samples contain different observations:
The sample means may differ.
The sample proportions may differ.
The models fit to the samples may differ.
These differences occur due to sampling error.
Suppose the average height of all students in a school is 66 inches.
A researcher takes a random sample and finds an average height of 65 inches.
Sampling Error = 65 − 66 = −1
The sample estimate differs from the population value by 1 inch.
Sampling error is a natural consequence of using a sample instead of observing the entire population.
Sampling error can occur even when data collection is done perfectly.
Random sampling helps reduce bias, but it does not eliminate sampling error.
Two researchers could:
Draw different random samples from the same data generating process.
Obtain slightly different results.
Both results could be correct given the samples they observed.
Larger samples tend to produce less sampling error.
As sample size increases:
Sample estimates become more stable.
Different samples become more similar.
Estimates tend to be closer to population values.
This is one reason researchers often prefer larger samples when possible.
When building models, different random samples may produce:
Different parameter estimates
Different PRE values
Different F ratios
Different predictions
Some variation in model results is due simply to sampling error.
Understanding sampling error helps researchers avoid treating every difference between samples as a meaningful pattern.
Sampling error reminds us that:
Samples are not identical to populations.
Different samples can produce different results.
Statistical conclusions always involve some uncertainty.
Recognizing sampling error is an important part of statistical reasoning.