Sampling variation is the natural tendency for different random samples from the same population or data generating process (DGP) to produce different results.
Because each random sample contains different observations, sample statistics and model results will vary from one sample to another.
Imagine repeatedly drawing random samples from the same population or DGP.
Each sample will contain a different collection of observations, so measures such as:
Means
Proportions
Correlations
PRE values
Model coefficients
will not be exactly the same every time.
These differences are called sampling variation.
Suppose the average test score in a school is 80.
A researcher draws three different random samples of students and calculates the average score for each sample:
The averages differ because of sampling variation.
These concepts are closely related but not identical.
For example:
Sample averages of 78, 82, and 80 demonstrate sampling variation.
The difference between an average of 78 and the population average of 80 is sampling error.
Sampling variation creates sampling error.
Sampling variation is not:
A mistake
Measurement error
A sign that the data are wrong
It is an expected consequence of taking samples from a larger population or DGP.
Even perfectly collected random samples will show sampling variation.
Larger samples tend to show less sampling variation.
As sample size increases:
Sample statistics become more stable.
Different samples produce more similar results.
Estimates tend to be closer to population values.
This is why larger samples often produce more precise estimates.
Suppose several researchers collect different random samples from the same DGP and fit the same model.
They may obtain:
Different regression coefficients
Different PRE values
Different F ratios
Different predictions
These differences may simply be the result of sampling variation rather than meaningful differences in the underlying process.
Understanding sampling variation helps researchers recognize that models built from samples contain uncertainty.
Sampling variation is the foundation of statistical inference.
It helps explain:
Why different studies can produce different results
Why estimates are uncertain
Why confidence intervals are needed
Why replication is important in science
Without sampling variation, statistical inference would not be necessary.