Sample size determination is one of the most consequential decisions a researcher makes before collecting a single data point. Get it wrong in either direction – too small or too large – and the entire study pays a price. A sample that is too small produces unreliable results that can’t be generalized, while an oversized sample wastes time, money, and participant goodwill. The goal, then, is to find the optimal number – the minimum that still delivers results you can trust. Understanding what drives that number is the first step.
Table of Contents
- Why sample size matters in research
- Key factors that determine sample size
- Confidence level
- Margin of error
- Population heterogeneity and variability
- Frequency of the attribute being studied
- Statistical power
- Sampling theory and standard error: the theoretical backbone
- Calculating sample size for estimating means
- Calculating sample size for estimating proportions
- Adjusting for finite populations
- Practical considerations before finalizing sample size
- Tools for calculating sample size
Why sample size matters in research
Sample size is a crucial aspect of research methodology that directly shapes the reliability and validity of study findings. When researchers study a sample rather than an entire population, the results will always differ slightly from the true population value. That difference is called sampling error. A larger sample reduces sampling error because the sample contains a larger proportion of the population, making the estimate more precise. As sample size increases, the standard error – which measures the spread of possible sample estimates around the true value – decreases proportionally.
This relationship between sample size and precision is not just theoretical. Standard error is calculated by dividing the sample’s standard deviation by the square root of the sample size, which means every increase in sample size tightens the range within which we can be confident our estimate is accurate. Researchers then use the standard error to build confidence intervals – the ranges that are statistically likely to contain the true population value.
Key factors that determine sample size
There is no single universal formula for the right sample size. Several factors interact to determine the appropriate number, and each must be carefully considered before the research begins.
Confidence level
The confidence level expresses how certain a researcher wants to be that the true population parameter falls within the estimated range. A 95% confidence level, for example, means you can be 95% certain the results lie between two specific values. The higher the confidence level required, the larger the sample must be. For the same margin of error, a higher confidence level always demands a larger sample size. The most common confidence levels used in social research are 90%, 95%, and 99%, corresponding to Z-scores of 1.65, 1.96, and 2.58 respectively.
Margin of error
The margin of error (also called the confidence interval) defines how much imprecision the researcher is willing to accept. The smaller the allowed margin of error, the larger the required sample size – because tighter precision demands more observations. A margin of error of ±2% requires a far bigger sample than one of ±5%, given the same confidence level. Most social science surveys use a margin of error between 2% and 5%.
Population heterogeneity and variability
How different members of a population are from each other – their heterogeneity – is one of the most important but sometimes underestimated drivers of sample size. A more heterogeneous population implies a larger standard deviation, and therefore requires a larger sample to produce accurate results. In contrast, a homogenous population requires a smaller sample.
Population heterogeneity can pose challenges for sample size determination, as it may require larger samples to adequately capture the diversity of the target population – affecting the precision and generalizability of research findings. When researchers lack prior data on variability, the conservative approach is to assume maximum variability (p = 0.5 in proportion-based studies), which produces the largest – and therefore safest – sample size estimate.
Frequency of the attribute being studied
How common or rare the trait or behavior under study is also affects sample size. When a characteristic is very rare in the population – say, a particular health condition affecting only 3% of people – a larger sample is needed to capture enough cases for meaningful analysis. Conversely, if around half the population holds the attribute, variability is at its maximum, and the sample size needed to estimate it precisely is also at its peak. This is why researchers who have no prior data on the proportion of interest default to p = 0.5 as a conservative planning estimate.
Statistical power
Statistical power refers to the probability of detecting a true difference if one exists, and is heavily dependent on sample size. A statistical power of 80% is common in practice, and 70% is typically considered the minimum acceptable threshold. An underpowered study – one with too small a sample – risks a Type II error: failing to detect a real effect. Adequate power ensures that when a meaningful difference exists in the population, the study has a strong chance of finding it.
Sampling theory and standard error: the theoretical backbone
The theoretical basis for sample size calculation rests on sampling theory and the concept of the standard error. Sampling theory tells us that if we repeatedly drew samples of the same size from a population, the distribution of sample means would approximate a normal (bell-curve) distribution – a principle known as the central limit theorem. This is what makes statistical inference possible: it allows researchers to use a single sample to make probabilistic statements about the entire population.
The standard error quantifies how much a sample statistic (like a mean or proportion) is expected to vary from the true population parameter just by chance. Standard error provides a quantitative measure of how far an estimate is expected to differ from the true population parameter, and is used to calculate the margin of error. When a researcher sets a desired margin of error and confidence level before the study begins, they are essentially working backward from the acceptable standard error to find the minimum sample size that achieves it.
Calculating sample size for estimating means
When the goal is to estimate a population mean – such as the average income, average test score, or average hours worked – the sample size formula is built around the standard deviation of the population and the acceptable margin of error.
The formula is: n = (Z² × σ²) / e²
Where n is the required sample size, Z is the Z-score corresponding to the desired confidence level, σ is the estimated standard deviation of the population, and e is the acceptable margin of error.
For example, if a researcher wants to estimate the population mean with an estimated standard deviation of 10, a margin of error of 2, and a 95% confidence level (Z = 1.96), substituting these values yields a required sample size of approximately 96. If no estimate of the standard deviation is available beforehand, a pilot study can be conducted to obtain one.
It is important to always round up to the next whole number – since a sample size must be an integer, rounding down would put the study below the minimum required threshold and compromise its precision.
Calculating sample size for estimating proportions
When the research question involves a proportion – such as the percentage of voters supporting a policy, or the share of consumers preferring a product – the approach shifts slightly. The most widely used method here is Cochran’s formula.
Cochran’s formula allows a researcher to calculate an ideal sample size given a desired level of precision, confidence level, and the estimated proportion of the attribute present in the population. The formula is: n₀ = (Z² × p × q) / e²
Where n₀ is the required sample size, Z is the Z-score for the chosen confidence level, p is the estimated proportion with the attribute, q is 1 − p, and e is the acceptable margin of error.
A practical example: to estimate the proportion of people at a supermarket who identify as vegan, with 95% confidence and a 5% margin of error, assuming p = 0.5 and an unlimited population size, the formula yields a required sample of at least 385 people. This is a figure that appears repeatedly across social science and public health research as the benchmark for large-population surveys.
Adjusting for finite populations
The standard Cochran formula assumes a very large or unknown population. When the population being studied is smaller and finite, continuing to use the unadjusted formula can result in unnecessary oversampling. The finite population correction (FPC) adjusts the sample size downward for smaller populations, ensuring accurate representation without oversampling.
The corrected formula is: n = n₀ / [1 + (n₀ − 1) / N]
Practical considerations before finalizing sample size
Beyond formulas, several practical factors shape the final number a researcher commits to. Researchers need to determine acceptable precision levels, decide on study power, specify the confidence level, determine the magnitude of practical significance, and engage in an open and realistic dialogue about the appropriateness of the calculated sample size given the research question, available data records, research timeline, and cost.
Non-response is another key consideration. In survey research, not everyone selected will respond. A standard practice is to inflate the calculated sample size to account for expected non-response. If 20% of selected participants typically don’t respond, a researcher planning to collect data from 385 people should initially select around 480. Optimal sample size must take into account total population size, effect size, statistical power, confidence level, and margin of error to ensure reliability, validity, and empirical rigor.
Finally, when population standard deviation is unknown and no prior studies are available, a small pilot study – typically involving 30 to 50 participants – can provide the variance estimate needed to calculate a reliable sample size before the main study begins.
Tools for calculating sample size
Researchers today have access to a range of free and accessible tools that automate sample size calculations. OpenEpi (an open-source online calculator) and G*Power (a statistical software package) are among the most commonly used for sample size calculations across different study designs. Online calculators from platforms like Creative Research Systems allow researchers to input confidence level, margin of error, and population size to get instant results. However, it is worth noting that these tools work best when the researcher already understands the underlying parameters – they automate the calculation, but not the judgment.
What do you think? If two studies examine the same research question but one uses a sample of 100 and another uses 400, how much should you trust any difference in their conclusions? And when resources are limited, how should a researcher decide whether to prioritize a higher confidence level or a smaller margin of error?
References
- https://pmc.ncbi.nlm.nih.gov/articles/PMC10000262/
- https://en.wikipedia.org/wiki/Sample_size_determination
- https://www.britannica.com/science/sampling-error
- https://docmckee.com/cj/docs-research-glossary/sampling-error-definition/
- https://www.surveymonkey.com/mp/sample-size-calculator/
- https://ecampusontario.pressbooks.pub/introstats/chapter/7-5-calculating-the-sample-size-for-a-confidence-interval/
- https://ihopejournalofophthalmology.com/sample-size-and-its-evolution-in-research/
- https://kuey.net/index.php/kuey/article/download/6040/4340/12261
- https://hsij.anandafound.com/journal/article/download/63/43
- https://sawtoothsoftware.com/resources/blog/posts/determining-sample-size-for-survey-research
- https://www.ijbmi.org/papers/Vol(13)7/1307152167.pdf
- https://www.statisticshowto.com/probability-and-statistics/find-sample-size/
- https://www.calculator.net/sample-size-calculator.html
- https://dissertationdataanalysishelp.com/cochrans-sample-size-calculator/
- https://www.sciencedirect.com/science/article/pii/S2772906024005089
- https://www.surveysystem.com/sscalc.htm
Leave a Reply