When researchers collect data – whether on income levels, test scores, or health outcomes – knowing the average tells only part of the story. What matters just as much is how spread out that data is. That’s where measures of dispersion come in. The simplest of these is the range: a quick, direct way to see how far apart your data points stretch. But simplicity comes at a price. Understanding what the range can and cannot tell you is essential for anyone working with data in social research.
Table of Contents
- What is the range?
- Why the range is considered a “crude” measure
- The range as a biased estimator
- How sample size affects the range
- Where the range is genuinely useful
- The range in context: comparing it to other measures
- Interquartile range (IQR)
- Standard deviation
- Mean deviation
- Avoiding over-reliance on the range
- Calculating the range: a quick example
What is the range?
The range is defined as the difference between the highest and lowest values in a dataset. The formula is straightforward:
Range = Highest value − Lowest value
For example, if a researcher surveys household incomes in a community and finds the lowest is $20,000 and the highest is $200,000, the range is $180,000. That single number immediately signals significant economic disparity. The key point is that the range reflects the values actually recorded in the study – not the highest and lowest scores theoretically possible.
Why the range is considered a “crude” measure
Statisticians regularly describe the range as a crude measure of dispersion, and there are clear reasons for that label. While the range is easy to calculate, it is very sensitive to outliers and does not use all the observations in a dataset. It only registers the two most extreme data points while completely ignoring every value in between.
Consider two datasets: Dataset A has values 10, 50, 50, 50, 90, and Dataset B has values 10, 20, 30, 80, 90. Both have an identical range of 80. Yet their distributions tell completely different stories – Dataset A clusters tightly around 50, while Dataset B is more evenly spread across the range. The range alone cannot reveal this distinction.
In social research, this limitation has real consequences. The range is the simplest measure of dispersion, but it ignores variability between the extremes and is heavily influenced by outliers. A single extremely wealthy household in a community survey, for instance, can make an entire neighborhood appear far more economically diverse than it actually is.
The range as a biased estimator
One of the more technically important criticisms of the range is that it is a biased estimator. In statistics, a biased estimator is one whose expected value does not equal the population parameter it aims to represent – meaning it systematically overestimates or underestimates the true value of the parameter, even with an infinite sample size.
The range is specifically a downward-biased estimator of the population range. Because the sample range can never exceed the population range, it will almost always underestimate the true spread of the full population. When you draw a sample from a larger population, it is very unlikely that your sample will capture the absolute maximum and minimum of the entire population. As a result, the sample range consistently falls short of the true population range.
A biased estimator systematically misses the true parameter in one direction – and this is exactly what happens with the range. The smaller the sample, the more pronounced this underestimation tends to be. This is why researchers working with small or skewed samples should be especially cautious about relying on the range as a measure of population variability.
How sample size affects the range
There is an important relationship between sample size and the range that researchers must keep in mind. Increased sample size is associated with an increase in the range, because larger samples have greater potential to capture more extreme values. Conversely, a small sample is less likely to include the true minimum or maximum of the population, making the sample range a poor reflection of the full population’s spread.
This means that when comparing ranges across two groups, any observed differences could partly be an artifact of different sample sizes rather than genuine differences in variability. A researcher comparing income ranges between a sample of 20 households and a sample of 200 households should not treat those ranges as directly comparable.
Where the range is genuinely useful
Despite its limitations, the range is not without purpose. The range is a very useful measure in statistical quality control of products in industries where the interest lies in getting a quick rather than an accurate measure of variability. In social and public policy research, it serves similar quick-scan functions.
Some practical applications include:
Preliminary data exploration: Before applying more advanced analysis, the range gives researchers a first impression of how broadly the data spreads.
Communicating with non-technical audiences: When presenting to non-technical audiences, the range is far more accessible than standard deviations or interquartile ranges. Telling a community that household incomes range from $20,000 to $200,000 is immediately understood without any statistical background.
Highlighting extreme disparities: In social research, the range can highlight extreme disparities – for example, if years of schooling in a community range from 0 to 20, the range quickly reveals the extent of educational inequality.
Detecting data entry errors: In survey research, an unusually large range might indicate data entry errors or outliers that need investigation.
Ordinal data analysis: When data is ordinal and a mean cannot be established, the range is commonly used to calculate the measure of dispersion in the dataset.
The range in context: comparing it to other measures
The range sits at the entry point of a family of dispersion measures, each more sophisticated than the last. Understanding where it stands relative to alternatives helps researchers decide when to use it.
Interquartile range (IQR)
The interquartile range describes the middle 50% of observations and has the advantage of being usable when extreme values are not recorded exactly, as in open-ended class intervals. Unlike the range, it is not affected by outliers, making it a more reliable measure when data is skewed.
Standard deviation
Standard deviation is the most commonly used measure of dispersion – it measures spread around the mean and takes into account all values in the distribution. In sociological research, a high standard deviation in income data points to greater inequality, while a low standard deviation in survey responses indicates consensus across respondents. Unlike the range, it uses every data point and is far less sensitive to single extreme values.
Mean deviation
Mean deviation is the sum of absolute differences between observations and their mean, divided by the number of observations. It is a more complete measure than the range because it considers all values, though it is used less frequently than standard deviation in practice.
Avoiding over-reliance on the range
A common mistake in data analysis is reporting only a single measure of dispersion – especially the range – and treating it as a complete picture of variability. Researchers also sometimes choose inappropriate measures of dispersion for their data type. Using only the range with large datasets provides insufficient information about the true nature of data distribution.
The sound approach is to use the range as a starting point, not a conclusion. Pair it with the median for a sense of center and spread together. If outliers are present, shift to the interquartile range. When the full distribution needs to be described, standard deviation is the more appropriate tool. The range is helpful in calculating preliminary variability, but it is not worthy for thorough analysis because it is affected by the extreme values of the sample distribution.
In social research specifically, where data on income, education, health, and behavior often contain extreme values, relying on the range alone can produce genuinely misleading conclusions. A community where one billionaire lives among low-income households will show an enormous income range – but that number tells us almost nothing about the economic reality facing the majority of residents.
Calculating the range: a quick example
The mechanics are simple. Suppose a researcher collects data on the number of years of formal education completed by 7 participants: 4, 8, 10, 12, 14, 16, 18.
Range = 18 − 4 = 14 years
The range is the simplest measure of variability to calculate, and one encountered frequently – but it can be a very unreliable measure when not all the scores between the extremes are considered. In this example, the range tells us the spread is 14 years, but it says nothing about whether most participants clustered around 10-12 years or were evenly distributed across the full span.
This is the essential tension at the heart of the range as a statistical tool: it is maximally easy to compute and interpret, but minimally informative about the actual distribution of data. It provides a boundary, not a picture.
What do you think? If two communities show the same income range but very different distributions of wealth within that range, does the range still serve a useful purpose in comparing them? And given that the sample range consistently underestimates the true population range, how should researchers adjust their interpretation when working with small datasets?
References
- https://www.qualityresearchinternational.com/socialresearch/dispersion.htm
- https://www.dummies.com/article/body-mind-spirit/emotional-health-psychology/psychology/research/choosing-the-right-measure-of-dispersion-in-psychology-statistics-169544/
- https://pmc.ncbi.nlm.nih.gov/articles/PMC3198538/
- https://pubadmin.institute/research-methodologies/understanding-range-measure-dispersion
- https://hubsociology.com/measures-of-dispersion-in-social-research-a-comp/
- https://www.bizmanualz.com/library/what-does-biased-estimator-mean
- https://www.johndcook.com/blog/2014/10/24/sample-range/
- https://www.examples.com/ap-statistics/biased-and-unbiased-point-estimates
- https://open.maricopa.edu/psy230mm/chapter/chapter-5-measures-of-dispersion/
- https://www.sociologyguide.com/research-methods&statistics/measures-of-dispersion.php
- https://www.vaia.com/en-us/explanations/psychology/data-handling-and-analysis/measures-of-dispersion/
- https://psychology.town/statistics/measures-of-dispersion-researchers-guide/
- https://www.analyticssteps.com/blogs/measure-dispersion-definition-methods-and-examples
Leave a Reply