Interpreting data
Describe the shape of distributions, identify and explain outliers, and compare two datasets using measures of centre and spread.
Worked examples
Describing the shape of a distribution
Straightforward
Problem
A frequency histogram shows the number of hours students spend on homework each week. The bars are tallest on the left (around 1–3 hours) and become shorter towards the right, with a few students recording 8–10 hours. Describe the shape of this distribution and explain what it means in context.
1
Identify where most of the data is concentrated.
The tallest bars are on the left side of the histogram, around – hours. This is where most students' data lies.
2
Identify the direction of the tail.
The bars gradually decrease in height as we move to the right, with a few students in the – hour range. The tail extends to the right.
3
Name the shape of the distribution.
Because most data is on the left and the tail extends to the right, this distribution is positively skewed (right-skewed).
4
Interpret the shape in context.
In context, most students spend only – hours on homework each week, while a small number spend considerably more time. The mean will be pulled to the right of the median because of these higher values.
Answer
The distribution is positively skewed. Most students spend – hours on homework, with a few students spending much more time, creating a tail on the right.
Identifying and explaining an outlier
Moderate
Problem
The monthly salaries (in dollars) of eight employees at a small business are:
Identify the outlier, then compare the mean salary with and without the outlier. Explain what effect the outlier has.
Identify the outlier, then compare the mean salary with and without the outlier. Explain what effect the outlier has.
1
Identify the outlier.
The salaries range from to , except for the last value of , which is far above the others. The outlier is (likely the owner's salary).
2
Calculate the mean of all eight values.
3
Calculate the mean without the outlier.
4
Compare and explain.
With the outlier, the mean is . Without it, the mean is . The outlier raises the mean by over , making it a poor representation of what a typical employee earns. The median of the seven non-outlier values is , which better represents the typical salary.
Answer
The outlier is . The mean with the outlier is ; without it the mean is . The outlier inflates the mean by over , making the median a more appropriate measure of centre.
Comparing two datasets using centre and spread
Challenging
Problem
Two brands of batteries were tested for their lifetimes (in hours). The results are summarised below:
Compare the two brands and make a recommendation for a customer who wants reliable batteries.
Compare the two brands and make a recommendation for a customer who wants reliable batteries.
1
Compare the measures of centre.
Brand X has a mean of hours and Brand Y has a mean of hours. Brand X lasts longer on average by hours. The medians ( vs ) also confirm Brand X performs better at the centre of the distribution.
2
Compare the measures of spread.
Brand X has an IQR of hours compared to Brand Y's IQR of hours. Brand X also has a much smaller range ( vs hours). The smaller spread for Brand X means its battery lifetimes are much more consistent.
3
Note the difference between mean and median for each brand.
For Brand X, mean () and median () are very close — the distribution is roughly symmetric. For Brand Y, mean and median are both — also symmetric, but with much greater spread.
4
Make a recommendation.
Brand X lasts longer on average and is far more consistent (smaller IQR and range). A customer wanting reliable batteries should choose Brand X, as they can expect most batteries to last close to – hours with little variation.
Answer
Brand X lasts longer on average (mean h vs h) and is more consistent (IQR h vs h). Brand X is the better choice for reliability.
Practise
Q1·Straightforward
A dot plot shows most values clustered near the left, with a long tail stretching to the right. How would you describe the shape of this distribution?
Explanation
When the tail extends to the right (towards higher values), the distribution is called positively skewed (or right-skewed). Most data is bunched on the left, with a few large values pulling the tail to the right.
Q2·Straightforward
A histogram is roughly the same shape on both sides of the centre, with no obvious long tail on either side. How would you describe this distribution?
Explanation
If the left and right sides of the histogram are roughly mirror images of each other, the distribution is symmetric. In a symmetric distribution the mean and median are close to each other.
Q3·Straightforward
A stem-and-leaf plot for test scores shows most data near the high end, with a long tail stretching to the left (low scores). Which description fits best?
Explanation
When the tail extends to the left (towards lower values), the distribution is negatively skewed (or left-skewed). Most scores are high, but a few low scores create the tail on the left.
Q4·Straightforward
The ages of people at a community event are: . Which value is an outlier?
Explanation
The value is an outlier because it is much larger than all other values in the dataset, which are clustered between and . Outliers are data points that lie far from the main group.
Q5·Moderate
The daily maximum temperatures (°C) for two weeks are: . The value is an outlier. Which statement best explains the effect of this outlier on the mean?
Explanation
The mean of the 13 non-outlier values is about °C. Including pushes the mean of all 14 values to about °C — higher than most of the actual temperatures. Outliers pull the mean towards them, which is why the median is often a better measure of centre for skewed data.
Q6·Moderate
A dataset of house prices (in s) is: . The value is an outlier. What is the mean of the seven non-outlier values? Give your answer to the nearest whole number.
Explanation
Sum of the seven values: . Mean , which rounds to (in s).
Q7·Moderate
A dataset has a large outlier at the high end. Which measure of centre is least affected by this outlier?
Explanation
The median is resistant to outliers because it is determined by the middle position — a single extreme value cannot move the middle very much. The mean is pulled towards the outlier because every value contributes to the sum. The range is also heavily affected because it uses the maximum and minimum values.
Q8·Moderate
A student earns these weekly amounts from casual work: . The is an outlier (an unusually busy week). Which statement correctly describes the effect of removing the outlier on the range?
Explanation
With the outlier: range . Without the outlier: range . Removing the outlier greatly reduces the range, showing how sensitive the range is to extreme values.
Q9·Challenging
Two classes sat the same maths test. Class A has a mean of and an IQR of . Class B has a mean of and an IQR of . Which conclusion is best supported by the data?
Explanation
Class A has a higher mean (), so it performed better on average. Class A also has a smaller IQR (), meaning its results are more tightly clustered — more consistent. Class B's wider IQR indicates more variation in individual scores.
Q10·Challenging
Dataset A: . Dataset B: . Both datasets have the same median. By how much does the mean of Dataset B exceed the mean of Dataset A?
Explanation
Mean of A . Mean of B . Difference . Both datasets have median (the middle value), showing that the outlier raised the mean but did not change the median.
Q11·Challenging
The points scored per game by two basketball players over a season are summarised below:
Which statement best describes the comparison between the two players?
Which statement best describes the comparison between the two players?
Explanation
Both players have the same mean ( points), so their average scoring is equal. However, Player A's IQR is compared to Player B's IQR of . A smaller IQR means Player A's scores are more tightly clustered — more consistent from game to game. Player B's larger IQR indicates more variable performance.
Q12·Challenging
A dataset of values has a mean of . A new value of is added. What happens to the mean and median after adding this outlier?
Explanation
Original sum . New mean , an increase of . The median of ordered values is the 6th value. Adding at the top shifts the 6th position by one place, typically producing a much smaller change in the median. The mean increases more because it is directly pulled towards the large outlier.