Predictions and limitations
Use lines of best fit to interpolate and extrapolate, and identify limitations of statistical models.
Worked examples
Interpolating from a line of best fit
Straightforward
Problem
A scatter plot shows the relationship between the number of hours a student studies per week () and their exam score out of 100 (). The line of best fit has equation . The data were collected for values between 2 and 12.
Use the line to predict the exam score for a student who studies 7 hours per week.
Use the line to predict the exam score for a student who studies 7 hours per week.
1
Check whether the value is within the data range.
The data range is to . Since is within this range, the prediction is an interpolation.
2
Substitute into the equation of the line of best fit.
3
Calculate the predicted score.
Answer
The line of best fit predicts an exam score of 65 for a student who studies 7 hours per week.
Extrapolating and explaining reduced reliability
Moderate
Problem
A researcher records the number of years of work experience () and annual salary ($000s, ) for employees at a company. The data were collected for employees with 1 to 15 years of experience. The line of best fit is .
(a) Calculate the predicted salary for an employee with 25 years of experience.
(b) Give one reason why this prediction is less reliable than a prediction made within the data range.
(a) Calculate the predicted salary for an employee with 25 years of experience.
(b) Give one reason why this prediction is less reliable than a prediction made within the data range.
1
Identify whether is within the data range.
The data range is to . Since is outside this range, the prediction is extrapolation.
2
Substitute into the equation.
3
The model predicts an annual salary of $122\,000.
Calculate the predicted salary.
The model predicts an annual salary of $122\,000.
4
Explain why this extrapolated prediction is less reliable.
The data only covers employees with up to 15 years of experience. Beyond this range, the rate of salary growth may slow down (e.g. due to salary caps or reduced promotions), so the linear pattern may not continue. Predictions outside the data range are less reliable because the relationship between variables may change.
Answer
(a) $122\,000. (b) The prediction is extrapolation — the linear trend observed up to 15 years may not continue at the same rate beyond the observed data range.
Identifying limitations of a statistical model
Challenging
Problem
A Year 10 student surveys 11 of her classmates about their daily screen time (hours) and their score on a recent maths test (%). She finds a moderate negative association and draws a line of best fit. She concludes: 'Screen time causes lower maths scores. This model can predict the maths score of any student in Australia from their screen time.'
Identify TWO specific limitations of this conclusion.
Identify TWO specific limitations of this conclusion.
1
Consider the sample size and how the sample was selected.
The sample consists of only 11 classmates — this is a very small sample. A reliable model typically requires a much larger sample. Additionally, the classmates were not randomly selected, so the sample is not representative of all Australian students. This is a convenience sample, which may share similar habits, school environments, and demographics.
2
Consider whether a statistical association implies causation.
A moderate negative association means that higher screen time tends to occur alongside lower test scores in this data set. However, association does not imply causation. Other variables — such as time management, study habits, or sleep — may influence both screen time and test scores. The model cannot justify the claim that screen time causes lower scores.
3
State the two limitations clearly.
Limitation 1: The sample is too small (11 students) and is a convenience sample, making it unrepresentative of all Australian students.
Limitation 2: Association does not imply causation — the negative association does not prove that screen time causes lower maths scores.
Limitation 2: Association does not imply causation — the negative association does not prove that screen time causes lower maths scores.
Answer
Two limitations: (1) The sample is too small and unrepresentative — it cannot be generalised to all Australian students. (2) Association does not imply causation — the model shows a statistical relationship only, not a causal one.
Practise
Q1·Straightforward
A scatter plot shows the relationship between daily temperature (, in °C) and the number of visitors to a beach (). The line of best fit has equation . The data were collected for temperatures between 15°C and 35°C.
Use the line to predict the number of visitors on a day when the temperature is 25°C.
Use the line to predict the number of visitors on a day when the temperature is 25°C.
Explanation
Since is within the data range (15 to 35), this is interpolation.
Substituting :
The line of best fit predicts 140 visitors.
Substituting :
The line of best fit predicts 140 visitors.
Q2·Straightforward
A line of best fit for a data set has equation . The data were collected for values between 2 and 10.
Use the line to predict when .
Use the line to predict when .
Explanation
Since is within the data range (2 to 10), this is interpolation.
Substituting :
Substituting :
Q3·Straightforward
A scatter plot comparing a student's hours of practice per week () and their music performance score (, out of 100) has a line of best fit passing through the points and . The data range is to .
Use the line to predict the performance score for a student who practises 5 hours per week.
Use the line to predict the performance score for a student who practises 5 hours per week.
Explanation
Gradient:
Using the point :
The line predicts a performance score of 60 for 5 hours of practice per week.
Using the point :
The line predicts a performance score of 60 for 5 hours of practice per week.
Q4·Straightforward
A scatter plot comparing the age of a car (, in years) and its resale value (, in thousands of dollars) has a line of best fit with equation . The data were collected for cars aged 1 to 8 years.
Use the line to estimate the resale value of a 5-year-old car.
Use the line to estimate the resale value of a 5-year-old car.
Explanation
Substituting :
The line of best fit estimates a resale value of $15\,500 for a 5-year-old car.
The line of best fit estimates a resale value of $15\,500 for a 5-year-old car.
Q5·Moderate
A researcher records the foot length (, in cm) and height (, in cm) of adults aged 20–40. The data range is cm to cm. The line of best fit is .
A student uses this equation to predict the height of a person with a foot length of 38 cm. Calculate the predicted height.
A student uses this equation to predict the height of a person with a foot length of 38 cm. Calculate the predicted height.
Explanation
Substituting :
The equation predicts a height of 235 cm. However, is well outside the observed data range (22 to 30 cm), so this is extrapolation. The prediction is unreliable because the linear relationship may not hold beyond the observed range.
The equation predicts a height of 235 cm. However, is well outside the observed data range (22 to 30 cm), so this is extrapolation. The prediction is unreliable because the linear relationship may not hold beyond the observed range.
Q6·Moderate
A line of best fit is based on data collected within a certain range of values. A researcher extends the line beyond the data range to make a prediction.
Which statement best explains why this extrapolated prediction is less reliable than an interpolated one?
Which statement best explains why this extrapolated prediction is less reliable than an interpolated one?
Explanation
When we extrapolate, we assume the linear pattern continues beyond the observed data. In practice, the relationship may level off, change direction, or behave differently outside the data range. This makes extrapolated predictions less reliable than interpolated ones.
Q7·Moderate
A line of best fit for a data set is . The data were collected for values from 10 to 40.
A scientist uses the equation to predict when . Calculate the predicted value of .
A scientist uses the equation to predict when . Calculate the predicted value of .
Explanation
Substituting :
Note that is outside the data range (10 to 40), so this is extrapolation. The predicted value of may not be reliable.
Note that is outside the data range (10 to 40), so this is extrapolation. The predicted value of may not be reliable.
Q8·Moderate
A meteorologist records daily maximum temperature and electricity usage in a city for the months of January to June. She uses the line of best fit to predict electricity usage in August.
Which statement about this prediction is most accurate?
Which statement about this prediction is most accurate?
Explanation
The data were collected from January to June. August falls outside this range, so any prediction for August is extrapolation. Extrapolation is less reliable because conditions (and the relationship between variables) may differ outside the observed period.
Q9·Challenging
A student surveys 9 of her friends to investigate whether the number of hours spent on homework each week is associated with their average assessment mark. She finds a strong positive association and concludes that her model can be used to predict marks for all NSW Year 10 students.
Which limitation is most relevant to her conclusion?
Which limitation is most relevant to her conclusion?
Explanation
A sample of 9 friends is very small and is not randomly selected — it is a convenience sample. Surveying only her friends means the sample is likely to share similar backgrounds, habits, and school environments. This makes the model unreliable for predicting marks across all NSW Year 10 students.
Q10·Challenging
A researcher plots a scatter plot of ice cream sales () versus the number of drownings at beaches () across different months. The plot shows a strong positive association. The researcher concludes that eating ice cream causes drowning.
Which statement best identifies the flaw in this conclusion?
Which statement best identifies the flaw in this conclusion?
Explanation
A strong association between two variables does not mean one causes the other. In this case, both ice cream sales and drownings tend to increase during hot weather — hot weather is a confounding variable that influences both. This is a classic example of a spurious correlation.
Q11·Challenging
A line of best fit is drawn through a scatter plot of 25 data points. Four of the points lie very far from the line and from the main cluster of data.
How do these outliers affect the line of best fit?
How do these outliers affect the line of best fit?
Explanation
Outliers are data points that lie far from the general trend. A line of best fit attempts to minimise the distance to all data points, so outliers can shift the line toward themselves, distorting its gradient or position. This reduces the line's accuracy as a model for the majority of the data.
Q12·Challenging
A researcher collects data from 200 Year 10 students at a single selective school to investigate the relationship between study time and academic performance. She builds a linear model and uses it to predict academic performance for all Australian secondary school students.
Which limitation is most relevant?
Which limitation is most relevant?
Explanation
Students at a selective school tend to be higher achieving and may have different study habits compared with students at comprehensive schools. The sample is not representative of all Australian secondary students, so applying the model to the broader population is not valid. This is a sampling bias limitation.