Linear Regression (AA HL)
Once a scatter diagram shows a roughly linear trend, the regression line of \(y\) on \(x\) gives you a single equation for predicting \(y\) from \(x\). This page covers how to read off and use that line - finding its equation, substituting values, and knowing when a prediction is trustworthy - with worked examples and the mistakes examiners see most. It's part of the broader Bivariate Statistics topic.
19 questions on this sub-topic.
The regression line
Covered under IB syllabus reference SL4.4: the equation of the regression line of \(y\) on \(x\), and its use for prediction, including the dangers of extrapolation.
Regression line of y on x
\(y=ax+b,\ a=\dfrac{S_{xy}}{S_{xx}},\ b=\bar y-a\bar x\)
Given on the formula booklet, but on paper you normally read \(a\) and \(b\) straight off your GDC's linear regression output rather than computing \(S_{xy}\) and \(S_{xx}\) by hand.
Using it for prediction
Substitute the \(x\)-value into \(y=ax+b\).
Only reliable for \(x\)-values inside the range of the original data - predicting outside that range is extrapolation, and the trend may not hold there.
Need the full syllabus wording and the wider statistics formula table? See Bivariate Statistics. For calculator steps, see the parent topic's GDC guidance.
Worked examples
A regression line is \(y = 3x + 5.\) Estimate \(y\) when \(x = 4.\)
Worked solution
\(y = 3(4) + 5\) M1
\(= 17.\) A1
A data set gives \(r = 0.88\) and regression line \(y = 1.5x + 2\).
(a) Describe the correlation.
(b) Estimate \(y\) when \(x = 10\).
Worked solution
(a) \(r = 0.88\) is close to \(+1\), giving a strong positive linear correlation. A1
(b) Estimate. \(y = 1.5(10) + 2\) M1
\(= 17.\) A1
The regression line of \(y\) on \(x\) is \(y = 2x + 5.\) A new variable \(X = x + 10\) is used (a shift).
(a) Express the regression line of \(y\) on \(X.\)
(b) State the effect of the shift on the correlation coefficient.
Worked solution
(a) \(x = X - 10\), so \(y\) M1
\(= 2(X - 10) + 5\) A1
\(= 2X - 15.\) A1
(b) A shift (adding a constant) does not change \(r\) M1
- correlation is invariant under translation. A1
Common mistakes
- Extrapolating beyond the data range. Using the regression line to predict far outside the \(x\)-values that were actually collected is unreliable, since there's no evidence the linear trend continues.
- Using the y on x line to predict x from y. The two regression lines are generally different - predicting \(x\) from a given \(y\) needs the \(x\) on \(y\) line, not a rearrangement of the \(y\) on \(x\) line.
- Rounding \(a\) and \(b\) too early. If you round the gradient and intercept before substituting, small errors compound - carry extra decimal places (or use exact GDC values) until the final answer.
Ready to practise properly?
19 linear-regression questions, marked instantly like the real exam.
Quick answers
What is the equation of the regression line of y on x?
\(y = ax + b\), where \(a = \dfrac{S_{xy}}{S_{xx}}\) and \(b = \bar y - a\bar x\). In practice you read \(a\) and \(b\) directly off your GDC's linear regression output.
Why shouldn't I use the regression line to extrapolate?
The line only models the trend across the \(x\)-values actually collected. Outside that range there's no evidence the linear pattern continues, so a prediction there can be unreliable or meaningless.