Linear Regression (AI HL)
When a scatter diagram shows points clustered around a straight-line trend, the line of best fit is the \(y\) on \(x\) regression line, found by technology so that it minimises the vertical distances to the data. This page covers the equation of that line, Pearson's correlation coefficient \(r\), and why the direction you fit the line matters. It's part of the broader Bivariate & Non-linear Regression topic.
13 questions on this sub-topic.
The regression line
Covered under IB syllabus reference SL4.4: linear correlation of bivariate data and Pearson's product-moment correlation coefficient \(r\), found using technology and meaningful only for linear relationships, together with the equation of the \(y\) on \(x\) regression line and the dangers of extrapolation.
Regression line of \(y\) on \(x\)
\(y=ax+b\)
Not in the formula booklet - your GDC generates \(a\) and \(b\) directly from the data. The line always passes through the mean point \((\bar{x},\bar{y})\).
Residual
\(\text{residual} = y_{\text{observed}} - y_{\text{predicted}}\)
How far a real data point sits above or below the line - a large residual flags an outlier or a point the line fits poorly.
Need the full syllabus wording and formula-booklet reference table? See Bivariate & Non-linear Regression.
Worked examples
A regression line is \(y = 3x + 2.\) A data point is \((4, 16).\)
(a) Find the predicted value of \(y\) when \(x = 4.\)
(b) Find the residual for this point.
Worked solution
(a) Predicted \(y = 3(4) + 2\) M1
\(= 14.\) A1
(b) Residual \(= \text{observed} - \text{predicted} = 16 - 14\) M1
\(= 2.\) A1
For a data set, the regression line of \(y\) on \(x\) is \(y = 1.8x + 4.\) The mean point is \((\bar{x}, \bar{y}) = (10, 22).\)
(a) Verify that the mean point lies on the line.
(b) Explain why the line of \(y\) on \(x\) should be used to predict \(y\) from \(x\), not \(x\) from \(y\).
Worked solution
(a) \(1.8(10) + 4 = 22\) M1
\(= \bar y\), so \((10,22)\) lies on the line. A1 AG
(b) The \(y\) on \(x\) line minimises vertical (\(y\)) errors, so it is for predicting \(y\) from \(x\); M1 A1
using it in reverse would not minimise the relevant errors. A1
Common mistakes
- Fitting the line in the wrong direction. The regression line of \(y\) on \(x\) minimises errors in \(y\), so it should only be used to predict \(y\) from a given \(x\) - not rearranged to predict \(x\) from \(y\).
- Getting the residual sign backwards. A residual is observed minus predicted, not the other way round - a point above the line gives a positive residual.
- Trusting a prediction far outside the data range. Extrapolating well beyond the smallest and largest \(x\)-values used to fit the line assumes the same linear trend continues, which often isn't true.
Ready to practise properly?
13 linear-regression questions, marked instantly like the real exam.
Quick answers
What is the equation of the linear regression line?
The regression line of \(y\) on \(x\) has equation \(y = ax + b\), where \(a\) and \(b\) are found using technology from the data. The line always passes through the mean point \((\bar{x}, \bar{y})\).
What does a residual measure?
A residual is the difference between an observed \(y\)-value and the value the regression line predicts for that \(x\), calculated as observed minus predicted. See using your GDC for how to generate the line itself.