Linear Regression (AI HL)

When a scatter diagram shows points clustered around a straight-line trend, the line of best fit is the \(y\) on \(x\) regression line, found by technology so that it minimises the vertical distances to the data. This page covers the equation of that line, Pearson's correlation coefficient \(r\), and why the direction you fit the line matters. It's part of the broader Bivariate & Non-linear Regression topic.

13 questions on this sub-topic.

Practise linear regression → Try exam-style questions

The regression line

Covered under IB syllabus reference SL4.4: linear correlation of bivariate data and Pearson's product-moment correlation coefficient \(r\), found using technology and meaningful only for linear relationships, together with the equation of the \(y\) on \(x\) regression line and the dangers of extrapolation.

Regression line of \(y\) on \(x\)

\(y=ax+b\)

Not in the formula booklet - your GDC generates \(a\) and \(b\) directly from the data. The line always passes through the mean point \((\bar{x},\bar{y})\).

Residual

\(\text{residual} = y_{\text{observed}} - y_{\text{predicted}}\)

How far a real data point sits above or below the line - a large residual flags an outlier or a point the line fits poorly.

Need the full syllabus wording and formula-booklet reference table? See Bivariate & Non-linear Regression.

Worked examples

1
Medium
GDC
[4 marks]

A regression line is \(y = 3x + 2.\) A data point is \((4, 16).\)

(a) Find the predicted value of \(y\) when \(x = 4.\)
(b) Find the residual for this point.

Worked solution

(a) Predicted \(y = 3(4) + 2\) M1
\(= 14.\) A1

(b) Residual \(= \text{observed} - \text{predicted} = 16 - 14\) M1
\(= 2.\) A1

M1 Substitute A1 Correct answer of \(14\) M1 Obs − pred A1 Correct answer of \(2\)
2
Hard
GDC
[5 marks]

For a data set, the regression line of \(y\) on \(x\) is \(y = 1.8x + 4.\) The mean point is \((\bar{x}, \bar{y}) = (10, 22).\)

(a) Verify that the mean point lies on the line.
(b) Explain why the line of \(y\) on \(x\) should be used to predict \(y\) from \(x\), not \(x\) from \(y\).

Worked solution

(a) \(1.8(10) + 4 = 22\) M1
\(= \bar y\), so \((10,22)\) lies on the line. A1 AG

(b) The \(y\) on \(x\) line minimises vertical (\(y\)) errors, so it is for predicting \(y\) from \(x\); M1 A1
using it in reverse would not minimise the relevant errors. A1

M1 Substitute A1 \(=\bar y\) M1 Minimises \(y\)-errors A1 Predict \(y\) from \(x\) A1 Not reverse

Common mistakes

Ready to practise properly?

13 linear-regression questions, marked instantly like the real exam.

Quick answers

What is the equation of the linear regression line?

The regression line of \(y\) on \(x\) has equation \(y = ax + b\), where \(a\) and \(b\) are found using technology from the data. The line always passes through the mean point \((\bar{x}, \bar{y})\).

What does a residual measure?

A residual is the difference between an observed \(y\)-value and the value the regression line predicts for that \(x\), calculated as observed minus predicted. See using your GDC for how to generate the line itself.

← Back to Applications & Interpretation HL topics