Calculating Correlation & Regression (AI HL)

When you have two numerical variables measured on the same individuals, you can ask how strongly they're linearly related and use that relationship to make predictions. This topic covers finding Pearson's product-moment correlation coefficient \(r\) on your GDC, fitting the regression line of \(y\) on \(x\) (and, less commonly, \(x\) on \(y\)), and using that line responsibly - including knowing when a prediction is trustworthy and when correlation is being mistaken for causation.

What the syllabus says

This topic maps onto one point in the official IB Applications & Interpretation syllabus.

CodeSyllabus content
SL4.4Linear correlation of bivariate data. Pearson's product-moment correlation coefficient \(r\), calculated using technology - meaningful only for linear relationships. Scatter diagrams and lines of best fit by eye, passing through the mean point \((\bar{x},\bar{y})\); distinguishing correlation from causation. The equation of the regression line of \(y\) on \(x\), found using technology, used for prediction with awareness of the dangers of extrapolation and that this line cannot always reliably predict \(x\) from a value of \(y\); interpretation of the parameters \(a\) and \(b\) in \(y=ax+b\).

This is core AI SL syllabus content that is also examinable at AI HL.

Key terms

Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.

What is Pearson's correlation coefficient?

Pearson's product-moment correlation coefficient \(r\) measures how strongly two numerical variables are linearly related. It ranges from \(-1\) to \(1\): values near \(\pm1\) show a strong linear relationship, values near \(0\) show a weak or no linear relationship. It's always found using technology.

e.g. For revision hours vs exam score data, a GDC gives \(r = 0.997\) - an almost perfect positive linear relationship.

What's the difference between correlation and causation?

Correlation just measures how strongly two variables move together. Causation means one variable directly brings about the change in the other. A strong \(r\) never proves causation on its own - a hidden third variable can be driving both.

e.g. Sunshine hours and ice cream sales have \(r = 0.999\), but warmer weather (a confounding variable) likely causes both, rather than sunshine directly causing sales.

What is the regression line of y on x?

The regression line of \(y\) on \(x\) is the line \(y = ax+b\) that best predicts \(y\) from a given \(x\), found by technology using the least-squares method. It always passes through the mean point \((\bar{x},\bar{y})\).

e.g. For revision hours \(x\) vs exam score \(y\), the GDC gives \(y = 7.09x + 38.5\).

What is interpolation vs extrapolation?

Interpolation is predicting a value for \(x\) that lies inside the range of the original data - generally safe. Extrapolation is predicting for an \(x\) outside that range, which is much riskier because the linear pattern might not continue.

e.g. With screen-time data from \(1\) to \(6\) hours, \(y \approx -0.503(3.5)+9.03 \approx 7.27\) hours of sleep is interpolation, since \(x=3.5\) lies inside that range.

What is a residual?

A residual is the difference between an actual data value and the value the regression line predicts for it: residual \(= y_{\text{actual}} - y_{\text{predicted}}\). Residuals show how well the line fits an individual point.

e.g. For \((15,25)\) on the line \(y \approx 1.114x+7\), the predicted value is \(\hat{y} \approx 23.71\), so the residual is \(25 - 23.71 = 1.29\).

Key formulas

Everything here is computed with your GDC rather than by hand - the tables below summarise what each piece means and where it comes from.

Formula reference

Pearson's \(r\) has a formula in the booklet, but the regression line itself is found directly by technology - there's no formula for \(a\) and \(b\) to memorise.

FormulaUsed forBooklet?
\(r = \dfrac{\sum(x-\bar{x})(y-\bar{y})}{\sqrt{\sum(x-\bar{x})^2\sum(y-\bar{y})^2}}\)Pearson's correlation coefficient✓ Yes
\(y = ax + b\)Regression line of \(y\) on \(x\)Not in the booklet - found using GDC
\(x = a'y + b'\)Regression line of \(x\) on \(y\)Not in the booklet - found using GDC
residual \(= y_{\text{actual}} - y_{\text{predicted}}\)Measuring fit at one pointNot in the formula booklet - prior knowledge

y on x vs x on y

These are two different lines fitted to the same data - swapping which variable you predict changes the calculation, not just the labels.

Featurey on xx on y
Predicts\(y\) from a given \(x\)\(x\) from a given \(y\)
GDC setup\(x\) as explanatory (Xlist)\(y\) as explanatory (swap lists)
Example\(y = 0.140x - ...\) style form\(x = 0.140y - 5.39\)
Passes through\((\bar{x},\bar{y})\)\((\bar{x},\bar{y})\)

Interpreting r

The sign and size of \(r\) describe the strength and direction of the linear relationship only.

Strong positive

\(r\) close to \(1\)

As \(x\) increases, \(y\) tends to increase in a clear, near-linear pattern.

Strong negative

\(r\) close to \(-1\)

As \(x\) increases, \(y\) tends to decrease in a clear, near-linear pattern.

Weak / no correlation

\(r\) close to \(0\)

Little to no linear relationship - though a strong non-linear pattern could still exist.

Correlation vs causation

These three ideas keep showing up together in exam questions asking you to "comment on" a correlation.

Correlation

Measures association

Just describes how strongly two variables move together - says nothing about why.

Causation

One variable drives the other

A change in \(x\) directly produces a change in \(y\). Never assume this from \(r\) alone.

Confounding variable

The hidden third factor

A variable that independently affects both \(x\) and \(y\), producing a correlation with no direct causal link.

Using the regression line

The safety of a prediction depends entirely on where the \(x\)-value sits relative to the original data.

Interpolation

Predicting inside the data range

Generally reliable, since the model is being applied where it was actually tested.

Extrapolation

Predicting outside the data range

Riskier - the linear trend might not hold, and the further outside the range, the less trustworthy the prediction.

Gradient in context

The rate of change

Always state the gradient's units and meaning in context, e.g. "sales rise by 1.36 hundred per extra hour of sunshine."

Worked examples

Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.

1
Easy
Calculator
[4 marks]

A café records the outdoor temperature \(x\) (\(^\circ\)C) and the number of hot coffees sold \(y\) on 6 days:

\(x\)101520253035
\(y\)20018015013010080

(a)  Find Pearson's product-moment correlation coefficient \(r\).

(b)  Find the equation of the regression line of \(y\) on \(x\).

(c)  Show that the mean point \((\bar{x},\bar{y})\) lies on this regression line.

Worked solution

(a) \(r\approx-0.998.\) A1

(b) M1
\(y\approx-4.91x+251.\) A1

(c) \(\bar{x}=22.5,\ \bar{y}=140.\) Substituting: \(-4.91(22.5)+251=-110.5+251\approx140=\bar{y}.\) A1 AG

🖩 On the GDC.
TI-84 Plus CE: STAT → CALC → LinReg(ax+b), DiagnosticOn for \(r\).
Casio fx-CG50: Statistics → CALC → LinearReg.
TI-Nspire CX: menu → Statistics → Stat Calculations → Linear Regression.

A1 Correct value of r (GDC output) M1 Attempt to run linear regression of y on x A1 Correct equation y=-4.91x+251 A1 Correct mean point substituted AG Matches the stated result
2
Medium
Calculator
[6 marks]

The table shows hours of revision \(x\) and exam score \(y\) (marks out of 100) for six students.

\(x\)123456
\(y\)455261687480

(a)  Use your GDC to find Pearson's \(r\).

(b)  Find the regression line of \(y\) on \(x\) and interpret the gradient in context.

(c)  Find the regression line of \(x\) on \(y\).

Worked solution

(a)   \(r = 0.997\) A1

(b)   Using GDC: \(y\) M1
\(= 7.09x + 38.5\) A1

The gradient 7.09 means that for each additional hour of revision, the predicted exam score increases by approximately 7.09 marks. A1

(c)   Swap lists on GDC (enter \(y\) as explanatory): \(x\) M1
\(= 0.140y - 5.39\) A1

A1 Correct value of Pearson's r, 0.997 (GDC output) M1 Attempt to run linear regression of y on x A1 Correct regression equation \(y=7.09x+38.5\) A1 Correct contextual interpretation of the gradient M1 Attempt to swap the explanatory and response variables for x on y A1 Correct regression equation \(x=0.140y-5.39\)

Common mistakes

The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.

  • Entering the lists the wrong way round for x on y. To fit \(x\) on \(y\), \(y\) must be entered as the explanatory list on the GDC - simply swapping the letters in the answer without re-running the regression gives the wrong line.
  • Extrapolating without comment. Using the regression line for an \(x\)-value well outside the data range without flagging that this is extrapolation and may be unreliable loses marks even if the arithmetic is correct.
  • Treating a strong r as proof of causation. A high \(|r|\) only shows a strong linear association - concluding one variable "causes" the other without considering confounding variables is a common reasoning error.
  • Rounding the equation before predicting. Substituting the rounded 3 s.f. gradient and intercept into a further calculation compounds rounding error - use the GDC's full-precision values for predictions, then round only the final answer.

Using your GDC

Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.

Show steps for:
Linear regression (line of best fit) and r

Core AI skill - gives the regression line and the correlation coefficient in one step.

  1. Enter x-values in one list and y-values in another.
  2. Turn on DiagnosticOn once (2nd → CATALOG) so r appears; then STAT → CALC → 4:LinReg(ax+b), set Xlist and Ylist, Calculate.TI-84
  3. Calculator page → menu → Statistics → Stat Calculations → Linear Regression (mx+b).Nspire
  4. Statistics menu → CALC (F2) → REG (F3) → X (linear); r shows automatically.Casio
  5. Read a (gradient), b (intercept), r (correlation) and r² (coefficient of determination).
  6. Use the equation to predict - but only within the data range (interpolation).

Tip: r near ±1 means a strong linear fit; near 0 means weak. r² is the proportion of variation explained.

One-variable statistics (mean, median, standard deviation)

Instant summary statistics from a list - useful for finding the mean point \((\bar{x},\bar{y})\) that the regression line always passes through.

  1. Enter the data into a list.
  2. STAT → Edit → type values into L1. Then STAT → CALC → 1:1-Var Stats, choose L1, Calculate.TI-84
  3. Add a Lists & Spreadsheet page, name a column and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
  4. Statistics menu → enter data in List 1 → CALC (F2) → 1-VAR.Casio
  5. Read x̄ (mean), Sx (sample sd) or σx (population sd), and the five-number summary.

Tip: Run 1-Var Stats on both the x-list and the y-list to get \(\bar{x}\) and \(\bar{y}\) separately when verifying the mean point lies on the regression line.

See the full GDC guide for more calculator models and topics.

Ready to practise properly?

Correlation & regression questions, marked instantly like the real exam.

Quick answers

The questions students on this topic ask most often.

How do I find Pearson's r on my GDC?

Enter the x-values and y-values into two lists, then run the linear regression command (LinReg on a TI-84, Linear Regression on an Nspire, REG on a Casio). Make sure diagnostics/r is switched on so the calculator displays r alongside the equation.

What's the difference between the regression line of y on x and x on y?

The y on x line predicts y from a given x and is the one you use by default. The x on y line predicts x from a given y and is a different line - you get it by swapping which list is explanatory and which is response before running the regression again.

Can I predict y for any x using my regression line?

Only safely within the range of the original data (interpolation). Using the line for an x well outside that range is extrapolation, and the prediction becomes increasingly unreliable the further you go.

Does a strong correlation mean one variable causes the other?

No. A high value of r only shows a strong linear association - it never proves that one variable causes the other. A third, confounding variable can drive both at once, so correlation never implies causation on its own.