Calculating Correlation & Regression (AA HL)

When two variables might be linked, a scatter diagram shows the pattern and Pearson's correlation coefficient \(r\) puts a number on how strong and in which direction that linear pattern runs. This topic covers finding \(r\) and \(r^2\) on your GDC, fitting the regression lines of \(y\) on \(x\) and of \(x\) on \(y\), and using those lines carefully to predict one variable from the other - without falling into the traps of extrapolation or mistaking correlation for causation.

What the syllabus says

This topic maps onto two points in the official IB Analysis & Approaches syllabus.

CodeSyllabus content
SL4.4Linear correlation of bivariate data - technology should be used to calculate Pearson's product-moment correlation coefficient \(r\), although hand calculation may enhance understanding. Scatter diagrams; lines of best fit passing through the mean point. Distinction between correlation and causation - correlation does not imply causation. Equation of the regression line of \(y\) on \(x\), found using technology, used for prediction, interpreting the parameters \(a\) and \(b\) in \(y=ax+b\), and awareness of the dangers of extrapolation.
SL4.10Equation of the regression line of \(x\) on \(y\), and its use for prediction - with awareness that this line cannot always reliably be used to predict \(y\) from a given \(x\) value.

These are core AA SL syllabus points that are also examinable at AA HL.

Key terms

Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.

What is Pearson's correlation coefficient \(r\)?

\(r\) is a number between \(-1\) and \(1\) that measures how closely a set of bivariate data follows a straight line. A value near \(1\) means a strong positive linear relationship, near \(-1\) means a strong negative one, and near \(0\) means little or no linear relationship. It says nothing about non-linear patterns.

e.g. For \((1,2), (2,4), (3,6), (4,8)\) every point lies exactly on \(y=2x\), so \(r = 1\).

What is the regression line of \(y\) on \(x\)?

The regression line of \(y\) on \(x\) is the straight line \(y=ax+b\) that best fits the data by minimising the vertical distances from each point to the line. It always passes through the mean point \((\bar x, \bar y)\), and is the line used when predicting \(y\) from a given value of \(x\).

e.g. If \(\bar x = 5\), \(\bar y = 20\) and the gradient is \(2\), then \(b = \bar y - a\bar x = 20 - 2(5) = 10\), giving \(y = 2x+10\).

What is the regression line of \(x\) on \(y\)?

The regression line of \(x\) on \(y\) is a different straight line, found by minimising the horizontal distances instead of the vertical ones. It's used only for predicting \(x\) from a given value of \(y\) - swapping the \(y\) on \(x\) line around algebraically does not give the same line.

e.g. If the \(y\) on \(x\) gradient is \(2\) and \(r=0.9\), the \(x\) on \(y\) gradient is \(r^2/2 = 0.81/2 = 0.405\), not \(1/2\).

What is \(r^2\)?

\(r^2\), the coefficient of determination, is the square of the correlation coefficient. It gives the proportion of the variation in one variable that can be explained by a linear relationship with the other, usually expressed as a percentage.

e.g. If \(r = 0.95\), then \(r^2 = 0.9025\), so \(90.25\%\) of the variation in \(y\) is explained by \(x\).

What's the difference between interpolation and extrapolation?

Interpolation is using a regression line to predict a value that falls within the range of the original data, which is generally reliable. Extrapolation is predicting outside that range, which is risky because there's no evidence the linear pattern continues beyond the data collected.

e.g. If \(x\) ranges from \(10\) to \(50\) in the data, predicting at \(x=30\) is interpolation; predicting at \(x=100\) is extrapolation.

Key formulas

A handful of formulas cover this topic, and every one of them is normally evaluated on your GDC rather than substituted into by hand. The tables below summarise them - the explanations underneath go into more depth.

Formula reference

Pearson's \(r\) and both regression line equations are given in the formula booklet - the coefficient of determination \(r^2\) is not listed separately since it's simply \(r\) squared.

FormulaUsed forBooklet?
\(r = \dfrac{S_{xy}}{\sqrt{S_{xx}S_{yy}}}\)Pearson's product-moment correlation coefficient✓ Yes
\(y-\bar y = \dfrac{S_{xy}}{S_{xx}}(x-\bar x)\)Regression line of \(y\) on \(x\)✓ Yes
\(x-\bar x = \dfrac{S_{xy}}{S_{yy}}(y-\bar y)\)Regression line of \(x\) on \(y\)✓ Yes
\(r^2\)Coefficient of determinationNot in booklet - square the value of r

Regression line of y on x vs x on y

These are two genuinely different lines, and mixing them up is one of the most common ways marks are lost on this topic.

FeatureLine of y on xLine of x on y
Minimisesvertical distances to the linehorizontal distances to the line
Form\(y=ax+b\)\(x=cy+d\)
Use it to predict\(y\) from a given \(x\)\(x\) from a given \(y\)
Passes through\((\bar x,\bar y)\)\((\bar x,\bar y)\)
Gradient relationshipproduct of the two gradients \(=r^2\)

Correlation and its strength

Sign of r

\[-1 \le r \le 1\]

Positive \(r\): as \(x\) increases, \(y\) tends to increase. Negative \(r\): as \(x\) increases, \(y\) tends to decrease.

✓ In the formula booklet

Strength of r

\[|r|\text{ near }1 = \text{strong}\]

\(|r|\) close to \(1\) means a strong linear relationship; \(|r|\) close to \(0\) means a weak or absent one.

Not in the formula booklet - interpretation only

Correlation vs causation

\[r \ne \text{proof of cause}\]

A strong \(r\) shows association, not that one variable causes the other - a third, confounding variable could explain both.

Not in the formula booklet - conceptual point

Using the regression line for prediction

Interpolation

\[x_{\min} \le x \le x_{\max}\]

Predicting within the range of the collected data - generally reliable, since the line is supported by nearby data.

Not in the formula booklet - definition

Extrapolation

\[x < x_{\min}\text{ or }x > x_{\max}\]

Predicting outside the data range - the linear pattern is not guaranteed to continue, so the prediction may be unreliable.

Not in the formula booklet - definition

Interpreting the gradient

\[y = ax+b\]

In context, \(a\) is the change in \(y\) per unit increase in \(x\), and \(b\) is the value of \(y\) predicted when \(x=0\) (if that makes sense in context).

Not in the formula booklet - interpretation

Worked examples

Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.

1
Medium
Calc
[7 marks]

A farmer records fertiliser applied \(x\) (kg/hectare) and crop yield \(y\) (tonnes/hectare) for 6 plots:

\(x\)203550658095
\(y\)3.13.94.65.46.06.7

(a) Calculate the Pearson's product-moment correlation coefficient \(r,\) and interpret its strength and direction.

(b) Find the equation of the regression line of \(y\) on \(x.\)

(c) Predict the yield for \(70\) kg/hectare of fertiliser.

Worked solution

(a) \(r \approx 0.999,\) a very strong positive linear correlation. M1
A1
R1

(b) \(y \approx 0.0478x + 2.20.\) M1 A1

(c) \(y \approx 0.0478(70) + 2.20\) M1
\(\approx 5.55\) tonnes/hectare. A1

M1 Attempt at computing r using the GDC A1 Correct value \(r\approx0.999\) R1 Reasoning identifying the correlation as very strong and positive M1 Attempt at linear regression of y on x using the GDC A1 Correct equation \(y\approx0.0478x+2.20\) M1 Substituting x=70 into the regression equation from (b) A1 Correct value \(\approx5.55\) tonnes/hectare
2
Hard
No calc
[5 marks]

Data on \(x\) and \(y\) gives the regression line \(y = 3x + 1\) with \(r = 0.95\).
New variables are defined: \(u = 2x + 1\) and \(v = y - 4\).

(a) Express the regression line of \(v\) on \(u\) in the form \(v = au + b\).

(b) State the value of the correlation coefficient between \(u\) and \(v\).

(c) State the gradient of the regression line of \(u\) on \(v\).

Worked solution

(a) \(v = y-4 = (3x+1)-4 = 3x-3\) M1 Substituting \(y=3x+1\) into \(v=y-4\) to express \(v\) in terms of \(x\). Since \(x=(u-1)/2\), \(v = 3\cdot\dfrac{u-1}{2}-3 = 1.5u-1.5-3\) M1 Substitute \(x\) in terms of \(u\). \(=1.5u-4.5\) A1 Correct simplified equation \(v=1.5u-4.5\).

(b) Linear transformations do not change \(|r|\), so \(r = 0.95\). A1 Correct value \(r=0.95\), stated directly from the fact that linear transformations preserve \(|r|\).

(c) \(a_2 = r^2/a_1 = 0.9025/1.5 = 0.602\) (3 significant figures) A1 Correct value \(\approx0.602\) for the gradient of \(u\) on \(v\), using \(a_2=r^2/a_1\).

M1 Method - substituting the transformation into the given regression line A1 Accuracy - correct equation or numerical value

Common mistakes

The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.

  • Confusing the sign of \(r\) with its strength. \(r=-0.9\) is a stronger relationship than \(r=0.3\) - the sign only tells you the direction, the magnitude \(|r|\) tells you the strength.
  • Using the \(y\) on \(x\) line to predict \(x\) from \(y\). Rearranging \(y=ax+b\) algebraically does not give the true \(x\) on \(y\) regression line - you need to run the regression the other way on your GDC.
  • Extrapolating with confidence. A prediction far outside the range of \(x\)-values in the data is unreliable, even with a strong \(r\) - always check whether you're interpolating or extrapolating before trusting the answer.
  • Treating a strong \(r\) as proof of causation. A high correlation shows association, not that one variable causes the other - a confounding factor could be driving both.

Using your GDC

Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.

Show steps for:
Linear regression (line of best fit) and r

Core AA skill - gives the regression line and the correlation coefficient in one step.

  1. Enter x-values in one list and y-values in another.
  2. Turn on DiagnosticOn once (2nd → CATALOG) so r appears; then STAT → CALC → 4:LinReg(ax+b), set Xlist and Ylist, Calculate.TI-84
  3. Calculator page → menu → Statistics → Stat Calculations → Linear Regression (mx+b).Nspire
  4. Statistics menu → CALC (F2) → REG (F3) → X (linear); r shows automatically.Casio
  5. Read a (gradient), b (intercept), r (correlation) and r² (coefficient of determination).
  6. Use the equation to predict - but only within the data range (interpolation).

Tip: r near ±1 means a strong linear fit; near 0 means weak. r² is the proportion of variation explained.

One-variable statistics (mean, median, standard deviation)

Instant summary statistics from a list - no formulas to compute by hand. Useful for finding \(\bar x\) and \(\bar y\) before working with the regression line.

  1. Enter the data into a list.
  2. STAT → Edit → type values into L1. Then STAT → CALC → 1:1-Var Stats, choose L1, Calculate.TI-84
  3. Add a Lists & Spreadsheet page, name a column and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
  4. Statistics menu → enter data in List 1 → CALC (F2) → 1-VAR.Casio
  5. Read x̄ (mean), Sx (sample sd) or σx (population sd), and the five-number summary (min, Q1, median, Q3, max).
  6. For frequency data, put values in one list and frequencies in another and set the frequency list.

Tip: Sx vs σx: use σx (population) for a complete data set, Sx (sample) for a sample. IB usually wants σx.

See the full GDC guide for more calculator models and topics.

Ready to practise properly?

Correlation & regression questions, marked instantly like the real exam.

Quick answers

The questions students on this topic ask most often.

What does a correlation coefficient close to 1 or -1 mean?

A value of \(r\) close to \(1\) or \(-1\) means the points lie close to a straight line - close to \(1\) is a strong positive relationship (as \(x\) increases, \(y\) increases), close to \(-1\) is a strong negative relationship (as \(x\) increases, \(y\) decreases). A value near \(0\) means little or no linear relationship.

What's the difference between the regression line of y on x and the line of x on y?

The line of \(y\) on \(x\) minimises the vertical distances from the points to the line, and is used to predict \(y\) from a given \(x\). The line of \(x\) on \(y\) minimises the horizontal distances instead, and is used to predict \(x\) from a given \(y\). They are two different lines unless the correlation is perfect.

Why can't I use the y on x line to predict x from a y value?

The \(y\) on \(x\) line was built to minimise vertical errors when predicting \(y\), not horizontal errors when predicting \(x\) - rearranging it algebraically gives a different, less accurate line than the true \(x\) on \(y\) regression. Use the \(x\) on \(y\) line whenever you're predicting \(x\) from \(y\).

Is Pearson's r formula given in the exam formula booklet?

Yes - the formula for \(r\), and the equations of both regression lines, are in the topic 4 section of the AA formula booklet. In practice you'll almost always find \(r\) and the regression line directly from your GDC rather than substituting into the formula by hand.