Calculating Correlation & Regression (AA HL)
When two variables might be linked, a scatter diagram shows the pattern and Pearson's correlation coefficient \(r\) puts a number on how strong and in which direction that linear pattern runs. This topic covers finding \(r\) and \(r^2\) on your GDC, fitting the regression lines of \(y\) on \(x\) and of \(x\) on \(y\), and using those lines carefully to predict one variable from the other - without falling into the traps of extrapolation or mistaking correlation for causation.
What the syllabus says
This topic maps onto two points in the official IB Analysis & Approaches syllabus.
| Code | Syllabus content |
|---|---|
| SL4.4 | Linear correlation of bivariate data - technology should be used to calculate Pearson's product-moment correlation coefficient \(r\), although hand calculation may enhance understanding. Scatter diagrams; lines of best fit passing through the mean point. Distinction between correlation and causation - correlation does not imply causation. Equation of the regression line of \(y\) on \(x\), found using technology, used for prediction, interpreting the parameters \(a\) and \(b\) in \(y=ax+b\), and awareness of the dangers of extrapolation. |
| SL4.10 | Equation of the regression line of \(x\) on \(y\), and its use for prediction - with awareness that this line cannot always reliably be used to predict \(y\) from a given \(x\) value. |
These are core AA SL syllabus points that are also examinable at AA HL.
Key terms
Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.
What is Pearson's correlation coefficient \(r\)?
\(r\) is a number between \(-1\) and \(1\) that measures how closely a set of bivariate data follows a straight line. A value near \(1\) means a strong positive linear relationship, near \(-1\) means a strong negative one, and near \(0\) means little or no linear relationship. It says nothing about non-linear patterns.
e.g. For \((1,2), (2,4), (3,6), (4,8)\) every point lies exactly on \(y=2x\), so \(r = 1\).
What is the regression line of \(y\) on \(x\)?
The regression line of \(y\) on \(x\) is the straight line \(y=ax+b\) that best fits the data by minimising the vertical distances from each point to the line. It always passes through the mean point \((\bar x, \bar y)\), and is the line used when predicting \(y\) from a given value of \(x\).
e.g. If \(\bar x = 5\), \(\bar y = 20\) and the gradient is \(2\), then \(b = \bar y - a\bar x = 20 - 2(5) = 10\), giving \(y = 2x+10\).
What is the regression line of \(x\) on \(y\)?
The regression line of \(x\) on \(y\) is a different straight line, found by minimising the horizontal distances instead of the vertical ones. It's used only for predicting \(x\) from a given value of \(y\) - swapping the \(y\) on \(x\) line around algebraically does not give the same line.
e.g. If the \(y\) on \(x\) gradient is \(2\) and \(r=0.9\), the \(x\) on \(y\) gradient is \(r^2/2 = 0.81/2 = 0.405\), not \(1/2\).
What is \(r^2\)?
\(r^2\), the coefficient of determination, is the square of the correlation coefficient. It gives the proportion of the variation in one variable that can be explained by a linear relationship with the other, usually expressed as a percentage.
e.g. If \(r = 0.95\), then \(r^2 = 0.9025\), so \(90.25\%\) of the variation in \(y\) is explained by \(x\).
What's the difference between interpolation and extrapolation?
Interpolation is using a regression line to predict a value that falls within the range of the original data, which is generally reliable. Extrapolation is predicting outside that range, which is risky because there's no evidence the linear pattern continues beyond the data collected.
e.g. If \(x\) ranges from \(10\) to \(50\) in the data, predicting at \(x=30\) is interpolation; predicting at \(x=100\) is extrapolation.
Key formulas
A handful of formulas cover this topic, and every one of them is normally evaluated on your GDC rather than substituted into by hand. The tables below summarise them - the explanations underneath go into more depth.
Formula reference
Pearson's \(r\) and both regression line equations are given in the formula booklet - the coefficient of determination \(r^2\) is not listed separately since it's simply \(r\) squared.
| Formula | Used for | Booklet? |
|---|---|---|
| \(r = \dfrac{S_{xy}}{\sqrt{S_{xx}S_{yy}}}\) | Pearson's product-moment correlation coefficient | ✓ Yes |
| \(y-\bar y = \dfrac{S_{xy}}{S_{xx}}(x-\bar x)\) | Regression line of \(y\) on \(x\) | ✓ Yes |
| \(x-\bar x = \dfrac{S_{xy}}{S_{yy}}(y-\bar y)\) | Regression line of \(x\) on \(y\) | ✓ Yes |
| \(r^2\) | Coefficient of determination | Not in booklet - square the value of r |
Regression line of y on x vs x on y
These are two genuinely different lines, and mixing them up is one of the most common ways marks are lost on this topic.
| Feature | Line of y on x | Line of x on y |
|---|---|---|
| Minimises | vertical distances to the line | horizontal distances to the line |
| Form | \(y=ax+b\) | \(x=cy+d\) |
| Use it to predict | \(y\) from a given \(x\) | \(x\) from a given \(y\) |
| Passes through | \((\bar x,\bar y)\) | \((\bar x,\bar y)\) |
| Gradient relationship | product of the two gradients \(=r^2\) | |
Correlation and its strength
Sign of r
\[-1 \le r \le 1\]
Positive \(r\): as \(x\) increases, \(y\) tends to increase. Negative \(r\): as \(x\) increases, \(y\) tends to decrease.
✓ In the formula bookletStrength of r
\[|r|\text{ near }1 = \text{strong}\]
\(|r|\) close to \(1\) means a strong linear relationship; \(|r|\) close to \(0\) means a weak or absent one.
Not in the formula booklet - interpretation onlyCorrelation vs causation
\[r \ne \text{proof of cause}\]
A strong \(r\) shows association, not that one variable causes the other - a third, confounding variable could explain both.
Not in the formula booklet - conceptual pointUsing the regression line for prediction
Interpolation
\[x_{\min} \le x \le x_{\max}\]
Predicting within the range of the collected data - generally reliable, since the line is supported by nearby data.
Not in the formula booklet - definitionExtrapolation
\[x < x_{\min}\text{ or }x > x_{\max}\]
Predicting outside the data range - the linear pattern is not guaranteed to continue, so the prediction may be unreliable.
Not in the formula booklet - definitionInterpreting the gradient
\[y = ax+b\]
In context, \(a\) is the change in \(y\) per unit increase in \(x\), and \(b\) is the value of \(y\) predicted when \(x=0\) (if that makes sense in context).
Not in the formula booklet - interpretationWorked examples
Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.
A farmer records fertiliser applied \(x\) (kg/hectare) and crop yield \(y\) (tonnes/hectare) for 6 plots:
| \(x\) | 20 | 35 | 50 | 65 | 80 | 95 |
|---|---|---|---|---|---|---|
| \(y\) | 3.1 | 3.9 | 4.6 | 5.4 | 6.0 | 6.7 |
(a) Calculate the Pearson's product-moment correlation coefficient \(r,\) and interpret its strength and direction.
(b) Find the equation of the regression line of \(y\) on \(x.\)
(c) Predict the yield for \(70\) kg/hectare of fertiliser.
Worked solution
(a) \(r \approx 0.999,\) a very strong positive linear correlation. M1
A1
R1
(b) \(y \approx 0.0478x + 2.20.\) M1 A1
(c) \(y \approx 0.0478(70) + 2.20\) M1
\(\approx 5.55\) tonnes/hectare. A1
Data on \(x\) and \(y\) gives the regression line \(y = 3x + 1\) with \(r = 0.95\).
New variables are defined: \(u = 2x + 1\) and \(v = y - 4\).
(a) Express the regression line of \(v\) on \(u\) in the form \(v = au + b\).
(b) State the value of the correlation coefficient between \(u\) and \(v\).
(c) State the gradient of the regression line of \(u\) on \(v\).
Worked solution
(a) \(v = y-4 = (3x+1)-4 = 3x-3\) M1 Substituting \(y=3x+1\) into \(v=y-4\) to express \(v\) in terms of \(x\). Since \(x=(u-1)/2\), \(v = 3\cdot\dfrac{u-1}{2}-3 = 1.5u-1.5-3\) M1 Substitute \(x\) in terms of \(u\). \(=1.5u-4.5\) A1 Correct simplified equation \(v=1.5u-4.5\).
(b) Linear transformations do not change \(|r|\), so \(r = 0.95\). A1 Correct value \(r=0.95\), stated directly from the fact that linear transformations preserve \(|r|\).
(c) \(a_2 = r^2/a_1 = 0.9025/1.5 = 0.602\) (3 significant figures) A1 Correct value \(\approx0.602\) for the gradient of \(u\) on \(v\), using \(a_2=r^2/a_1\).
Common mistakes
The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.
- Confusing the sign of \(r\) with its strength. \(r=-0.9\) is a stronger relationship than \(r=0.3\) - the sign only tells you the direction, the magnitude \(|r|\) tells you the strength.
- Using the \(y\) on \(x\) line to predict \(x\) from \(y\). Rearranging \(y=ax+b\) algebraically does not give the true \(x\) on \(y\) regression line - you need to run the regression the other way on your GDC.
- Extrapolating with confidence. A prediction far outside the range of \(x\)-values in the data is unreliable, even with a strong \(r\) - always check whether you're interpolating or extrapolating before trusting the answer.
- Treating a strong \(r\) as proof of causation. A high correlation shows association, not that one variable causes the other - a confounding factor could be driving both.
Using your GDC
Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.
Core AA skill - gives the regression line and the correlation coefficient in one step.
- Enter x-values in one list and y-values in another.
- Turn on DiagnosticOn once (2nd → CATALOG) so r appears; then STAT → CALC → 4:LinReg(ax+b), set Xlist and Ylist, Calculate.TI-84
- Calculator page → menu → Statistics → Stat Calculations → Linear Regression (mx+b).Nspire
- Statistics menu → CALC (F2) → REG (F3) → X (linear); r shows automatically.Casio
- Read a (gradient), b (intercept), r (correlation) and r² (coefficient of determination).
- Use the equation to predict - but only within the data range (interpolation).
Tip: r near ±1 means a strong linear fit; near 0 means weak. r² is the proportion of variation explained.
Instant summary statistics from a list - no formulas to compute by hand. Useful for finding \(\bar x\) and \(\bar y\) before working with the regression line.
- Enter the data into a list.
- STAT → Edit → type values into L1. Then STAT → CALC → 1:1-Var Stats, choose L1, Calculate.TI-84
- Add a Lists & Spreadsheet page, name a column and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
- Statistics menu → enter data in List 1 → CALC (F2) → 1-VAR.Casio
- Read x̄ (mean), Sx (sample sd) or σx (population sd), and the five-number summary (min, Q1, median, Q3, max).
- For frequency data, put values in one list and frequencies in another and set the frequency list.
Tip: Sx vs σx: use σx (population) for a complete data set, Sx (sample) for a sample. IB usually wants σx.
See the full GDC guide for more calculator models and topics.
Ready to practise properly?
Correlation & regression questions, marked instantly like the real exam.
Quick answers
The questions students on this topic ask most often.
What does a correlation coefficient close to 1 or -1 mean?
A value of \(r\) close to \(1\) or \(-1\) means the points lie close to a straight line - close to \(1\) is a strong positive relationship (as \(x\) increases, \(y\) increases), close to \(-1\) is a strong negative relationship (as \(x\) increases, \(y\) decreases). A value near \(0\) means little or no linear relationship.
What's the difference between the regression line of y on x and the line of x on y?
The line of \(y\) on \(x\) minimises the vertical distances from the points to the line, and is used to predict \(y\) from a given \(x\). The line of \(x\) on \(y\) minimises the horizontal distances instead, and is used to predict \(x\) from a given \(y\). They are two different lines unless the correlation is perfect.
Why can't I use the y on x line to predict x from a y value?
The \(y\) on \(x\) line was built to minimise vertical errors when predicting \(y\), not horizontal errors when predicting \(x\) - rearranging it algebraically gives a different, less accurate line than the true \(x\) on \(y\) regression. Use the \(x\) on \(y\) line whenever you're predicting \(x\) from \(y\).
Is Pearson's r formula given in the exam formula booklet?
Yes - the formula for \(r\), and the equations of both regression lines, are in the topic 4 section of the AA formula booklet. In practice you'll almost always find \(r\) and the regression line directly from your GDC rather than substituting into the formula by hand.
Related topics
More Statistics & Probability topics from the same AA HL syllabus unit, in case you want to keep going.