Calculating Correlation & Regression (AA SL)
When you have two connected sets of data, correlation tells you how strongly they move together, and regression gives you the actual line for predicting one from the other. This topic covers Pearson's correlation coefficient \(r\), fitting and interpreting the regression line of \(y\) on \(x\), using that line to make predictions, and knowing when a prediction can and can't be trusted.
What the syllabus says
This topic maps onto one main point in the official IB Analysis & Approaches syllabus, with a related SL point on regression of x on y.
| Code | Syllabus content |
|---|---|
| SL4.4 | Linear correlation of bivariate data. Pearson's product-moment correlation coefficient \(r\) (found using technology). Scatter diagrams and lines of best fit by eye, passing through the mean point. Equation of the regression line of \(y\) on \(x\) (found using technology), used for prediction, interpreting the parameters \(a\) and \(b\) in \(y=ax+b\), being aware of the dangers of extrapolation and that a \(y\) on \(x\) line cannot always reliably predict \(x\) from a value of \(y\). |
| SL4.10 | Equation of the regression line of \(x\) on \(y\), and its use for prediction. Students should be aware that this cannot always reliably predict \(y\) from a value of \(x\). |
These are core AA SL syllabus points, also examinable at AA HL.
Key terms
Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.
What is Pearson's r?
Pearson's product-moment correlation coefficient \(r\) is a number between \(-1\) and \(1\) that measures how strongly two variables follow a straight-line pattern. Values near \(\pm1\) mean a strong linear relationship; values near 0 mean little or no linear relationship. Your GDC calculates it directly from the data.
e.g. \(r = 0.999\) for data on hours of sunlight against plant height means an almost perfect positive linear relationship.
What is the regression line of y on x?
The regression line of \(y\) on \(x\) is the "line of best fit" that minimises the total squared vertical distance between the line and each data point. It's written \(y = ax+b\), and it's the line you use whenever you're predicting \(y\) from a given \(x\).
e.g. For temperature \(x\) and sales \(y\), \(y = 10.2x - 76.4\) predicts sales for any given temperature within the data range.
What is the mean point?
The mean point \((\bar{x}, \bar{y})\) is the point formed from the mean of the \(x\)-values and the mean of the \(y\)-values. Every regression line of \(y\) on \(x\) passes exactly through this point, which makes it a useful check on a calculated regression equation.
e.g. For \(x: 3,4,5,6,7,8\) and \(y: 8,11,14,17,20,24\), \(\bar{x} = 33/6 = 5.5\) and \(\bar{y} = 94/6 = 15.67\).
What is extrapolation?
Extrapolation is using a regression line to predict a value of \(y\) for an \(x\) that lies outside the range of the original data. This is unreliable because there's no evidence the same linear relationship continues beyond the data actually collected.
e.g. If \(x\) ranges from \(3\) to \(17\) in the data, predicting at \(x=50\) is extrapolation and the result should be treated with caution.
What does the gradient of a regression line mean?
The gradient \(a\) in \(y=ax+b\) tells you, in context, how much \(y\) changes for every one-unit increase in \(x\). Interpreting it correctly means stating both the size of the change and the units of \(x\) and \(y\) from the original question.
e.g. For \(y = 2.91x + 3.41\) with \(x\) in $100s of marketing spend and \(y\) in new customers, each extra $100 spent is associated with about 3 more new customers.
Key formulas
This topic is built around two GDC outputs - and knowing exactly what each one is used for. The tables below summarise everything at a glance - the explanations underneath go into more depth.
Formula reference
Neither Pearson's \(r\) nor the regression equation needs to be memorised as a hand formula - both are found using technology, so the syllabus lists them as concepts to interpret rather than formulas to recall.
| Formula | Used for | Booklet? |
|---|---|---|
| \(r\) (Pearson's coefficient) | Strength and direction of linear correlation | Not in booklet - found by GDC |
| \(y = ax+b\) | Regression line of \(y\) on \(x\), for predicting \(y\) | Not in booklet - found by GDC |
| \(x = cy+d\) | Regression line of \(x\) on \(y\), for predicting \(x\) | Not in booklet - found by GDC |
| Mean point \((\bar{x}, \bar{y})\) | The point every \(y\) on \(x\) regression line passes through | Not in booklet - prior knowledge |
y on x vs x on y
These are two different lines that both pass through the mean point but are otherwise not the same - picking the wrong one is a common source of lost marks.
| Feature | Regression of y on x | Regression of x on y |
|---|---|---|
| Form | \(y = ax+b\) | \(x = cy+d\) |
| Minimises errors in | the \(y\)-direction | the \(x\)-direction |
| Use to predict | \(y\) from a given \(x\) | \(x\) from a given \(y\) |
| Example | \(y=3.14x-1.62\) | \(x=-0.773y+80.0\) |
Correlation and its interpretation
Reading r
\(-1 \le r \le 1\)
Closer to \(\pm1\) means a stronger linear relationship; closer to 0 means weaker.
Not in the formula booklet - found by GDCSign of r
Positive \(r\): as \(x\) increases, \(y\) tends to increase. Negative \(r\): as \(x\) increases, \(y\) tends to decrease.
The sign of \(r\) always matches the sign of the regression gradient \(a\).
Not in the formula booklet - prior knowledgeCorrelation vs causation
A strong \(r\) shows association, not cause and effect.
Never claim that \(x\) causes \(y\) just because \(r\) is close to \(\pm1\).
Not in the formula booklet - key exam ideaUsing the regression line for prediction
Interpolation
Predicting within the range of the given \(x\)-data.
Generally reliable, since the model is only tested inside this range.
Not in the formula booklet - key exam ideaExtrapolation
Predicting outside the range of the given data.
Unreliable - always flag this in your answer if the value of \(x\) falls outside the data range.
Not in the formula booklet - key exam ideaInterpreting the gradient
\(a\) in \(y=ax+b\) is the change in \(y\) per unit increase in \(x\).
Always state the interpretation using the actual units and context from the question.
Not in the formula booklet - key exam ideaWorked examples
Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.
Data on hours of sunlight \(x\) and plant height \(y\) (cm) after 4 weeks:
| \(x\) | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|
| \(y\) | 8 | 11 | 14 | 17 | 20 | 24 |
(a)(i) Calculate \(\bar{x}.\)
(a)(ii) Calculate \(\bar{y}.\)
(b) Find the regression line of \(y\) on \(x\) using your GDC.
(c) Verify algebraically that \((\bar{x}, \bar{y})\) lies on your regression line.
(d) Find \(r\) and describe the correlation.
Worked solution
(a) \(\bar{x} = 33/6 = 5.5,\quad \bar{y} = 94/6 = 15.67\) A1 A1
(b)
\(y = 3.14x - 1.62\) M1
\(y = 3.14x - 1.62\) A1
(c) \(3.14(5.5)-1.62 = 17.27-1.62 = 15.65\approx\bar{y}=15.67\) ✓ (small rounding difference from using the 3 s.f. line) A1 AG
(d)
\(r = 0.999\) A1
very strong positive correlation. A1
A teacher records the number of hours \(x\) spent revising and the test score \(y\) (%) for 6 students:
| \(x\) | 2 | 4 | 6 | 8 | 10 | 12 |
|---|---|---|---|---|---|---|
| \(y\) | 45 | 52 | 61 | 68 | 75 | 85 |
(a) State which variable is the explanatory (independent) variable and which is the response (dependent) variable.
(b)(i) Find the value of \(r.\)
(b)(ii) Find the equation of the regression line of \(y\) on \(x.\)
(c) A student revises for 9 hours. Predict their test score.
(d) A different student scores 95%. Explain why the regression line found in (b)(ii) should not be used to estimate how many hours they revised.
Worked solution
(a) Hours revising \((x)\) is the explanatory variable and test score \((y)\) is the response variable. A1
(b)(i) \(r \approx 0.999\) M1 A1
(b)(ii) \(y \approx 3.94x+36.7\) M1 A1
(c) \(y \approx 3.9429(9)+36.733\) M1
\(\approx 72.2\%.\) A1
(d) The regression line of \(y\) on \(x\) is constructed to predict \(y\) from a given \(x,\) minimizing errors in the \(y\)-direction; R1 using it in reverse to estimate \(x\) from a given \(y\) does not give a reliable estimate - a separate regression line of \(x\) on \(y\) would be needed. A1
Common mistakes
The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.
- Using the y on x line to predict x. The regression line of \(y\) on \(x\) is designed to predict \(y\) from a given \(x\) - to predict \(x\) from a given \(y\), you need the separate regression line of \(x\) on \(y\).
- Ignoring extrapolation. Predicting a value far outside the range of the collected data assumes the linear trend continues, which is often not a safe assumption - always check the \(x\)-value against the data range before trusting a prediction.
- Assuming a high r means a linear model is appropriate. \(r\) only measures linear association and can still be high for a clearly curved relationship - always check the scatter diagram, not just the value of \(r\).
- Treating correlation as causation. A strong \(r\) shows two variables move together, not that one causes the other - other factors could explain the relationship.
Using your GDC
Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.
Core skill for this topic - gives the regression line and the correlation coefficient in one step.
- Enter \(x\)-values in one list and \(y\)-values in another.
- Turn on DiagnosticOn once (2nd → CATALOG) so \(r\) appears; then STAT → CALC → 4:LinReg(ax+b), set Xlist and Ylist, Calculate.TI-84
- Calculator page → menu → Statistics → Stat Calculations → Linear Regression (mx+b).Nspire
- Statistics menu → CALC (F2) → REG (F3) → X (linear); \(r\) shows automatically.Casio
- Read \(a\) (gradient), \(b\) (intercept), \(r\) (correlation) and \(r^2\) (coefficient of determination).
- Use the equation to predict - but only within the data range (interpolation).
Tip: \(r\) near \(\pm1\) means a strong linear fit; near 0 means weak. \(r^2\) is the proportion of variation explained.
Used to find the mean point \((\bar{x}, \bar{y})\) quickly, for verifying a regression line or drawing a line of best fit by eye.
- Enter the data into a list.
- STAT → Edit → type values into L1. Then STAT → CALC → 1:1-Var Stats, choose L1, Calculate.TI-84
- Add a Lists & Spreadsheet page, name a column and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
- Statistics menu → enter data in List 1 → CALC (F2) → 1-VAR.Casio
- Read \(\bar{x}\) (mean), \(S_x\) (sample sd) or \(\sigma_x\) (population sd), and the five-number summary.
Tip: Run 1-Var Stats on the \(x\)-list and the \(y\)-list separately to get \(\bar{x}\) and \(\bar{y}\), then check they satisfy your regression equation.
See the full GDC guide for more calculator models and topics.
Ready to practise properly?
Correlation and regression questions, marked instantly like the real exam.
Quick answers
The questions students on this topic ask most often.
What does Pearson's r actually measure?
Pearson's r measures the strength and direction of a LINEAR relationship between two variables, on a scale from -1 to 1. It says nothing about curved relationships - a strong curve can still give a high r, which is why you should always look at the scatter diagram too.
Why is extrapolation dangerous?
A regression line is only fitted to the data you have. Using it to predict far outside that range (extrapolation) assumes the same linear pattern continues, which often isn't true - the relationship can change, level off, or reverse beyond the data you actually collected.
Why can't I always use the y on x line to predict x?
The regression line of y on x is built to minimise errors in the y-direction, so it's optimised for predicting y from a given x. To reliably predict x from a given y, you need the separate regression line of x on y, which minimises errors in the x-direction instead.
Do I need to calculate r by hand?
No - the syllabus expects you to use your GDC to find r and the regression equation. Hand calculation is only there to help you understand what the value means, not something you'll be examined on directly. See the GDC guide for model-specific instructions.
Related topics
More Statistics & Probability topics from the same AA SL syllabus unit, in case you want to keep going.