Correlation & Regression (AI SL)
Correlation and regression describe how two variables move together and let you predict one from the other. This topic covers Pearson's product-moment correlation coefficient \(r\) as a measure of linear relationship strength, scatter diagrams, the regression line of \(y\) on \(x\), and how to interpret what the gradient and intercept mean in context - plus the crucial distinction between correlation and causation.
What the syllabus says
This topic maps onto one point in the official IB Applications & Interpretation syllabus.
| Code | Syllabus content |
|---|---|
| SL4.4 | Linear correlation of bivariate data. Pearson's product-moment correlation coefficient \(r\), calculated using technology. Scatter diagrams; lines of best fit by eye, passing through the mean point. The equation of the regression line of \(y\) on \(x\), found using technology, and interpretation of the parameters \(a\) and \(b\) in \(y=ax+b\). Awareness of the dangers of extrapolation, and that correlation does not imply causation. |
Critical values of \(r\) are given where needed - Pearson's \(r\) is only meaningful for linear relationships.
Key terms
Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.
What is Pearson's correlation coefficient?
Pearson's \(r\) is a number between \(-1\) and \(1\) that measures the strength and direction of a linear relationship between two variables. Values near \(\pm1\) mean a strong linear relationship; values near 0 mean little or no linear relationship. It's calculated by your GDC from bivariate data.
e.g. \(r=0.992\) for advertising spend vs revenue means a very strong positive linear relationship.
What is the regression line of y on x?
The regression line of \(y\) on \(x\) is the line \(y=ax+b\) that best fits a set of bivariate data, minimising the vertical distances between the points and the line. It's used to predict \(y\) for a given \(x\), and always passes through the mean point \((\bar x,\bar y)\).
e.g. For \(y=3.77x+3.25\), predicting at \(x=6\) gives \(y=3.77(6)+3.25=25.9\).
What does the gradient of a regression line mean?
The gradient \(a\) in \(y=ax+b\) tells you how much \(y\) is predicted to change for every one-unit increase in \(x\). Interpreting it in context turns an abstract number into a real statement about the relationship being modelled.
e.g. For \(y=8x+30\) (score vs hours studied), each extra hour of study is associated with about 8 more marks.
What is the difference between interpolation and extrapolation?
Interpolation means predicting a \(y\)-value for an \(x\) that lies within the range of the original data - generally reliable. Extrapolation means predicting outside that range, which is risky because the linear pattern might not continue beyond the data collected.
e.g. With data for \(x\) between 10 and 30, predicting at \(x=22\) is interpolation; predicting at \(x=50\) would be extrapolation.
Why doesn't correlation imply causation?
A strong correlation shows two variables move together, but it doesn't prove one causes the other - a third, confounding factor could be driving both, or the link could be coincidental. Concluding causation needs evidence beyond a correlation coefficient.
e.g. Advertising spend and sales might both rise together in December simply because of the season, not because one directly causes the other.
Key formulas
This topic has one central equation and a lot of interpretation - the tables below summarise what to know, and the sections underneath explain how to use it.
Formula reference
Both \(r\) and the regression line's coefficients \(a\) and \(b\) are always found using your GDC's statistics mode - you're never expected to compute them from raw sums by hand in this topic.
| Formula | Used for | Booklet? |
|---|---|---|
| \(y=ax+b\) | Regression line of \(y\) on \(x\) | ✓ Yes |
| \(-1\le r\le 1\) | Range of Pearson's correlation coefficient | Not in booklet - GDC output |
| Line passes through \((\bar x,\bar y)\) | Check on a computed regression line | Not in booklet - property of the line |
Interpreting the strength of r
There's no single official cutoff, but this is the general scale examiners expect you to use when describing correlation strength in words.
| Value of \(|r|\) | Description | Direction (sign of \(r\)) |
|---|---|---|
| 0.8 - 1.0 | Strong correlation | Positive if \(r>0\), negative if \(r<0\) |
| 0.5 - 0.8 | Moderate correlation | Positive if \(r>0\), negative if \(r<0\) |
| 0 - 0.5 | Weak correlation | Positive if \(r>0\), negative if \(r<0\) |
| 0 | No linear correlation | N/A |
Reading a scatter diagram
Before calculating anything, a scatter diagram tells you what kind of relationship (if any) you're dealing with.
Positive correlation
As \(x\) increases, \(y\) tends to increase too - points trend upward left to right, giving a positive gradient and \(r>0\).
Negative correlation
As \(x\) increases, \(y\) tends to decrease - points trend downward left to right, giving a negative gradient and \(r<0\).
No correlation
The points show no clear upward or downward trend - \(r\) is close to zero, and fitting a line wouldn't be meaningful.
Using the regression line responsibly
A regression equation is only as trustworthy as the data and range it came from.
Predict, don't extrapolate
Only trust predictions for \(x\)-values within (or close to) the range of the original data. Far outside that range, the linear pattern may simply not hold.
y on x predicts y, not x
The \(y\) on \(x\) line is built to predict \(y\) from a given \(x\). To predict \(x\) from a given \(y\), you need the separate \(x\) on \(y\) regression line instead.
Correlation is not causation
Even a very high \(|r|\) only shows association, not cause and effect - always consider whether a confounding variable could explain the pattern.
Worked examples
Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.
Hours studied \(x\) and test score \(y\) give the regression line \(y=8x+30.\)
(a) Predict the score for 5 hours.
(b) Interpret the gradient.
(c) Interpret the \(y\)-intercept.
Worked solution
(a) Prediction. Substitute \(x=5:\) M1
\(y=8(5)+30=70.\) A1
(b) Gradient. The gradient \(a=8\) is the rate of change: R1
each extra hour of study is associated with about 8 more marks. A1
(c) Intercept. The intercept \(b=30\) is the predicted \(y\) when \(x=0:\) a student who does no study is predicted to score 30 R1
Temperature \(x\)°C and energy use \(y\) (kWh):
| \(x\) | 10 | 15 | 20 | 25 | 30 |
|---|---|---|---|---|---|
| \(y\) | 80 | 72 | 60 | 55 | 42 |
(a) Find Pearson's \(r\) (3 significant figures).
(b) Find the regression line \(y=ax+b\) (3 significant figures).
(c) Predict the energy use at 22°C.
Worked solution
(a) Pearson’s \(r.\) A1
Enter the five points and run linear regression. M1
(b) Regression line. The GDC returns \(a\) A1
and \(b\) for \(y=ax+b.\) A1
(c) Prediction at 22°C. \(y=-1.86(22)+99.0\approx 58.1\) M1
kWh A1
(interpolation, since 22 is within 10–30). A1
Common mistakes
The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.
- Claiming causation from correlation. A strong \(r\) shows association, not cause and effect - always consider whether a confounding factor could explain the relationship instead.
- Extrapolating without comment. Using the regression line to predict far outside the data range and treating the answer as equally reliable as an interpolated one - always flag extrapolation as less trustworthy.
- Predicting \(x\) from \(y\) using the \(y\) on \(x\) line. The \(y\) on \(x\) regression line is built to predict \(y\), not to be rearranged to predict \(x\) - that needs the separate \(x\) on \(y\) line.
- Forgetting to interpret \(a\) and \(b\) in context. A question asking to "interpret the gradient" wants a sentence about what it means for the real quantities involved, not just a restatement of the number.
Using your GDC
Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.
Core AI skill - gives the regression line and the correlation coefficient in one step.
- Enter \(x\)-values in one list and \(y\)-values in another.
- Turn on DiagnosticOn once (2nd → CATALOG) so \(r\) appears; then STAT → CALC → 4:LinReg(ax+b), set Xlist and Ylist, Calculate.TI-84
- Calculator page → menu → Statistics → Stat Calculations → Linear Regression (mx+b).Nspire
- Statistics menu → CALC (F2) → REG (F3) → X (linear); \(r\) shows automatically.Casio
- Read \(a\) (gradient), \(b\) (intercept), \(r\) (correlation) and \(r^2\) (coefficient of determination).
- Use the equation to predict - but only within the data range (interpolation).
Tip: \(r\) near \(\pm1\) means a strong linear fit; near 0 means weak. \(r^2\) is the proportion of variation explained.
Instant summary statistics from a list - useful for finding the mean point \((\bar x,\bar y)\) that a regression line always passes through.
- Enter the data into a list.
- STAT → Edit → type values into L1. Then STAT → CALC → 1:1-Var Stats, choose L1, Calculate.TI-84
- Add a Lists & Spreadsheet page, name a column and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
- Statistics menu → enter data in List 1 → CALC (F2) → 1-VAR.Casio
- Read \(\bar x\) (mean) for each list - the mean point \((\bar x,\bar y)\) always lies on the regression line, useful for checking your answer.
Tip: Run 1-Var Stats separately on your \(x\)-list and \(y\)-list to get \(\bar x\) and \(\bar y\) for a quick verification of the regression line.
See the full GDC guide for more calculator models and topics.
Ready to practise properly?
Correlation & regression questions, marked instantly like the real exam.
Quick answers
The questions students on this topic ask most often.
What does Pearson's r actually tell me?
\(r\) measures the strength and direction of a linear relationship between two variables. It ranges from -1 (perfect negative correlation) to +1 (perfect positive correlation), with 0 meaning no linear relationship. It's only meaningful for relationships that are actually linear.
Does a strong correlation mean one variable causes the other?
No. Correlation measures how closely two variables move together, not whether one causes the other. A third, confounding factor can drive both variables at once, or the relationship could be coincidental - correlation never proves causation.
Why can't I always predict x from a y on x regression line?
The \(y\) on \(x\) line is built to minimise errors in the \(y\)-direction, so it's optimised for predicting \(y\) from a given \(x\). Using it backwards to predict \(x\) from \(y\) gives a biased result - you'd need the separate \(x\) on \(y\) regression line for that direction.
What's the danger of extrapolation?
Extrapolation means using a regression line to predict a value outside the range of the original data. The linear pattern you fitted may not continue beyond that range, so predictions there are far less reliable than predictions within the data range (interpolation).
Sub-topics
Correlation & Regression broken down into its individual skills, each with its own focused page.
Related topics
More Statistics & Probability topics from the same AI SL syllabus unit, in case you want to keep going.