Correlation & Regression (AI SL)

Correlation and regression describe how two variables move together and let you predict one from the other. This topic covers Pearson's product-moment correlation coefficient \(r\) as a measure of linear relationship strength, scatter diagrams, the regression line of \(y\) on \(x\), and how to interpret what the gradient and intercept mean in context - plus the crucial distinction between correlation and causation.

What the syllabus says

This topic maps onto one point in the official IB Applications & Interpretation syllabus.

CodeSyllabus content
SL4.4Linear correlation of bivariate data. Pearson's product-moment correlation coefficient \(r\), calculated using technology. Scatter diagrams; lines of best fit by eye, passing through the mean point. The equation of the regression line of \(y\) on \(x\), found using technology, and interpretation of the parameters \(a\) and \(b\) in \(y=ax+b\). Awareness of the dangers of extrapolation, and that correlation does not imply causation.

Critical values of \(r\) are given where needed - Pearson's \(r\) is only meaningful for linear relationships.

Key terms

Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.

What is Pearson's correlation coefficient?

Pearson's \(r\) is a number between \(-1\) and \(1\) that measures the strength and direction of a linear relationship between two variables. Values near \(\pm1\) mean a strong linear relationship; values near 0 mean little or no linear relationship. It's calculated by your GDC from bivariate data.

e.g. \(r=0.992\) for advertising spend vs revenue means a very strong positive linear relationship.

What is the regression line of y on x?

The regression line of \(y\) on \(x\) is the line \(y=ax+b\) that best fits a set of bivariate data, minimising the vertical distances between the points and the line. It's used to predict \(y\) for a given \(x\), and always passes through the mean point \((\bar x,\bar y)\).

e.g. For \(y=3.77x+3.25\), predicting at \(x=6\) gives \(y=3.77(6)+3.25=25.9\).

What does the gradient of a regression line mean?

The gradient \(a\) in \(y=ax+b\) tells you how much \(y\) is predicted to change for every one-unit increase in \(x\). Interpreting it in context turns an abstract number into a real statement about the relationship being modelled.

e.g. For \(y=8x+30\) (score vs hours studied), each extra hour of study is associated with about 8 more marks.

What is the difference between interpolation and extrapolation?

Interpolation means predicting a \(y\)-value for an \(x\) that lies within the range of the original data - generally reliable. Extrapolation means predicting outside that range, which is risky because the linear pattern might not continue beyond the data collected.

e.g. With data for \(x\) between 10 and 30, predicting at \(x=22\) is interpolation; predicting at \(x=50\) would be extrapolation.

Why doesn't correlation imply causation?

A strong correlation shows two variables move together, but it doesn't prove one causes the other - a third, confounding factor could be driving both, or the link could be coincidental. Concluding causation needs evidence beyond a correlation coefficient.

e.g. Advertising spend and sales might both rise together in December simply because of the season, not because one directly causes the other.

Key formulas

This topic has one central equation and a lot of interpretation - the tables below summarise what to know, and the sections underneath explain how to use it.

Formula reference

Both \(r\) and the regression line's coefficients \(a\) and \(b\) are always found using your GDC's statistics mode - you're never expected to compute them from raw sums by hand in this topic.

FormulaUsed forBooklet?
\(y=ax+b\)Regression line of \(y\) on \(x\)✓ Yes
\(-1\le r\le 1\)Range of Pearson's correlation coefficientNot in booklet - GDC output
Line passes through \((\bar x,\bar y)\)Check on a computed regression lineNot in booklet - property of the line

Interpreting the strength of r

There's no single official cutoff, but this is the general scale examiners expect you to use when describing correlation strength in words.

Value of \(|r|\)DescriptionDirection (sign of \(r\))
0.8 - 1.0Strong correlationPositive if \(r>0\), negative if \(r<0\)
0.5 - 0.8Moderate correlationPositive if \(r>0\), negative if \(r<0\)
0 - 0.5Weak correlationPositive if \(r>0\), negative if \(r<0\)
0No linear correlationN/A

Reading a scatter diagram

Before calculating anything, a scatter diagram tells you what kind of relationship (if any) you're dealing with.

Positive correlation

As \(x\) increases, \(y\) tends to increase too - points trend upward left to right, giving a positive gradient and \(r>0\).

Negative correlation

As \(x\) increases, \(y\) tends to decrease - points trend downward left to right, giving a negative gradient and \(r<0\).

No correlation

The points show no clear upward or downward trend - \(r\) is close to zero, and fitting a line wouldn't be meaningful.

Using the regression line responsibly

A regression equation is only as trustworthy as the data and range it came from.

Predict, don't extrapolate

Only trust predictions for \(x\)-values within (or close to) the range of the original data. Far outside that range, the linear pattern may simply not hold.

y on x predicts y, not x

The \(y\) on \(x\) line is built to predict \(y\) from a given \(x\). To predict \(x\) from a given \(y\), you need the separate \(x\) on \(y\) regression line instead.

Correlation is not causation

Even a very high \(|r|\) only shows association, not cause and effect - always consider whether a confounding variable could explain the pattern.

Worked examples

Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.

1
Medium
GDC
[5 marks]

Hours studied \(x\) and test score \(y\) give the regression line \(y=8x+30.\)

(a) Predict the score for 5 hours.
(b) Interpret the gradient.
(c) Interpret the \(y\)-intercept.

Worked solution

(a) Prediction. Substitute \(x=5:\) M1
\(y=8(5)+30=70.\) A1

(b) Gradient. The gradient \(a=8\) is the rate of change: R1
each extra hour of study is associated with about 8 more marks. A1

(c) Intercept. The intercept \(b=30\) is the predicted \(y\) when \(x=0:\) a student who does no study is predicted to score 30 R1

M1 Substitute A1 Prediction R1 Gradient meaning A1 Gradient value R1 Intercept
2
Hard
GDC
[7 marks]

Temperature \(x\)°C and energy use \(y\) (kWh):

\(x\)1015202530
\(y\)8072605542

(a) Find Pearson's \(r\) (3 significant figures).
(b) Find the regression line \(y=ax+b\) (3 significant figures).
(c) Predict the energy use at 22°C.

Worked solution

(a) Pearson’s \(r.\) A1
Enter the five points and run linear regression. M1

(b) Regression line. The GDC returns \(a\) A1
and \(b\) for \(y=ax+b.\) A1

(c) Prediction at 22°C. \(y=-1.86(22)+99.0\approx 58.1\) M1
kWh A1
(interpolation, since 22 is within 10–30). A1

A1 Obtain \(r\) M1 Run regression A1 Gradient \(a\) A1 Intercept \(b\) M1 Substitute \(x=22\) A1 Prediction A1 Interpolation justified

Common mistakes

The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.

  • Claiming causation from correlation. A strong \(r\) shows association, not cause and effect - always consider whether a confounding factor could explain the relationship instead.
  • Extrapolating without comment. Using the regression line to predict far outside the data range and treating the answer as equally reliable as an interpolated one - always flag extrapolation as less trustworthy.
  • Predicting \(x\) from \(y\) using the \(y\) on \(x\) line. The \(y\) on \(x\) regression line is built to predict \(y\), not to be rearranged to predict \(x\) - that needs the separate \(x\) on \(y\) line.
  • Forgetting to interpret \(a\) and \(b\) in context. A question asking to "interpret the gradient" wants a sentence about what it means for the real quantities involved, not just a restatement of the number.

Using your GDC

Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.

Show steps for:
Linear regression (line of best fit) and r

Core AI skill - gives the regression line and the correlation coefficient in one step.

  1. Enter \(x\)-values in one list and \(y\)-values in another.
  2. Turn on DiagnosticOn once (2nd → CATALOG) so \(r\) appears; then STAT → CALC → 4:LinReg(ax+b), set Xlist and Ylist, Calculate.TI-84
  3. Calculator page → menu → Statistics → Stat Calculations → Linear Regression (mx+b).Nspire
  4. Statistics menu → CALC (F2) → REG (F3) → X (linear); \(r\) shows automatically.Casio
  5. Read \(a\) (gradient), \(b\) (intercept), \(r\) (correlation) and \(r^2\) (coefficient of determination).
  6. Use the equation to predict - but only within the data range (interpolation).

Tip: \(r\) near \(\pm1\) means a strong linear fit; near 0 means weak. \(r^2\) is the proportion of variation explained.

One-variable statistics (mean, median, standard deviation)

Instant summary statistics from a list - useful for finding the mean point \((\bar x,\bar y)\) that a regression line always passes through.

  1. Enter the data into a list.
  2. STAT → Edit → type values into L1. Then STAT → CALC → 1:1-Var Stats, choose L1, Calculate.TI-84
  3. Add a Lists & Spreadsheet page, name a column and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
  4. Statistics menu → enter data in List 1 → CALC (F2) → 1-VAR.Casio
  5. Read \(\bar x\) (mean) for each list - the mean point \((\bar x,\bar y)\) always lies on the regression line, useful for checking your answer.

Tip: Run 1-Var Stats separately on your \(x\)-list and \(y\)-list to get \(\bar x\) and \(\bar y\) for a quick verification of the regression line.

See the full GDC guide for more calculator models and topics.

Ready to practise properly?

Correlation & regression questions, marked instantly like the real exam.

Quick answers

The questions students on this topic ask most often.

What does Pearson's r actually tell me?

\(r\) measures the strength and direction of a linear relationship between two variables. It ranges from -1 (perfect negative correlation) to +1 (perfect positive correlation), with 0 meaning no linear relationship. It's only meaningful for relationships that are actually linear.

Does a strong correlation mean one variable causes the other?

No. Correlation measures how closely two variables move together, not whether one causes the other. A third, confounding factor can drive both variables at once, or the relationship could be coincidental - correlation never proves causation.

Why can't I always predict x from a y on x regression line?

The \(y\) on \(x\) line is built to minimise errors in the \(y\)-direction, so it's optimised for predicting \(y\) from a given \(x\). Using it backwards to predict \(x\) from \(y\) gives a biased result - you'd need the separate \(x\) on \(y\) regression line for that direction.

What's the danger of extrapolation?

Extrapolation means using a regression line to predict a value outside the range of the original data. The linear pattern you fitted may not continue beyond that range, so predictions there are far less reliable than predictions within the data range (interpolation).

Sub-topics

Correlation & Regression broken down into its individual skills, each with its own focused page.