Calculating Correlation & Regression (AI SL)

This topic is the hands-on companion to correlation and regression theory: given raw pairs of data, you enter them into your GDC and calculate Pearson's \(r\), the regression line \(y=ax+b\), the mean point, and residuals from scratch. Where the theory topic hands you an equation to interpret, here you build that equation yourself and check it against the data.

What the syllabus says

This topic maps onto one point in the official IB Applications & Interpretation syllabus - the same point as Correlation & Regression, but focused on the calculation process itself.

CodeSyllabus content
SL4.4Linear correlation of bivariate data. Pearson's product-moment correlation coefficient \(r\), calculated using technology. Scatter diagrams; lines of best fit by eye, passing through the mean point. The equation of the regression line of \(y\) on \(x\), found using technology, and interpretation of the parameters \(a\) and \(b\) in \(y=ax+b\). Use of the regression line for prediction, being aware of the dangers of extrapolation.

Hand calculations of \(r\) and the regression coefficients may enhance understanding, but technology should always be used for the actual result.

Key terms

Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.

What is bivariate data?

Bivariate data consists of paired observations of two variables measured on the same subjects - one \(x\)-value and one \(y\)-value per data point. It's the raw material for correlation and regression: you enter it as two matched lists into your GDC.

e.g. Advertising spend and revenue for six months: \((2,12), (4,18), (5,22),\dots\)

What is a residual?

A residual is the vertical distance between an actual data point and the regression line's prediction at that \(x\)-value: \(\text{residual}=y_{\text{actual}}-y_{\text{predicted}}.\) Small residuals mean the line fits well at that point; large ones flag a poor fit or an unusual observation.

e.g. For \((150,11)\) with predicted \(y=0.0909(150)-2.30=11.3,\) the residual is \(11-11.3=-0.3.\)

What is the mean point?

The mean point is \((\bar x,\bar y)\), the point formed by the mean of all \(x\)-values and the mean of all \(y\)-values. Every least-squares regression line passes exactly through this point, which makes it a handy check on a GDC-calculated equation.

e.g. With \(\bar x=6.17\) and regression line \(y=3.77x+3.25,\) substituting gives \(3.77(6.17)+3.25=26.5=\bar y.\)

What is the regression line of x on y?

The regression line of \(x\) on \(y\) is a separate least-squares line that predicts \(x\) from a given \(y\), minimising horizontal rather than vertical distances. It's different from the \(y\) on \(x\) line and is needed whenever you want to predict \(x\) from a known \(y\)-value.

e.g. Given a plant's height, you'd need the \(x\) on \(y\) line - not the \(y\) on \(x\) line - to estimate its sunlight hours.

What does 3 significant figures mean for a and b?

IB regression answers are usually rounded to 3 significant figures for the gradient \(a\) and intercept \(b\), unless told otherwise. Using the rounded coefficients (rather than full GDC precision) in later parts can introduce small but noticeable rounding errors.

e.g. A GDC value of \(a=-1.2236\) is reported as \(a=-1.22\) to 3 significant figures.

Key formulas

The calculation workflow matters more than any single formula here - the tables below summarise what each quantity is and where it comes from, and the sections underneath walk through the process.

Formula reference

Every quantity below is read straight from your GDC's regression output rather than computed from raw sums by hand.

FormulaUsed forBooklet?
\(y=ax+b\)Regression line of \(y\) on \(x\)✓ Yes
\(\bar x=\dfrac{\Sigma x}{n},\ \bar y=\dfrac{\Sigma y}{n}\)Mean point, verified to lie on the lineNot in booklet - prior knowledge
\(\text{residual}=y_{\text{actual}}-y_{\text{predicted}}\)Measures fit at one data pointNot in booklet - defined quantity

Calculating a residual: worked micro-example

Given the regression line \(y=0.0909x-2.30\) and the data point \((150,11)\): predicted \(y=0.0909(150)-2.30=13.6-2.30=11.3\); residual \(=11-11.3=-0.3.\) The negative sign shows the actual value was slightly below what the line predicted.

The full calculation workflow

Every question in this topic follows roughly the same sequence of GDC steps, in this order.

1. Enter the data

Type the \(x\)-values into one list and the matching \(y\)-values into a second list, keeping the pairing consistent row by row.

2. Run the regression

Use the GDC's linear regression function on the two lists to get \(r\), the gradient \(a\), and the intercept \(b\) in one calculation.

3. Predict and verify

Substitute a given \(x\) into \(y=ax+b\) for a prediction, or substitute \(\bar x\) to verify the line passes through the mean point.

Checking your regression line

Two quick sanity checks catch most calculation errors before you move on to using the line.

Sign of the gradient matches the trend

If the scatter of points clearly rises left to right, \(a\) should be positive; if it falls, \(a\) should be negative. A mismatch usually means the \(x\)- and \(y\)-lists were swapped.

Mean point lies on the line

Substituting \(\bar x\) into your regression equation should return \(\bar y\) (to rounding). If it doesn't, re-check the coefficients you copied down from the GDC.

Worked examples

Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.

1
Medium
GDC
[6 marks]

A salesperson's weekly distance \(x\) (km) and number of sales \(y\):

\(x\)80120150180200220250300
\(y\)59111416172125

(a) Find \(r.\)
(b) Find the regression line of \(y\) on \(x.\)
(c) Estimate sales when the salesperson travels 170 km.
(d) Find the residual for the data point \((150, 11).\)

Worked solution

(a)   \(r=0.998\) A1

(b)   \(y\) M1
\(=0.0909x-2.30\) A1

(c)   \(y=0.0909(170)-2.30=15.5-2.30=13.2\) A1

(d)   Predicted: \(0.0909(150)-2.30=13.6-2.30=11.3\)
Residual \(=11-11.3\) M1
\(=-0.3\) A1

A1 Correct value of r M1 Attempt to find the regression line of y on x A1 Correct equation y=0.0909x-2.30 A1 Correct estimate of sales 13.2 at x=170 M1 Substituting x=150 into the regression equation to find the predicted value A1 Correct residual -0.3
2
Hard
GDC
[8 marks]

Background noise level \(x\) (dB) and worker productivity score \(y\) in an office:

\(x\)3040455055606570
\(y\)8880756861554840

(a) Find \(\bar x\) and \(\bar y.\)
(b) Find \(r\) and the regression line of \(y\) on \(x.\)
(c) Verify that the mean point lies on the regression line.
(d) Estimate productivity when noise level is 52 dB.
(e) State the residual for the data point \((65, 48).\)

Worked solution

(a)   \(\bar{x}=51.875,\; \bar{y}=64.375\) A1

(b)   \(r=-0.994\) A1
\(y=-1.22x+128\) (3 significant figures) M1 A1

(c)   \(-1.22(51.875)+128=-63.3+128=64.7\approx 64.4\) ✓ A1 AG

(d)   \(y=-1.2236(52)+127.85\approx64.2\) A1

(e)   Predicted for \(x=65\): \(-1.2236(65)+127.85\approx48.3\)
Residual \(=48-48.3\) M1
\(\approx-0.3\) A1

A1 Correct mean values \(\bar{x}=51.875\) and \(\bar{y}=64.375\) A1 Correct value of r=-0.994 M1 Attempt to find the regression line of y on x A1 Correct equation y=-1.22x+128 A1 Verifying the mean point lies on the regression line A1 Correct estimate of productivity 64.2 at x=52 M1 Finding the predicted value at x=65 A1 Correct residual -0.3

Common mistakes

The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.

  • Swapping the \(x\)- and \(y\)-lists. Entering the wrong variable into the wrong list flips the sign of the gradient and gives a regression line for the wrong direction entirely - always double check which quantity is \(x\) and which is \(y\) before running LinReg.
  • Using rounded coefficients for later calculations. Rounding \(a\) and \(b\) to 3 significant figures before using them in a prediction or residual compounds the rounding error - where possible, keep full GDC precision until the final step.
  • Getting the sign of a residual backwards. A residual is always actual minus predicted, not the other way round - a positive residual means the actual value was above the line, not below it.
  • Forgetting DiagnosticOn (TI-84). Without turning DiagnosticOn, the TI-84's LinReg output won't display \(r\) at all, only \(a\) and \(b\) - an easy thing to miss under exam pressure.

Using your GDC

Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.

Show steps for:
Linear regression (line of best fit) and r

Core AI skill - gives the regression line and the correlation coefficient in one step.

  1. Enter \(x\)-values in one list and \(y\)-values in another.
  2. Turn on DiagnosticOn once (2nd → CATALOG) so \(r\) appears; then STAT → CALC → 4:LinReg(ax+b), set Xlist and Ylist, Calculate.TI-84
  3. Calculator page → menu → Statistics → Stat Calculations → Linear Regression (mx+b).Nspire
  4. Statistics menu → CALC (F2) → REG (F3) → X (linear); \(r\) shows automatically.Casio
  5. Read \(a\) (gradient), \(b\) (intercept), \(r\) (correlation) and \(r^2\) (coefficient of determination).
  6. Use the equation to predict - but only within the data range (interpolation).

Tip: \(r\) near \(\pm1\) means a strong linear fit; near 0 means weak. \(r^2\) is the proportion of variation explained.

One-variable statistics (mean, median, standard deviation)

Run this separately on your \(x\)-list and \(y\)-list to get \(\bar x\) and \(\bar y\) for the mean point check.

  1. Enter the data into a list.
  2. STAT → Edit → type values into L1. Then STAT → CALC → 1:1-Var Stats, choose L1, Calculate.TI-84
  3. Add a Lists & Spreadsheet page, name a column and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
  4. Statistics menu → enter data in List 1 → CALC (F2) → 1-VAR.Casio
  5. Read \(\bar x\) (mean) for each list, then substitute \(\bar x\) into your regression line and check it returns \(\bar y\).

Tip: If substituting \(\bar x\) doesn't return \(\bar y\) (to rounding), re-check the coefficients you copied down from the LinReg screen.

See the full GDC guide for more calculator models and topics.

Ready to practise properly?

Calculating correlation & regression questions, marked instantly like the real exam.

Quick answers

The questions students on this topic ask most often.

What's the difference between this topic and Correlation & Regression?

Correlation & Regression focuses on interpreting \(r\) and a regression line you're given. This topic is about the calculation itself - entering raw bivariate data into your GDC and computing \(r\), the regression line, the mean point, and residuals from scratch.

What is a residual?

A residual is the difference between an actual data value and the value the regression line predicts for the same \(x\): residual = actual \(y\) - predicted \(y\). A small residual means the line fits that point well; a large one means it doesn't.

Why does the regression line always pass through the mean point?

The least-squares method that generates the regression line is constructed so the positive and negative residuals balance out around the average \(x\) and \(y\) values - this makes the point \((\bar x,\bar y)\) always sit exactly on the line, which is a useful check on your GDC output.

Do I need to enter data by hand every time?

Yes - unlike some other topics, calculating correlation and regression starts from raw data pairs, so you always type the \(x\)-values into one list and the \(y\)-values into another before running the regression, rather than being handed the equation directly.