Calculating Correlation & Regression (AI SL)
This topic is the hands-on companion to correlation and regression theory: given raw pairs of data, you enter them into your GDC and calculate Pearson's \(r\), the regression line \(y=ax+b\), the mean point, and residuals from scratch. Where the theory topic hands you an equation to interpret, here you build that equation yourself and check it against the data.
What the syllabus says
This topic maps onto one point in the official IB Applications & Interpretation syllabus - the same point as Correlation & Regression, but focused on the calculation process itself.
| Code | Syllabus content |
|---|---|
| SL4.4 | Linear correlation of bivariate data. Pearson's product-moment correlation coefficient \(r\), calculated using technology. Scatter diagrams; lines of best fit by eye, passing through the mean point. The equation of the regression line of \(y\) on \(x\), found using technology, and interpretation of the parameters \(a\) and \(b\) in \(y=ax+b\). Use of the regression line for prediction, being aware of the dangers of extrapolation. |
Hand calculations of \(r\) and the regression coefficients may enhance understanding, but technology should always be used for the actual result.
Key terms
Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.
What is bivariate data?
Bivariate data consists of paired observations of two variables measured on the same subjects - one \(x\)-value and one \(y\)-value per data point. It's the raw material for correlation and regression: you enter it as two matched lists into your GDC.
e.g. Advertising spend and revenue for six months: \((2,12), (4,18), (5,22),\dots\)
What is a residual?
A residual is the vertical distance between an actual data point and the regression line's prediction at that \(x\)-value: \(\text{residual}=y_{\text{actual}}-y_{\text{predicted}}.\) Small residuals mean the line fits well at that point; large ones flag a poor fit or an unusual observation.
e.g. For \((150,11)\) with predicted \(y=0.0909(150)-2.30=11.3,\) the residual is \(11-11.3=-0.3.\)
What is the mean point?
The mean point is \((\bar x,\bar y)\), the point formed by the mean of all \(x\)-values and the mean of all \(y\)-values. Every least-squares regression line passes exactly through this point, which makes it a handy check on a GDC-calculated equation.
e.g. With \(\bar x=6.17\) and regression line \(y=3.77x+3.25,\) substituting gives \(3.77(6.17)+3.25=26.5=\bar y.\)
What is the regression line of x on y?
The regression line of \(x\) on \(y\) is a separate least-squares line that predicts \(x\) from a given \(y\), minimising horizontal rather than vertical distances. It's different from the \(y\) on \(x\) line and is needed whenever you want to predict \(x\) from a known \(y\)-value.
e.g. Given a plant's height, you'd need the \(x\) on \(y\) line - not the \(y\) on \(x\) line - to estimate its sunlight hours.
What does 3 significant figures mean for a and b?
IB regression answers are usually rounded to 3 significant figures for the gradient \(a\) and intercept \(b\), unless told otherwise. Using the rounded coefficients (rather than full GDC precision) in later parts can introduce small but noticeable rounding errors.
e.g. A GDC value of \(a=-1.2236\) is reported as \(a=-1.22\) to 3 significant figures.
Key formulas
The calculation workflow matters more than any single formula here - the tables below summarise what each quantity is and where it comes from, and the sections underneath walk through the process.
Formula reference
Every quantity below is read straight from your GDC's regression output rather than computed from raw sums by hand.
| Formula | Used for | Booklet? |
|---|---|---|
| \(y=ax+b\) | Regression line of \(y\) on \(x\) | ✓ Yes |
| \(\bar x=\dfrac{\Sigma x}{n},\ \bar y=\dfrac{\Sigma y}{n}\) | Mean point, verified to lie on the line | Not in booklet - prior knowledge |
| \(\text{residual}=y_{\text{actual}}-y_{\text{predicted}}\) | Measures fit at one data point | Not in booklet - defined quantity |
Calculating a residual: worked micro-example
Given the regression line \(y=0.0909x-2.30\) and the data point \((150,11)\): predicted \(y=0.0909(150)-2.30=13.6-2.30=11.3\); residual \(=11-11.3=-0.3.\) The negative sign shows the actual value was slightly below what the line predicted.
The full calculation workflow
Every question in this topic follows roughly the same sequence of GDC steps, in this order.
1. Enter the data
Type the \(x\)-values into one list and the matching \(y\)-values into a second list, keeping the pairing consistent row by row.
2. Run the regression
Use the GDC's linear regression function on the two lists to get \(r\), the gradient \(a\), and the intercept \(b\) in one calculation.
3. Predict and verify
Substitute a given \(x\) into \(y=ax+b\) for a prediction, or substitute \(\bar x\) to verify the line passes through the mean point.
Checking your regression line
Two quick sanity checks catch most calculation errors before you move on to using the line.
Sign of the gradient matches the trend
If the scatter of points clearly rises left to right, \(a\) should be positive; if it falls, \(a\) should be negative. A mismatch usually means the \(x\)- and \(y\)-lists were swapped.
Mean point lies on the line
Substituting \(\bar x\) into your regression equation should return \(\bar y\) (to rounding). If it doesn't, re-check the coefficients you copied down from the GDC.
Worked examples
Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.
A salesperson's weekly distance \(x\) (km) and number of sales \(y\):
| \(x\) | 80 | 120 | 150 | 180 | 200 | 220 | 250 | 300 |
|---|---|---|---|---|---|---|---|---|
| \(y\) | 5 | 9 | 11 | 14 | 16 | 17 | 21 | 25 |
(a) Find \(r.\)
(b) Find the regression line of \(y\) on \(x.\)
(c) Estimate sales when the salesperson travels 170 km.
(d) Find the residual for the data point \((150, 11).\)
Worked solution
(a) \(r=0.998\) A1
(b) \(y\) M1
\(=0.0909x-2.30\) A1
(c) \(y=0.0909(170)-2.30=15.5-2.30=13.2\) A1
(d) Predicted: \(0.0909(150)-2.30=13.6-2.30=11.3\)
Residual \(=11-11.3\) M1
\(=-0.3\) A1
Background noise level \(x\) (dB) and worker productivity score \(y\) in an office:
| \(x\) | 30 | 40 | 45 | 50 | 55 | 60 | 65 | 70 |
|---|---|---|---|---|---|---|---|---|
| \(y\) | 88 | 80 | 75 | 68 | 61 | 55 | 48 | 40 |
(a) Find \(\bar x\) and \(\bar y.\)
(b) Find \(r\) and the regression line of \(y\) on \(x.\)
(c) Verify that the mean point lies on the regression line.
(d) Estimate productivity when noise level is 52 dB.
(e) State the residual for the data point \((65, 48).\)
Worked solution
(a) \(\bar{x}=51.875,\; \bar{y}=64.375\) A1
(b) \(r=-0.994\) A1
\(y=-1.22x+128\) (3 significant figures) M1 A1
(c) \(-1.22(51.875)+128=-63.3+128=64.7\approx 64.4\) ✓ A1 AG
(d) \(y=-1.2236(52)+127.85\approx64.2\) A1
(e) Predicted for \(x=65\): \(-1.2236(65)+127.85\approx48.3\)
Residual \(=48-48.3\) M1
\(\approx-0.3\) A1
Common mistakes
The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.
- Swapping the \(x\)- and \(y\)-lists. Entering the wrong variable into the wrong list flips the sign of the gradient and gives a regression line for the wrong direction entirely - always double check which quantity is \(x\) and which is \(y\) before running LinReg.
- Using rounded coefficients for later calculations. Rounding \(a\) and \(b\) to 3 significant figures before using them in a prediction or residual compounds the rounding error - where possible, keep full GDC precision until the final step.
- Getting the sign of a residual backwards. A residual is always actual minus predicted, not the other way round - a positive residual means the actual value was above the line, not below it.
- Forgetting DiagnosticOn (TI-84). Without turning DiagnosticOn, the TI-84's LinReg output won't display \(r\) at all, only \(a\) and \(b\) - an easy thing to miss under exam pressure.
Using your GDC
Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.
Core AI skill - gives the regression line and the correlation coefficient in one step.
- Enter \(x\)-values in one list and \(y\)-values in another.
- Turn on DiagnosticOn once (2nd → CATALOG) so \(r\) appears; then STAT → CALC → 4:LinReg(ax+b), set Xlist and Ylist, Calculate.TI-84
- Calculator page → menu → Statistics → Stat Calculations → Linear Regression (mx+b).Nspire
- Statistics menu → CALC (F2) → REG (F3) → X (linear); \(r\) shows automatically.Casio
- Read \(a\) (gradient), \(b\) (intercept), \(r\) (correlation) and \(r^2\) (coefficient of determination).
- Use the equation to predict - but only within the data range (interpolation).
Tip: \(r\) near \(\pm1\) means a strong linear fit; near 0 means weak. \(r^2\) is the proportion of variation explained.
Run this separately on your \(x\)-list and \(y\)-list to get \(\bar x\) and \(\bar y\) for the mean point check.
- Enter the data into a list.
- STAT → Edit → type values into L1. Then STAT → CALC → 1:1-Var Stats, choose L1, Calculate.TI-84
- Add a Lists & Spreadsheet page, name a column and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
- Statistics menu → enter data in List 1 → CALC (F2) → 1-VAR.Casio
- Read \(\bar x\) (mean) for each list, then substitute \(\bar x\) into your regression line and check it returns \(\bar y\).
Tip: If substituting \(\bar x\) doesn't return \(\bar y\) (to rounding), re-check the coefficients you copied down from the LinReg screen.
See the full GDC guide for more calculator models and topics.
Ready to practise properly?
Calculating correlation & regression questions, marked instantly like the real exam.
Quick answers
The questions students on this topic ask most often.
What's the difference between this topic and Correlation & Regression?
Correlation & Regression focuses on interpreting \(r\) and a regression line you're given. This topic is about the calculation itself - entering raw bivariate data into your GDC and computing \(r\), the regression line, the mean point, and residuals from scratch.
What is a residual?
A residual is the difference between an actual data value and the value the regression line predicts for the same \(x\): residual = actual \(y\) - predicted \(y\). A small residual means the line fits that point well; a large one means it doesn't.
Why does the regression line always pass through the mean point?
The least-squares method that generates the regression line is constructed so the positive and negative residuals balance out around the average \(x\) and \(y\) values - this makes the point \((\bar x,\bar y)\) always sit exactly on the line, which is a useful check on your GDC output.
Do I need to enter data by hand every time?
Yes - unlike some other topics, calculating correlation and regression starts from raw data pairs, so you always type the \(x\)-values into one list and the \(y\)-values into another before running the regression, rather than being handed the equation directly.
Related topics
More Statistics & Probability topics from the same AI SL syllabus unit, in case you want to keep going.