Bivariate Statistics (AA HL)
Bivariate statistics looks at the relationship between two variables measured on the same subjects - how closely they move together, and how to model that relationship with a straight line. This topic covers scatter diagrams, Pearson's correlation coefficient \(r\), the regression line of \(y\) on \(x\), and how to interpret \(r\) and \(r^2\) without falling into the trap of confusing correlation with causation.
What the syllabus says
This topic maps onto two points in the official IB Analysis & Approaches syllabus.
| Code | Syllabus content |
|---|---|
| SL4.4 | Linear correlation of bivariate data; Pearson's product-moment correlation coefficient \(r\), found using technology. Scatter diagrams; lines of best fit by eye, through the mean point. Distinction between correlation and causation. Equation of the regression line of \(y\) on \(x\), and use for prediction, including the dangers of extrapolation. Interpreting the parameters \(a\) and \(b\) in \(y=ax+b\). |
| SL4.10 | Equation of the regression line of \(x\) on \(y\), and its use for prediction. Awareness that a prediction of \(y\) from \(x\) cannot always reliably be made using an \(x\) on \(y\) line. |
These are core AA SL syllabus points that are also examinable at AA HL.
Key terms
Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.
What is bivariate data?
Bivariate data is a set of paired measurements taken on the same subjects, such as each person's height and weight. It's studied using a scatter diagram, which plots each pair as a point to reveal whether a pattern or relationship exists.
e.g. Recording each day's temperature \(x\) alongside that day's ice-cream sales \(y\) produces a set of paired \((x,y)\) points.
What is Pearson's correlation coefficient, r?
\(r\) measures the strength and direction of a linear relationship between two variables, always between \(-1\) and \(1\). Values near \(\pm1\) mean a strong linear relationship; values near \(0\) mean little to no linear relationship. It's only meaningful when the underlying relationship is roughly linear.
e.g. \(r=-0.92\) between hours of screen time and hours of sleep indicates a strong negative linear relationship.
What is the regression line of y on x?
The regression line of \(y\) on \(x\), \(y=ax+b\), is the line of best fit that minimises vertical distances from the data points, always passing through the mean point \((\bar x,\bar y)\). It's used to predict \(y\) from a given \(x\).
e.g. If \(y=2.5x+8\), then at \(x=4\): \(y=2.5(4)+8=10+8=18\).
What is r² (coefficient of determination)?
\(r^2\) is the square of the correlation coefficient, giving the proportion of the variation in \(y\) that's explained by the linear relationship with \(x\). It's usually expressed as a percentage, and it's always between 0 and 1.
e.g. \(r=0.7 \Rightarrow r^2 = 0.49\), so 49% of the variation in \(y\) is explained by the linear relationship.
What's the difference between interpolation and extrapolation?
Interpolation is predicting a value within the range of the collected data, which is generally reliable. Extrapolation is predicting outside that range, which is risky because there's no evidence the linear pattern continues beyond where data was actually gathered.
e.g. For data with \(x\) between \(5\) and \(20\), predicting at \(x=12\) is interpolation; predicting at \(x=50\) is extrapolation.
Key formulas
A handful of formulas cover this whole topic - the tables below summarise them, and the sections underneath explain each in more depth.
Formula reference
The correlation and regression formulas are on the official formula booklet; \(r^2\) itself is simply the square of \(r\) and isn't listed as a separate formula.
| Formula | Used for | Booklet? |
|---|---|---|
| \(S_{xy}=\dfrac{\sum xy}{n}-\bar x\bar y\) | Covariance-style sum used in \(r\) and \(b\) | ✓ Yes |
| \(S_{xx}=\dfrac{\sum x^2}{n}-\bar x^2\) | Spread of \(x\), used in \(r\) and \(b\) | ✓ Yes |
| \(r=\dfrac{S_{xy}}{\sqrt{S_{xx}S_{yy}}}\) | Pearson's correlation coefficient | ✓ Yes |
| \(y=ax+b,\ a=\dfrac{S_{xy}}{S_{xx}},\ b=\bar y-a\bar x\) | Regression line of \(y\) on \(x\) | ✓ Yes |
| \(r^2\) | Coefficient of determination | Not in booklet - simply \(r\) squared |
Regression line of y on x vs x on y
There are two different regression lines for the same data, and using the wrong one gives an unreliable prediction.
| Feature | y on x line | x on y line |
|---|---|---|
| Minimises | Vertical distances from the points | Horizontal distances from the points |
| Gradient | \(a=\dfrac{S_{xy}}{S_{xx}}\) | Uses \(S_{xy}\) and \(S_{yy}\) instead of \(S_{xx}\) |
| Use it to predict | \(y\), given a value of \(x\) | \(x\), given a value of \(y\) |
| Passes through | Both lines pass through the mean point \((\bar x,\bar y)\) - and coincide only when \(r=\pm1\). | |
Correlation
Correlation describes how closely two variables move together in a straight-line pattern - nothing more.
Pearson's r
\[r=\dfrac{S_{xy}}{\sqrt{S_{xx}S_{yy}}}\]
Ranges from \(-1\) (perfect negative) to \(1\) (perfect positive); only meaningful for linear relationships.
✓ In the formula bookletInterpreting r
The sign gives the direction (positive/negative); the size gives the strength - close to \(\pm1\) is strong, close to \(0\) is weak or absent.
Correlation vs causation
A strong \(r\) never proves one variable causes the other - a third factor could be driving both, or the link could be coincidental.
The regression line
The regression line summarises the linear trend and lets you predict one variable from the other, within limits.
Equation & parameters
\[y=ax+b\]
\(a\) is the gradient - the change in \(y\) per unit increase in \(x\). \(b\) is the \(y\)-intercept - the predicted \(y\) when \(x=0\).
✓ In the formula bookletInterpolation vs extrapolation
Predicting within the data range (interpolation) is generally safe; predicting outside it (extrapolation) is unreliable, since the pattern might not continue.
Worked examples
Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.
A regression of yield \(y\) on rainfall \(x\) gives \(r=-0.6.\)
(a) Find \(r^2.\)
(b) Interpret \(r^2\) in context.
(c) State the percentage of variation NOT explained by the linear relationship.
Worked solution
(a) \(r^2=(-0.6)^2\) M1
\(=0.36.\) A1
(b) About \(36\%\) of the variation in yield is explained by the linear relationship with rainfall. A1
(c) \(100\%-36\%=64\%.\) A1
Data \((x,y)\) have \(r=-0.8.\) New variables are \(u=3x\) and \(v=-2y.\)
(a) State the value of the correlation coefficient between \(u\) and \(v.\)
(b) Explain the effect of the negative scale factor.
(c) State the correlation coefficient if instead \(u=-3x\) (both variables now negatively scaled).
Worked solution
(a) Scaling by a positive factor leaves \(r\) unchanged; multiplying \(y\) by a negative factor reverses the sign. M1
So \(r_{uv}=+0.8.\) A1
(b) \(v=-2y\) reflects the relationship, turning a negative correlation into a positive one of the same strength. R1
(c) Two sign reversals cancel, so \(r_{uv}=-0.8\) (the original sign is restored). A1
Common mistakes
The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.
- Treating correlation as causation. A strong \(r\) shows two variables move together, not that one causes the other - a third factor, or coincidence, could explain the pattern.
- Extrapolating beyond the data range. Using the regression line to predict far outside the \(x\)-values that were actually collected is unreliable, since there's no evidence the linear trend continues.
- Using the y on x line to predict x from y. The two regression lines are generally different - predicting \(x\) from a given \(y\) needs the \(x\) on \(y\) line, not a rearrangement of the \(y\) on \(x\) line.
- Reporting r when the question asks for r². \(r\) measures strength and direction; \(r^2\) is the percentage of variation explained. Mixing these up is a common way to lose an accuracy mark.
Using your GDC
Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.
Gives the regression line and the correlation coefficient in one step - the core skill for this whole topic.
- Enter x-values in one list and y-values in another.
- Turn on DiagnosticOn once (2nd → CATALOG) so \(r\) appears; then STAT → CALC → 4:LinReg(ax+b), set Xlist and Ylist, Calculate.TI-84
- Calculator page → menu → Statistics → Stat Calculations → Linear Regression (mx+b).Nspire
- Statistics menu → CALC (F2) → REG (F3) → X (linear); \(r\) shows automatically.Casio
- Read \(a\) (gradient), \(b\) (intercept), \(r\) (correlation) and \(r^2\) (coefficient of determination).
- Use the equation to predict - but only within the data range (interpolation).
Tip: \(r\) near \(\pm1\) means a strong linear fit; near \(0\) means weak. \(r^2\) is the proportion of variation explained.
Lists let you store paired data sets without re-entering values - essential for setting up \(x\) and \(y\) columns before running a regression.
- Enter data manually: go to the list editor and type \(x\)-values and \(y\)-values into two separate lists.
- 2nd → STAT (LIST) → OPS lets you sort or clear a list. Use L1 for \(x\), L2 for \(y\), and reference both together in LinReg.TI-84
- In a Lists & Spreadsheet page, use two named columns for \(x\) and \(y\) so both are available together for regression.Nspire
- In Statistics list editor, enter \(x\)-values in List 1 and \(y\)-values in List 2 before running REG.Casio
- Keep the \(x\) and \(y\) lists the same length and in matching pairs - a mismatched row shifts every pair after it.
Tip: Double-check the lists are paired correctly before running a regression - a single mistyped value can shift \(r\) noticeably.
See the full GDC guide for more calculator models and topics.
Ready to practise properly?
Bivariate statistics questions, marked instantly like the real exam.
Quick answers
The questions students on this topic ask most often.
Does a strong correlation mean one variable causes the other?
No. Correlation only measures how closely two variables move together in a straight-line pattern - it says nothing about cause and effect. A third, unmeasured variable could be driving both, or the relationship could be coincidental.
Can I use the regression line to predict outside the data range?
Predicting within the range of the collected data (interpolation) is generally reliable. Predicting outside that range (extrapolation) is risky, because there's no evidence the same linear pattern continues beyond where you actually have data.
What's the difference between the regression line of y on x and of x on y?
The \(y\) on \(x\) line minimises vertical distances and should be used to predict \(y\) from a given \(x\). The \(x\) on \(y\) line minimises horizontal distances and should be used to predict \(x\) from a given \(y\). The two lines are different (unless \(r=\pm1\)) and shouldn't be used in the wrong direction.
Do I need to calculate r by hand?
No. Technology should be used to calculate \(r\) and the regression line equation, although a hand calculation can help you understand what the value represents. Critical values of \(r\), where needed, are given to you in the exam.
Sub-topics
Bivariate Statistics broken down into its individual skills, each with its own focused page.
Related topics
More Statistics & Probability topics from the same AA HL syllabus unit, in case you want to keep going.