Non-linear Regression (AI HL)
Not every relationship between two variables is a straight line. When a scatter diagram curves, technology can fit a quadratic, cubic, exponential, power or sine model instead, and the fit is judged using \(R^2\) rather than Pearson's \(r\). This page covers picking a sensible model shape, reading off the coefficients your GDC gives you, and the traps in comparing models by \(R^2\) alone. It's part of the broader Bivariate & Non-linear Regression topic.
19 questions on this sub-topic.
Choosing a model
Covered under IB syllabus reference AHL4.13: evaluating least-squares regression curves (linear, quadratic, cubic, exponential, power and sine) using technology, and using the coefficient of determination \(R^2\) to gauge fit while remembering it isn't the whole story.
Quadratic / cubic
\(y=ax^2+bx+c\)
Choose these when the scatter diagram turns once (quadratic) or twice (cubic) - a single peak or trough, or an S-shaped curve.
Exponential
\(y=ab^x\)
Use for growth or decay that speeds up (or slows towards zero) at a constant percentage rate rather than a constant amount.
Power
\(y=ax^b\)
Use when the rate of change itself scales with \(x\), for example area against side length, or wind resistance against speed.
None of these regression equations appear in the formula booklet - your GDC's STAT/CALC menu generates the coefficients for whichever model you select. Need the full syllabus wording? See Bivariate & Non-linear Regression.
Worked examples
The table shows distance \(d\) (km) braked in time \(t\) (s).
| \(t\) | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|
| \(d\) | 1.2 | 4.5 | 9.1 | 14.8 | 21.0 |
Use technology to fit a cubic model \(d = at^3 + bt^2 + ct + e\).
(a) State \(a,\) \(b,\) \(c\) and \(e.\)
(b) State \(R^2.\)
Worked solution
(a) Enter data and run cubic regression. M1
Output: \(a \approx 0.015,\ b \approx 0.108,\ c \approx 0.867,\ e \approx 0.210.\) A1
(b) \(R^2 \approx 1.000\) (excellent fit). A1
GDC: STAT → CALC → CubicReg with DiagnosticOn for \(R^2.\)
Two models are fitted to the same data: a linear model with \(R^2=0.78\) and a quadratic with \(R^2=0.95\).
(a) State which model fits better and why.
(b) Explain a danger of always choosing the model with the higher \(R^2\).
Worked solution
(a) The quadratic, because a higher \(R^2\) means it explains more of the variation. M1
The quadratic. A1
(b) Adding parameters can artificially raise \(R^2\) M1
and cause overfitting - fitting noise A1
and predicting poorly outside the data. A1
Common mistakes
- Quoting \(r\) for a non-linear model. Pearson's \(r\) only measures linear correlation - for quadratic, exponential or power models, compare fits using \(R^2\) instead.
- Treating the highest \(R^2\) as automatically "correct". A model with more parameters (cubic over quadratic, say) can chase individual data points rather than the underlying trend - always sanity-check the shape against the scatter diagram too.
- Extrapolating a curved model far beyond the data. An exponential or power curve can shoot off to unrealistic values outside the range it was fitted on, even when \(R^2\) inside that range is excellent.
Ready to practise properly?
15 non-linear regression questions, marked instantly like the real exam.
Quick answers
What does R-squared tell you about a non-linear model?
\(R^2\) gives the proportion of the variability in the data accounted for by the chosen model. A value close to 1 means a close fit, but a high \(R^2\) alone doesn't guarantee the model is the right choice - see using your GDC for how to check it.
Which non-linear regression models can I fit on my GDC?
Most IB-approved GDCs fit quadratic, cubic, exponential, power and sine regression curves from the STAT menu, alongside the usual linear regression option.