Bivariate & Non-linear Regression (AI HL)

Once you can see a relationship between two variables in a scatter diagram, regression lets you fit a model to it and use that model to predict values, interpret the real-world meaning of its parameters, and judge how good a fit it actually is. This topic covers linear regression and its interpretation, non-linear models (quadratic, exponential, power and sine), and using \(R^2\) to compare how well different models fit the same data - all found using your GDC.

What the syllabus says

This topic maps onto two points in the official IB Applications & Interpretation syllabus.

CodeSyllabus content
SL4.4Linear correlation of bivariate data and Pearson's product-moment correlation coefficient \(r\), found using technology - meaningful only for linear relationships. Scatter diagrams with lines of best fit by eye through the mean point, the equation of the regression line of \(y\) on \(x\), and interpreting the parameters \(a\) and \(b\) in \(y=ax+b\) in context, including the dangers of extrapolation.
AHL4.13Non-linear regression: evaluating least squares regression curves (linear, quadratic, cubic, exponential, power and sine) using technology. The coefficient of determination \(R^2\), which gives the proportion of variability accounted for by the chosen model, and the awareness that \(R^2\) alone is not a good way to decide between models.

SL4.4 is core AI syllabus content examined at both SL and HL; AHL4.13 extends regression to non-linear models at HL.

Key terms

Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.

What is Pearson's \(r\)?

Pearson's product-moment correlation coefficient \(r\) measures the strength and direction of a linear relationship between two variables, from \(-1\) (perfect negative) to \(+1\) (perfect positive). A value near 0 means little or no linear relationship.

e.g. \(r=-0.87\) is a strong, negative linear correlation.

What is the regression line?

The regression line of \(y\) on \(x\), \(y=ax+b\), is the line of best fit found by minimising the sum of squared vertical distances from the data points. Your GDC generates it directly from the raw data.

e.g. For hours studied vs test score: \(y\approx4.23x+47.5\).

What is \(R^2\)?

The coefficient of determination \(R^2\) is the proportion of the variation in \(y\) explained by the model, as a value between 0 and 1. It applies to any regression model, not just linear ones, which is why it's used to compare non-linear fits.

e.g. \(R^2=0.96\) means 96% of the variation in \(y\) is explained by the model.

What is extrapolation?

Extrapolation means using a regression model to predict a \(y\)-value for an \(x\)-value outside the range of the original data. It's unreliable because there's no evidence the same pattern continues beyond what was actually measured.

e.g. Data spans \(15\)-\(30°\)C; predicting sales at \(5°\)C is unreliable extrapolation.

What is a non-linear model?

A non-linear model fits a curve rather than a straight line - quadratic, exponential, power or sine - when a scatter diagram shows the data doesn't follow a linear pattern. The IB GDC menu offers each of these as a regression option.

e.g. Power model \(y=2.4x^{1.6}\): at \(x=5\), \(y=2.4(5)^{1.6}\approx31.5\).

Key formulas

Six formulas and forms cover almost every question on this topic. The two tables below summarise all of them at a glance - the explanations underneath go into more depth on each one.

Formula reference

None of these are listed separately in the formula booklet - the regression equation, \(r\) and \(R^2\) are all generated directly by your GDC's statistics menu rather than substituted by hand.

FormUsed forBooklet?
\(y=ax+b\)Linear regression line of \(y\) on \(x\)Not in booklet
\(y=ax^2+bx+c\)Quadratic regressionNot in booklet
\(y=ab^x\)Exponential regressionNot in booklet
\(y=ax^b\)Power regressionNot in booklet
\(-1\le r\le1\)Pearson's correlation coefficient (linear only)Not in booklet
\(0\le R^2\le1\)Coefficient of determination (any model)Not in booklet

Linear vs non-linear regression

The choice of model should always start from what the scatter diagram looks like, not just from which gives the highest \(R^2\).

FeatureLinear regressionNon-linear regression
Shape fittedStraight lineQuadratic, exponential, power or sine curve
Strength measure\(r\) (only meaningful for lines)\(R^2\) (works for any model)
Parameter meaning\(a\) = gradient, \(b\) = interceptDepends on the model - not always a rate
GDC menuLinReg(ax+b)QuadReg, ExpReg, PwrReg, SinReg

Interpreting a linear model

A regression equation is only useful once you can explain what its parameters mean in the context of the question.

The gradient

In \(y=ax+b\), \(a\) is the change in \(y\) for every one-unit increase in \(x\) - always state it in context and in the original units.

Not in the formula booklet - prior knowledge

The intercept

\(b\) is the predicted value of \(y\) when \(x=0\) - sometimes a genuine "fixed" starting value (like a base charge), sometimes not meaningful if \(x=0\) is outside the data range.

Not in the formula booklet - prior knowledge

Prediction

Substitute the given \(x\) into the regression equation. Reliable only for interpolation - \(x\)-values within the range of the original data.

Not in the formula booklet - prior knowledge

Choosing and evaluating a model

With several regression options available on the GDC, \(R^2\) is the main tool for comparing them - but it isn't the only thing that matters.

Comparing models with \(R^2\)

Fit each candidate model and compare \(R^2\) values - closer to 1 means more variation explained. An exponential model with \(R^2=0.98\) beats a linear model with \(R^2=0.91\) on the same data.

Not in the formula booklet - prior knowledge

\(R^2\) isn't everything

A high \(R^2\) on a model that doesn't suit the context (e.g. a sine curve for steadily growing data) is meaningless - always check the scatter diagram's shape matches the model chosen.

Not in the formula booklet - prior knowledge

Linearising an exponential model

Taking \(\ln\) of \(y=ab^x\) gives \(\ln y = (\ln b)x + \ln a\) - a straight line in \(\ln y\) vs \(x\), whose gradient is \(\ln b\) and intercept is \(\ln a\).

Not in the formula booklet - prior knowledge

Worked examples

Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.

1
Easy
GDC
[4 marks]

A regression line for monthly electricity cost \(y\) ($) against units used \(x\) is \(y = 0.18x + 12.\)

(a) Interpret the \(y\)-intercept.

(b) Interpret the gradient.

(c) Predict the cost for \(50\) units.

Worked solution

(a) The fixed monthly charge occurs when no units are used, i.e. $12. R1

(b) Each additional unit costs $0.18. R1

🖩 Regression on the GDC - TI‑84 STAT▸CALC▸LinReg/ExpReg/PwrReg · Casio STAT▸CALC · Nspire Stat Calculations; turn DiagnosticOn for \(r\).

(c) \(y=0.18(50)+12\) M1
\(=$21.\) A1

R1 Intercept in context, $12 fixed R1 Gradient in context, $0.18/unit M1 Substitute \(x=50\) A1 $21
2
Hard
GDC
[8 marks]

Temperature \(x\) (°C) and ice cream sales \(y\) ($100s) are recorded:

\(x\)151820242830
\(y\)46791214

(a) Find \(r\) and describe the correlation.

(b) Find the regression line of \(y\) on \(x\).

(c) Predict sales when the temperature is 22°C.

(d) State whether predicting sales at 5°C is reliable.

Worked solution

(a) \(r \approx 0.995\) - a very strong positive linear correlation. M1 A1

(b) \(y \approx 0.641x - 5.76.\) M1 A1

(c) \(y = 0.641(22) - 5.76\) M1
\(\approx 8.35\) ($835). A1

(d) No - 5°C is far below the data range (15–30°C): unreliable extrapolation. M1 A1

M1 GDC \(r\) A1 Strong positive M1 LinReg A1 Equation M1 Substitute A1 Correct answer of \(\approx8.35\) M1 Outside range A1 Unreliable

Common mistakes

The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.

  • Extrapolating without comment. Predicting for an \(x\)-value outside the data range is unreliable - always check the data range before trusting (or defending) a prediction.
  • Quoting \(r\) for a non-linear model. Pearson's \(r\) only measures linear correlation - for quadratic, exponential or power models, compare fits using \(R^2\) instead.
  • Interpreting the gradient without units or context. "\(a=4.23\)" alone earns little - state what it means, e.g. "each extra hour of study is associated with about a 4.23-point increase in score".
  • Confusing correlation with causation. A strong \(r\) or \(R^2\) shows the two variables move together, not that one causes the other - a third factor could be driving both.

Using your GDC

Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.

Show steps for:
Linear regression (line of best fit) and r

Core AI skill - gives the regression line and the correlation coefficient in one step.

  1. Enter \(x\)-values in one list and \(y\)-values in another.
  2. Turn on DiagnosticOn once (2nd → CATALOG) so \(r\) appears; then STAT → CALC → 4:LinReg(ax+b), set Xlist and Ylist, Calculate.TI-84
  3. Calculator page → menu → Statistics → Stat Calculations → Linear Regression (mx+b).Nspire
  4. Statistics menu → CALC (F2) → REG (F3) → X (linear); \(r\) shows automatically.Casio
  5. Read \(a\) (gradient), \(b\) (intercept), \(r\) (correlation) and \(r^2\) (coefficient of determination).
  6. Use the equation to predict - but only within the data range (interpolation).

Tip: \(r\) near \(\pm1\) means a strong linear fit; near 0 means weak. \(r^2\) is the proportion of variation explained.

Compare regression models using R²

After fitting several models (linear, quadratic, exponential...), you need to decide which fits the data best - \(R^2\) is the key tool.

  1. Fit each candidate model in turn and note the \(R^2\) value each time.
  2. Turn DiagnosticOn first (2nd → 0, scroll to DiagnosticOn, ENTER) - then \(R^2\) appears after every regression.TI-84
  3. \(R^2\) is shown automatically after each regression calculation in the Statistics menu.Nspire
  4. \(R^2\) (displayed as \(r^2\)) appears in the regression output; run CALC → REG for each model type and compare.Casio
  5. The model with \(R^2\) closest to 1 explains the most variation in \(y\) - but also consider whether the model makes sense for the context.
  6. An exponential model with \(R^2 = 0.98\) is better than a linear model with \(R^2 = 0.91\) for the same data.

Tip: \(R^2\) alone doesn't tell you whether the model is appropriate - always look at the scatter plot too. A high \(R^2\) on a model that shouldn't apply (e.g. sinusoidal for steadily growing data) is meaningless.

See the full GDC guide for more calculator models and topics.

Ready to practise properly?

Bivariate & non-linear regression questions, marked instantly like the real exam.

Quick answers

The questions students on this topic ask most often.

What's the difference between linear and non-linear regression?

Linear regression fits a straight line, \(y=ax+b\), and is measured by Pearson's \(r\). Non-linear regression fits a curve - quadratic, exponential, power or sine - to data that doesn't follow a straight-line pattern, and is measured by \(R^2\) instead, since \(r\) only applies to linear relationships.

What does R-squared actually tell me?

\(R^2\) is the proportion of the variation in \(y\) that is explained by the model, as a value between 0 and 1. An \(R^2\) of \(0.96\) means 96% of the variation is explained by the model - close to 1 is a very good fit, but a high \(R^2\) alone doesn't prove the model type is the right one.

Why is extrapolation dangerous?

A regression model is only fitted to, and only reliable within, the range of \(x\)-values in your data. Predicting far outside that range assumes the same pattern continues, which often isn't true - always check whether the \(x\)-value you're predicting for falls inside the original data range.

Can I use my GDC for this topic?

Yes, and you're expected to. Every regression equation, correlation coefficient and \(R^2\) value in the exam is generated by your GDC's regression menu - you're never expected to compute these from raw data by hand. See the GDC guide for model-specific instructions.

Sub-topics

Bivariate & Non-linear Regression broken down into its individual skills, each with its own focused page.