Regression Line (AI SL)
Once bivariate data shows a linear trend, the regression line of \(y\) on \(x\) lets you predict a \(y\)-value for a given \(x\) - found on your GDC, never derived by hand. This page focuses on reading and using that equation, including the difference between a safe prediction and an unreliable one. It's part of the broader Correlation & Regression topic.
26 questions on this sub-topic.
The regression equation
Covered under IB syllabus reference SL4.4: linear correlation of bivariate data, the equation of the regression line of \(y\) on \(x\) found using technology, and awareness that extrapolating beyond the data is unsafe.
Regression line of \(y\) on \(x\)
\(y=ax+b\)
Not in the formula booklet - your GDC's LinReg function returns \(a\) and \(b\) directly from raw data. Use it to predict \(y\) for a given \(x\), never the reverse.
Check on a computed line
Line passes through \((\bar x,\bar y)\)
Not in the booklet either - it's a property of the least-squares line. If you're given the mean point and part of the equation, this lets you solve for a missing coefficient.
Need the full syllabus wording and formula-booklet reference table? See Correlation & Regression. For calculator steps, see the parent topic's GDC guidance.
Worked examples
Data for \(x\) ranges from 2 to 12, with regression line \(y=3x+5.\)
(a) Predict \(y\) at \(x=7\) and state whether this is reliable.
(b) Predict \(y\) at \(x=20\) and state whether this is reliable.
Worked solution
(a) At \(x=7.\) \(y=3(7)+5=26.\) A1
Since \(7\) lies inside the data range \([2,12],\) this is interpolation - reliable. R1
(b) At \(x=20.\) \(y=3(20)+5=65.\) A1
But \(20\) is well beyond the largest data value (12), so this is extrapolation - unreliable, R1
because the linear pattern may not continue outside the observed range. R1
A regression line \(y=2.5x+1\) passes through the mean point \((\bar x,\bar y),\) where \(\bar x=6.\) Find \(\bar y.\)
Worked solution
Key fact. The least-squares regression line always passes through the mean point A1
\((\bar x,\bar y).\) R1
Step - substitute \(\bar x=6:\) M1
\(\bar y=2.5(6)+1=16.\) A1
Common mistakes
- Extrapolating without comment. Using the regression line to predict far outside the data range and treating the answer as equally reliable as an interpolated one - always flag extrapolation as less trustworthy.
- Predicting \(x\) from \(y\) using the \(y\) on \(x\) line. The \(y\) on \(x\) regression line is built to predict \(y\), not to be rearranged to predict \(x\) - that needs the separate \(x\) on \(y\) line.
- Forgetting the mean-point property. When a question gives \(\bar x\) or \(\bar y\) alongside a partial equation, missing that the line must pass through \((\bar x,\bar y)\) loses an easy substitution mark.
Ready to practise properly?
28 regression-line questions, marked instantly like the real exam.
Quick answers
What does the regression line of y on x predict?
\(y=ax+b\), found by technology, predicts \(y\) from a given \(x\) value. It cannot be rearranged to predict \(x\) from \(y\) - that needs a separate \(x\) on \(y\) line.
Why is extrapolation unreliable?
The regression line is only fitted to the range of \(x\)-values in the data. Outside that range the linear pattern may not continue, so predictions there (extrapolation) are far less trustworthy than predictions inside the range (interpolation).