Descriptive Statistics & Correlation (AA SL)
Descriptive statistics summarise a data set with a handful of numbers - where it's centred, how spread out it is, and whether there's a linear relationship with a second variable. This topic covers the mean, median, mode, quartiles and standard deviation, reading and comparing box and whisker diagrams, and using Pearson's correlation coefficient and the regression line to describe bivariate data. Almost everything here is calculated on your GDC rather than by hand.
What the syllabus says
This topic maps onto three points in the official IB Analysis & Approaches syllabus.
| Code | Syllabus content |
|---|---|
| SL4.2 | Presentation of data (discrete and continuous): frequency distributions, histograms, cumulative frequency graphs. Production and understanding of box and whisker diagrams, including comparing two distributions and identifying outliers. |
| SL4.3 | Measures of central tendency (mean, median, mode) and estimation of the mean from grouped data. Modal class. Measures of dispersion (interquartile range, standard deviation and variance). Effect of constant changes on the original data. |
| SL4.4 | Linear correlation of bivariate data. Pearson's product-moment correlation coefficient, \(r\). Scatter diagrams; lines of best fit. Equation of the regression line of \(y\) on \(x\), found using technology, used for prediction and interpreting the parameters \(a\) and \(b\) in \(y=ax+b\), with awareness of the dangers of extrapolation and that a prediction of \(x\) from a value of \(y\) cannot always be made reliably using a \(y\) on \(x\) line. |
These are core AA syllabus points, examined mainly on Paper 2 where a calculator is allowed.
Key terms
Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.
What is the interquartile range?
The interquartile range (IQR) is the spread of the middle 50% of a data set, found by subtracting the lower quartile from the upper quartile. It's a measure of dispersion that isn't affected by extreme values, unlike the range.
e.g. If \(Q_1=18\) and \(Q_3=27\), then \(\text{IQR}=27-18=9\).
What is standard deviation?
Standard deviation measures how spread out a data set is around its mean - a small standard deviation means the values are clustered close to the mean, while a large one means they're spread widely. It's found using technology from the raw data.
e.g. A class with mean score 65 and standard deviation 4 has scores clustered close to 65; one with standard deviation 15 does not.
What is an outlier?
An outlier is a data value that lies unusually far from the rest of the data set. On the AA syllabus, a value is treated as a potential outlier if it lies more than \(1.5\times\text{IQR}\) beyond the nearer quartile.
e.g. With \(Q_1=18\), \(Q_3=27\), \(\text{IQR}=9\), the upper boundary is \(27+1.5(9)=40.5\), so any value above 40.5 is a potential outlier.
What is Pearson's correlation coefficient?
Pearson's product-moment correlation coefficient \(r\) measures the strength and direction of a linear relationship between two variables. Values close to \(+1\) or \(-1\) indicate a strong linear relationship; values close to \(0\) indicate little or no linear relationship.
e.g. \(r=0.95\) indicates a strong positive linear correlation; \(r=-0.1\) indicates almost none.
What is the regression line?
The regression line of \(y\) on \(x\) is the line of best fit found using technology, used to estimate \(y\) for a given value of \(x\). It always passes through the mean point \((\bar{x},\bar{y})\) of the data.
e.g. With \(y=2.4x+1.5\), at \(x=6\): \(y=2.4(6)+1.5=14.4+1.5=15.9\).
Key formulas
The core measures below are all computed with technology on the exam, but you still need to know what each one means and how to interpret it. The two tables summarise them at a glance - the explanations underneath go into more depth.
Formula reference
Standard deviation and Pearson's \(r\) are not given as standalone formulas in the booklet - both are found using your GDC directly from the data.
| Quantity | Used for | Booklet? |
|---|---|---|
| \(\text{IQR}=Q_3-Q_1\) | Spread of the middle 50% | Not in booklet - definition |
| Outlier boundary: \(Q_1-1.5\,\text{IQR}\), \(Q_3+1.5\,\text{IQR}\) | Identifying potential outliers | Not in booklet - convention |
| \(\bar x=\dfrac{\sum fx}{\sum f}\) | Mean of grouped/frequency data | Not in booklet - use technology |
| Pearson's \(r\) | Strength/direction of linear correlation | Not in booklet - found using GDC |
| \(y=mx+c\) (regression line) | Estimating \(y\) from \(x\) | Not in booklet - found using GDC |
Centre vs spread
A full description of a data set needs one measure of where it's centred and one measure of how spread out it is - this table lines the common choices up side by side.
| Feature | Measure of centre | Measure of spread |
|---|---|---|
| Common choice | Mean or median | Standard deviation or IQR |
| Affected by outliers? | Mean: yes. Median: no | Standard deviation: yes. IQR: no |
| Needs full data set? | Mean: yes. Median: no (just order) | Both need the full data set |
| Read from a box plot? | Median: yes (the line in the box) | IQR: yes (the width of the box) |
Reading and comparing distributions
These are the three main statistical skills tested on this topic beyond straightforward calculation.
Box and whisker diagrams
A box plot displays the five-number summary: minimum, \(Q_1\), median, \(Q_3\), maximum. The box spans the IQR, and the whiskers extend to the most extreme values that aren't outliers.
Outliers are plotted as separate crosses or dots, not included in the whisker.
Comparing two data sets
To compare two distributions, comment on both centre (which median is higher) and spread (which IQR or range is larger), always linking back to the context of the question.
A comparison needs both a numerical statement and a written comment - neither alone earns full marks.
Correlation and causation
A strong correlation coefficient shows two variables move together, but it never proves that one causes the other - a third factor could be driving both.
Using a regression line far outside the data range (extrapolation) makes any estimate unreliable, since the linear trend isn't guaranteed to continue.
Worked examples
Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.
For the data set \(4, 7, 8, 11, 13, 15, 18\):
(a) Find the median.
(b)(i) Find the lower quartile.
(b)(ii) Find the upper quartile.
(c) Find the interquartile range.
Worked solution
(a) Median. With \(n=7\) ordered values the median is the \(\tfrac{7+1}{2}=4\)th value, so \(Q_2=11.\) A1
(b)(i) Lower half \(4,7,8\Rightarrow Q_1=7.\) A1
(b)(ii) Upper half \(13,15,18\Rightarrow Q_3=15.\) A1
(c) IQR. The interquartile range is the spread of the middle 50%: \(\text{IQR}=Q_3-Q_1=15-7=8.\) A1
Bivariate data gives \(r = 0.95\) and regression line \(y = 2.4x + 1.5\).
(a) Describe the correlation.
(b) Estimate \(y\) when \(x = 6\).
(c) Comment on the reliability of estimating \(y\) when \(x = 50\).
Worked solution
(a) Describe the correlation. \(r=0.95\) is very close to \(+1,\) so there is a strong positive linear correlation: as \(x\) increases, \(y\) tends to increase. A1 R1
(b) Estimate \(y\) at \(x=6\). Substitute into the regression line: \(y=2.4(6)+1.5=14.4+1.5\) M1
\(=15.9.\) A1
(c) Reliability at \(x=50\). \(x=50\) lies far outside the range of the data, so using the line there is extrapolation; the linear trend need not continue, making the estimate unreliable. A1R1
Common mistakes
The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.
- Confusing standard deviation with variance. Variance is the square of the standard deviation - if your GDC gives \(\sigma_x=4\), the variance is \(16\), not \(4\).
- Using \(S_x\) instead of \(\sigma_x\) (or the reverse). \(S_x\) is the sample standard deviation and \(\sigma_x\) is the population standard deviation - the IB usually wants \(\sigma_x\) unless the question specifically describes a sample drawn from a larger population.
- Claiming correlation proves causation. A strong \(r\) value shows two variables move together; it never establishes that one causes the other, since a third factor could explain both.
- Extrapolating a regression line beyond the data range. Using \(y=2.4x+1.5\) to estimate \(y\) at \(x=50\) when the data only covers \(x=1\) to \(9\) produces an estimate with no guarantee of validity - the linear pattern might not continue.
Using your GDC
Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.
Instant summary statistics from a list - no formulas to compute by hand.
- Enter the data into a list.
- STAT → Edit → type values into L1. Then STAT → CALC → 1:1-Var Stats, choose L1, Calculate.TI-84
- Add a Lists & Spreadsheet page, name a column and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
- Statistics menu → enter data in List 1 → CALC (F2) → 1-VAR.Casio
- Read \(\bar x\) (mean), \(S_x\) (sample sd) or \(\sigma_x\) (population sd), and the five-number summary (min, \(Q_1\), median, \(Q_3\), max).
- For frequency data, put values in one list and frequencies in another and set the frequency list.
Tip: \(S_x\) vs \(\sigma_x\): use \(\sigma_x\) (population) for a complete data set, \(S_x\) (sample) for a sample. IB usually wants \(\sigma_x\).
Lists let you store data sets, generate sequences, and compute statistics without re-entering values - essential for grouped data and frequency tables.
- Enter data manually: go to the list editor and type values one by one.
- 2nd → STAT (LIST) → OPS → 5:seq(expression, variable, start, end) generates a list from a formula, e.g. seq(X², X, 1, 10). For grouped data, put midpoints in L1 and frequencies in L2, then run 1-Var Stats L1, L2.TI-84
- In a Lists & Spreadsheet page, type a formula in the header row (e.g. = seq(x², x, 1, 10)) to fill the column automatically. Use two columns for midpoints and frequencies.Nspire
- In Statistics list editor, enter data directly; OPTN → LIST → Seq( builds a list from a formula. For frequency data, put midpoints in List 1 and frequencies in List 2.Casio
- For frequency data (grouped), enter midpoints in one list and frequencies in another, then reference both lists when running 1-Var Stats.
Tip: For cumulative frequency or working out \(\sum fx\) by hand - don't. Store midpoints in L1, frequencies in L2, and let 1-Var Stats do it all.
See the full GDC guide for more calculator models and topics.
Ready to practise properly?
Descriptive statistics & correlation questions, marked instantly like the real exam.
Quick answers
The questions students on this topic ask most often.
What's the difference between the mean and the median?
The mean is the arithmetic average of all the data values, so it's pulled around by extreme values. The median is the middle value when the data is ordered, so it's a more reliable measure of centre when there are outliers or skewed data.
How do I know if there's an outlier?
Use the \(1.5\times\text{IQR}\) rule: a value is a potential outlier if it lies more than \(1.5\times\text{IQR}\) below \(Q_1\) or above \(Q_3\). Work out the two boundaries first, then check whether any data value falls outside them.
What does a correlation coefficient of \(r = 0.91\) actually tell me?
It tells you the linear relationship between the two variables is strong and positive - as one variable increases, the other tends to increase too, with the points lying close to a straight line. It says nothing about whether one variable causes the other.
Can I use my GDC for this topic?
Yes, on both papers where a calculator is allowed - your GDC computes the mean, standard deviation, quartiles, correlation coefficient, and regression line directly from a list of data, which is exactly what most exam questions expect you to do. See the GDC guide for model-specific instructions.
Sub-topics
Descriptive Statistics & Correlation broken down into its individual skills, each with its own focused page.
Related topics
More Statistics & Probability topics from the same AA SL syllabus unit, in case you want to keep going.