Statistics & Sampling (AI HL)
Statistics turns a mass of raw numbers into a handful of numbers you can actually use - a typical value, a measure of how spread out the data is, and a way to spot the values that don't fit the pattern. This topic covers the measures of central tendency and dispersion, how to summarise and compare data sets with box-and-whisker diagrams, and the 1.5 IQR rule for identifying outliers. Almost everything here leans on your GDC's 1-Var Stats function rather than hand calculation.
What the syllabus says
This topic maps onto two points in the official IB Applications & Interpretation syllabus.
| Code | Syllabus content |
|---|---|
| SL4.2 | Presentation of discrete and continuous data (frequency distributions, histograms with equal class intervals). Cumulative frequency and cumulative frequency graphs, used to find the median, quartiles, percentiles, range and interquartile range (IQR). Production and interpretation of box-and-whisker diagrams, including comparing two distributions by symmetry, median, IQR or range, with outliers marked by a cross. |
| SL4.3 | Measures of central tendency (mean, median, mode), including estimating the mean of grouped data using mid-interval values and identifying the modal class. Measures of dispersion (interquartile range, standard deviation, variance) and quartiles of discrete data, calculated using technology, with awareness that different methods may give slightly different quartile values. The effect of constant changes (adding a constant, or scaling) on the mean and standard deviation of a data set. |
These are core AI SL syllabus points that are also examinable at AI HL.
Key terms
Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.
What is the mean?
The mean is the sum of all data values divided by the number of values - the everyday "average". It's sensitive to every value in the data set, including unusually large or small ones, which is why a single outlier can shift it noticeably.
e.g. For \(2,4,4,4,5,5,7,9\): mean \(= \dfrac{40}{8} = 5\).
What is the median?
The median is the middle value when the data is ordered - or the mean of the two middle values if there's an even number of items. Unlike the mean, it's resistant to outliers, which makes it the better choice for skewed data.
e.g. For \(3,5,7,9,11\): the median is \(7\), the middle value.
What is standard deviation?
Standard deviation measures how spread out the data is around the mean - a small value means the data clusters tightly, a large value means it's spread widely. It's the square root of the variance, and on the IB you calculate it using your GDC's 1-Var Stats.
e.g. For \(2,4,4,4,5,5,7,9\) (mean \(5\)): variance \(= \dfrac{32}{8}=4\), so \(\sigma = \sqrt4 = 2\).
What is interquartile range (IQR)?
The interquartile range is the width of the middle 50% of the data: \(\text{IQR} = Q_3 - Q_1\), where \(Q_1\) and \(Q_3\) are the lower and upper quartiles. It's a measure of spread that, unlike the range, ignores extreme values at either end.
e.g. If \(Q_1=24\) and \(Q_3=40\), then \(\text{IQR} = 40-24 = 16\).
What counts as an outlier?
A value is an outlier if it falls more than 1.5 times the IQR below \(Q_1\) or above \(Q_3\). On a box-and-whisker diagram, outliers are marked with a cross rather than being included in the whisker.
e.g. With \(Q_1=24, Q_3=40, \text{IQR}=16\): upper boundary \(= 40+1.5(16)=64\), so \(66\) is an outlier.
Key formulas
Six formulas cover almost every question on this topic. The two tables below summarise all of them at a glance - the explanations underneath go into more depth on each one.
Formula reference
None of these are listed separately in the formula booklet - they're treated as prior knowledge, and IB questions expect you to use your GDC's statistics mode rather than substitute by hand.
| Formula | Used for | Booklet? |
|---|---|---|
| \(\bar x = \dfrac{\Sigma x}{n}\) | Mean of ungrouped data | Not in booklet |
| \(\bar x = \dfrac{\Sigma fx}{\Sigma f}\) | Estimated mean of grouped data | Not in booklet |
| \(\sigma = \sqrt{\dfrac{\Sigma(x-\bar x)^2}{n}}\) | Population standard deviation | Not in booklet |
| \(\text{IQR} = Q_3 - Q_1\) | Interquartile range | Not in booklet |
| \(Q_1 - 1.5\,\text{IQR},\ \ Q_3 + 1.5\,\text{IQR}\) | Outlier boundaries | Not in booklet |
| New mean \(=a\bar x+b\); new sd \(=|a|\sigma\) | Effect of a linear data transformation | Not in booklet |
Population vs sample standard deviation
Your GDC's 1-Var Stats screen shows both \(\sigma_x\) and \(s_x\) side by side. Knowing which one the question wants matters.
| Feature | Population (\(\sigma_x\)) | Sample (\(s_x\)) |
|---|---|---|
| Divides sum of squares by | \(n\) | \(n-1\) |
| Use when | the data set is the entire population | the data is a sample used to estimate a wider population |
| Value | always the smaller of the two | always the larger of the two |
| IB default | \(\sigma_x\), unless the question says otherwise | only when explicitly told it's a sample |
Measures of central tendency
These three describe a "typical" value in the data set - each answers a slightly different question.
Mean
\[\bar x = \frac{\Sigma x}{n}\]
Sum every value and divide by how many there are. Uses all the data, so it's sensitive to outliers.
Not in the formula booklet - prior knowledgeMedian
The middle value once the data is ordered (or the mean of the two middle values for an even-sized set). Resistant to outliers and skew.
Not in the formula booklet - prior knowledgeMode
The most frequently occurring value. For grouped data, the modal class is the class interval with the highest frequency.
Not in the formula booklet - prior knowledgeMeasures of spread
These describe how tightly or widely the data is scattered around its centre.
Range
Maximum value minus minimum value. Simple, but distorted enormously by a single extreme value.
Not in the formula booklet - prior knowledgeInterquartile range
\[\text{IQR} = Q_3 - Q_1\]
The spread of the middle 50% of the data. Ignores the most extreme quarter at each end, so it's resistant to outliers.
Not in the formula booklet - prior knowledgeVariance
The mean of the squared deviations from the mean. Equal to standard deviation squared - found on the GDC, never expected by hand.
Not in the formula booklet - prior knowledgeStandard deviation
\[\sigma = \sqrt{\frac{\Sigma(x-\bar x)^2}{n}}\]
Square root of the variance - back in the original units of the data, which makes it easier to interpret than variance.
Not in the formula booklet - prior knowledgeIdentifying outliers
The 1.5 IQR rule gives an objective boundary for flagging unusual values, rather than relying on a judgement call.
1.5 x IQR rule
\[Q_1-1.5\,\text{IQR}, \quad Q_3+1.5\,\text{IQR}\]
Any value beyond either boundary is an outlier and is plotted separately with a cross on a box plot.
Not in the formula booklet - prior knowledgeEffect of an outlier
A single extreme value drags the mean toward it but barely moves the median - which is exactly why the median is preferred for skewed or outlier-heavy data.
Not in the formula booklet - prior knowledgeWorked examples
Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.
A data set has \(Q_1 = 24\), \(Q_3 = 40\).
(a) Find the interquartile range.
(b)(i) Find the lower outlier boundary.
(b)(ii) Find the upper outlier boundary.
(c) Determine whether the value 66 is an outlier.
Worked solution
(a) \(\text{IQR} = 40 - 24\) M1
\(= 16.\) A1
(b)(i) Lower \(= 24 - 1.5(16)\) M1
\(= 0.\) A1
(b)(ii) Upper \(= 40 + 1.5(16) = 64.\) A1
(c) \(66 > 64\), so 66 is an outlier. A1
Group A has 12 values with mean 18. Group B has 8 values with mean 23.
(a) Find the combined mean of all 20 values.
(b) The combined mean is then required to be exactly 21 by adding one extra value to the 20. Find that value.
Worked solution
(a) Total \(= 12(18) + 8(23)\) M1
\(= 400.\) A1
Mean \(= \dfrac{400}{20} = 20.\) A1
(b) \(\dfrac{400 + v}{21} = 21 \Rightarrow v\) M1
\(= 41.\) A1
Common mistakes
The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.
- Reading \(s_x\) instead of \(\sigma_x\) off the GDC. Your calculator's 1-Var Stats screen shows both - the IB almost always wants the population value \(\sigma_x\) unless the question explicitly calls the data a sample.
- Getting the outlier boundary backwards. The lower boundary is \(Q_1 - 1.5\,\text{IQR}\) and the upper boundary is \(Q_3 + 1.5\,\text{IQR}\) - subtracting from \(Q_3\) or adding to \(Q_1\) gives the wrong side entirely.
- Using class boundaries instead of mid-interval values for grouped means. \(\bar x = \Sigma fx/\Sigma f\) needs the midpoint of each class, not its lower or upper limit - using the wrong value skews the whole estimate.
- Assuming the mean is as resistant to outliers as the median. A single extreme value can shift the mean substantially while barely moving the median - always check which measure the question is actually asking for.
Using your GDC
Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.
Instant summary statistics from a list - no formulas to compute by hand.
- Enter the data into a list.
- STAT → Edit → type values into L1. Then STAT → CALC → 1:1-Var Stats, choose L1, Calculate.TI-84
- Add a Lists & Spreadsheet page, name a column and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
- Statistics menu → enter data in List 1 → CALC (F2) → 1-VAR.Casio
- Read \(\bar x\) (mean), \(S_x\) (sample sd) or \(\sigma_x\) (population sd), and the five-number summary (min, \(Q_1\), median, \(Q_3\), max).
- For frequency data, put values in one list and frequencies in another and set the frequency list.
Tip: \(S_x\) vs \(\sigma_x\): use \(\sigma_x\) (population) for a complete data set, \(S_x\) (sample) for a sample. IB usually wants \(\sigma_x\).
Lists let you store data sets, generate sequences, and compute statistics without re-entering values - essential for grouped data and frequency tables.
- Enter data manually: go to the list editor and type values one by one.
- 2nd → STAT (LIST) → OPS → 5:seq(expression, variable, start, end) generates a list from a formula. For grouped data, put midpoints in L1 and frequencies in L2, then run 1-Var Stats L1, L2.TI-84
- In a Lists & Spreadsheet page, type a formula in the header row to fill the column automatically. Use two columns for midpoints and frequencies.Nspire
- In Statistics list editor, enter data directly; OPTN → LIST → Seq( builds a list from a formula. For frequency data, put midpoints in List 1 and frequencies in List 2.Casio
- For frequency data (grouped), enter midpoints in one list and frequencies in another, then reference both lists when running 1-Var Stats.
Tip: For cumulative frequency or working out \(\Sigma fx\) by hand - don't. Store midpoints in L1, frequencies in L2, and let 1-Var Stats do it all.
See the full GDC guide for more calculator models and topics.
Ready to practise properly?
Statistics & sampling questions, marked instantly like the real exam.
Quick answers
The questions students on this topic ask most often.
What's the difference between population and sample standard deviation?
Population standard deviation (\(\sigma_x\) on your GDC) treats your data as the entire population. Sample standard deviation (\(S_x\)) divides by \(n-1\) instead of \(n\), giving a slightly larger value that estimates the spread of a wider population from a sample. IB questions usually want \(\sigma_x\) unless told the data is a sample.
How do I know if a value is an outlier?
Find the interquartile range, \(\text{IQR} = Q_3 - Q_1\). A value is an outlier if it lies below \(Q_1 - 1.5\,\text{IQR}\) or above \(Q_3 + 1.5\,\text{IQR}\). Box-and-whisker diagrams mark outliers with a cross rather than extending the whisker to them.
Should I use the mean or median to summarise skewed data?
The median, because it is resistant to outliers and skew - a few very large or small values barely move it. The mean is pulled toward extreme values, so for skewed data (like income) it can give a misleading picture of a typical value.
Can I use my GDC for this topic?
Yes - almost every question on this topic expects it. 1-Var Stats gives you the mean, standard deviation and five-number summary in one step, so you rarely need to compute these by hand. See the GDC guide for model-specific instructions.
Sub-topics
Statistics & Sampling broken down into its individual skills, each with its own focused page.
Related topics
More Statistics & Probability topics from the same AI HL syllabus unit, in case you want to keep going.