Descriptive Statistics (AI SL)
Descriptive statistics summarise a data set with a handful of numbers instead of listing every value. This topic covers where the "middle" of a data set is (mean, median, mode) and how spread out it is (range, interquartile range, standard deviation, variance), plus how to spot and describe outliers and read a box-and-whisker plot. Almost every calculation here is done with your GDC's one-variable statistics function.
What the syllabus says
This topic maps onto three points in the official IB Applications & Interpretation syllabus.
| Code | Syllabus content |
|---|---|
| SL4.1 | Concepts of population, sample, random sample, and discrete and continuous data. Reliability of data sources and bias in sampling. Interpretation of outliers - a data item more than \(1.5\times\) IQR from the nearest quartile - and awareness that some outliers are valid while others may be recording errors. Sampling techniques (simple random, convenience, systematic, quota, stratified) and their effectiveness. |
| SL4.2 | Presentation of data in frequency tables, histograms and cumulative frequency graphs; using cumulative frequency to find the median, quartiles, percentiles, range and IQR. Production and understanding of box and whisker diagrams, including comparing two distributions using symmetry, median, IQR or range, and assessing possible normality from the symmetry of the box and whiskers. |
| SL4.3 | Measures of central tendency (mean, median, mode), including estimating the mean from grouped data. Measures of dispersion (IQR, standard deviation and variance), calculated using technology. Effect of constant changes (e.g. adding or scaling every value) on the mean and standard deviation. |
Variance and standard deviation of a sample are calculated using technology only - hand calculation is for understanding, not for the exam.
Key terms
Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.
What is the interquartile range?
The interquartile range (IQR) is the spread of the middle 50% of a data set: the difference between the upper quartile \(Q_3\) and the lower quartile \(Q_1\). It ignores the most extreme quarter of values on each end, which makes it resistant to outliers.
e.g. If \(Q_1=7\) and \(Q_3=16.5\), then \(\text{IQR}=16.5-7=9.5\).
What is standard deviation?
Standard deviation measures how far, on average, data values sit from the mean. A small standard deviation means the data is tightly clustered; a large one means it's spread out. It's calculated using every value in the data set, via your GDC.
e.g. For \(8,9,10,11,12\) (mean 10), the population standard deviation is \(\sigma=\sqrt{2}\approx1.41\).
What is an outlier?
An outlier is a data item that lies more than \(1.5\times\)IQR from the nearest quartile - unusually far from the bulk of the data. Some outliers are genuine extreme values; others are recording errors, so context matters when deciding what to do with one.
e.g. With \(Q_1=20,\ Q_3=32\) (IQR \(=12\)), the upper fence is \(32+1.5(12)=50\), so a value of 55 is an outlier.
What is a box and whisker diagram?
A box and whisker diagram displays the five-number summary of a data set - minimum, \(Q_1\), median, \(Q_3\), maximum - as a box (the middle 50%) with whiskers stretching to the extremes. It's a fast way to compare the spread and skew of two data sets side by side.
e.g. A data set with min 3, \(Q_1=6\), median 10, \(Q_3=14\), max 19 draws a box from 6 to 14 with a line at 10, and whiskers to 3 and 19.
Why is the median resistant to outliers?
The median only depends on the position of values once ordered, not their size - so one extreme value can't drag it far. The mean, by contrast, uses every value's actual size, so a single huge or tiny value can shift it substantially.
e.g. Adding 40 to \(8,9,10,11,12\) (mean 10) pushes the mean to 15, but the median only moves from 10 to 10.5.
Key formulas
A handful of formulas cover this topic - the two tables below summarise them at a glance, and the sections underneath go into more depth.
Formula reference
Range and IQR are simple arithmetic on values you read off directly; standard deviation and variance are calculated by technology, and their defining relationship is worth knowing even though you won't compute it by hand.
| Formula | Used for | Booklet? |
|---|---|---|
| \(\text{Range} = \text{max} - \text{min}\) | Total spread of the data | Not in booklet - prior knowledge |
| \(\text{IQR} = Q_3 - Q_1\) | Spread of the middle 50% | Not in booklet - prior knowledge |
| Lower fence \(=Q_1-1.5\times\text{IQR}\) | Outlier test (below) | Not in booklet - prior knowledge |
| Upper fence \(=Q_3+1.5\times\text{IQR}\) | Outlier test (above) | Not in booklet - prior knowledge |
| \(\text{Variance} = \sigma^2\) | Relationship between variance and standard deviation | ✓ Yes |
Mean vs median: which to trust
The two measures of centre behave very differently once a data set has extreme values - this table lines up when each one is the better choice.
| Feature | Mean | Median |
|---|---|---|
| Uses | Every value in the data set | Only the position of values once ordered |
| Affected by outliers? | Yes - can shift a lot | No - barely moves |
| Best for | Symmetric data with no extreme values | Skewed data or data with outliers |
| Paired dispersion measure | Standard deviation | Interquartile range |
| Example (add 40 to \(8,9,10,11,12\)) | Rises from 10 to 15 | Rises from 10 to 10.5 |
Measures of central tendency
These describe where the "middle" of a data set sits, using three different ideas of what "middle" means.
Mean
The sum of all values divided by how many there are. Calculated using technology from raw or grouped data (using midpoints for grouped data).
Not in the formula booklet - GDC computationMedian
The middle value when the data is ordered. For \(n\) values, it's the \(\tfrac{n+1}{2}\)th value if \(n\) is odd, or the mean of the two middle values if \(n\) is even.
Not in the formula booklet - prior knowledgeMode
The most frequently occurring value. For grouped data, the modal class is the class interval with the highest frequency (only meaningful for equal class widths).
Not in the formula booklet - prior knowledgeMeasures of dispersion
These describe how spread out the data is - each with a different sensitivity to extreme values.
Range and IQR
\[\text{IQR}=Q_3-Q_1\]
Range uses the extremes; IQR uses only the middle 50%, so it's resistant to outliers.
Standard deviation and variance
\[\text{Variance}=\sigma^2\]
Standard deviation measures typical distance from the mean; variance is its square. Both calculated by GDC.
Effect of transforming data
Adding or subtracting a constant shifts the mean but leaves the standard deviation unchanged. Multiplying every value by a constant scales both the mean and the standard deviation by that constant.
Worked examples
Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.
Data: \(2, 5, 6, 6, 7, 8, 9, 30.\)
(a) Find the range.
(b) Find the IQR.
(c) State which better describes the spread and why.
Worked solution
(a) Range. \(30-2=28.\) A1
(b) IQR. Lower half \(2,5,6,6\Rightarrow Q_1=\tfrac{5+6}{2}=5.5;\) M1
upper half \(7,8,9,30\Rightarrow Q_3=\tfrac{8+9}{2}=8.5.\) So \(\text{IQR}=8.5-5.5=3.\) A1
(c) Which is better. The single value 30 is an outlier that inflates the range to 28, but the IQR (3) ignores the extreme quarter of the data. So the IQR better describes the spread R1
A data set has \(Q_1=20,\ Q_3=32.\)
(a) Find the IQR.
(b)(i) Find the lower outlier boundary using the \(1.5\times\)IQR rule.
(b)(ii) Find the upper outlier boundary using the \(1.5\times\)IQR rule.
(c) State whether a value of 55 is an outlier.
Worked solution
(a) IQR. \(\text{IQR}=Q_3-Q_1=32-20=12.\) A1
(b)(i) Lower fence \(=Q_1-1.5\times\text{IQR}=20-1.5(12)=20-18\) M1
\(=2.\) A1
(b)(ii) Upper fence \(=Q_3+1.5\times\text{IQR}=32+18\) M1
\(=50.\) A1 Any value outside \([2,50]\) is flagged as an outlier.
(c) Is 55 an outlier? Since \(55>50\) (above the upper fence), yes, 55 is an outlier. A1
Common mistakes
The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.
- Using the mean when the median is more appropriate. If a data set has an outlier or is heavily skewed, the mean can be misleading - the median (paired with the IQR) usually gives a fairer picture.
- Mixing up \(\sigma_x\) (population) and \(s_x\) (sample). Your GDC's one-variable statistics screen shows both - IB generally wants \(\sigma_x\), the population standard deviation, so check which one you're reading off.
- Forgetting the effect of constant changes. Adding a constant to every value shifts the mean but does NOT change the standard deviation; only scaling (multiplying) every value changes the standard deviation.
- Misreading a box plot. The line inside the box is the median, not the mean, and the box edges are \(Q_1\) and \(Q_3,\) not the minimum and maximum - those are the whisker ends.
Using your GDC
Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.
Instant summary statistics from a list - no formulas to compute by hand.
- Enter the data into a list.
- STAT → Edit → type values into L1. Then STAT → CALC → 1:1-Var Stats, choose L1, Calculate.TI-84
- Add a Lists & Spreadsheet page, name a column and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
- Statistics menu → enter data in List 1 → CALC (F2) → 1-VAR.Casio
- Read \(\bar x\) (mean), \(S_x\) (sample sd) or \(\sigma_x\) (population sd), and the five-number summary (min, \(Q_1\), median, \(Q_3\), max).
- For frequency data, put values in one list and frequencies in another and set the frequency list.
Tip: \(S_x\) vs \(\sigma_x\): use \(\sigma_x\) (population) for a complete data set, \(S_x\) (sample) for a sample. IB usually wants \(\sigma_x\).
Lists let you store data sets, generate sequences, and compute statistics without re-entering values - essential for grouped data and frequency tables.
- Enter data manually: go to the list editor and type values one by one.
- 2nd → STAT (LIST) → OPS → 5:seq(expression, variable, start, end) generates a list from a formula, e.g. seq(X², X, 1, 10). For grouped data, put midpoints in L1 and frequencies in L2, then run 1-Var Stats L1, L2.TI-84
- In a Lists & Spreadsheet page, type a formula in the header row (e.g. = seq(x², x, 1, 10)) to fill the column automatically. Use two columns for midpoints and frequencies.Nspire
- In Statistics list editor, enter data directly; OPTN → LIST → Seq( builds a list from a formula. For frequency data, put midpoints in List 1 and frequencies in List 2.Casio
- For frequency data (grouped), enter midpoints in one list and frequencies in another, then reference both lists when running 1-Var Stats.
- To sort a list ascending: SortA(L1) on TI-84; menu → List → Sort Ascending on Nspire.
Tip: For cumulative frequency or working out \(\Sigma fx\) by hand - don't. Store midpoints in L1, frequencies in L2, and let 1-Var Stats do it all.
See the full GDC guide for more calculator models and topics.
Ready to practise properly?
Descriptive statistics questions, marked instantly like the real exam.
Quick answers
The questions students on this topic ask most often.
What's the difference between the mean and the median?
The mean is the sum of all values divided by how many there are - it uses every value and is pulled around by extreme ones. The median is the middle value once the data is ordered - it only depends on position, so it's far more resistant to outliers.
When should I use standard deviation instead of IQR?
Standard deviation uses every data value and suits data without extreme outliers, giving a precise measure of spread around the mean. The IQR only uses the middle 50% of the data, so it's the better choice when outliers are present or the data is skewed.
How do I know if a value is an outlier?
Use the \(1.5\times\)IQR rule: find the lower fence (\(Q_1-1.5\times\text{IQR}\)) and upper fence (\(Q_3+1.5\times\text{IQR}\)). Any value below the lower fence or above the upper fence is flagged as an outlier.
Do I need to calculate standard deviation by hand?
No. The syllabus says variance and standard deviation of a sample are calculated using technology only, though hand calculation may help you understand what the formula is doing. In exams you'll always use your GDC's one-variable statistics function.
Sub-topics
Descriptive Statistics broken down into its individual skills, each with its own focused page.
Related topics
More Statistics & Probability topics from the same AI SL syllabus unit, in case you want to keep going.