Descriptive Statistics (AA HL)

Descriptive statistics summarise a data set with a handful of numbers: where its centre lies, how spread out it is, and whether any values sit unusually far from the rest. This topic covers the mean, median and mode, quartiles and the interquartile range, standard deviation and variance, and how to spot an outlier using the 1.5 IQR rule. Most of the calculation is done on your GDC, so the emphasis is on reading and interpreting the numbers it gives you.

What the syllabus says

This topic maps onto three points in the official IB Analysis & Approaches syllabus.

CodeSyllabus content
SL4.1Concepts of population, sample and discrete/continuous data. Reliability of data sources and bias in sampling. An outlier is a data item more than 1.5 × IQR from the nearest quartile; some outliers are valid data, others are recording errors.
SL4.2Presentation of data in frequency tables and histograms. Cumulative frequency graphs, used to find the median, quartiles, percentiles and IQR. Production and interpretation of box and whisker diagrams, including comparing two distributions.
SL4.3Measures of central tendency (mean, median, mode), including estimating the mean of grouped data from midpoints. Measures of dispersion (IQR, standard deviation, variance). The effect of constant changes (adding a constant, scaling) on the mean and standard deviation. Quartiles of discrete data, found using technology.

These are core AA SL syllabus points that are also examinable at AA HL.

Key terms

Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.

What is the mean?

The mean is the sum of every value in a data set divided by how many values there are. It uses every data point, which makes it sensitive to unusually large or small values (outliers) pulling it away from where most of the data actually sits.

e.g. For \(4, 7, 8, 10, 11\): \(\bar x = \dfrac{4+7+8+10+11}{5} = \dfrac{40}{5} = 8\).

What is the median?

The median is the middle value when the data is arranged in order. For an odd number of values it's the single middle term; for an even number it's the mean of the two middle terms. Unlike the mean, it isn't affected by extreme values.

e.g. For \(2, 5, 7, 7, 9, 12, 14\) (7 values, already ordered): the median is the 4th value, \(7\).

What are the quartiles and the IQR?

The quartiles \(Q_1\), \(Q_2\) (the median) and \(Q_3\) split ordered data into four equal parts. The interquartile range \(\text{IQR} = Q_3 - Q_1\) measures the spread of the middle 50% of the data, ignoring the extremes at either end.

e.g. For \(1, 3, 5, 6, 8, 9, 11, 13\): lower half \(1,3,5,6\) gives \(Q_1=4\); upper half \(8,9,11,13\) gives \(Q_3=10\); \(\text{IQR}=10-4=6\).

What is standard deviation?

Standard deviation measures the typical distance of each data value from the mean. A small standard deviation means the data is tightly clustered around the mean; a large one means it's widely spread out. Variance is just the standard deviation squared.

e.g. For \(2, 4, 4, 6\) (mean \(4\)): \(\sigma = \sqrt{\dfrac{(2{-}4)^2+(4{-}4)^2+(4{-}4)^2+(6{-}4)^2}{4}} = \sqrt{\dfrac{8}{4}} = \sqrt2 \approx 1.41\).

What is an outlier?

An outlier is a value more than \(1.5\times\text{IQR}\) beyond the nearer quartile - below \(Q_1-1.5\,\text{IQR}\) or above \(Q_3+1.5\,\text{IQR}\). It might be a genuine extreme observation or a recording error, so it needs checking in context rather than automatically discarding.

e.g. If \(Q_1=6\) and \(Q_3=10\), then \(\text{IQR}=4\) and the upper fence is \(10+1.5(4)=16\); a value of \(18\) would count as an outlier.

Key formulas

Descriptive statistics leans heavily on your GDC's statistics mode - the tables below show which formulas are on the booklet and which are just definitions you apply directly.

Formula reference

The mean and standard deviation formulas are on the official formula booklet; the IQR and outlier rule are definitions applied directly, with no formula given.

FormulaUsed forBooklet?
\(\bar x = \dfrac{\sum f_i x_i}{\sum f_i}\)Mean (including grouped/frequency data)✓ Yes
\(\sigma = \sqrt{\dfrac{\sum f_i(x_i-\bar x)^2}{\sum f_i}}\)Population standard deviation✓ Yes
\(\text{IQR} = Q_3 - Q_1\)Interquartile rangeNot in booklet - prior knowledge
\(Q_1-1.5\,\text{IQR},\ Q_3+1.5\,\text{IQR}\)Outlier fencesNot in booklet - prior knowledge

Mean vs median: which is more useful?

Both measure the "centre" of a data set, but they respond very differently when the data is skewed or contains outliers.

FeatureMeanMedian
Uses every value?Yes - all data points enter the calculationNo - only the position in the ordered list matters
Affected by outliers?Yes, heavilyNo, largely resistant
Best forSymmetric data with no extreme valuesSkewed data or data with outliers
Example\(1,2,3,4,50 \Rightarrow \bar x = 12\)\(1,2,3,4,50 \Rightarrow \text{median} = 3\)

Measures of central tendency

These describe where the "middle" of the data lies - each answers a slightly different question.

Mean

\[\bar x = \dfrac{\sum x_i}{n}\]

The arithmetic average - sum every value, divide by how many there are.

✓ In the formula booklet

Median

The middle value of the ordered data - the value at position \(\tfrac{n+1}{2}\) for odd \(n\), or the mean of the two middle values for even \(n\).

Not in the formula booklet - prior knowledge

Mode

The most frequently occurring value. For grouped data, the modal class is the class interval with the highest frequency (equal class widths only).

Not in the formula booklet - prior knowledge

Measures of dispersion

These describe how spread out the data is - a single "centre" number is never enough on its own.

Range

Largest value minus smallest value. Simple, but entirely determined by the two most extreme points.

Not in the formula booklet - prior knowledge

Interquartile range

\[\text{IQR} = Q_3 - Q_1\]

The spread of the middle 50% of the data - resistant to outliers, unlike the range.

Not in the formula booklet - prior knowledge

Standard deviation & variance

\[\sigma^2 = \dfrac{\sum(x_i-\bar x)^2}{n}\]

Variance is the mean squared distance from the mean; standard deviation is its square root, back in the original units.

✓ In the formula booklet

Effect of constant changes on the data

Adding a constant shifts every value, but doesn't change how spread out they are relative to each other. Scaling stretches both the centre and the spread.

Adding a constant

If every value has \(k\) added, the mean increases by \(k\) but the standard deviation is unchanged - the shape of the data doesn't stretch, it just slides.

Scaling by a constant

If every value is multiplied by \(k\), both the mean and the standard deviation are multiplied by \(|k|\) - the whole data set stretches (or shrinks) proportionally.

Worked examples

Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.

1
Easy
No calc
[4 marks]

The data set is \(3, 6, 6, 9, 11.\)

(a) Find the mean.

(b) Find the median.

Worked solution

(b) Sorted: \(3,6,6,9,11\); median \(=6.\) M1
Median \(=6.\) A1

M1 Sum \(\div n\) A1 \(\bar x=7\) M1 Sorts the data \(3,6,6,9,11\) A1 Median \(=6\)
2
Medium
No calc
[5 marks]

Ordered data: \(4, 6, 9, 10, 13, 15, 19.\)

(a) Find the median.

(b)(i) Find \(Q_1.\)

(b)(ii) Find \(Q_3.\)

(c) Find the interquartile range.

Worked solution

(a) 7 values; median is the 4th \(= 10.\) M1
\(\text{Median} = 10.\) A1

(b)(i) Lower half \(4,6,9\): \(Q_1 = 6.\) M1

(b)(ii) Upper half \(13,15,19\): \(Q_3 = 15.\) A1

(c) \(\text{IQR} = 15 - 6 = 9.\) A1

M1 Median position A1 Correct answer of \(10\) M1 Lower half \(4,6,9\) gives \(Q_1=6\) A1 Upper half \(13,15,19\) gives \(Q_3=15\) A1 Correct answer of \(9\)

Common mistakes

The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.

  • Using \(S_x\) instead of \(\sigma_x\). Your calculator gives both a sample standard deviation (\(S_x\), divides by \(n-1\)) and a population one (\(\sigma_x\), divides by \(n\)). IB treats the data set as the population unless told otherwise, so \(\sigma_x\) is almost always the one to quote.
  • Forgetting to use midpoints for grouped data. When data is given in class intervals like \(35 \le x < 50\), you must use the midpoint (\(42.5\)) as the representative value - using the interval endpoints instead gives the wrong mean.
  • Splitting the data incorrectly to find quartiles. Different methods (whether the median is included in each half) give slightly different quartile values, and technology may not match a hand calculation - IB expects the GDC's method to be used and trusted.
  • Assuming every extreme value is an error. The \(1.5\times\text{IQR}\) rule flags a value as an outlier, but that doesn't automatically mean it should be deleted - some outliers are a genuine part of the data.

Using your GDC

Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.

Show steps for:
One-variable statistics (mean, median, standard deviation)

Instant summary statistics from a list - no formulas to compute by hand.

  1. Enter the data into a list.
  2. STAT → Edit → type values into L1. Then STAT → CALC → 1:1-Var Stats, choose L1, Calculate.TI-84
  3. Add a Lists & Spreadsheet page, name a column and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
  4. Statistics menu → enter data in List 1 → CALC (F2) → 1-VAR.Casio
  5. Read \(\bar x\) (mean), \(S_x\) (sample sd) or \(\sigma_x\) (population sd), and the five-number summary (min, \(Q_1\), median, \(Q_3\), max).
  6. For frequency data, put values in one list and frequencies in another and set the frequency list.

Tip: \(S_x\) vs \(\sigma_x\): use \(\sigma_x\) (population) for a complete data set, \(S_x\) (sample) for a sample. IB usually wants \(\sigma_x\).

Build and manipulate lists

Lists let you store data sets, generate sequences, and compute statistics without re-entering values - essential for grouped data and frequency tables.

  1. Enter data manually: go to the list editor and type values one by one.
  2. 2nd → STAT (LIST) → OPS → 5:seq(expression, variable, start, end) generates a list from a formula, e.g. seq(X², X, 1, 10). For grouped data, put midpoints in L1 and frequencies in L2, then run 1-Var Stats L1, L2.TI-84
  3. In a Lists & Spreadsheet page, type a formula in the header row (e.g. = seq(x², x, 1, 10)) to fill the column automatically. Use two columns for midpoints and frequencies.Nspire
  4. In Statistics list editor, enter data directly; OPTN → LIST → Seq( builds a list from a formula. For frequency data, put midpoints in List 1 and frequencies in List 2.Casio
  5. For frequency data (grouped), enter midpoints in one list and frequencies in another, then reference both lists when running 1-Var Stats.
  6. To sort a list ascending: SortA(L1).TI-84
  7. To sort a list ascending: menu → List → Sort Ascending.Nspire

Tip: For cumulative frequency or working out \(\sum fx\) by hand - don't. Store midpoints in L1, frequencies in L2, and let 1-Var Stats do it all.

See the full GDC guide for more calculator models and topics.

Ready to practise properly?

Descriptive statistics questions, marked instantly like the real exam.

Quick answers

The questions students on this topic ask most often.

Do I need to memorise the standard deviation formula?

No. The syllabus expects you to calculate standard deviation using your GDC's statistics mode, not by hand. Knowing the formula still helps you understand what the calculator is doing and answer "explain" questions.

What's the difference between \(\sigma_x\) and \(S_x\) on my calculator?

\(\sigma_x\) is the population standard deviation (dividing by \(n\)); \(S_x\) is the sample standard deviation (dividing by \(n-1\)). At SL and HL the data set is treated as the population unless told otherwise, so IB almost always wants \(\sigma_x\).

How do I find the mean of grouped data?

Use the midpoint of each class interval as a representative x-value, multiply each midpoint by its frequency, sum those products, then divide by the total frequency: \(\bar x = \dfrac{\sum fx}{\sum f}\).

What counts as an outlier?

IB defines an outlier as any value more than \(1.5\times\text{IQR}\) beyond the nearer quartile - below \(Q_1-1.5\,\text{IQR}\) or above \(Q_3+1.5\,\text{IQR}\). Not every extreme value is an error, so outliers should be checked in context rather than automatically removed.

Sub-topics

Descriptive Statistics broken down into its individual skills, each with its own focused page.