Measures of Central Tendency (AI HL)
Mean, median and mode each try to summarise a data set with a single "typical" value, but they don't always agree - especially once the data is skewed or grouped into classes. This page covers how to find and choose between them, including the mid-interval trick for grouped data, with worked examples and the mistakes that lose the most marks. It's part of the broader Statistics & Sampling topic.
26 questions on this sub-topic.
Mean, median and mode
Covered under IB syllabus reference SL4.3: mean, median and mode for a data set, including estimating the mean of grouped data from mid-interval values and identifying the modal class.
Mean of raw data
\(\bar{x} = \dfrac{\sum x}{n}\)
Sum every value and divide by how many values there are. Not in the formula booklet - your GDC's one-variable statistics screen gives \(\bar{x}\) directly once the data is entered.
Mean of grouped data
\(\bar{x} \approx \dfrac{\sum fx}{\sum f}\)
Replace each class by its mid-interval value \(x\), weight by the frequency \(f\), then divide by the total frequency. This is an estimate, since the raw values inside each class are unknown.
Need standard deviation, IQR, or the outlier rule too? See Statistics & Sampling.
Worked examples
Find the mean of \(6, 9, 4, 7, 9\).
Worked solution
\(6+9+4+7+9=35.\) M1
\(\bar x=\dfrac{35}{5}=7.\) A1
The mean of 6 numbers is 14. A seventh number is added and the mean becomes 15.
Find the seventh number.
Worked solution
Original total \(= 6(14)\) M1
\(= 84.\) A1
New total \(= 7(15) = 105.\) M1
Seventh number \(= 105 - 84 = 21.\) A1
Common mistakes
- Using the mean when the median is what's asked for (or vice versa). Skewed data or a single extreme outlier can pull the mean well away from where most of the data actually sits - the median is often the fairer "typical value" in that case, and examiners expect you to know which one a question wants.
- Using the class width instead of the mid-interval value. For grouped data, \(x\) in \(\sum fx / \sum f\) is the midpoint of each class (lower bound plus upper bound, divided by 2), not the width of the class - mixing the two gives a completely wrong mean.
- Forgetting a reversed-mean question is still just "total \(\div\) count". When a mean is given and you need to recover a total or a missing value, multiply back up first (\(\text{total} = n\bar{x}\)) rather than trying to manipulate the mean formula directly.
Ready to practise properly?
29 central-tendency questions, marked instantly like the real exam.
Quick answers
What is the difference between mean, median and mode?
The mean is the sum of the values divided by how many there are; the median is the middle value once the data is ordered; the mode is the value (or class) that occurs most often. Outliers or skew can make them disagree.
How do you estimate the mean of grouped data?
Replace each class with its mid-interval value, multiply each mid-interval value by its frequency, sum those products, then divide by the total frequency: \(\bar{x} \approx \dfrac{\sum fx}{\sum f}\). It's an estimate because the raw values inside each class are unknown. See using your GDC for entering frequency data directly.