Measures of Spread (AA SL)
Two data sets can share the same mean and still look completely different once you ask how spread out they are. This page covers the range, interquartile range and standard deviation - what each one measures, when the GDC gives you two different standard deviations to choose from, and how a linear transformation of the data affects spread. It's part of the broader Descriptive Statistics & Correlation topic.
17 questions on this sub-topic.
The key measures
Covered under IB syllabus reference SL4.3: measures of dispersion (interquartile range, standard deviation and variance), and the effect of constant changes on the original data.
Interquartile range
\(\text{IQR}=Q_3-Q_1\)
Not in the formula booklet - it's a definition. Measures the spread of the middle 50% of the data, so it isn't pulled around by extreme values the way the range is.
Effect of a linear transformation
If \(y=ax+b\), the new mean is \(a\bar x+b\) and the new standard deviation is \(|a|\sigma_x\).
Adding or subtracting a constant \(b\) shifts every value equally, so it never changes the spread - only the multiplier \(a\) does.
Need the full syllabus wording and formula-booklet reference table? See Descriptive Statistics & Correlation.
Worked examples
A data set in order is \(3, 5, 6, 8, 9, 11, 14, 18\).
(a) Find the range.
(b) Find the interquartile range.
Worked solution
(a) \(\text{Range}=18-3=15.\) A1
(b) With \(n=8,\) lower half \(3,5,6,8\Rightarrow Q_1=\dfrac{5+6}{2}=5.5;\) upper half \(9,11,14,18\Rightarrow Q_3=\dfrac{11+14}{2}=12.5.\) M1
\(Q_1=5.5,\ Q_3=12.5.\) A1
\(\text{IQR}=12.5-5.5=7.\) A1
The number of goals scored in 10 matches is \(1, 0, 2, 3, 1, 2, 4, 1, 0, 2\).
(a) Find the mean.
(b) Find the standard deviation, to 3 significant figures.
Worked solution
(a) Add the ten values: \(1+0+2+3+1+2+4+1+0+2=16,\) M1
\(\bar x=\dfrac{16}{10}=1.6\) goals. A1
(b) For the spread of this whole data set use the population SD \(\sigma\) (the IB default for a given list), not the sample \(s_x\). M1
The GDC returns \(\sigma\approx 1.20\) (to 3 significant figures). A1
Common mistakes
- Reading \(Sx\) instead of \(\sigma x\) off the GDC. The calculator's 1-Var Stats screen shows both the sample SD \(Sx\) (dividing by \(n-1\)) and the population SD \(\sigma x\) (dividing by \(n\)) - for a data set treated as the whole population, take \(\sigma x\).
- Forgetting that shifting data doesn't change its spread. Under \(y=ax+b\), the constant \(b\) moves every value by the same amount, so gaps between values - and hence the standard deviation - are unaffected; only \(a\) rescales the spread.
- Splitting an odd-sized data set incorrectly for the quartiles. When \(n\) is odd, whether the median itself is included in the lower and upper halves changes \(Q_1\) and \(Q_3\) - check which convention the question or GDC is using.
Ready to practise properly?
17 measures-of-spread questions, marked instantly like the real exam.
Quick answers
What is the difference between the range and the interquartile range?
The range is the maximum minus the minimum, so it uses only the two most extreme values. The interquartile range, \(\text{IQR}=Q_3-Q_1\), is the spread of the middle 50% of the data, so it isn't distorted by outliers the way the range is.
Which standard deviation should I use on the GDC, sigma x or Sx?
For a given list of data treated as the whole population, use the population standard deviation \(\sigma x\), which divides by \(n\). The sample standard deviation \(Sx\), dividing by \(n-1\), is only used when the data is a sample used to estimate a wider population. See the parent topic's GDC guidance for calculator steps.