Spearman's Rank (AI SL)
Spearman's rank correlation coefficient \(r_s\) measures how closely two variables move together in rank order, even when the underlying relationship isn't a straight line. Because it works on ranks rather than raw values, it's less thrown off by a single extreme outlier than Pearson's \(r\). It's part of the broader Correlation & Regression topic.
21 questions on this sub-topic.
The formula
Covered under IB syllabus reference SL4.10: Spearman's rank correlation coefficient \(r_s\), found using technology in examinations, together with awareness of when it's more appropriate than Pearson's \(r\) and how it responds to outliers.
Spearman's rank coefficient
\(r_s = 1 - \dfrac{6\sum d^2}{n(n^2-1)}\)
In the formula booklet - rank both variables, find the difference \(d\) between each pair of ranks, then apply the formula (or let the GDC compute it from the ranked lists directly).
Reading the value
\(-1\le r_s\le 1\)
Interpreted exactly like Pearson's \(r\): close to \(+1\) is a strong positive monotonic relationship, close to \(-1\) strong negative, near \(0\) little or no monotonic relationship.
Need the full syllabus wording and formula-booklet reference table? See Correlation & Regression. For calculator steps, see the parent topic's GDC guidance.
Worked examples
Two judges rank 5 cakes. Judge A: 1,2,3,4,5; Judge B: 2,1,4,3,5.
Find Spearman's rank correlation coefficient \(r_s.\)
Worked solution
Differences \(d=-1,1,-1,1,0\), so \(\sum d^2=4.\) M1
\(r_s=1-\dfrac{6\sum d^2}{n(n^2-1)}=1-\dfrac{6(4)}{5(24)}=1-0.2=0.8.\) A1
A study finds \(r_s=-0.85\) between hours of TV and exam rank. Interpret this result.
Worked solution
The negative sign shows a strong inverse monotonic relationship: A1
more TV hours are associated with worse (higher-numbered) exam ranks. A1
Two coaches rank 6 athletes. The rank differences give \(\sum d^2=8.\) Find \(r_s\) and comment on the level of agreement.
Worked solution
\(r_s=1-\dfrac{6(8)}{6(35)}\) M1
\(=1-\dfrac{48}{210}.\) A1
\(\approx0.771.\) A1
The coaches show fairly strong agreement. A1
Five towns are ranked by rainfall and by number of umbrellas sold. The rank differences are \(d=1,0,-2,1,0.\) Find \(r_s.\)
Worked solution
\(\sum d^2=1+0+4+1+0=6.\) M1
\(r_s=1-\dfrac{6(6)}{5(24)}=1-0.3=0.7.\) A1
Common mistakes
- Ranking the wrong direction. Whether rank 1 means "largest" or "smallest" only matters that it's applied consistently to both variables - a mixed convention scrambles every \(d\) value.
- Forgetting to average tied ranks. When two data items are equal, both should get the mean of the ranks they'd otherwise occupy, not two different whole-number ranks.
- Treating \(r_s\) and Pearson's \(r\) as interchangeable. \(r_s\) detects any monotonic trend (including curved ones), while \(r\) only detects a linear one - quoting the wrong coefficient for what the question describes loses marks.
Ready to practise properly?
21 Spearman's-rank questions, marked instantly like the real exam.
Quick answers
What is the formula for Spearman's rank correlation coefficient?
\(r_s = 1 - \dfrac{6\sum d^2}{n(n^2-1)}\), where \(d\) is the difference between the ranks of each pair and \(n\) is the number of pairs. In examinations it is normally found using technology instead.
How is Spearman's rank different from Pearson's r?
Pearson's \(r\) measures the strength of a linear relationship in the raw data. Spearman's rank works on the ranks of the data instead, so it measures any monotonic relationship (consistently increasing or decreasing, not necessarily in a straight line) and is less affected by outliers.