Data Types and Collection (AI SL)
Before you can pick a sampling method or judge whether a data source can be trusted, you need the vocabulary this sub-topic covers: population versus sample, discrete versus continuous data, and what actually makes a source reliable or a result biased. It's part of the broader Sampling & Data Collection topic.
27 questions on this sub-topic.
Outliers and data types
Covered under IB syllabus reference SL4.1. This part of the syllabus sets up the population/sample vocabulary, the discrete/continuous split, and the reliability and bias language used across the whole statistics course - plus a precise rule for spotting outliers.
The outlier rule
\(x < Q_1 - 1.5(IQR)\) or \(x > Q_3 + 1.5(IQR)\)
This is the exact boundary the IB uses to define an outlier. It flags a value as unusual - it doesn't automatically tell you to remove it.
Not in the formula booklet - syllabus definitionDiscrete vs continuous
Discrete data can only take separate, countable values - number of siblings, shoe size, goals scored. Continuous data can take any value in a range, usually from measuring - height, time, mass. The distinction decides which charts and calculations make sense.
Looking for sampling techniques instead? See Sampling methods, or the full Sampling & Data Collection page for GDC guidance.
Worked examples
Two reports give different figures for a town's population.
State two factors that affect the reliability of a data source.
Worked solution
Any two of: how recently the data were collected; A1
the sample size and method, or whether the source is independent/unbiased. A1
A reporter interviews people leaving a gym about how much they exercise.
State the sampling method and explain a likely source of bias.
Worked solution
This is convenience sampling. A1
People leaving a gym exercise more than average, R1
so the sample is biased and overestimates exercise levels. A1
A survey on internet use is run only through a website. Explain why results may be biased.
Worked solution
People without internet cannot respond. R1
So the sample is unrepresentative. A1
Common mistakes
- Treating every outlier as an error. The \(1.5\times IQR\) rule flags a value as unusual, not as wrong - some outliers are a genuine, valid part of the data and should stay in unless there's a clear reason to exclude them.
- Confusing discrete with continuous. Data you count - number of pets, shoe size - is discrete even though it looks numerical. Only data that comes from measuring belongs on a continuous scale.
- Blurring "small" with "biased". A small sample gives an imprecise estimate; a biased sample gives a wrong one regardless of size. Be clear about which criticism actually applies to a scenario.
Ready to practise properly?
29 data-types-and-collection questions, marked instantly like the real exam.
Quick answers
What counts as an outlier in IB AI SL statistics?
A data item more than \(1.5\times\) the interquartile range (IQR) from the nearest quartile - so below \(Q_1 - 1.5(IQR)\) or above \(Q_3 + 1.5(IQR)\).
What is the difference between discrete and continuous data?
Discrete data can only take separate, countable values (e.g. number of goals scored); continuous data can take any value in a range and usually comes from measuring (e.g. height, time).