Sampling Methods (AI HL)
Before you can analyse a data set you have to know how it was collected, and a badly-chosen sample can make even the best statistics misleading. This page covers the five named sampling techniques and the kinds of bias examiners like to test, with worked examples and the mistakes that lose marks. It's part of the broader Statistics & Sampling topic.
11 questions on this sub-topic.
Sampling techniques
Covered under IB syllabus reference SL4.1. There's no formula to learn here - you need to be able to describe each method, know its strengths and weaknesses, and spot bias in a described sampling process.
The five methods
Simple random, convenience, systematic, quota and stratified sampling. Stratified splits the population into groups (strata) first, then samples randomly and proportionally within each - it's the one examiners ask you to calculate.
Bias to spot
Selection bias (the method itself excludes part of the population) and non-response/self-selection bias (only certain people choose to reply). Name the bias, then explain who is under- or over-represented as a result.
Need the full syllabus wording and formula-booklet reference table? See Statistics & Sampling.
Worked examples
A school has 1200 students in 4 year groups. A researcher wants a sample of 60.
(a) Describe how to take a stratified sample by year group, given the groups have 360, 300, 270 and 270 students.
(b) State one advantage of stratified sampling over simple random sampling here.
Worked solution
(a) Sampling fraction \(= \dfrac{60}{1200} = \dfrac{1}{20}.\) M1
Take \(\tfrac{1}{20}\) of each stratum: \(18, 15, 14, 13\) (total 60), selecting at random within each. A1
(b) It guarantees each year group is represented in proportion to its size, R1
reducing sampling bias. A1
A council emails an online survey to all residents and uses only the replies.
(a) State the type of bias most likely present.
(b) Suggest one improvement to reduce this bias.
Worked solution
(a) Non-response (self-selection) bias. R1
Those with strong opinions or internet access are over-represented. A1
(b) Take a random sample and follow up non-respondents. R1
Or offer multiple channels. A1
Common mistakes
- Naming the wrong bias. "Selection bias" comes from how the sample is chosen (e.g. only surveying commuters at one station); "non-response bias" comes from who chooses to reply once selected. Mixing these up costs the identification mark even when the explanation is otherwise sound.
- Confusing stratified with quota sampling. Both split the population into groups first, but stratified sampling then selects randomly within each group, while quota sampling lets the interviewer pick anyone until the quota is met - which reintroduces bias.
- Stopping at "it's biased" without saying who is affected. A full-mark answer names the bias, then states which group is over- or under-represented as a result - a bare label rarely earns both marks.
Ready to practise properly?
10 sampling-methods questions, marked instantly like the real exam.
Quick answers
What sampling methods do I need to know for IB AI?
Simple random, convenience, systematic, quota and stratified sampling. You should be able to describe each method and comment on when it is appropriate.
What is the difference between quota and stratified sampling?
Both split the population into groups (strata), but stratified sampling then chooses members randomly within each group, while quota sampling lets the researcher choose non-randomly until each quota is filled - so quota sampling is more prone to bias.