Sampling & Data Collection (AI SL)
Before any statistics can be trusted, the data behind it has to be collected sensibly. This topic covers the difference between a population and a sample, the five named sampling methods on the syllabus - simple random, convenience, systematic, quota and stratified - how to spot bias in a sampling design, and how the syllabus defines an outlier.
What the syllabus says
This topic maps onto one point in the official IB Applications & Interpretation syllabus.
| Code | Syllabus content |
|---|---|
| SL4.1 | Concepts of population, sample, random sample, discrete and continuous data. Reliability of data sources and bias in sampling. Interpretation of outliers, defined as a data item more than 1.5\(\times\) the interquartile range (IQR) from the nearest quartile. Sampling techniques and their effectiveness: simple random, convenience, systematic, quota and stratified sampling methods. |
Some outliers are a valid part of the sample; others may be errors in the data - the syllabus expects awareness of both possibilities before deciding what to do with one.
Key terms
Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.
What is a population vs a sample?
The population is every individual you could possibly measure; a sample is the smaller subset you actually collect data from. A random sample gives every member of the population a known, non-zero chance of being chosen, which is what makes it possible to trust the sample to represent the whole.
e.g. Population: 5000 customers. Sample: 100 customers surveyed by picking random numbers 1-5000.
What is stratified sampling?
Stratified sampling splits the population into groups (strata) - such as year groups or departments - and samples from each group in proportion to its size, using the same sampling fraction throughout. It guarantees every group is represented, unlike a simple random sample which could miss a small group by chance.
e.g. Sampling fraction \(\dfrac{75}{900}=\dfrac{1}{12}\); a stratum of 300 contributes \(300\times\dfrac{1}{12}=25\).
What is systematic sampling?
Systematic sampling selects every \(k\)th member of a list, where the sampling interval \(k\) is the population size divided by the desired sample size, after a random starting point. It's quick to apply to an ordered list, but can introduce bias if the list has a hidden repeating pattern.
e.g. For 600 records and a sample of 20: interval \(=\dfrac{600}{20}=30\).
What is convenience sampling, and why can it be biased?
Convenience sampling selects whoever is easiest to reach - the first people you meet, rather than a randomly chosen group. It's fast but risks systematic bias, since the people who happen to be convenient to sample are often not representative of the whole population.
e.g. Surveying the first 30 shoppers at 9am misses everyone who shops later in the day.
What is an outlier?
The syllabus defines an outlier as a data item more than 1.5 times the interquartile range (IQR) from the nearest quartile. Some outliers are genuine, valid data; others are recording errors - the definition flags a value as unusual, but doesn't decide which it is.
e.g. If \(Q_1=10\), \(Q_3=20\), then \(\text{IQR}=10\), so any value above \(20+1.5(10)=35\) is an outlier.
Key formulas
This topic is mostly about identifying the right method and its weaknesses, but three genuine calculations come up repeatedly. The tables below summarise them - the explanations underneath go into more depth on each one.
Formula reference
None of these are officially listed in the formula booklet - they're direct-proportion reasoning and a syllabus-stated convention, not booklet entries.
| Formula | Used for | Booklet? |
|---|---|---|
| Stratum sample size \(= \dfrac{\text{sample size}}{\text{population size}}\times\) stratum size | Stratified sampling | Not in booklet - direct proportion |
| Sampling interval \(= \dfrac{\text{population size}}{\text{sample size}}\) | Systematic sampling | Not in booklet - direct proportion |
| Outlier if more than \(1.5\times\text{IQR}\) from nearest quartile | Interpreting outliers | Not in booklet - syllabus-defined convention |
Probability sampling vs non-probability sampling
The five named methods split naturally into two groups, depending on whether every member of the population has a known chance of selection.
| Feature | Probability sampling | Non-probability sampling |
|---|---|---|
| Methods | Simple random, systematic, stratified | Convenience, quota |
| Selection | Every member has a known chance of being chosen | Chosen by ease of access or a fixed target count |
| Bias risk | Low, if applied correctly | Higher - depends on who happens to be available |
| Example | Every 20th shopper all day (systematic) | First 30 shoppers at 9am (convenience) |
Sampling techniques
All five methods appear on the syllabus by name - you should be able to identify each one from a short description.
Simple random sampling
Number every member of the population, then use random numbers to pick the sample - every individual and every combination has an equal chance of selection.
Not in the formula booklet - selection procedureStratified sampling
Split the population into strata, then sample from each using the same sampling fraction, so every group is represented proportionally.
Not in the formula booklet - direct proportionSystematic sampling
Pick every \(k\)th item from an ordered list after a random start, spreading the sample evenly across the whole population.
Not in the formula booklet - direct proportionConvenience & quota sampling
Convenience takes whoever is easiest to reach; quota fixes target numbers per group but still lets the interviewer choose anyone willing - both are non-random and can be biased.
Not in the formula booklet - selection procedureBias, reliability and outliers
Choosing a method is only half the job - you also need to judge how trustworthy the resulting data is.
Sample size and reliability
A larger sample reduces sampling variability and generally gives a more reliable estimate of the true population value, all else being equal.
Not in the formula booklet - general principleSources of bias
Watch for non-response bias, self-selection (voluntary-response) bias, and time-of-day or location bias - name the specific mechanism, not just "it could be biased".
Not in the formula booklet - identified in contextInterpreting an outlier
An outlier flagged by the \(1.5\times\text{IQR}\) rule may be a genuine extreme value or a recording error - investigate before removing it from the data set.
Not in the formula booklet - syllabus-defined conventionWorked examples
Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.
A stratified sample uses fraction 0.05. A stratum has 640 members.
(a) Find the number sampled.
(b) If the fraction were instead 0.08, find the new number sampled.
(c) Find the total population size if this stratum represents 40% of the total population.
Worked solution
(a) \(0.05\times640=32.\) M1 A1
(b) \(0.08\times640=51.2\) M1
\(\approx51.\) A1
(c) \(640\div0.4=1600.\) M1 A1
A college has 1200 students in four years: 360, 320, 280, 240.
(a)(i) A stratified sample of 60 is required. Find the number from year 1 (largest group).
(a)(ii) Find the number from year 2.
(a)(iii) Find the number from year 3.
(a)(iv) Find the number from year 4 (smallest group).
(b) Verify that the four sample sizes sum to 60.
(c) The college instead wants a total sample of 90. Find the new number of students required from the largest year group.
Worked solution
(a)(i) \(18,\)
(a)(ii) \(16,\)
(a)(iii) \(14,\)
(a)(iv) \(12.\) M1 A1 A1
(b) \(18+16+14+12=60.\) M1
Total \(=60.\) ✓ A1 AG
(c) \(\dfrac{90}{1200}=0.075.\) M1
\(360\times0.075=27.\) A1
Common mistakes
The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.
- Confusing "population" with "sample". The population is everyone you could measure; the sample is who you actually measured. Mixing these up flips the direction of an estimate.
- Treating convenience or voluntary-response results as representative. A sample that wasn't chosen randomly can look fine but still be systematically biased - always name the specific group it's likely to over- or under-represent.
- Rounding each stratum's sample size without checking the total. Rounded stratum sizes don't always sum exactly to the target sample size - check the total, and adjust the largest stratum if needed.
- Removing outliers automatically. The \(1.5\times\text{IQR}\) rule flags a value as unusual, not as wrong - some outliers are a genuine, valid part of the data and should stay in.
Using your GDC
This topic is mostly reasoning about method and bias, but when raw sample data is given - like pilot survey results - your GDC's list and statistics tools save you from adding everything by hand. Pick your model to filter down to just the steps that apply to you.
Lists let you store data sets, generate sequences, and compute statistics without re-entering values - essential for grouped data and pilot samples.
- Enter data manually: go to the list editor and type values one by one.
- 2nd → STAT (LIST) → OPS → 5:seq(expression, variable, start, end) generates a list from a formula. For frequency data, put midpoints in L1 and frequencies in L2, then run 1-Var Stats L1, L2.TI-84
- In a Lists & Spreadsheet page, type a formula in the header row (e.g. = seq(x², x, 1, 10)) to fill the column automatically. Use two columns for midpoints and frequencies.Nspire
- In Statistics list editor, enter data directly; OPTN → LIST → Seq( builds a list from a formula. For frequency data, put midpoints in List 1 and frequencies in List 2.Casio
- To sort a list ascending: SortA(L1) on TI-84; menu → List → Sort Ascending on Nspire.
Tip: For cumulative frequency or working out \(\Sigma fx\) by hand - don't. Store midpoints in L1, frequencies in L2, and let 1-Var Stats do it all.
Instant summary statistics from a list - no formulas to compute by hand once your sample data is entered.
- Enter the data into a list.
- STAT → Edit → type values into L1. Then STAT → CALC → 1:1-Var Stats, choose L1, Calculate.TI-84
- Add a Lists & Spreadsheet page, name a column and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
- Statistics menu → enter data in List 1 → CALC (F2) → 1-VAR.Casio
- Read \(\bar{x}\) (mean), \(S_x\) (sample sd) or \(\sigma_x\) (population sd), and the five-number summary (min, \(Q_1\), median, \(Q_3\), max).
Tip: \(S_x\) vs \(\sigma_x\): use \(S_x\) (sample) when your data is a sample used to estimate the population, which is the usual case for a pilot survey.
See the full GDC guide for more calculator models and topics.
Ready to practise properly?
Sampling & data collection questions, marked instantly like the real exam.
Quick answers
The questions students on this topic ask most often.
What's the difference between a population and a sample?
The population is every individual you could possibly measure - all 5000 customers, every student in a school. A sample is the smaller group you actually collect data from, used to estimate something about the whole population without measuring everyone.
Which sampling method is "best"?
There's no single best method - it depends on the situation. Simple random and stratified sampling are generally the most representative when they're practical. Convenience and voluntary-response sampling are quick but carry a high risk of bias, so examiners usually expect you to name their specific weakness in context.
How do I find the sample size for one stratum in stratified sampling?
Multiply the stratum's population size by the overall sampling fraction (total sample size divided by total population size). Each stratum then contributes a share of the sample proportional to its share of the population.
Should I always remove outliers from a data set?
No. An outlier is any data item more than 1.5 times the interquartile range from the nearest quartile, but some outliers are genuine, valid data points, not errors. You should investigate an outlier before deciding whether to exclude it, rather than removing it automatically.
Sub-topics
Sampling & Data Collection broken down into its individual skills, each with its own focused page.
Related topics
More Statistics & Probability topics from the same AI SL syllabus unit, in case you want to keep going.