Sampling & Data Collection (AA SL)
Almost every statistical claim is based on a sample, not the whole population, so how that sample is chosen matters enormously. This topic covers the difference between a population and a sample, the main sampling techniques (simple random, systematic, stratified, quota and convenience), how to spot bias in a sampling method or a survey question, and classifying data as discrete or continuous.
What the syllabus says
This topic maps onto one point in the official IB Analysis & Approaches syllabus.
| Code | Syllabus content |
|---|---|
| SL4.1 | Concepts of population, sample, random sample, discrete and continuous data. Reliability of data sources and bias in sampling. Interpretation of outliers. Sampling techniques and their effectiveness: simple random, convenience, systematic, quota and stratified sampling methods. |
This is a core AA SL syllabus point that is also examinable at AA HL.
Key terms
Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.
What is the difference between a population and a sample?
A population is the entire group you want to draw conclusions about. A sample is a smaller subset of the population that is actually studied, chosen because surveying the whole population is usually too slow, costly or impossible.
e.g. All 800 students at a school form the population; a stratified sample of 40 of them is the sample.
What is stratified sampling?
Stratified sampling divides the population into subgroups (strata) and samples from each in proportion to its size, guaranteeing that every subgroup is represented fairly. The proportion sampled from each stratum equals the overall sampling fraction.
e.g. For 240 staff (120 production, 80 sales, 40 admin) and a sample of 30, the fraction is \(\tfrac{30}{240}=\tfrac18\), so production contributes \(\tfrac18\times120=15.\)
What is systematic sampling?
Systematic sampling selects every \(k\)th member of a list after a random starting point, where \(k\) is the sampling interval. It's fast and spreads the sample evenly, but can be biased if the list has a hidden pattern matching the interval.
e.g. For a population of 800 and a sample of 40, the interval is \(k=800/40=20\), so every 20th person is chosen.
What is convenience sampling?
Convenience sampling selects whoever is easiest to reach, such as the first people who happen to walk past. It's quick to carry out but is likely to be biased, since it only reaches a particular kind of person rather than a representative cross-section.
e.g. Surveying people leaving a gym about exercise habits over-represents people who already exercise regularly.
What is sampling bias?
Sampling bias occurs when the way a sample is chosen makes it systematically unrepresentative of the population, so conclusions drawn from it are unreliable. It can come from the sampling method itself, or from how survey questions are worded.
e.g. "Don't you agree that taxes are too high?" is a leading question that pushes respondents toward one answer.
Key formulas
This topic is mostly conceptual, but stratified and systematic sampling both rely on one simple calculation. The tables below summarise the sampling methods at a glance - the explanations underneath go into more depth.
Formula reference
The sampling fraction is not a listed formula booklet entry - it's a straightforward ratio you're expected to know and apply.
| Formula | Used for | Booklet? |
|---|---|---|
| Sampling fraction \(=\dfrac{\text{sample size}}{\text{population size}}\) | Proportion of the population sampled from each stratum | Not in booklet - prior knowledge |
| Stratum sample \(=\) fraction \(\times\) stratum size | Number sampled from a given subgroup | Not in booklet - prior knowledge |
| Systematic interval \(k=\dfrac{\text{population size}}{\text{sample size}}\) | Gap between selected members in systematic sampling | Not in booklet - prior knowledge |
Sampling methods compared
The five named sampling methods trade off representativeness against ease of collection - knowing which is which, and what their strengths and weaknesses are, is the core of this topic.
| Method | How it works | Main weakness |
|---|---|---|
| Simple random | Every member has an equal chance of selection | Can still miss small subgroups by chance |
| Systematic | Every \(k\)th member selected after a random start | Biased if the list has periodicity matching \(k\) |
| Stratified | Sample taken from each subgroup, proportional to size | Needs the population divided into known strata first |
| Quota | Interviewer fills fixed quotas for each subgroup | Selection within each quota is not random |
| Convenience | Whoever is easiest to reach is sampled | Usually unrepresentative of the wider population |
Data and reliability
Discrete vs continuous
Discrete data is counted (whole numbers); continuous data is measured (any value in a range).
Number of cars is discrete; height or time is continuous.
Not in the formula booklet - prior knowledgeReliability of data
Consider missing data, recording errors, and whether the sampling method itself introduces bias.
A large sample doesn't fix bias - a biased method stays biased no matter how many people you ask.
Not in the formula booklet - key exam ideaOutliers
An outlier is a data item more than \(1.5\times\) IQR from the nearest quartile.
Some outliers are genuine and should stay in the data; others are recording errors and should be investigated.
Not in the formula booklet - definition to knowWorked examples
Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.
Consider the following variables.
(a) Classify the number of cars in a car park as discrete or continuous.
(b) Classify the height of a tree as discrete or continuous.
(c) Classify the time to run 100 m as discrete or continuous.
Worked solution
(a) Discrete (counted). A1
(b) Continuous (measured). A1
(c) Continuous (measured). A1
A school has 600 students in three year groups in the ratio \(5:4:3\). A stratified sample of size \(n\) is taken; 20 students are chosen from the largest year group.
(a)(i) Find the number of students in the largest year group.
(a)(ii) Find the number of students in the middle year group.
(a)(iii) Find the number of students in the smallest year group.
(b) Find the sampling fraction.
(c) Hence find \(n\).
Worked solution
(a)(i) Parts \(=5+4+3=12;\) largest group \(=\tfrac{5}{12}\times600\) M1
\(=250.\) A1
(a)(ii) Middle group \(=\tfrac{4}{12}\times600=200.\) A1
(a)(iii) Smallest group \(=\tfrac{3}{12}\times600=150.\) A1
(b) Largest group \(=250,\) and \(20\) chosen: fraction \(=\tfrac{20}{250}\) M1
\(=\tfrac{2}{25}.\) A1
(c) \(n=\tfrac{2}{25}\times600\) M1
\(=48.\) A1
Common mistakes
The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.
- Mixing up stratified and quota sampling. Stratified sampling selects randomly within each proportional subgroup; quota sampling only fixes the quota size and lets the interviewer choose who fills it, so it isn't random within each group.
- Applying the population's sampling fraction to the wrong total. The fraction (sample size ÷ population size) must be multiplied by each stratum's own size, not by the total sample size or population size again.
- Assuming a bigger sample automatically fixes bias. A biased sampling method - like convenience sampling - stays biased no matter how large the sample gets; only changing the method fixes it.
- Classifying data by its numerical appearance rather than how it arises. A quantity like shoe size is discrete because it's counted from a fixed set of values, even though it can include halves - the test is whether it's measured (continuous) or counted (discrete), not whether it "looks like a whole number".
Using your GDC
Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.
Once a sample is collected, lists let you store and summarise it without re-entering values - essential once you move from "how was this sampled" to "what does the sample tell us".
- Enter data manually: go to the list editor and type values one by one.
- 2nd → STAT (LIST) → OPS → 5:seq(expression, variable, start, end) generates a list from a formula. For grouped data, put midpoints in L1 and frequencies in L2, then run 1-Var Stats L1, L2.TI-84
- In a Lists & Spreadsheet page, type a formula in the header row to fill a column automatically. Use two columns for midpoints and frequencies.Nspire
- In Statistics list editor, enter data directly; OPTN → LIST → Seq( builds a list from a formula. For frequency data, put midpoints in List 1 and frequencies in List 2.Casio
- Sum a list to check totals: sum(L1) on TI-84; sum() on Nspire; OPTN → LIST → Sum on Casio.
Tip: For cumulative frequency or working out totals by hand - don't. Store midpoints in one list, frequencies in another, and let 1-Var Stats do it all.
Once a proper sample has been collected, this gives instant summary statistics - and lets you compare how representative different sampling methods turn out to be.
- Enter the sample data into a list.
- STAT → Edit → type values into L1. Then STAT → CALC → 1:1-Var Stats, choose L1, Calculate.TI-84
- Add a Lists & Spreadsheet page, name a column and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
- Statistics menu → enter data in List 1 → CALC (F2) → 1-VAR.Casio
- Read \(\bar{x}\) (mean), \(S_x\) (sample sd) or \(\sigma_x\) (population sd).
Tip: Sx vs σx: use σx (population) for a complete data set, Sx (sample) for a sample. IB usually wants σx.
See the full GDC guide for more calculator models and topics.
Ready to practise properly?
Sampling and data collection questions, marked instantly like the real exam.
Quick answers
The questions students on this topic ask most often.
What's the difference between a population and a sample?
A population is the entire group you're interested in studying. A sample is a smaller subset of the population that is actually surveyed or measured. Statistics is largely about using what you find in a sample to draw reliable conclusions about the whole population.
When should I use stratified sampling instead of simple random sampling?
Use stratified sampling when the population has distinct subgroups (like year groups or departments) and you want each one represented in proportion to its size. Simple random sampling doesn't guarantee this - by chance it could under- or over-represent a subgroup, especially with a small sample.
Why is convenience sampling usually unreliable?
Convenience sampling only reaches people who happen to be present or easy to reach at a particular time and place, which is rarely representative of the wider population. Surveying gym-goers about exercise habits, for example, over-represents people who already exercise a lot.
How do I calculate a stratified sample size for one group?
Find the sampling fraction - the total sample size divided by the total population - then multiply that fraction by the size of the subgroup. For example, sampling 30 from a population of 240 gives a fraction of 1/8, so a subgroup of 120 contributes 1/8 × 120 = 15 to the sample.
Sub-topics
Sampling & Data Collection broken down into its individual skills, each with its own focused page.
Related topics
More Statistics & Probability topics from the same AA SL syllabus unit, in case you want to keep going.