Sampling & Data Collection (AA SL)

Almost every statistical claim is based on a sample, not the whole population, so how that sample is chosen matters enormously. This topic covers the difference between a population and a sample, the main sampling techniques (simple random, systematic, stratified, quota and convenience), how to spot bias in a sampling method or a survey question, and classifying data as discrete or continuous.

What the syllabus says

This topic maps onto one point in the official IB Analysis & Approaches syllabus.

CodeSyllabus content
SL4.1Concepts of population, sample, random sample, discrete and continuous data. Reliability of data sources and bias in sampling. Interpretation of outliers. Sampling techniques and their effectiveness: simple random, convenience, systematic, quota and stratified sampling methods.

This is a core AA SL syllabus point that is also examinable at AA HL.

Key terms

Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.

What is the difference between a population and a sample?

A population is the entire group you want to draw conclusions about. A sample is a smaller subset of the population that is actually studied, chosen because surveying the whole population is usually too slow, costly or impossible.

e.g. All 800 students at a school form the population; a stratified sample of 40 of them is the sample.

What is stratified sampling?

Stratified sampling divides the population into subgroups (strata) and samples from each in proportion to its size, guaranteeing that every subgroup is represented fairly. The proportion sampled from each stratum equals the overall sampling fraction.

e.g. For 240 staff (120 production, 80 sales, 40 admin) and a sample of 30, the fraction is \(\tfrac{30}{240}=\tfrac18\), so production contributes \(\tfrac18\times120=15.\)

What is systematic sampling?

Systematic sampling selects every \(k\)th member of a list after a random starting point, where \(k\) is the sampling interval. It's fast and spreads the sample evenly, but can be biased if the list has a hidden pattern matching the interval.

e.g. For a population of 800 and a sample of 40, the interval is \(k=800/40=20\), so every 20th person is chosen.

What is convenience sampling?

Convenience sampling selects whoever is easiest to reach, such as the first people who happen to walk past. It's quick to carry out but is likely to be biased, since it only reaches a particular kind of person rather than a representative cross-section.

e.g. Surveying people leaving a gym about exercise habits over-represents people who already exercise regularly.

What is sampling bias?

Sampling bias occurs when the way a sample is chosen makes it systematically unrepresentative of the population, so conclusions drawn from it are unreliable. It can come from the sampling method itself, or from how survey questions are worded.

e.g. "Don't you agree that taxes are too high?" is a leading question that pushes respondents toward one answer.

Key formulas

This topic is mostly conceptual, but stratified and systematic sampling both rely on one simple calculation. The tables below summarise the sampling methods at a glance - the explanations underneath go into more depth.

Formula reference

The sampling fraction is not a listed formula booklet entry - it's a straightforward ratio you're expected to know and apply.

FormulaUsed forBooklet?
Sampling fraction \(=\dfrac{\text{sample size}}{\text{population size}}\)Proportion of the population sampled from each stratumNot in booklet - prior knowledge
Stratum sample \(=\) fraction \(\times\) stratum sizeNumber sampled from a given subgroupNot in booklet - prior knowledge
Systematic interval \(k=\dfrac{\text{population size}}{\text{sample size}}\)Gap between selected members in systematic samplingNot in booklet - prior knowledge

Sampling methods compared

The five named sampling methods trade off representativeness against ease of collection - knowing which is which, and what their strengths and weaknesses are, is the core of this topic.

MethodHow it worksMain weakness
Simple randomEvery member has an equal chance of selectionCan still miss small subgroups by chance
SystematicEvery \(k\)th member selected after a random startBiased if the list has periodicity matching \(k\)
StratifiedSample taken from each subgroup, proportional to sizeNeeds the population divided into known strata first
QuotaInterviewer fills fixed quotas for each subgroupSelection within each quota is not random
ConvenienceWhoever is easiest to reach is sampledUsually unrepresentative of the wider population

Data and reliability

Discrete vs continuous

Discrete data is counted (whole numbers); continuous data is measured (any value in a range).

Number of cars is discrete; height or time is continuous.

Not in the formula booklet - prior knowledge

Reliability of data

Consider missing data, recording errors, and whether the sampling method itself introduces bias.

A large sample doesn't fix bias - a biased method stays biased no matter how many people you ask.

Not in the formula booklet - key exam idea

Outliers

An outlier is a data item more than \(1.5\times\) IQR from the nearest quartile.

Some outliers are genuine and should stay in the data; others are recording errors and should be investigated.

Not in the formula booklet - definition to know

Worked examples

Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.

1
Easy
No calc
[3 marks]

Consider the following variables.

(a) Classify the number of cars in a car park as discrete or continuous.

(b) Classify the height of a tree as discrete or continuous.

(c) Classify the time to run 100 m as discrete or continuous.

Worked solution

(a) Discrete (counted). A1

(b) Continuous (measured). A1

(c) Continuous (measured). A1

A1 Discrete A1 Continuous
2
Hard
No calc
[8 marks]

A school has 600 students in three year groups in the ratio \(5:4:3\). A stratified sample of size \(n\) is taken; 20 students are chosen from the largest year group.

(a)(i) Find the number of students in the largest year group.

(a)(ii) Find the number of students in the middle year group.

(a)(iii) Find the number of students in the smallest year group.

(b) Find the sampling fraction.

(c) Hence find \(n\).

Worked solution

(a)(i) Parts \(=5+4+3=12;\) largest group \(=\tfrac{5}{12}\times600\) M1
\(=250.\) A1

(a)(ii) Middle group \(=\tfrac{4}{12}\times600=200.\) A1

(a)(iii) Smallest group \(=\tfrac{3}{12}\times600=150.\) A1

(b) Largest group \(=250,\) and \(20\) chosen: fraction \(=\tfrac{20}{250}\) M1
\(=\tfrac{2}{25}.\) A1

(c) \(n=\tfrac{2}{25}\times600\) M1
\(=48.\) A1

M1 Find the ratio fraction and apply it to the total of 600 students A1 Correct largest group of 250 A1 Correct middle group of 200 A1 Correct smallest group of 150 M1 Form the sampling fraction using 20 out of the largest group of 250 A1 Correct sampling fraction of 2/25 M1 Apply the sampling fraction to the total population of 600 A1 Correct value n=48 (FT)

Common mistakes

The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.

  • Mixing up stratified and quota sampling. Stratified sampling selects randomly within each proportional subgroup; quota sampling only fixes the quota size and lets the interviewer choose who fills it, so it isn't random within each group.
  • Applying the population's sampling fraction to the wrong total. The fraction (sample size ÷ population size) must be multiplied by each stratum's own size, not by the total sample size or population size again.
  • Assuming a bigger sample automatically fixes bias. A biased sampling method - like convenience sampling - stays biased no matter how large the sample gets; only changing the method fixes it.
  • Classifying data by its numerical appearance rather than how it arises. A quantity like shoe size is discrete because it's counted from a fixed set of values, even though it can include halves - the test is whether it's measured (continuous) or counted (discrete), not whether it "looks like a whole number".

Using your GDC

Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.

Show steps for:
Build and manipulate lists

Once a sample is collected, lists let you store and summarise it without re-entering values - essential once you move from "how was this sampled" to "what does the sample tell us".

  1. Enter data manually: go to the list editor and type values one by one.
  2. 2nd → STAT (LIST) → OPS → 5:seq(expression, variable, start, end) generates a list from a formula. For grouped data, put midpoints in L1 and frequencies in L2, then run 1-Var Stats L1, L2.TI-84
  3. In a Lists & Spreadsheet page, type a formula in the header row to fill a column automatically. Use two columns for midpoints and frequencies.Nspire
  4. In Statistics list editor, enter data directly; OPTN → LIST → Seq( builds a list from a formula. For frequency data, put midpoints in List 1 and frequencies in List 2.Casio
  5. Sum a list to check totals: sum(L1) on TI-84; sum() on Nspire; OPTN → LIST → Sum on Casio.

Tip: For cumulative frequency or working out totals by hand - don't. Store midpoints in one list, frequencies in another, and let 1-Var Stats do it all.

One-variable statistics (mean, median, standard deviation)

Once a proper sample has been collected, this gives instant summary statistics - and lets you compare how representative different sampling methods turn out to be.

  1. Enter the sample data into a list.
  2. STAT → Edit → type values into L1. Then STAT → CALC → 1:1-Var Stats, choose L1, Calculate.TI-84
  3. Add a Lists & Spreadsheet page, name a column and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
  4. Statistics menu → enter data in List 1 → CALC (F2) → 1-VAR.Casio
  5. Read \(\bar{x}\) (mean), \(S_x\) (sample sd) or \(\sigma_x\) (population sd).

Tip: Sx vs σx: use σx (population) for a complete data set, Sx (sample) for a sample. IB usually wants σx.

See the full GDC guide for more calculator models and topics.

Ready to practise properly?

Sampling and data collection questions, marked instantly like the real exam.

Quick answers

The questions students on this topic ask most often.

What's the difference between a population and a sample?

A population is the entire group you're interested in studying. A sample is a smaller subset of the population that is actually surveyed or measured. Statistics is largely about using what you find in a sample to draw reliable conclusions about the whole population.

When should I use stratified sampling instead of simple random sampling?

Use stratified sampling when the population has distinct subgroups (like year groups or departments) and you want each one represented in proportion to its size. Simple random sampling doesn't guarantee this - by chance it could under- or over-represent a subgroup, especially with a small sample.

Why is convenience sampling usually unreliable?

Convenience sampling only reaches people who happen to be present or easy to reach at a particular time and place, which is rarely representative of the wider population. Surveying gym-goers about exercise habits, for example, over-represents people who already exercise a lot.

How do I calculate a stratified sample size for one group?

Find the sampling fraction - the total sample size divided by the total population - then multiply that fraction by the size of the subgroup. For example, sampling 30 from a population of 240 gives a fraction of 1/8, so a subgroup of 120 contributes 1/8 × 120 = 15 to the sample.

Sub-topics

Sampling & Data Collection broken down into its individual skills, each with its own focused page.