Sampling & Data Collection (AI SL)

Before any statistics can be trusted, the data behind it has to be collected sensibly. This topic covers the difference between a population and a sample, the five named sampling methods on the syllabus - simple random, convenience, systematic, quota and stratified - how to spot bias in a sampling design, and how the syllabus defines an outlier.

What the syllabus says

This topic maps onto one point in the official IB Applications & Interpretation syllabus.

CodeSyllabus content
SL4.1Concepts of population, sample, random sample, discrete and continuous data. Reliability of data sources and bias in sampling. Interpretation of outliers, defined as a data item more than 1.5\(\times\) the interquartile range (IQR) from the nearest quartile. Sampling techniques and their effectiveness: simple random, convenience, systematic, quota and stratified sampling methods.

Some outliers are a valid part of the sample; others may be errors in the data - the syllabus expects awareness of both possibilities before deciding what to do with one.

Key terms

Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.

What is a population vs a sample?

The population is every individual you could possibly measure; a sample is the smaller subset you actually collect data from. A random sample gives every member of the population a known, non-zero chance of being chosen, which is what makes it possible to trust the sample to represent the whole.

e.g. Population: 5000 customers. Sample: 100 customers surveyed by picking random numbers 1-5000.

What is stratified sampling?

Stratified sampling splits the population into groups (strata) - such as year groups or departments - and samples from each group in proportion to its size, using the same sampling fraction throughout. It guarantees every group is represented, unlike a simple random sample which could miss a small group by chance.

e.g. Sampling fraction \(\dfrac{75}{900}=\dfrac{1}{12}\); a stratum of 300 contributes \(300\times\dfrac{1}{12}=25\).

What is systematic sampling?

Systematic sampling selects every \(k\)th member of a list, where the sampling interval \(k\) is the population size divided by the desired sample size, after a random starting point. It's quick to apply to an ordered list, but can introduce bias if the list has a hidden repeating pattern.

e.g. For 600 records and a sample of 20: interval \(=\dfrac{600}{20}=30\).

What is convenience sampling, and why can it be biased?

Convenience sampling selects whoever is easiest to reach - the first people you meet, rather than a randomly chosen group. It's fast but risks systematic bias, since the people who happen to be convenient to sample are often not representative of the whole population.

e.g. Surveying the first 30 shoppers at 9am misses everyone who shops later in the day.

What is an outlier?

The syllabus defines an outlier as a data item more than 1.5 times the interquartile range (IQR) from the nearest quartile. Some outliers are genuine, valid data; others are recording errors - the definition flags a value as unusual, but doesn't decide which it is.

e.g. If \(Q_1=10\), \(Q_3=20\), then \(\text{IQR}=10\), so any value above \(20+1.5(10)=35\) is an outlier.

Key formulas

This topic is mostly about identifying the right method and its weaknesses, but three genuine calculations come up repeatedly. The tables below summarise them - the explanations underneath go into more depth on each one.

Formula reference

None of these are officially listed in the formula booklet - they're direct-proportion reasoning and a syllabus-stated convention, not booklet entries.

FormulaUsed forBooklet?
Stratum sample size \(= \dfrac{\text{sample size}}{\text{population size}}\times\) stratum sizeStratified samplingNot in booklet - direct proportion
Sampling interval \(= \dfrac{\text{population size}}{\text{sample size}}\)Systematic samplingNot in booklet - direct proportion
Outlier if more than \(1.5\times\text{IQR}\) from nearest quartileInterpreting outliersNot in booklet - syllabus-defined convention

Probability sampling vs non-probability sampling

The five named methods split naturally into two groups, depending on whether every member of the population has a known chance of selection.

FeatureProbability samplingNon-probability sampling
MethodsSimple random, systematic, stratifiedConvenience, quota
SelectionEvery member has a known chance of being chosenChosen by ease of access or a fixed target count
Bias riskLow, if applied correctlyHigher - depends on who happens to be available
ExampleEvery 20th shopper all day (systematic)First 30 shoppers at 9am (convenience)

Sampling techniques

All five methods appear on the syllabus by name - you should be able to identify each one from a short description.

Simple random sampling

Number every member of the population, then use random numbers to pick the sample - every individual and every combination has an equal chance of selection.

Not in the formula booklet - selection procedure

Stratified sampling

Split the population into strata, then sample from each using the same sampling fraction, so every group is represented proportionally.

Not in the formula booklet - direct proportion

Systematic sampling

Pick every \(k\)th item from an ordered list after a random start, spreading the sample evenly across the whole population.

Not in the formula booklet - direct proportion

Convenience & quota sampling

Convenience takes whoever is easiest to reach; quota fixes target numbers per group but still lets the interviewer choose anyone willing - both are non-random and can be biased.

Not in the formula booklet - selection procedure

Bias, reliability and outliers

Choosing a method is only half the job - you also need to judge how trustworthy the resulting data is.

Sample size and reliability

A larger sample reduces sampling variability and generally gives a more reliable estimate of the true population value, all else being equal.

Not in the formula booklet - general principle

Sources of bias

Watch for non-response bias, self-selection (voluntary-response) bias, and time-of-day or location bias - name the specific mechanism, not just "it could be biased".

Not in the formula booklet - identified in context

Interpreting an outlier

An outlier flagged by the \(1.5\times\text{IQR}\) rule may be a genuine extreme value or a recording error - investigate before removing it from the data set.

Not in the formula booklet - syllabus-defined convention

Worked examples

Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.

1
Easy
[6 marks]

A stratified sample uses fraction 0.05. A stratum has 640 members.

(a) Find the number sampled.

(b) If the fraction were instead 0.08, find the new number sampled.

(c) Find the total population size if this stratum represents 40% of the total population.

Worked solution

(a) \(0.05\times640=32.\) M1 A1

(b) \(0.08\times640=51.2\) M1
\(\approx51.\) A1

(c) \(640\div0.4=1600.\) M1 A1

M1 Attempt at the stratified sample size 0.05×640 A1 Correct value 32 M1 Attempt at the recalculated sample size 0.08×640 A1 Correct value 51 (rounded) M1 Attempt at the total population 640÷0.4 A1 Correct value 1600
2
Medium
[7 marks]

A college has 1200 students in four years: 360, 320, 280, 240.

(a)(i) A stratified sample of 60 is required. Find the number from year 1 (largest group).

(a)(ii) Find the number from year 2.

(a)(iii) Find the number from year 3.

(a)(iv) Find the number from year 4 (smallest group).

(b) Verify that the four sample sizes sum to 60.

(c) The college instead wants a total sample of 90. Find the new number of students required from the largest year group.

Worked solution

(a)(i) \(18,\)

(a)(ii) \(16,\)

(a)(iii) \(14,\)

(a)(iv) \(12.\) M1 A1 A1

Multiply each by 0.05.

(b) \(18+16+14+12=60.\) M1
Total \(=60.\) ✓ A1 AG

(c) \(\dfrac{90}{1200}=0.075.\) M1
\(360\times0.075=27.\) A1

M1 Attempt sampling fraction \(60/1200=0.05\) A1 At least two of \(18,16,14,12\) correct A1 All four values correct M1 Sum the four values A1 Confirms total \(60\) M1 New sampling fraction \(90/1200=0.075\) A1 Correct answer of \(27\)

Common mistakes

The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.

  • Confusing "population" with "sample". The population is everyone you could measure; the sample is who you actually measured. Mixing these up flips the direction of an estimate.
  • Treating convenience or voluntary-response results as representative. A sample that wasn't chosen randomly can look fine but still be systematically biased - always name the specific group it's likely to over- or under-represent.
  • Rounding each stratum's sample size without checking the total. Rounded stratum sizes don't always sum exactly to the target sample size - check the total, and adjust the largest stratum if needed.
  • Removing outliers automatically. The \(1.5\times\text{IQR}\) rule flags a value as unusual, not as wrong - some outliers are a genuine, valid part of the data and should stay in.

Using your GDC

This topic is mostly reasoning about method and bias, but when raw sample data is given - like pilot survey results - your GDC's list and statistics tools save you from adding everything by hand. Pick your model to filter down to just the steps that apply to you.

Show steps for:
Build and manipulate lists

Lists let you store data sets, generate sequences, and compute statistics without re-entering values - essential for grouped data and pilot samples.

  1. Enter data manually: go to the list editor and type values one by one.
  2. 2nd → STAT (LIST) → OPS → 5:seq(expression, variable, start, end) generates a list from a formula. For frequency data, put midpoints in L1 and frequencies in L2, then run 1-Var Stats L1, L2.TI-84
  3. In a Lists & Spreadsheet page, type a formula in the header row (e.g. = seq(x², x, 1, 10)) to fill the column automatically. Use two columns for midpoints and frequencies.Nspire
  4. In Statistics list editor, enter data directly; OPTN → LIST → Seq( builds a list from a formula. For frequency data, put midpoints in List 1 and frequencies in List 2.Casio
  5. To sort a list ascending: SortA(L1) on TI-84; menu → List → Sort Ascending on Nspire.

Tip: For cumulative frequency or working out \(\Sigma fx\) by hand - don't. Store midpoints in L1, frequencies in L2, and let 1-Var Stats do it all.

One-variable statistics (mean, median, standard deviation)

Instant summary statistics from a list - no formulas to compute by hand once your sample data is entered.

  1. Enter the data into a list.
  2. STAT → Edit → type values into L1. Then STAT → CALC → 1:1-Var Stats, choose L1, Calculate.TI-84
  3. Add a Lists & Spreadsheet page, name a column and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
  4. Statistics menu → enter data in List 1 → CALC (F2) → 1-VAR.Casio
  5. Read \(\bar{x}\) (mean), \(S_x\) (sample sd) or \(\sigma_x\) (population sd), and the five-number summary (min, \(Q_1\), median, \(Q_3\), max).

Tip: \(S_x\) vs \(\sigma_x\): use \(S_x\) (sample) when your data is a sample used to estimate the population, which is the usual case for a pilot survey.

See the full GDC guide for more calculator models and topics.

Ready to practise properly?

Sampling & data collection questions, marked instantly like the real exam.

Quick answers

The questions students on this topic ask most often.

What's the difference between a population and a sample?

The population is every individual you could possibly measure - all 5000 customers, every student in a school. A sample is the smaller group you actually collect data from, used to estimate something about the whole population without measuring everyone.

Which sampling method is "best"?

There's no single best method - it depends on the situation. Simple random and stratified sampling are generally the most representative when they're practical. Convenience and voluntary-response sampling are quick but carry a high risk of bias, so examiners usually expect you to name their specific weakness in context.

How do I find the sample size for one stratum in stratified sampling?

Multiply the stratum's population size by the overall sampling fraction (total sample size divided by total population size). Each stratum then contributes a share of the sample proportional to its share of the population.

Should I always remove outliers from a data set?

No. An outlier is any data item more than 1.5 times the interquartile range from the nearest quartile, but some outliers are genuine, valid data points, not errors. You should investigate an outlier before deciding whether to exclude it, rather than removing it automatically.

Sub-topics

Sampling & Data Collection broken down into its individual skills, each with its own focused page.