Sampling & Data Collection (AA HL)

Before you can calculate anything about a data set, you need to know how it was collected - and whether that collection method could have introduced bias. This topic covers the five sampling methods on the AA syllabus, what makes a sampling frame reliable or biased, the difference between discrete and continuous data, and how to spot an outlier using the \(1.5\times\text{IQR}\) rule.

What the syllabus says

This topic maps onto one point in the official IB Analysis & Approaches syllabus.

CodeSyllabus content
SL4.1Concepts of population, sample, random sample, discrete and continuous data. Reliability of data sources and bias in sampling - dealing with missing data and errors in recording. Interpretation of outliers, defined as a data item more than \(1.5\times\) the interquartile range (IQR) from the nearest quartile - with awareness that some outliers are valid while others may be recording errors. Sampling techniques and their effectiveness: simple random, convenience, systematic, quota and stratified sampling methods.

This is a core AA SL syllabus point that is also examinable at AA HL. Some worked examples below also draw on the basic mean/median/quartile calculations from SL4.2, since finding these is often the first step in checking whether a sample is reasonable.

Key terms

Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.

What's the difference between a population and a sample?

The population is the entire group you want to know about. A sample is a smaller subset of that population that is actually measured or surveyed, chosen so that conclusions about the sample can reasonably be generalised back to the whole population.

e.g. A school has \(800\) students (the population); a survey of \(50\) randomly chosen students is a sample.

What is a sampling frame?

A sampling frame is the actual list of individuals from which a sample is drawn. It should cover the whole population as closely as possible - if it leaves people out, the sample drawn from it can be biased, however carefully the selection itself is done.

e.g. The school register listing all \(800\) students is the sampling frame for a survey of that school.

What is simple random sampling?

Simple random sampling is a method where every member of the population has an equal, independent chance of being chosen, typically using a random-number generator on a numbered list. It's the fairest method, but needs a complete sampling frame to work.

e.g. Numbering all \(800\) students \(1\) to \(800\) and using a random-number generator to pick \(50\) of them.

What is stratified sampling?

Stratified sampling splits the population into subgroups (strata) - such as year groups or genders - then samples from each stratum in proportion to its size. This guarantees every subgroup is represented, which pure random sampling doesn't promise.

e.g. A school of \(480\) boys and \(320\) girls needs a sample of \(50\): boys \(=\dfrac{480}{800}\times50=30\), girls \(=\dfrac{320}{800}\times50=20\).

What is systematic sampling?

Systematic sampling picks every \(k\)th item from a list after a random starting point, where the interval \(k\) is the population size divided by the sample size. It's quick to carry out, but can introduce bias if the list has a hidden pattern that matches the interval.

e.g. For \(800\) students and a sample of \(50\): \(k=800/50=16\), so pick a random start from \(1\) to \(16\), then every \(16\)th student.

Key formulas

This topic is more conceptual than most, but there are two calculations you're expected to carry out directly, plus the outlier rule from SL4.1. The tables below summarise them - the explanations underneath go into more depth.

Formula reference

None of the calculations below appear in the formula booklet as a named formula - they're direct applications of proportional reasoning and the outlier definition given in the syllabus itself.

FormulaUsed forBooklet?
\(\text{stratum sample size} = \dfrac{\text{stratum size}}{\text{population size}}\times n\)Stratified samplingNot in the formula booklet - proportional reasoning
\(k = \dfrac{N}{n}\)Systematic sampling intervalNot in the formula booklet - direct calculation
outlier if \(Q_3+1.5\times\text{IQR}\)Outlier detectionNot in the formula booklet - definition given in the syllabus guidance

Probability sampling vs non-probability sampling

Every method on the syllabus falls into one of these two categories, and knowing which one you're dealing with tells you a lot about how trustworthy the sample is likely to be.

FeatureProbability samplingNon-probability sampling
Selectionevery member has a known, non-zero chance of selectionselection is not random - based on convenience or fixed quotas
Examplessimple random, systematic, stratifiedconvenience, quota
Bias risklow, provided the sampling frame is completehigher - can easily be unrepresentative
Generalising to the populationstatistically justifiedriskier, even with a large sample

The sampling methods

The IB syllabus names five sampling techniques - you need to be able to name, describe and evaluate each one.

Simple random sampling

Every member of the population has an equal, independent chance of selection - usually via a random-number generator applied to a numbered sampling frame.

Not in the formula booklet - definition

Systematic sampling

\[k = N/n\]

Pick a random start within the first interval, then select every \(k\)th item afterwards.

Not in the formula booklet - definition

Stratified sampling

Divide the population into subgroups, then sample from each in proportion to its size, so every subgroup is represented fairly.

Not in the formula booklet - definition

Quota and convenience sampling

Quota sampling fills fixed subgroup targets with whoever is willing; convenience sampling simply takes whoever is easiest to reach. Both are non-random and can be biased.

Not in the formula booklet - definition

Data quality and outliers

Sampling frame and bias

A sampling frame that excludes part of the population (e.g. a telephone directory missing people without a landline) produces a biased sample, however random the selection itself is.

Not in the formula booklet - definition

Discrete vs continuous data

Discrete data can only take specific, separate values (number of pets, goals scored). Continuous data can take any value in a range (mass, time, height).

Not in the formula booklet - definition

Outliers

\[Q_3+1.5\,\text{IQR}\]

A data item this far from the nearest quartile is flagged as an outlier - but check whether it's a genuine (valid) extreme value or a recording error before deciding what to do with it.

Not in the formula booklet - syllabus guidance definition

Worked examples

Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.

1
Easy
No calc
[4 marks]

Name the sampling method used in each case.

(a) Every 8th item on a production line is checked.

(b) The first 30 people leaving a shop are surveyed.

(c) A researcher numbers the population and uses a random-number generator to select 30 people. Name this method, and state one advantage it has over method (a).

Worked solution

(a) Systematic sampling. A1

(b) Convenience sampling. A1

(c) Simple random sampling. A1
Every member has an equal, independent chance of selection, avoiding any periodicity bias that systematic sampling can suffer. A1

A1 Systematic A1 Convenience A1 Names simple random sampling A1 Advantage over systematic sampling
2
Medium
No calc
[4 marks]

A college has 450 arts, 300 science and 150 sport students. A stratified sample of 90 is taken.

Find the number selected from each group.

(a) Find the number from arts.

(b) Find the number from science.

(c) Find the number from sport.

Worked solution

(b) science \(=30,\) A1

(c) sport \(=15.\) A1

M1 Fraction A1 Correct answer of 45 A1 Correct answer of 30 A1 Correct answer of 15

Common mistakes

The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.

  • Mixing up systematic and stratified sampling. Systematic sampling picks every \(k\)th item from one list; stratified sampling splits the population into subgroups first and samples proportionally from each - they're testing different things and aren't interchangeable answers.
  • Forgetting that a sample can be biased even with perfect randomness. If the sampling frame itself excludes part of the population (like a directory missing unlisted numbers), no amount of careful random selection afterwards fixes that.
  • Rounding stratum sizes without checking they sum to the total. When \(\dfrac{\text{stratum}}{\text{population}}\times n\) doesn't give a whole number, round sensibly and check the parts still add up to the required total sample size.
  • Assuming every outlier should be removed. An outlier flagged by the \(1.5\times\text{IQR}\) rule might be a genuine, valid data point rather than a recording error - the syllabus expects you to comment on which is more likely in context, not to delete it automatically.

Using your GDC

Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.

Show steps for:
Build and manipulate lists

Lists let you store data sets, generate sequences, and compute statistics without re-entering values - essential for grouped data and frequency tables when checking a sample.

  1. Enter data manually: go to the list editor and type values one by one.
  2. 2nd → STAT (LIST) → OPS → 5:seq(expression, variable, start, end) generates a list from a formula, e.g. seq(X², X, 1, 10). For grouped data, put midpoints in L1 and frequencies in L2, then run 1-Var Stats L1, L2.TI-84
  3. In a Lists & Spreadsheet page, type a formula in the header row (e.g. = seq(x², x, 1, 10)) to fill the column automatically. Use two columns for midpoints and frequencies.Nspire
  4. In Statistics list editor, enter data directly; OPTN → LIST → Seq( builds a list from a formula. For frequency data, put midpoints in List 1 and frequencies in List 2.Casio
  5. For frequency data (grouped), enter midpoints in one list and frequencies in another, then reference both lists when running 1-Var Stats.
  6. To sort a list ascending: SortA(L1) on TI-84; menu → List → Sort Ascending on Nspire.
  7. Sum a list: sum(L1) on TI-84; sum() on Nspire; OPTN → LIST → Sum on Casio.

Tip: Building your sample data as a list first makes every later calculation (mean, quartiles, outlier check) a single menu command rather than a re-entry.

One-variable statistics (mean, median, standard deviation)

Instant summary statistics from a list - no formulas to compute by hand. Useful once your sample is entered, to check its mean, median or quartiles.

  1. Enter the data into a list.
  2. STAT → Edit → type values into L1. Then STAT → CALC → 1:1-Var Stats, choose L1, Calculate.TI-84
  3. Add a Lists & Spreadsheet page, name a column and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
  4. Statistics menu → enter data in List 1 → CALC (F2) → 1-VAR.Casio
  5. Read x̄ (mean), Sx (sample sd) or σx (population sd), and the five-number summary (min, Q1, median, Q3, max).
  6. For frequency data, put values in one list and frequencies in another and set the frequency list.

Tip: Sx vs σx: use σx (population) for a complete data set, Sx (sample) for a sample. IB usually wants σx.

See the full GDC guide for more calculator models and topics.

Ready to practise properly?

Sampling & data collection questions, marked instantly like the real exam.

Quick answers

The questions students on this topic ask most often.

What's the difference between a population and a sample?

The population is the entire group you're interested in - every student in a school, every item off a production line. A sample is a smaller subset of that population that is actually surveyed or measured, chosen so it (hopefully) represents the whole population well.

Which sampling method is 'best'?

Simple random sampling is the fairest, since every member of the population has an equal, independent chance of selection - but stratified sampling is often more practical when the population has distinct subgroups, since it guarantees each subgroup is represented in proportion.

How do I find the sample size for each stratum in stratified sampling?

Multiply the total sample size by the fraction of the population that stratum makes up: sample size for a stratum \(=\dfrac{\text{stratum size}}{\text{population size}}\times n.\) Round to the nearest whole number if needed.

What makes a sampling frame biased?

A sampling frame is biased if it systematically excludes part of the population - for example, a telephone directory misses anyone without a listed landline. Always check whether everyone in the population had a genuine chance of appearing on the list the sample was drawn from.

Sub-topics

Sampling & Data Collection broken down into its individual skills, each with its own focused page.