Data Types and Collection (AA HL)
Before you can sample or summarise anything, you need to know what kind of data you're holding and whether you can trust where it came from. This page covers discrete vs continuous data, how questionnaires and data sources can go wrong, and how the IB defines an outlier - with worked examples and the mistakes markers see most often. It's part of the broader Sampling & Data Collection topic.
18 questions on this sub-topic.
Key ideas
Covered under IB syllabus reference SL4.1: population, sample and random sample, discrete and continuous data, reliability and bias in data sources, and the interpretation of outliers.
Discrete vs continuous data
Discrete data can only take specific, separate values (number of pets, goals scored). Continuous data can take any value in a range (mass, time, height).
Not in the formula booklet - definitionOutliers
A data item is an outlier if it lies more than \(1.5\times\) the interquartile range (IQR) from the nearest quartile. Some outliers are genuine values; others are recording errors, and exam questions expect you to judge which given the context.
Not in the formula booklet - definitionNeed the full syllabus wording and sampling methods table? See Sampling & Data Collection.
Worked examples
State two features of a well-designed questionnaire question.
Worked solution
Any two of: it is clear and unambiguous; A1
it is not leading, and offers appropriate, non-overlapping response options. A1
A survey about internet habits is conducted only through an online form. Explain why the results may be biased.
Worked solution
Only people who already use the internet can respond, R1
so the sample is not representative of the whole population. A1
Common mistakes
- Assuming every outlier should be removed. An outlier flagged by the \(1.5\times\text{IQR}\) rule might be a genuine, valid data point rather than a recording error - the syllabus expects you to comment on which is more likely in context, not to delete it automatically.
- Mixing up discrete and continuous. "Number of siblings" is discrete even though it could theoretically grow large; "height measured to the nearest mm" is continuous because it's still an approximation of an underlying value that could fall anywhere in a range.
- Treating unreliability as only about sample size. A data source can be untrustworthy for several separate reasons at once - how recent it is, how it was collected, and whether the source has an interest in the result - not just how many people were asked.
Ready to practise properly?
18 data-types-and-collection questions, marked instantly like the real exam.
Quick answers
What is the difference between discrete and continuous data?
Discrete data can only take specific, separate values, such as the number of pets someone owns. Continuous data can take any value within a range, such as mass, time or height.
What counts as an outlier in IB Maths?
A data item that is more than \(1.5\times\) the interquartile range (IQR) from the nearest quartile. Some outliers are genuine, valid values; others are recording errors, and you're expected to comment on which is more likely given the context.