Hypothesis Testing (AI HL)
Hypothesis testing gives you a formal way to decide whether a pattern in a sample of data is real, or just what you'd expect from random chance. This topic covers writing null and alternative hypotheses, the chi-squared test for goodness-of-fit and independence, the t-test for a population mean, and how to compare a p-value or test statistic against a significance level to reach a defensible conclusion.
What the syllabus says
This topic maps onto two points in the official IB Applications & Interpretation syllabus.
| Code | Syllabus content |
|---|---|
| SL4.11 | Formulation of null and alternative hypotheses (\(H_0\) and \(H_1\)), significance levels and p-values. The \(\chi^2\) test for independence (contingency tables, degrees of freedom, critical value) and the \(\chi^2\) goodness-of-fit test - only upper-tail tests at common significance levels (1%, 5%, 10%) are set, and expected frequencies must be greater than 5. |
| AHL4.18 | Critical values and critical regions. Test for a population mean, using the normal distribution when \(\sigma\) is known and the t-distribution when \(\sigma\) is unknown regardless of sample size, for paired or unpaired samples (matched pairs treated as a single-sample technique) - critical regions for t-tests are not required. |
SL4.11 is core AI syllabus content examined at both SL and HL; AHL4.18 extends the t-test to HL-only content.
Key terms
Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.
What are \(H_0\) and \(H_1\)?
\(H_0\), the null hypothesis, is the default assumption of "no effect" or "no difference" - it's what you assume true until the data gives strong evidence otherwise. \(H_1\), the alternative hypothesis, is the claim you're testing for.
e.g. Testing a die: \(H_0\): the die is fair; \(H_1\): the die is not fair.
What is a p-value?
The p-value is the probability of getting a result at least as extreme as the one observed, assuming \(H_0\) is true. A small p-value means the observed data would be unlikely if \(H_0\) were really true - evidence against it.
e.g. If \(p=0.011\) and the significance level is 5%: \(0.011<0.05\), so reject \(H_0\).
What is a significance level?
The significance level (often 1%, 5% or 10%) is the threshold you fix in advance for how small the p-value must be before you reject \(H_0\). It's the risk you're willing to accept of rejecting a true \(H_0\) by chance.
e.g. At the 5% level, \(p=0.090\) is not significant since \(0.090>0.05\).
What is a chi-squared test?
A \(\chi^2\) test compares observed frequencies to expected frequencies - either to check whether data fits a claimed distribution (goodness-of-fit) or whether two categorical variables are related (independence).
e.g. Degrees of freedom for a \(2\times3\) contingency table: \(\nu=(2-1)(3-1)=2\).
What is a t-test?
A t-test compares a sample mean to a claimed population mean (or compares two sample means) when the population standard deviation is unknown and must be estimated from the sample itself. The IB always runs it on the GDC.
e.g. Sample of 10 cartons, mean 496 ml, testing \(H_0:\mu=500\) against \(H_1:\mu<500\).
Key formulas
Six formulas and rules cover almost every question on this topic. The two tables below summarise all of them at a glance - the explanations underneath go into more depth on each one.
Formula reference
Degrees-of-freedom rules and the decision rule for comparing p-values are treated as syllabus content rather than formula-booklet entries - the \(\chi^2\) statistic and t-statistic themselves are always generated by technology, not substituted by hand.
| Rule | Used for | Booklet? |
|---|---|---|
| Reject \(H_0\) if \(p < \) significance level | Decision rule for any hypothesis test | Not in booklet |
| \(\nu = k-1\) | Degrees of freedom, \(\chi^2\) goodness-of-fit (\(k\) categories) | Not in booklet |
| \(\nu = (r-1)(c-1)\) | Degrees of freedom, \(\chi^2\) independence (\(r\) rows, \(c\) columns) | Not in booklet |
| Expected frequency \(= \dfrac{\text{total}}{\text{no. of categories}}\) | Uniform expected frequencies for \(\chi^2\) GOF | Not in booklet |
| \(\nu = n-1\) | Degrees of freedom, one-sample t-test | Not in booklet |
| Reject \(H_0\) if test statistic is beyond the critical value | Critical-value decision rule (alternative to p-value) | Not in booklet |
Chi-squared vs t-test
These are the two tests you'll meet at AI HL - one for categorical data, one for a numerical mean.
| Feature | Chi-squared test | t-test |
|---|---|---|
| Data type | Categorical (counts/frequencies) | Numerical (a mean) |
| Tests | Fit to a distribution, or independence of two variables | A claimed population mean, or two means |
| Tail | Always upper-tail only | One- or two-tailed, depending on \(H_1\) |
| Key condition | Every expected frequency \(>5\) | Underlying variable is (approximately) normal |
| GDC test | \(\chi^2\)GOF-Test or \(\chi^2\)-Test | T-Test or 2-SampTTest |
Setting up a test
Every hypothesis test starts the same way, regardless of which test you'll eventually run.
State the hypotheses
\(H_0\) is always the "no effect / no difference / independent" statement; \(H_1\) is what you're trying to find evidence for. State the significance level at the same time.
Not in the formula booklet - prior knowledgeOne-tailed vs two-tailed
\(H_1\) with \(<\) or \(>\) is one-tailed (testing a specific direction); \(H_1\) with \(\ne\) is two-tailed (testing for any difference at all).
Not in the formula booklet - prior knowledgeThe decision rule
Compare the GDC's p-value to the significance level: \(p<\) significance level means reject \(H_0\); \(p\ge\) significance level means do not reject \(H_0\).
Not in the formula booklet - prior knowledgeChi-squared tests
Both flavours of \(\chi^2\) test compare an observed pattern to an expected one - only the source of the expected values differs.
Goodness-of-fit
Tests whether observed frequencies match a claimed distribution (e.g. a fair die, or a genetic ratio). Degrees of freedom \(=k-1\).
Not in the formula booklet - prior knowledgeIndependence
Tests whether two categorical variables (e.g. gender and subject choice) are associated, using a contingency table. Degrees of freedom \(=(r-1)(c-1)\).
Not in the formula booklet - prior knowledgeThe expected-frequency rule
Every expected frequency must be greater than 5 for the \(\chi^2\) approximation to be valid. If one isn't, combine it with an adjacent category, which reduces \(\nu\) by 1.
Not in the formula booklet - prior knowledgeWorked examples
Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.
A one-tailed t-test on a sample of size 16 gives a test statistic \(t = 1.92.\) The critical value at the 5% level with 15 degrees of freedom is 1.753.
(a) State the degrees of freedom.
(b) State and justify the conclusion at the 5% level.
Worked solution
(a) \(\nu = n - 1.\) M1
\(\nu = 15.\) A1
(b) \(t = 1.92 > 1.753.\) R1
In the rejection region: reject \(H_0\). A1
Reject \(H_0\) at 5%. A1
A die is rolled 120 times. The observed frequencies for faces 1-6 are 14, 22, 18, 25, 21, 20. Test at the 5% level whether the die is fair.
(a) State the hypotheses.
(b)(i) Write down the expected frequency for each face.
(b)(ii) State the degrees of freedom.
(c) Given the calculated \(\chi^2 = 3.7\) and critical value 11.07, state the conclusion.
Worked solution
(a) \(H_0\): the die is fair; \(H_1\): the die is not fair. A1 A1
(b) Expected \(= \dfrac{120}{6} = 20\) per face; \(\nu = 5.\) A1 A1
(c) \(3.7 < 11.07\), so do not reject \(H_0\): R1
insufficient evidence the die is unfair. A1
Common mistakes
The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.
- Getting the decision rule backwards. A small p-value is evidence against \(H_0\), so \(p<\) significance level means reject \(H_0\) - not the other way round.
- Using the wrong degrees-of-freedom formula. Goodness-of-fit uses \(\nu=k-1\); independence on a contingency table uses \(\nu=(r-1)(c-1)\) - mixing these up on a contingency table gives the wrong critical value.
- Concluding \(H_0\) is "true" after not rejecting it. Failing to reject \(H_0\) only means there's insufficient evidence against it - it's never proof that \(H_0\) is correct.
- Forgetting to state the conclusion in context. "Reject \(H_0\)" alone rarely earns full marks - the final answer needs to say what that means for the actual situation, e.g. "there is evidence the machine underfills".
Using your GDC
Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.
Needed before most t-tests - to find or confirm the sample mean and standard deviation you'll feed into the test.
- Enter the data into a list.
- STAT → Edit → type values into L1. Then STAT → CALC → 1:1-Var Stats, choose L1, Calculate.TI-84
- Add a Lists & Spreadsheet page, name a column and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
- Statistics menu → enter data in List 1 → CALC (F2) → 1-VAR.Casio
- Read \(\bar x\) (mean), \(S_x\) (sample sd) or \(\sigma_x\) (population sd), and the five-number summary.
- For frequency data, put values in one list and frequencies in another and set the frequency list.
Tip: \(S_x\) vs \(\sigma_x\): use \(\sigma_x\) (population) for a complete data set, \(S_x\) (sample) for a sample. A t-test always uses the sample standard deviation.
Underpins normal-based hypothesis tests and critical-value questions - converts a significance level into a critical z-value.
- Work out the area to the LEFT of the value you want.
- 2nd → VARS (DISTR) → invNorm(area, μ, σ). Newer OS lets you pick the tail.TI-84
- menu → Probability → Distributions → Inverse Normal; enter the area, μ and σ.Nspire
- Main menu → Statistics → DIST → NORM → InvN; set the tail and enter area, σ, μ.Casio
Tip: invNorm needs the area to the LEFT. For a 5% upper-tail critical value, use area = 0.95.
See the full GDC guide for more calculator models and topics.
Ready to practise properly?
Hypothesis testing questions, marked instantly like the real exam.
Quick answers
The questions students on this topic ask most often.
What's the difference between a chi-squared test and a t-test?
A chi-squared test compares observed and expected frequencies - it's used for goodness-of-fit (does the data match a distribution?) or independence (are two categorical variables related?). A t-test compares a sample mean to a claimed population mean, or compares two sample means, when the population standard deviation is unknown.
How do I decide whether to reject H0?
Compare the p-value your GDC gives to the stated significance level. If \(p\) is less than the significance level, reject \(H_0\) - there's evidence for \(H_1\). If \(p\) is greater than or equal to the significance level, do not reject \(H_0\) - there's insufficient evidence against it.
How do I find the degrees of freedom for a chi-squared test?
For a goodness-of-fit test with \(k\) categories, degrees of freedom \(=k-1\). For an independence test on a contingency table with \(r\) rows and \(c\) columns, degrees of freedom \(=(r-1)(c-1)\).
Can I use my GDC for this topic?
Yes, and you're expected to. Every hypothesis test in the exam is run on your GDC, which returns the test statistic and the p-value directly - your job is to set up \(H_0\) and \(H_1\) correctly and interpret the output in context. See the GDC guide for model-specific instructions.
Sub-topics
Hypothesis Testing broken down into its individual skills, each with its own focused page.
Related topics
More Statistics & Probability topics from the same AI HL syllabus unit, in case you want to keep going.