Hypothesis Testing (AI SL)
Hypothesis testing gives a formal way to decide whether data provides enough evidence against an assumption. This topic covers writing null and alternative hypotheses, using significance levels and p-values to make a decision, the chi-squared test for independence and goodness of fit with categorical data, and the t-test for comparing two population means.
What the syllabus says
This topic maps onto one point in the official IB Applications & Interpretation syllabus.
| Code | Syllabus content |
|---|---|
| SL4.11 | Formulation of null and alternative hypotheses, \(H_0\) and \(H_1\). Significance levels, p-values, expected and observed frequencies. The \(\chi^2\) test for independence: contingency tables, degrees of freedom, critical value. The \(\chi^2\) goodness of fit test. The t-test, using the p-value to compare the means of two populations, one- and two-tailed tests. |
At SL, contingency tables have at most 4 rows or columns, expected frequencies must be greater than 5, and only upper-tail tests at 1%, 5% or 10% significance are set. Students use technology to find the \(\chi^2\) statistic, p-value and t-test results.
Key terms
Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.
What is a null hypothesis?
The null hypothesis, \(H_0\), is the "no effect" or "no difference" starting assumption - that variables are independent, or that a population parameter equals a stated value. It's always written as a statement about the population, never about the sample data.
e.g. For a fairness test: \(H_0\): the die is fair (all faces equally likely).
What is a p-value?
The p-value is the probability, assuming \(H_0\) is true, of getting a result at least as extreme as the one observed. A small p-value means the observed data would be unusual if \(H_0\) were true, giving evidence against it.
e.g. If \(p=0.032\) at the 5% level: since \(0.032<0.05\), reject \(H_0\).
What is a chi-squared test?
The chi-squared test compares observed and expected frequencies in categories, either to test whether two categorical variables are independent (a contingency table) or whether data fits a claimed distribution (goodness of fit). A GDC gives the \(\chi^2\) statistic and p-value directly.
e.g. With \(\chi^2_{calc}=5.42\) and \(p=0.0199<0.05\), reject \(H_0\): there's evidence of an association.
What is a significance level?
The significance level is the cut-off probability (commonly 1%, 5% or 10%) chosen before the test, below which the p-value counts as strong enough evidence to reject \(H_0\). It represents the risk of rejecting a true \(H_0\) (a Type I error) that you're willing to accept.
e.g. At the 1% level, reject \(H_0\) only if \(p<0.01\) - a stricter bar than the 5% level.
What is degrees of freedom (chi-squared)?
For a chi-squared test of independence on an \(m\times n\) contingency table, the degrees of freedom is \(\nu=(m-1)(n-1)\). It determines the shape of the chi-squared distribution used to find the critical value or p-value.
e.g. For a \(2\times3\) table, \(\nu=(2-1)(3-1)=2\).
Key formulas
This topic is mostly about setting up and interpreting a test correctly - the underlying statistic and p-value are always found using technology. The tables below cover what's genuinely formula-based.
Formula reference
Almost nothing here is calculated by hand in examinations - the decision rules and degrees-of-freedom formula are conventions you apply, not booklet formulas you look up.
| Formula / rule | Used for | Booklet? |
|---|---|---|
| \(\nu=(m-1)(n-1)\) | Degrees of freedom for an \(m\times n\) contingency table | Not in booklet - counting rule |
| Reject \(H_0\) if \(p<\) significance level | Decision rule by p-value | Not in booklet - standard convention |
| Reject \(H_0\) if \(\chi^2_{calc}>\chi^2_{crit}\) | Decision rule by critical value | Not in booklet - standard convention |
| Expected frequency \(>5\) in every cell | Validity condition for the \(\chi^2\) test | Not in booklet - syllabus requirement |
Chi-squared test vs t-test
These are the two named tests on the AI SL syllabus - the type of data you have decides which one applies.
| Feature | Chi-squared test | t-test |
|---|---|---|
| Data type | Categorical (counts in categories) | Continuous (measurements) |
| Question answered | Are two variables independent? Does data fit a distribution? | Do two populations have the same mean? |
| Key assumption | Expected frequency \(>5\) in every cell | Underlying variable is normally distributed; equal variances assumed at SL |
| Example | Is eye colour independent of handedness? | Do two teaching methods give different mean scores? |
Setting up a test
Every hypothesis test starts the same way, regardless of which test you eventually run.
State the hypotheses
\(H_0\) is always the "no difference / independent" statement; \(H_1\) is the claim being tested for, expressed in words or as an inequality.
Not in the formula booklet - stated in words each timeChoose the significance level
Fix the significance level (1%, 5% or 10% at SL) before looking at the data - it sets how much evidence is needed to reject \(H_0\).
Not in the formula booklet - stated in the questionCheck the validity condition
For the \(\chi^2\) test, every expected frequency must exceed 5; if not, combine categories until it does before running the test.
Not in the formula booklet - syllabus requirementMaking a decision
Once the GDC gives you a statistic and p-value, the decision itself follows one of two equivalent rules.
Compare with the p-value
Reject \(H_0\) if \(p<\) significance level; otherwise do not reject \(H_0\).
Not in the formula booklet - standard conventionCompare with the critical value
Reject \(H_0\) if \(\chi^2_{calc}>\chi^2_{crit}\); otherwise do not reject \(H_0\).
Not in the formula booklet - standard conventionInterpret in context
Always finish by restating the conclusion in the words of the original problem - "reject \(H_0\)" alone is not a complete answer.
Not in the formula booklet - examiner expectationWorked examples
Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.
A researcher tests whether eye colour and handedness are independent.
(a) State the null hypothesis.
(b) State the alternative hypothesis.
Worked solution
(a) A \(\chi^2\) test of independence always pits ‘independent’ against ‘associated’. M1
\(H_0\): eye colour and handedness are independent. A1
(b) \(H_1\): eye colour and handedness are not independent (they are associated). A1
Note. Hypotheses are always statements about the populations, never the sample.
A survey of 200 people tests whether exercise level (low/high) is independent of having a cold last winter (yes/no). A GDC gives \(\chi^2_{calc}=5.42\), p-value 0.0199, with 1 degree of freedom.
(a) State the hypotheses.
(b) State the degrees of freedom check (for a \(2\times2\) table).
(c) At the 5% level, state and interpret the conclusion.
Worked solution
(a) \(H_0\): exercise level and having a cold are independent. \(H_1\): they are not independent. M1 A1
(b) Degrees of freedom. For a \(2\times 2\) table, \(\nu=(2-1)(2-1)=1\) - matching the GDC. ✓ A1
(c) Conclusion. \(p=0.0199<0.05,\) so reject \(H_0\) A1 M1 A1 : there is significant evidence of an association between exercise level and catching a cold. R1
Common mistakes
The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.
- Writing \(H_0\) as a statement about the sample. Hypotheses describe the population, not the specific data collected - "the die is fair", not "these 120 rolls came out even".
- Treating "reject \(H_0\)" as proof that \(H_1\) is true. A hypothesis test only measures the strength of evidence at a chosen significance level - it never proves anything with certainty, in either direction.
- Forgetting the expected-frequency-at-least-5 rule. If any cell's expected frequency is 5 or below, the \(\chi^2\) approximation becomes unreliable - categories should be combined first.
- Comparing \(p\) the wrong way round. Reject \(H_0\) when \(p\) is less than the significance level, not greater - it's easy to flip this under exam pressure.
Using your GDC
The IB syllabus expects the chi-squared statistic, p-value and t-test results to be found using technology - your job is to enter the data correctly and read off the right numbers. Pick your model to filter down to just the steps that apply to you.
Enter the observed counts and let the GDC compute the \(\chi^2\) statistic, degrees of freedom, p-value and expected frequencies in one step.
- State \(H_0\) and \(H_1\) and the significance level before running the test.
- STAT → TESTS → χ²-Test (independence, enter counts as a matrix) or χ²GOF-Test (goodness of fit).TI-84
- Statistics menu → TEST → CHI → 2WAY (independence) or GOF (goodness of fit).Casio
- menu → Statistics → Stat Tests → χ² 2-way Test or χ² GOF Test.Nspire
- Read off \(\chi^2_{calc}\), the p-value and the degrees of freedom, then compare with the significance level or critical value.
Tip: Check the GDC's degrees of freedom against \(\nu=(m-1)(n-1)\) as a sanity check that the table was entered correctly.
Runs the two-sample t-test for the test statistic and p-value used to compare the means of two populations.
- Decide whether the test is one-tailed (a claimed increase or decrease) or two-tailed (just "different").
- STAT → TESTS → 2-SampTTest (two samples) or T-Test (one sample against a claimed mean).TI-84
- Statistics menu → TEST → t → 2-Sample.Casio
- menu → Statistics → Stat Tests → 2-Sample t Test.Nspire
- At SL, assume the two population variances are equal, so use the pooled two-sample t-test setting.
Tip: "Do not reject \(H_0\)" for a t-test does not prove the means are equal - only that the data give too little evidence to claim a difference.
See the full GDC guide for more calculator models and topics.
Ready to practise properly?
Hypothesis testing questions, marked instantly like the real exam.
Quick answers
The questions students on this topic ask most often.
What's the difference between the chi-squared test and the t-test?
The chi-squared test works with categorical data - it checks whether two categorical variables are independent, or whether observed counts fit an expected pattern. The t-test works with continuous data - it compares the means of two populations. Pick chi-squared for counts in categories, t-test for comparing averages.
Does "do not reject H0" mean H0 is true?
No. "Do not reject \(H_0\)" only means the data didn't give strong enough evidence against it - it never proves \(H_0\) is true. Likewise, "reject \(H_0\)" doesn't prove \(H_1\) is true, just that the evidence against \(H_0\) is strong enough at the chosen significance level.
How do I decide whether to reject the null hypothesis?
Compare the p-value to the significance level: reject \(H_0\) if \(p\) is less than the significance level, otherwise do not reject. Equivalently, for the chi-squared test, reject \(H_0\) if the calculated chi-squared statistic is greater than the critical value.
Why does the expected frequency need to be at least 5?
The chi-squared test statistic only approximately follows the chi-squared distribution, and that approximation becomes unreliable when expected frequencies are small. If a cell's expected frequency is below 5, categories are usually combined so every expected frequency reaches at least 5 before the test is run.
Sub-topics
Hypothesis Testing broken down into its individual skills, each with its own focused page.
Related topics
More Statistics & Probability topics from the same AI SL syllabus unit, in case you want to keep going.