Hypothesis Testing (AI SL)

Hypothesis testing gives a formal way to decide whether data provides enough evidence against an assumption. This topic covers writing null and alternative hypotheses, using significance levels and p-values to make a decision, the chi-squared test for independence and goodness of fit with categorical data, and the t-test for comparing two population means.

What the syllabus says

This topic maps onto one point in the official IB Applications & Interpretation syllabus.

CodeSyllabus content
SL4.11Formulation of null and alternative hypotheses, \(H_0\) and \(H_1\). Significance levels, p-values, expected and observed frequencies. The \(\chi^2\) test for independence: contingency tables, degrees of freedom, critical value. The \(\chi^2\) goodness of fit test. The t-test, using the p-value to compare the means of two populations, one- and two-tailed tests.

At SL, contingency tables have at most 4 rows or columns, expected frequencies must be greater than 5, and only upper-tail tests at 1%, 5% or 10% significance are set. Students use technology to find the \(\chi^2\) statistic, p-value and t-test results.

Key terms

Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.

What is a null hypothesis?

The null hypothesis, \(H_0\), is the "no effect" or "no difference" starting assumption - that variables are independent, or that a population parameter equals a stated value. It's always written as a statement about the population, never about the sample data.

e.g. For a fairness test: \(H_0\): the die is fair (all faces equally likely).

What is a p-value?

The p-value is the probability, assuming \(H_0\) is true, of getting a result at least as extreme as the one observed. A small p-value means the observed data would be unusual if \(H_0\) were true, giving evidence against it.

e.g. If \(p=0.032\) at the 5% level: since \(0.032<0.05\), reject \(H_0\).

What is a chi-squared test?

The chi-squared test compares observed and expected frequencies in categories, either to test whether two categorical variables are independent (a contingency table) or whether data fits a claimed distribution (goodness of fit). A GDC gives the \(\chi^2\) statistic and p-value directly.

e.g. With \(\chi^2_{calc}=5.42\) and \(p=0.0199<0.05\), reject \(H_0\): there's evidence of an association.

What is a significance level?

The significance level is the cut-off probability (commonly 1%, 5% or 10%) chosen before the test, below which the p-value counts as strong enough evidence to reject \(H_0\). It represents the risk of rejecting a true \(H_0\) (a Type I error) that you're willing to accept.

e.g. At the 1% level, reject \(H_0\) only if \(p<0.01\) - a stricter bar than the 5% level.

What is degrees of freedom (chi-squared)?

For a chi-squared test of independence on an \(m\times n\) contingency table, the degrees of freedom is \(\nu=(m-1)(n-1)\). It determines the shape of the chi-squared distribution used to find the critical value or p-value.

e.g. For a \(2\times3\) table, \(\nu=(2-1)(3-1)=2\).

Key formulas

This topic is mostly about setting up and interpreting a test correctly - the underlying statistic and p-value are always found using technology. The tables below cover what's genuinely formula-based.

Formula reference

Almost nothing here is calculated by hand in examinations - the decision rules and degrees-of-freedom formula are conventions you apply, not booklet formulas you look up.

Formula / ruleUsed forBooklet?
\(\nu=(m-1)(n-1)\)Degrees of freedom for an \(m\times n\) contingency tableNot in booklet - counting rule
Reject \(H_0\) if \(p<\) significance levelDecision rule by p-valueNot in booklet - standard convention
Reject \(H_0\) if \(\chi^2_{calc}>\chi^2_{crit}\)Decision rule by critical valueNot in booklet - standard convention
Expected frequency \(>5\) in every cellValidity condition for the \(\chi^2\) testNot in booklet - syllabus requirement

Chi-squared test vs t-test

These are the two named tests on the AI SL syllabus - the type of data you have decides which one applies.

FeatureChi-squared testt-test
Data typeCategorical (counts in categories)Continuous (measurements)
Question answeredAre two variables independent? Does data fit a distribution?Do two populations have the same mean?
Key assumptionExpected frequency \(>5\) in every cellUnderlying variable is normally distributed; equal variances assumed at SL
ExampleIs eye colour independent of handedness?Do two teaching methods give different mean scores?

Setting up a test

Every hypothesis test starts the same way, regardless of which test you eventually run.

State the hypotheses

\(H_0\) is always the "no difference / independent" statement; \(H_1\) is the claim being tested for, expressed in words or as an inequality.

Not in the formula booklet - stated in words each time

Choose the significance level

Fix the significance level (1%, 5% or 10% at SL) before looking at the data - it sets how much evidence is needed to reject \(H_0\).

Not in the formula booklet - stated in the question

Check the validity condition

For the \(\chi^2\) test, every expected frequency must exceed 5; if not, combine categories until it does before running the test.

Not in the formula booklet - syllabus requirement

Making a decision

Once the GDC gives you a statistic and p-value, the decision itself follows one of two equivalent rules.

Compare with the p-value

Reject \(H_0\) if \(p<\) significance level; otherwise do not reject \(H_0\).

Not in the formula booklet - standard convention

Compare with the critical value

Reject \(H_0\) if \(\chi^2_{calc}>\chi^2_{crit}\); otherwise do not reject \(H_0\).

Not in the formula booklet - standard convention

Interpret in context

Always finish by restating the conclusion in the words of the original problem - "reject \(H_0\)" alone is not a complete answer.

Not in the formula booklet - examiner expectation

Worked examples

Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.

1
Easy
[3 marks]

A researcher tests whether eye colour and handedness are independent.

(a) State the null hypothesis.

(b) State the alternative hypothesis.

Worked solution

(a) A \(\chi^2\) test of independence always pits ‘independent’ against ‘associated’. M1
\(H_0\): eye colour and handedness are independent. A1

(b) \(H_1\): eye colour and handedness are not independent (they are associated). A1
Note. Hypotheses are always statements about the populations, never the sample.

M1 Recognise the independence test structure A1 H₀ independent A1 H₁ not independent
2
Hard
[7 marks]

A survey of 200 people tests whether exercise level (low/high) is independent of having a cold last winter (yes/no). A GDC gives \(\chi^2_{calc}=5.42\), p-value 0.0199, with 1 degree of freedom.

(a) State the hypotheses.

(b) State the degrees of freedom check (for a \(2\times2\) table).

(c) At the 5% level, state and interpret the conclusion.

Worked solution

(a) \(H_0\): exercise level and having a cold are independent. \(H_1\): they are not independent. M1 A1

(b) Degrees of freedom. For a \(2\times 2\) table, \(\nu=(2-1)(2-1)=1\) - matching the GDC. ✓ A1

(c) Conclusion. \(p=0.0199<0.05,\) so reject \(H_0\) A1 M1 A1 : there is significant evidence of an association between exercise level and catching a cold. R1

M1 Hypotheses A1 H₀, H₁ A1 ν = 1 A1 Quote p = 0.0199 M1 Compare with 0.05 A1 Reject H₀ R1 Conclusion in context

Common mistakes

The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.

  • Writing \(H_0\) as a statement about the sample. Hypotheses describe the population, not the specific data collected - "the die is fair", not "these 120 rolls came out even".
  • Treating "reject \(H_0\)" as proof that \(H_1\) is true. A hypothesis test only measures the strength of evidence at a chosen significance level - it never proves anything with certainty, in either direction.
  • Forgetting the expected-frequency-at-least-5 rule. If any cell's expected frequency is 5 or below, the \(\chi^2\) approximation becomes unreliable - categories should be combined first.
  • Comparing \(p\) the wrong way round. Reject \(H_0\) when \(p\) is less than the significance level, not greater - it's easy to flip this under exam pressure.

Using your GDC

The IB syllabus expects the chi-squared statistic, p-value and t-test results to be found using technology - your job is to enter the data correctly and read off the right numbers. Pick your model to filter down to just the steps that apply to you.

Show steps for:
Chi-squared test (independence or goodness of fit)

Enter the observed counts and let the GDC compute the \(\chi^2\) statistic, degrees of freedom, p-value and expected frequencies in one step.

  1. State \(H_0\) and \(H_1\) and the significance level before running the test.
  2. STAT → TESTS → χ²-Test (independence, enter counts as a matrix) or χ²GOF-Test (goodness of fit).TI-84
  3. Statistics menu → TEST → CHI → 2WAY (independence) or GOF (goodness of fit).Casio
  4. menu → Statistics → Stat Tests → χ² 2-way Test or χ² GOF Test.Nspire
  5. Read off \(\chi^2_{calc}\), the p-value and the degrees of freedom, then compare with the significance level or critical value.

Tip: Check the GDC's degrees of freedom against \(\nu=(m-1)(n-1)\) as a sanity check that the table was entered correctly.

The t-test (comparing two means)

Runs the two-sample t-test for the test statistic and p-value used to compare the means of two populations.

  1. Decide whether the test is one-tailed (a claimed increase or decrease) or two-tailed (just "different").
  2. STAT → TESTS → 2-SampTTest (two samples) or T-Test (one sample against a claimed mean).TI-84
  3. Statistics menu → TEST → t → 2-Sample.Casio
  4. menu → Statistics → Stat Tests → 2-Sample t Test.Nspire
  5. At SL, assume the two population variances are equal, so use the pooled two-sample t-test setting.

Tip: "Do not reject \(H_0\)" for a t-test does not prove the means are equal - only that the data give too little evidence to claim a difference.

See the full GDC guide for more calculator models and topics.

Ready to practise properly?

Hypothesis testing questions, marked instantly like the real exam.

Quick answers

The questions students on this topic ask most often.

What's the difference between the chi-squared test and the t-test?

The chi-squared test works with categorical data - it checks whether two categorical variables are independent, or whether observed counts fit an expected pattern. The t-test works with continuous data - it compares the means of two populations. Pick chi-squared for counts in categories, t-test for comparing averages.

Does "do not reject H0" mean H0 is true?

No. "Do not reject \(H_0\)" only means the data didn't give strong enough evidence against it - it never proves \(H_0\) is true. Likewise, "reject \(H_0\)" doesn't prove \(H_1\) is true, just that the evidence against \(H_0\) is strong enough at the chosen significance level.

How do I decide whether to reject the null hypothesis?

Compare the p-value to the significance level: reject \(H_0\) if \(p\) is less than the significance level, otherwise do not reject. Equivalently, for the chi-squared test, reject \(H_0\) if the calculated chi-squared statistic is greater than the critical value.

Why does the expected frequency need to be at least 5?

The chi-squared test statistic only approximately follows the chi-squared distribution, and that approximation becomes unreliable when expected frequencies are small. If a cell's expected frequency is below 5, categories are usually combined so every expected frequency reaches at least 5 before the test is run.

Sub-topics

Hypothesis Testing broken down into its individual skills, each with its own focused page.