Hypothesis Testing (AI HL)

Hypothesis testing gives you a formal way to decide whether a pattern in a sample of data is real, or just what you'd expect from random chance. This topic covers writing null and alternative hypotheses, the chi-squared test for goodness-of-fit and independence, the t-test for a population mean, and how to compare a p-value or test statistic against a significance level to reach a defensible conclusion.

What the syllabus says

This topic maps onto two points in the official IB Applications & Interpretation syllabus.

CodeSyllabus content
SL4.11Formulation of null and alternative hypotheses (\(H_0\) and \(H_1\)), significance levels and p-values. The \(\chi^2\) test for independence (contingency tables, degrees of freedom, critical value) and the \(\chi^2\) goodness-of-fit test - only upper-tail tests at common significance levels (1%, 5%, 10%) are set, and expected frequencies must be greater than 5.
AHL4.18Critical values and critical regions. Test for a population mean, using the normal distribution when \(\sigma\) is known and the t-distribution when \(\sigma\) is unknown regardless of sample size, for paired or unpaired samples (matched pairs treated as a single-sample technique) - critical regions for t-tests are not required.

SL4.11 is core AI syllabus content examined at both SL and HL; AHL4.18 extends the t-test to HL-only content.

Key terms

Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.

What are \(H_0\) and \(H_1\)?

\(H_0\), the null hypothesis, is the default assumption of "no effect" or "no difference" - it's what you assume true until the data gives strong evidence otherwise. \(H_1\), the alternative hypothesis, is the claim you're testing for.

e.g. Testing a die: \(H_0\): the die is fair; \(H_1\): the die is not fair.

What is a p-value?

The p-value is the probability of getting a result at least as extreme as the one observed, assuming \(H_0\) is true. A small p-value means the observed data would be unlikely if \(H_0\) were really true - evidence against it.

e.g. If \(p=0.011\) and the significance level is 5%: \(0.011<0.05\), so reject \(H_0\).

What is a significance level?

The significance level (often 1%, 5% or 10%) is the threshold you fix in advance for how small the p-value must be before you reject \(H_0\). It's the risk you're willing to accept of rejecting a true \(H_0\) by chance.

e.g. At the 5% level, \(p=0.090\) is not significant since \(0.090>0.05\).

What is a chi-squared test?

A \(\chi^2\) test compares observed frequencies to expected frequencies - either to check whether data fits a claimed distribution (goodness-of-fit) or whether two categorical variables are related (independence).

e.g. Degrees of freedom for a \(2\times3\) contingency table: \(\nu=(2-1)(3-1)=2\).

What is a t-test?

A t-test compares a sample mean to a claimed population mean (or compares two sample means) when the population standard deviation is unknown and must be estimated from the sample itself. The IB always runs it on the GDC.

e.g. Sample of 10 cartons, mean 496 ml, testing \(H_0:\mu=500\) against \(H_1:\mu<500\).

Key formulas

Six formulas and rules cover almost every question on this topic. The two tables below summarise all of them at a glance - the explanations underneath go into more depth on each one.

Formula reference

Degrees-of-freedom rules and the decision rule for comparing p-values are treated as syllabus content rather than formula-booklet entries - the \(\chi^2\) statistic and t-statistic themselves are always generated by technology, not substituted by hand.

RuleUsed forBooklet?
Reject \(H_0\) if \(p < \) significance levelDecision rule for any hypothesis testNot in booklet
\(\nu = k-1\)Degrees of freedom, \(\chi^2\) goodness-of-fit (\(k\) categories)Not in booklet
\(\nu = (r-1)(c-1)\)Degrees of freedom, \(\chi^2\) independence (\(r\) rows, \(c\) columns)Not in booklet
Expected frequency \(= \dfrac{\text{total}}{\text{no. of categories}}\)Uniform expected frequencies for \(\chi^2\) GOFNot in booklet
\(\nu = n-1\)Degrees of freedom, one-sample t-testNot in booklet
Reject \(H_0\) if test statistic is beyond the critical valueCritical-value decision rule (alternative to p-value)Not in booklet

Chi-squared vs t-test

These are the two tests you'll meet at AI HL - one for categorical data, one for a numerical mean.

FeatureChi-squared testt-test
Data typeCategorical (counts/frequencies)Numerical (a mean)
TestsFit to a distribution, or independence of two variablesA claimed population mean, or two means
TailAlways upper-tail onlyOne- or two-tailed, depending on \(H_1\)
Key conditionEvery expected frequency \(>5\)Underlying variable is (approximately) normal
GDC test\(\chi^2\)GOF-Test or \(\chi^2\)-TestT-Test or 2-SampTTest

Setting up a test

Every hypothesis test starts the same way, regardless of which test you'll eventually run.

State the hypotheses

\(H_0\) is always the "no effect / no difference / independent" statement; \(H_1\) is what you're trying to find evidence for. State the significance level at the same time.

Not in the formula booklet - prior knowledge

One-tailed vs two-tailed

\(H_1\) with \(<\) or \(>\) is one-tailed (testing a specific direction); \(H_1\) with \(\ne\) is two-tailed (testing for any difference at all).

Not in the formula booklet - prior knowledge

The decision rule

Compare the GDC's p-value to the significance level: \(p<\) significance level means reject \(H_0\); \(p\ge\) significance level means do not reject \(H_0\).

Not in the formula booklet - prior knowledge

Chi-squared tests

Both flavours of \(\chi^2\) test compare an observed pattern to an expected one - only the source of the expected values differs.

Goodness-of-fit

Tests whether observed frequencies match a claimed distribution (e.g. a fair die, or a genetic ratio). Degrees of freedom \(=k-1\).

Not in the formula booklet - prior knowledge

Independence

Tests whether two categorical variables (e.g. gender and subject choice) are associated, using a contingency table. Degrees of freedom \(=(r-1)(c-1)\).

Not in the formula booklet - prior knowledge

The expected-frequency rule

Every expected frequency must be greater than 5 for the \(\chi^2\) approximation to be valid. If one isn't, combine it with an adjacent category, which reduces \(\nu\) by 1.

Not in the formula booklet - prior knowledge

Worked examples

Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.

1
Medium
GDC
[5 marks]

A one-tailed t-test on a sample of size 16 gives a test statistic \(t = 1.92.\) The critical value at the 5% level with 15 degrees of freedom is 1.753.

(a) State the degrees of freedom.

(b) State and justify the conclusion at the 5% level.

Worked solution

(a) \(\nu = n - 1.\) M1
\(\nu = 15.\) A1

(b) \(t = 1.92 > 1.753.\) R1
In the rejection region: reject \(H_0\). A1
Reject \(H_0\) at 5%. A1

M1 \(n-1\) A1 Correct answer of \(15\) R1 Compare A1 Reject A1 Justify
2
Hard
GDC
[6 marks]

A die is rolled 120 times. The observed frequencies for faces 1-6 are 14, 22, 18, 25, 21, 20. Test at the 5% level whether the die is fair.

(a) State the hypotheses.

(b)(i) Write down the expected frequency for each face.

(b)(ii) State the degrees of freedom.

(c) Given the calculated \(\chi^2 = 3.7\) and critical value 11.07, state the conclusion.

Worked solution

(a) \(H_0\): the die is fair; \(H_1\): the die is not fair. A1 A1

(b) Expected \(= \dfrac{120}{6} = 20\) per face; \(\nu = 5.\) A1 A1

(c) \(3.7 < 11.07\), so do not reject \(H_0\): R1
insufficient evidence the die is unfair. A1

A1 \(H_0\) A1 \(H_1\) A1 Expected = 20 A1 \(\nu=5\) R1 Compare A1 Conclusion

Common mistakes

The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.

  • Getting the decision rule backwards. A small p-value is evidence against \(H_0\), so \(p<\) significance level means reject \(H_0\) - not the other way round.
  • Using the wrong degrees-of-freedom formula. Goodness-of-fit uses \(\nu=k-1\); independence on a contingency table uses \(\nu=(r-1)(c-1)\) - mixing these up on a contingency table gives the wrong critical value.
  • Concluding \(H_0\) is "true" after not rejecting it. Failing to reject \(H_0\) only means there's insufficient evidence against it - it's never proof that \(H_0\) is correct.
  • Forgetting to state the conclusion in context. "Reject \(H_0\)" alone rarely earns full marks - the final answer needs to say what that means for the actual situation, e.g. "there is evidence the machine underfills".

Using your GDC

Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.

Show steps for:
One-variable statistics (mean, median, standard deviation)

Needed before most t-tests - to find or confirm the sample mean and standard deviation you'll feed into the test.

  1. Enter the data into a list.
  2. STAT → Edit → type values into L1. Then STAT → CALC → 1:1-Var Stats, choose L1, Calculate.TI-84
  3. Add a Lists & Spreadsheet page, name a column and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
  4. Statistics menu → enter data in List 1 → CALC (F2) → 1-VAR.Casio
  5. Read \(\bar x\) (mean), \(S_x\) (sample sd) or \(\sigma_x\) (population sd), and the five-number summary.
  6. For frequency data, put values in one list and frequencies in another and set the frequency list.

Tip: \(S_x\) vs \(\sigma_x\): use \(\sigma_x\) (population) for a complete data set, \(S_x\) (sample) for a sample. A t-test always uses the sample standard deviation.

Inverse normal (find the value for a given probability)

Underpins normal-based hypothesis tests and critical-value questions - converts a significance level into a critical z-value.

  1. Work out the area to the LEFT of the value you want.
  2. 2nd → VARS (DISTR) → invNorm(area, μ, σ). Newer OS lets you pick the tail.TI-84
  3. menu → Probability → Distributions → Inverse Normal; enter the area, μ and σ.Nspire
  4. Main menu → Statistics → DIST → NORM → InvN; set the tail and enter area, σ, μ.Casio

Tip: invNorm needs the area to the LEFT. For a 5% upper-tail critical value, use area = 0.95.

See the full GDC guide for more calculator models and topics.

Ready to practise properly?

Hypothesis testing questions, marked instantly like the real exam.

Quick answers

The questions students on this topic ask most often.

What's the difference between a chi-squared test and a t-test?

A chi-squared test compares observed and expected frequencies - it's used for goodness-of-fit (does the data match a distribution?) or independence (are two categorical variables related?). A t-test compares a sample mean to a claimed population mean, or compares two sample means, when the population standard deviation is unknown.

How do I decide whether to reject H0?

Compare the p-value your GDC gives to the stated significance level. If \(p\) is less than the significance level, reject \(H_0\) - there's evidence for \(H_1\). If \(p\) is greater than or equal to the significance level, do not reject \(H_0\) - there's insufficient evidence against it.

How do I find the degrees of freedom for a chi-squared test?

For a goodness-of-fit test with \(k\) categories, degrees of freedom \(=k-1\). For an independence test on a contingency table with \(r\) rows and \(c\) columns, degrees of freedom \(=(r-1)(c-1)\).

Can I use my GDC for this topic?

Yes, and you're expected to. Every hypothesis test in the exam is run on your GDC, which returns the test statistic and the p-value directly - your job is to set up \(H_0\) and \(H_1\) correctly and interpret the output in context. See the GDC guide for model-specific instructions.

Sub-topics

Hypothesis Testing broken down into its individual skills, each with its own focused page.