Probability & Distributions (AI HL)
A probability distribution is a model of how likely each outcome of a random variable is - and once you know which model fits a situation, your GDC does the rest of the work. This topic covers discrete random variables and expected value, the binomial distribution for counting successes in fixed trials, and the normal distribution for continuous data, including working backwards from a probability to find an unknown value, mean or standard deviation.
What the syllabus says
This topic maps onto two points in the official IB Applications & Interpretation syllabus.
| Code | Syllabus content |
|---|---|
| SL4.7 | Concept of discrete random variables and their probability distributions, given as a table or a formula. Expected value \(E(X)\) for discrete data, and its applications - including that \(E(X)=0\) indicates a fair game where \(X\) is a player's gain. |
| SL4.8 / SL4.9 | The binomial distribution, its mean and variance, and the situations where it's an appropriate model (found using technology). The normal distribution and its properties, including that approximately 68% of data lies within one standard deviation of the mean, 95% within two, and 99.7% within three - with normal and inverse normal probabilities found using technology. |
These are core AI SL syllabus points that are also examinable at AI HL.
Key terms
Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.
What is a discrete random variable?
A discrete random variable is a quantity whose outcomes can be listed - like the number on a die, or the number of heads in five coin tosses - each with its own probability. The probabilities in its distribution must sum to exactly 1.
e.g. If \(P(X=x)=kx\) for \(x=1,2,3,4\), then \(k(1+2+3+4)=1\), so \(k=0.1\).
What is expected value?
The expected value \(E(X) = \Sigma xP(X=x)\) is the long-run average outcome if the random experiment were repeated many times. It doesn't have to be a value \(X\) can actually take - it's a weighted average.
e.g. For \(X\) with \(P(0)=0.5,P(1)=0.3,P(2)=0.2\): \(E(X)=0(0.5)+1(0.3)+2(0.2)=0.7\).
What is the binomial distribution?
The binomial distribution \(X \sim B(n,p)\) models the number of successes in \(n\) independent trials, each with the same success probability \(p\). It applies whenever there's a fixed number of identical yes/no trials.
e.g. \(X\sim B(50,0.2)\): mean \(E(X)=np=50(0.2)=10\).
What is the normal distribution?
The normal distribution is a continuous, bell-shaped model for data that clusters symmetrically around a mean, described entirely by its mean \(\mu\) and standard deviation \(\sigma\). Probabilities are areas under the curve, found with your GDC.
e.g. About 68% of a normal distribution's values lie within one standard deviation of \(\mu\).
What is inverse normal?
Inverse normal works backwards from a normal cdf calculation: given a probability (the area to the left), it finds the corresponding \(x\)-value or \(z\)-value. It's how you find a cut-off mark, like "the score needed to be in the top 10%".
e.g. If \(P(X>30)=0.2\) for \(\sigma=5\): \(z=\text{invNorm}(0.8)\approx0.8416\).
Key formulas
Six formulas cover almost every question on this topic. The two tables below summarise all of them at a glance - the explanations underneath go into more depth on each one.
Formula reference
The binomial mean and variance formulas are not listed separately in the formula booklet - they follow from prior knowledge of expected value and are usually confirmed by GDC output rather than substituted by hand.
| Formula | Used for | Booklet? |
|---|---|---|
| \(E(X) = \Sigma xP(X=x)\) | Expected value of a discrete random variable | Not in booklet |
| \(\text{Var}(X) = E(X^2)-[E(X)]^2\) | Variance of a discrete random variable | Not in booklet |
| \(X\sim B(n,p)\): \(E(X)=np\) | Mean of the binomial distribution | Not in booklet |
| \(\text{Var}(X)=np(1-p)\) | Variance of the binomial distribution | Not in booklet |
| \(z = \dfrac{x-\mu}{\sigma}\) | Standardising a normal value | Not in booklet |
| \(z = \text{invNorm}(\text{area})\) | Inverse normal (find \(x\) from a probability) | Not in booklet |
Binomial vs normal
These are the two distributions you'll meet most often - one discrete, one continuous. Knowing which applies is usually the first thing a question tests.
| Feature | Binomial \(B(n,p)\) | Normal \(N(\mu,\sigma^2)\) |
|---|---|---|
| Type | Discrete | Continuous |
| Models | Number of successes in \(n\) fixed trials | Measurements clustering around a mean |
| Parameters | \(n\) (trials), \(p\) (success probability) | \(\mu\) (mean), \(\sigma\) (standard deviation) |
| GDC functions | binompdf, binomcdf | normalcdf, invNorm |
| Mean | \(np\) | \(\mu\) (given directly) |
Discrete random variables
Everything here starts from the requirement that the probabilities in the distribution sum to 1.
Probability distribution
A table or formula giving \(P(X=x)\) for every possible value of \(X\). The probabilities must always sum to exactly 1 - this is often how you find an unknown constant.
Not in the formula booklet - prior knowledgeExpected value
\[E(X) = \Sigma xP(X=x)\]
The long-run average outcome. In a game of chance, \(E(X)=0\) for the player's net gain means the game is fair.
Not in the formula booklet - prior knowledgeVariance
\[\text{Var}(X)=E(X^2)-[E(X)]^2\]
Find \(E(X^2) = \Sigma x^2P(X=x)\) first, then subtract the square of the mean.
Not in the formula booklet - prior knowledgeBinomial distribution
Use the binomial model whenever there's a fixed number of independent, identical trials with only two outcomes each.
When it applies
A fixed number of trials \(n\); each trial is independent; each has the same success probability \(p\); only two outcomes per trial (success/failure).
Not in the formula booklet - prior knowledgeMean and variance
\[E(X)=np, \quad \text{Var}(X)=np(1-p)\]
Both come straight from \(n\) and \(p\) - no need to build the full distribution first.
Not in the formula booklet - prior knowledgepdf vs cdf
binompdf gives \(P(X=k)\) exactly; binomcdf gives \(P(X\le k)\). "At least" needs \(1-P(X\le k-1)\).
Not in the formula booklet - prior knowledgeNormal distribution
The normal distribution is defined entirely by its mean and standard deviation - every probability calculation reduces to an area under this one curve.
The 68-95-99.7 rule
About 68% of values lie within 1 standard deviation of \(\mu\), 95% within 2, and 99.7% within 3 - a quick sanity check for GDC output.
Not in the formula booklet - prior knowledgeStandardising
\[z=\frac{x-\mu}{\sigma}\]
Converts any normal value to the standard normal scale - useful for finding an unknown \(\mu\) or \(\sigma\) algebraically.
Not in the formula booklet - prior knowledgeFinding an unknown parameter
Turn each given probability into a \(z\)-value with inverse normal, then substitute into \(z=(x-\mu)/\sigma\). Two unknowns need two equations solved simultaneously.
Not in the formula booklet - prior knowledgeWorked examples
Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.
\(X\sim\mathrm{B}(50,0.2)\).
(a) Find \(E(X)\).
(b) Find \(\mathrm{Var}(X)\).
(c) Find the standard deviation, to 3 s.f.
Worked solution
(a) \(E(X)=np\) M1
\(=50(0.2)=10.\) A1
(b) \(\text{Var}(X) = 50(0.2)(0.8)\) M1
\(= 8.\) A1
(c) \(\sigma = \sqrt 8 \approx 2.83.\) A1
For a normally distributed variable, \(P(X < 40) = 0.10\) and \(P(X < 70) = 0.95.\)
(a) Write down two equations linking \(\mu\) and \(\sigma\) using \(z\)-values.
(b) Solve to find \(\mu\) and \(\sigma\).
Worked solution
(a) \(z_1 = -1.2816\) M1
\(= \dfrac{40 - \mu}{\sigma}\) A1
and \(z_2 = 1.6449 = \dfrac{70 - \mu}{\sigma}.\) A1
(b) Subtract: \(-30 = -2.9265\sigma \Rightarrow \sigma\) M1
\(\approx 10.3.\) A1
Then \(\mu = 40 + 1.2816(10.3)\) M1
\(\approx 53.1.\) A1
Common mistakes
The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.
- Using \(P(X\le k)\) when the question means \(P(X
For a discrete variable, "fewer than 3" is \(X\le2\), not \(X\le3\) - always translate the wording into an inequality before reaching for binomcdf. - Forgetting inverse normal needs the area to the LEFT. For "the top 10%", the area to enter is \(0.90\), not \(0.10\) - misreading this flips the answer to the wrong tail entirely.
- Mixing up binompdf and binomcdf. binompdf gives an exact probability \(P(X=k)\); binomcdf gives a cumulative probability \(P(X\le k)\) - using the wrong one is a very common one-mark slip.
- Assuming a variable is binomial without checking independence. The binomial model needs a fixed number of independent trials with a constant success probability - sampling without replacement from a small population breaks this and the model no longer applies exactly.
Using your GDC
Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.
The reverse of a normal probability: given a percentage, find the cut-off value (e.g. the mark for the top 10%).
- Work out the area to the LEFT of the value you want.
- 2nd → VARS (DISTR) → invNorm(area, μ, σ). Newer OS lets you pick the tail.TI-84
- menu → Probability → Distributions → Inverse Normal; enter the area, μ and σ.Nspire
- Main menu → Statistics → DIST → NORM → InvN; set the tail and enter area, σ, μ.Casio
Tip: invNorm needs the area to the LEFT. For "top 10%", use area = 0.90; for "bottom 25%", use area = 0.25.
Given a probability and one parameter, work back to the missing \(\mu\) or \(\sigma\) - a standard exam twist.
- Turn the probability into a \(z\)-value using the inverse normal with \(\mu = 0\), \(\sigma = 1\).
- invNorm(area-to-left, 0, 1) gives \(z\); then solve \(z = (x-\mu)/\sigma\) for the unknown.TI-84
- Use Inverse Normal with \(\mu=0\), \(\sigma=1\) to get \(z\), then solve the standardising equation for \(\mu\) or \(\sigma\).Nspire
- DIST → NORM → InvN with \(\mu=0\), \(\sigma=1\) gives \(z\); substitute into \(z=(x-\mu)/\sigma\).Casio
- If two probabilities are given, form two equations and solve them simultaneously for \(\mu\) and \(\sigma\).
Tip: Always sketch the curve first so you can tell which side of the mean each value should fall on.
See the full GDC guide for more calculator models and topics.
Ready to practise properly?
Probability & distributions questions, marked instantly like the real exam.
Quick answers
The questions students on this topic ask most often.
What's the difference between binomial and normal distributions?
Binomial is discrete - it counts successes out of a fixed number of independent trials, each with the same success probability. Normal is continuous - it models measurements like height or time that can take any value in a range and cluster symmetrically around a mean.
When do I use inverse normal instead of normal cdf?
Use normal cdf when you're given a value and need a probability. Use inverse normal when you're given a probability (or percentage) and need to find the corresponding value - inverse normal always needs the area to the LEFT of the value you want.
How do I find an unknown mean or standard deviation?
Convert each given probability to a \(z\)-value using inverse normal with mean 0 and standard deviation 1, then substitute into \(z = (x-\mu)/\sigma\). With one unknown you solve directly; with two unknowns (\(\mu\) and \(\sigma\)) you form two equations and solve them simultaneously.
Can I use my GDC for this topic?
Yes, and you're expected to. Binomial and normal probabilities are found using your GDC's distribution menu (binompdf/binomcdf, normalcdf, invNorm) rather than by formula - hand calculation is only for setting up the problem. See the GDC guide for model-specific instructions.
Sub-topics
Probability & Distributions broken down into its individual skills, each with its own focused page.
Related topics
More Statistics & Probability topics from the same AI HL syllabus unit, in case you want to keep going.