Probability Distributions (AA HL)
A probability distribution assigns a probability to every possible value of a random variable, and this topic covers the three ways IB expects you to work with one: discrete distributions built from a formula or table, the binomial distribution for repeated trials, and the continuous normal distribution. You'll find expected values, variances, and probabilities in both directions - from a value to a probability, and from a probability back to a value.
What the syllabus says
This topic maps onto four points in the official IB Analysis & Approaches syllabus, including one that's HL-only.
| Code | Syllabus content |
|---|---|
| SL4.7 | Concept of discrete random variables and their probability distributions. Expected value (mean) for discrete data. \(E(X)=0\) indicates a fair game where \(X\) is a player's gain. |
| SL4.8 | The binomial distribution, and its mean and variance. Binomial probabilities are found using technology; situations where the binomial model is appropriate. |
| SL4.9 | The normal distribution and curve, and its properties (approximately 68% of data within 1 standard deviation, 95% within 2, 99.7% within 3). Normal and inverse normal probabilities are found using technology. |
| SL4.12 | Standardization of normal variables (z-values), giving the number of standard deviations from the mean. Inverse normal calculations where the mean or standard deviation is unknown. |
| AHL4.14 | Variance of a discrete random variable, using \(\text{Var}(X)=E(X^2)-[E(X)]^2\). Continuous random variables and their probability density functions, including mode and median. |
SL4.7-SL4.12 are core AA content also examinable at HL; AHL4.14 is additional HL-only content.
Key terms
Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.
What is a discrete random variable?
A discrete random variable takes a countable set of possible values, each with its own probability, and those probabilities must sum to exactly 1. A probability distribution is the full listing (or formula) that pairs each value with its probability.
e.g. \(X\) takes values \(1,2,3\) with \(P(X=1)=0.2\), \(P(X=2)=0.5\), \(P(X=3)=0.3\); check: \(0.2+0.5+0.3=1\).
What is expected value, \(E(X)\)?
The expected value is the long-run average of a random variable if the trial were repeated many times - each value weighted by how likely it is. It's the probability-distribution equivalent of the mean.
e.g. \(X\) takes \(0,1,2\) with \(P(X=0)=0.5\), \(P(X=1)=0.3\), \(P(X=2)=0.2\): \(E(X)=0(0.5)+1(0.3)+2(0.2)=0.7\).
What is the binomial distribution?
The binomial distribution \(X\sim B(n,p)\) models the number of successes in \(n\) fixed, independent trials, each with the same success probability \(p\). It only applies when there are exactly two outcomes per trial and \(p\) doesn't change between trials.
e.g. \(X\sim B(5,0.4)\): \(P(X=2)=\binom{5}{2}(0.4)^2(0.6)^3=10\times0.16\times0.216=0.3456\).
What is the normal distribution?
The normal distribution \(X\sim N(\mu,\sigma^2)\) is a continuous, symmetric, bell-shaped distribution centred on the mean \(\mu\). About 68% of values lie within 1 standard deviation of the mean, 95% within 2, and 99.7% within 3.
e.g. \(X\sim N(50,10^2)\): \(P(X<60)\) has \(z=\dfrac{60-50}{10}=1\), which a GDC evaluates to \(P(X<60)\approx0.8413\).
What is standardization?
Standardizing converts a normal variable's value into a z-value: the number of standard deviations it sits from the mean, using \(z=\dfrac{x-\mu}{\sigma}\). It puts any normal distribution onto the same standard scale.
e.g. \(X\sim N(20,4^2)\), \(x=26\): \(z=\dfrac{26-20}{4}=1.5\).
Key formulas
These formulas cover discrete random variables, the binomial distribution and the normal distribution - the tables below summarise them at a glance.
Formula reference
All the formulas below are on the official formula booklet, though almost every actual probability or inverse-probability value must still be found using your GDC.
| Formula | Used for | Booklet? |
|---|---|---|
| \(E(X)=\sum x\,P(X=x)\) | Expected value of a discrete RV | ✓ Yes |
| \(\text{Var}(X)=E(X^2)-[E(X)]^2\) | Variance of a discrete RV (HL) | ✓ Yes |
| \(P(X=r)=\binom{n}{r}p^r(1-p)^{n-r}\) | Binomial probability | ✓ Yes |
| \(E(X)=np,\ \text{Var}(X)=np(1-p)\) | Binomial mean & variance | ✓ Yes |
| \(z=\dfrac{x-\mu}{\sigma}\) | Standardizing a normal variable | ✓ Yes |
Binomial vs normal: discrete or continuous?
These are the two named distributions on the syllabus, and they model fundamentally different kinds of variable.
| Feature | Binomial \(B(n,p)\) | Normal \(N(\mu,\sigma^2)\) |
|---|---|---|
| Type of variable | Discrete - counts successes | Continuous - any value in a range |
| Parameters | \(n\) (trials), \(p\) (success probability) | \(\mu\) (mean), \(\sigma\) (standard deviation) |
| Shape | Can be skewed, especially for small \(n\) or extreme \(p\) | Always symmetric, bell-shaped |
| Typical use | Number of defective items, correct guesses, successes | Heights, masses, measurement errors |
Discrete random variables
Every discrete distribution must obey one basic rule before anything else can be calculated from it.
Valid probability distribution
\[\sum P(X=x) = 1\]
Every discrete distribution's probabilities must sum to exactly 1 - this is usually how an unknown constant \(k\) is found.
Not in the formula booklet - prior knowledgeExpected value
\[E(X)=\sum x\,P(X=x)\]
The probability-weighted average - the long-run mean outcome if the trial were repeated indefinitely.
✓ In the formula bookletVariance (HL)
\[\text{Var}(X)=E(X^2)-[E(X)]^2\]
Find \(E(X^2)=\sum x^2 P(X=x)\) first, then subtract the square of the mean.
✓ In the formula bookletThe binomial distribution
Binomial questions are recognisable from the setup: a fixed number of identical, independent trials with two outcomes each.
Conditions for binomial
A fixed number of trials \(n\); each trial has exactly two outcomes; trials are independent; the probability of success \(p\) is constant across every trial.
Mean & variance
\[E(X)=np \qquad \text{Var}(X)=np(1-p)\]
No need to sum the full distribution - these shortcuts come straight from the parameters \(n\) and \(p\).
✓ In the formula bookletThe normal distribution
Normal distribution questions run in two directions: value-to-probability, and probability-to-value.
Standardizing (z-value)
\[z=\dfrac{x-\mu}{\sigma}\]
Converts any normal variable onto the standard \(N(0,1)\) scale - the number of standard deviations from the mean.
✓ In the formula bookletInverse normal
Given a probability, find the corresponding \(x\)-value - always found using the GDC's inverse normal function, working from the area to the LEFT of the value.
Worked examples
Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.
The discrete random variable \(X\) has \(P(X=x)=kx\) for \(x=1,2,3,4\).
(a) Find \(k\).
(b) Find \(E(X)\).
Worked solution
(a) \(\sum P = k(1+2+3+4)=10k=1\) M1
so \(k=\tfrac{1}{10}\). A1
(b) \(E(X)=\tfrac{1}{10}(1\cdot1+2\cdot2+3\cdot3+4\cdot4)=\tfrac{30}{10}\) M1
\(=3\). A1
\(X\sim N(30,\sigma^2)\). The central 95% of values lie within \(z=\pm1.96\) of the mean and the interval is \((26.08,33.92)\). Find \(\sigma\).
Worked solution
Half-width \(=33.92-30=3.92\) M1
\(=1.96\sigma\). A1
\(\sigma=\dfrac{3.92}{1.96}=2\). A1
Common mistakes
The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.
- Forgetting probabilities must sum to 1. This is the standard way to find an unknown constant \(k\) in a discrete distribution - if the probabilities you've written down don't sum to 1, something upstream is wrong.
- Using binomial when the conditions don't hold. Binomial needs a fixed number of independent trials with a constant success probability - drawing without replacement, or a probability that changes between trials, rules it out.
- Getting "at least" and "at most" backwards on the GDC. \(P(X\ge k)\) and \(P(X\le k)\) are not the same calculation, and one is usually \(1\) minus the other - always sketch which region is being asked for.
- Reporting the z-value instead of the actual x-value. Inverse normal calculations are asking for a real value in context (a mass, a time, a mark), not the intermediate standardized z-value used to find it.
Using your GDC
Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.
The reverse of a normal probability: given a percentage, find the cut-off value (e.g. the mark for the top 10%).
- Work out the area to the LEFT of the value you want.
- 2nd → VARS (DISTR) → invNorm(area, μ, σ). Newer OS lets you pick the tail.TI-84
- menu → Probability → Distributions → Inverse Normal; enter the area, μ and σ.Nspire
- Main menu → Statistics → DIST → NORM → InvN; set the tail and enter area, σ, μ.Casio
Tip: invNorm needs the area to the LEFT. For "top 10%", use area = 0.90; for "bottom 25%", use area = 0.25.
Given a probability and one parameter, work back to the missing μ or σ - a standard exam twist.
- Turn the probability into a z-value using the inverse normal with μ = 0, σ = 1.
- invNorm(area-to-left, 0, 1) gives z; then solve \(z = (x - \mu)/\sigma\) for the unknown.TI-84
- Use Inverse Normal with μ = 0, σ = 1 to get z, then solve the standardising equation for μ or σ.Nspire
- DIST → NORM → InvN with μ = 0, σ = 1 gives z; substitute into \(z = (x - \mu)/\sigma\).Casio
- If two probabilities are given, form two equations and solve them simultaneously for μ and σ.
Tip: Sketch the curve and shade the area on the correct side before using invNorm.
See the full GDC guide for more calculator models and topics.
Ready to practise properly?
Probability distributions questions, marked instantly like the real exam.
Quick answers
The questions students on this topic ask most often.
How do I know when to use the binomial distribution?
Use the binomial distribution when there is a fixed number of independent trials, each with only two outcomes (success/failure), and the probability of success stays the same on every trial. If any of those conditions fail, binomial is the wrong model.
Do I need to memorise the normal distribution formula?
No. Normal probabilities and inverse normal values must be found using your GDC's distribution functions - there's no expectation you evaluate the normal density function by hand.
What does standardizing a normal variable actually do?
It converts an x-value into a z-value, the number of standard deviations that value is from the mean, using \(z = \dfrac{x - \mu}{\sigma}\). This puts any normal distribution onto the same standard scale, which is what inverse normal calculations work with internally.
What's the difference between a normal probability and an inverse normal calculation?
A normal probability calculation goes from a value to a probability (find \(P(X<60)\)). An inverse normal calculation goes the other way, from a probability to a value (find the \(x\) such that \(P(X
Sub-topics
Probability Distributions broken down into its individual skills, each with its own focused page.
Related topics
More Statistics & Probability topics from the same AA HL syllabus unit, in case you want to keep going.