Distributions (AI SL)
A probability distribution describes every possible value of a random variable and how likely each one is. This topic covers discrete random variables and expected value, the binomial distribution for counting successes in a fixed number of independent trials, and the normal distribution for continuous quantities that cluster symmetrically around a mean - including reading probabilities off the curve and working backwards with the inverse normal.
What the syllabus says
This topic maps onto three points in the official IB Applications & Interpretation syllabus.
| Code | Syllabus content |
|---|---|
| SL4.7 | Concept of discrete random variables and their probability distributions. Expected value (mean), \(E(X)\), for discrete data. Applications. \(E(X)=0\) indicates a fair game where \(X\) represents the gain of a player. |
| SL4.8 | Binomial distribution. Mean and variance of the binomial distribution. In examinations, binomial probabilities should be found using available technology. Formal proof of mean and variance is not required. |
| SL4.9 | The normal distribution and curve. Properties of the normal distribution, including that approximately 68% of the data lies between \(\mu\pm\sigma\), 95% between \(\mu\pm2\sigma\) and 99.7% between \(\mu\pm3\sigma\). Normal probability calculations and inverse normal calculations must be found using technology. |
For inverse normal calculations the mean and standard deviation are always given - you don't need to standardise to \(z\) by hand.
Key terms
Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.
What is a discrete random variable?
A discrete random variable takes a countable set of values, each with its own probability, and those probabilities sum to 1. It's the framework behind both the binomial distribution and simpler distributions given as a table.
e.g. If \(P(X=1)=0.2\), \(P(X=3)=0.5\), \(P(X=5)=0.3\), the probabilities sum to \(0.2+0.5+0.3=1\).
What is the binomial distribution?
The binomial distribution, \(X\sim B(n,p)\), models the number of successes in \(n\) independent trials, each with the same success probability \(p\). Its mean is \(np\) and its variance is \(np(1-p)\).
e.g. \(X\sim B(20,0.3)\) has mean \(E(X)=20(0.3)=6\).
What is the normal distribution?
The normal distribution is a symmetric, bell-shaped curve for continuous data, defined by its mean \(\mu\) and standard deviation \(\sigma\), written \(X\sim N(\mu,\sigma^2)\). Many natural measurements - heights, masses, exam scores - are approximately normal.
e.g. For \(X\sim N(500,5^2)\), about 68% of values lie between 495 and 505.
What is an inverse normal calculation?
An inverse normal calculation works backwards from a known probability to find the cut-off value of \(X\) - the reverse of a normal probability calculation. You always give the area to the LEFT of the unknown value.
e.g. For \(X\sim N(58,12^2)\), the mark for the top 10% is \(\text{invNorm}(0.90,58,12)\approx 73.4\).
What is the 68-95-99.7 rule?
For any normal distribution, about 68% of values lie within one standard deviation of the mean, 95% within two, and 99.7% within three. It's a fast sanity check for normal probability answers before you trust the GDC's decimal output.
e.g. For \(X\sim N(500,5^2)\), \(P(495
Key formulas
Two formulas cover the binomial side of this topic by hand; the normal distribution is technology-only. The tables below summarise both - the explanations underneath go into more depth on each.
Formula reference
The binomial mean and variance formulas are in the official formula booklet. Normal probabilities and inverse normal values have no hand formula - the syllabus requires you to find them using your GDC.
Binomial vs normal
These are the two distributions this topic covers - choosing the right one is often the hardest part of the question, before any calculation begins.
| Feature | Binomial | Normal |
|---|---|---|
| Type of data | Discrete count | Continuous measurement |
| Notation | \(X\sim B(n,p)\) | \(X\sim N(\mu,\sigma^2)\) |
| Mean | \(np\) | \(\mu\) (given directly) |
| Typical use | Number of successes in \(n\) trials | Heights, weights, exam marks |
| Example | Number of heads in 20 coin flips | Heights of adult women in a city |
Binomial distribution
The binomial distribution applies whenever you have a fixed number of independent trials, each with the same two-outcome probability.
Mean
\[E(X)=np\]
The average number of successes over many repeats of \(n\) trials.
✓ In the formula bookletVariance and standard deviation
\[\text{Var}(X)=np(1-p),\quad \sigma=\sqrt{np(1-p)}\]
Variance is always smaller than the mean for a binomial distribution, since \(1-p<1\).
✓ In the formula bookletProbabilities by GDC
Exact probabilities \(P(X=r)\) use the binomial pdf; "at most" or "at least" probabilities use the binomial cdf. Both must be found using technology in examinations.
Not in the formula booklet - GDC requiredNormal distribution
The normal distribution is defined entirely by its mean and standard deviation - every probability question on it is solved with the GDC's normal-distribution functions.
68-95-99.7 rule
About 68% of data lies within \(\mu\pm\sigma\), 95% within \(\mu\pm2\sigma\), and 99.7% within \(\mu\pm3\sigma\) - a quick check on any normal probability answer.
Not in the formula booklet - general propertyNormal probability
Find \(P(Xa)\) or \(P(a
Inverse normal
Given a probability (area to the left), find the corresponding value of \(X\) using your GDC's inverse normal function - \(\mu\) and \(\sigma\) are always given.
Not in the formula booklet - GDC requiredWorked examples
Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.
\(X\sim B(20, 0.3)\).
(a) Find the mean \(E(X)\).
(b) Find the variance.
Worked solution
(a) Mean. \(E(X)=np=20(0.3)\) M1 \(=6.\) A1
(b) Variance. \(\text{Var}(X)=np(1-p)=20(0.3)(0.7)\) M1 \(=4.2.\) A1
Check. The variance is always smaller than the mean for a binomial (since \(1-p<1\)); here \(4.2<6,\) as expected.
Exam marks are \(N(58, 12^2)\). The top 10% receive a distinction.
(a) Find the minimum mark for a distinction (3 s.f.).
(b) Interpret your answer.
(c) Find the probability a randomly selected student scores below 40.
Worked solution
(a) invNorm 90th percentile. M1 \(\approx73.4.\) A1 This means the top 10% of students need at least 73.4 marks for a distinction. R1
(c) \(z=-1.5;\ P(X<40)\approx0.0668.\) M1 A1
Common mistakes
The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.
- Using the normal distribution for a discrete count. A count of successes from a fixed number of trials (e.g. number of heads in 20 flips) is binomial, not normal - check whether the data is discrete or continuous before choosing a model.
- Getting "at least" and "at most" the wrong way round. \(P(X\ge r)=1-P(X\le r-1)\), not \(1-P(X\le r)\) - always subtract one from the count before complementing, or you'll be off by one probability term.
- Forgetting invNorm needs the area to the LEFT. For "the top 10%", the area to the left is \(0.90\), not \(0.10\) - sketch the curve and shade the region first so you enter the correct area.
- Reporting too few significant figures too early. Round only the final answer, not intermediate probabilities - rounding early can shift a 3 s.f. answer in the last digit.
Using your GDC
Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.
Finds an exact or cumulative binomial probability without expanding \(\binom{n}{r}p^r(1-p)^{n-r}\) by hand.
- Decide whether you need "exactly \(r\)" (pdf) or "at most / at least \(r\)" (cdf).
- 2nd → DISTR → binompdf(\(n,p,r\)) for exactly \(r\); binomcdf(\(n,p,r\)) for at most \(r\).TI-84
- DIST → BINM → Bpd for exactly \(r\); Bcd for at most \(r\).Casio
- menu → Probability/Statistics → Distributions → Binomial Pdf or Binomial Cdf.Nspire
- For "at least \(r\)", compute \(1-\)binomcdf\((n,p,r-1)\) - subtract one from the count before complementing.
Tip: \(\texttt{binompdf}(30,\ 0.3,\ 10)\approx 0.142\) - check the value is near the peak if \(r\) is close to the mean \(np\).
The reverse of a normal probability: given a percentage, find the cut-off value (e.g. the mark for the top 10%).
- Work out the area to the LEFT of the value you want.
- 2nd → VARS (DISTR) → invNorm(area, \(\mu\), \(\sigma\)). Newer OS lets you pick the tail.TI-84
- menu → Probability → Distributions → Inverse Normal; enter the area, \(\mu\) and \(\sigma\).Nspire
- Main menu → Statistics → DIST → NORM → InvN; set the tail and enter area, \(\sigma\), \(\mu\).Casio
Tip: invNorm needs the area to the LEFT. For "top 10%", use area = 0.90; for "bottom 25%", use area = 0.25.
See the full GDC guide for more calculator models and topics.
Ready to practise properly?
Distributions questions, marked instantly like the real exam.
Quick answers
The questions students on this topic ask most often.
How do I know whether to use a binomial or a normal distribution?
Binomial models a count of successes from a fixed number of independent yes/no trials (e.g. number of heads in 20 flips). Normal models a continuous measurement that clusters symmetrically around a mean (e.g. heights, weights). If you're counting discrete successes, it's binomial; if you're measuring a continuous quantity, it's normal.
Do I need to memorise the binomial or normal formulas?
The mean and variance formulas for the binomial distribution, \(E(X) = np\) and \(\text{Var}(X) = np(1-p)\), are in the formula booklet. Normal probabilities and inverse normal values are never calculated by hand - the syllabus requires you to find them using your GDC.
What does invNorm actually do?
invNorm reverses a normal probability calculation: instead of giving you a probability from a value, you give it the area to the LEFT of an unknown value and it returns that value. It's how you find things like "the mark needed for the top 10%".
Is this topic examined with a calculator?
Mostly yes. Binomial and normal probabilities must be found using technology, so most questions appear on the calculator papers. A few questions - like stating a distribution or finding the mean of a simple binomial - can appear without a calculator.
Sub-topics
Distributions broken down into its individual skills, each with its own focused page.
Related topics
More Statistics & Probability topics from the same AI SL syllabus unit, in case you want to keep going.