Random Variables & Distributions (AA SL)
A random variable assigns a number to the outcome of a random process, and a probability distribution describes how likely each possible value is. This topic covers discrete random variables and expected value, the binomial distribution for counting successes in repeated trials, and the normal distribution for continuous data that clusters symmetrically around a mean.
What the syllabus says
This topic maps onto three points in the official IB Analysis & Approaches syllabus.
| Code | Syllabus content |
|---|---|
| SL4.7 | Concept of discrete random variables and their probability distributions. Expected value (mean) for discrete data, \(E(X)=\sum xP(X=x)\). Applications, including \(E(X)=0\) indicating a fair game. |
| SL4.8 | Binomial distribution, including mean and variance of the binomial distribution. In examinations, binomial probabilities should be found using technology. |
| SL4.9 | The normal distribution and curve. Properties of the normal distribution: approximately 68% of the data lies between \(\mu\pm\sigma\), 95% between \(\mu\pm2\sigma\), 99.7% between \(\mu\pm3\sigma\). Normal probability calculations and inverse normal calculations, found using technology. |
These are core AA SL syllabus points that are also examinable at AA HL.
Key terms
Five words worth knowing cold before you touch the formulas below - each with a worked example showing exactly what it means.
What is a discrete random variable?
A discrete random variable takes a countable set of numerical values, each with its own probability, and all the probabilities must sum to 1. It's usually written as a table listing each value of \(X\) alongside \(P(X=x)\).
e.g. For \(P(X=0)=0.1,\ P(X=1)=0.4,\ P(X=2)=0.3,\ P(X=3)=0.2\), the probabilities sum to \(0.1+0.4+0.3+0.2=1.\)
What is expected value?
The expected value \(E(X) = \sum xP(X=x)\) is the long-run average value of \(X\) if the random process were repeated many times. For a game, \(E(X)=0\) means the game is fair on average.
e.g. For \(P(X=0)=0.1,\ P(X=1)=0.4,\ P(X=2)=0.3,\ P(X=3)=0.2\): \(E(X)=0(0.1)+1(0.4)+2(0.3)+3(0.2)=1.6.\)
What is the binomial distribution?
The binomial distribution \(X\sim B(n,p)\) models the number of successes in \(n\) independent trials, each with the same probability of success \(p\). Its mean is \(np\) and it's found on your GDC using binompdf and binomcdf.
e.g. For 25 items with an 8% defect rate, \(X\sim B(25,0.08)\) and \(E(X)=np=25(0.08)=2.\)
What is the normal distribution?
The normal distribution \(X\sim N(\mu,\sigma^2)\) models continuous data that clusters symmetrically around a mean \(\mu\), with spread controlled by the standard deviation \(\sigma\). About 68% of values lie within one standard deviation of the mean.
e.g. For heights \(X\sim N(170,6^2)\), \(P(164
What is an inverse normal calculation?
An inverse normal calculation works backwards from a known probability to find the corresponding boundary value of \(X\), instead of finding a probability from a given value. You always enter the area to the LEFT of the boundary.
e.g. For \(X\sim N(170,6^2)\), the height exceeded by only 10% of people uses \(P(X
Key formulas
This topic has one core formula for discrete variables and two GDC-based distributions to recognise. The tables below summarise everything at a glance - the explanations underneath go into more depth.
Formula reference
Expected value is on the official formula booklet as \(E(X)=\sum xP(X=x)\); binomial and normal probabilities are found using technology rather than a hand formula, so the syllabus does not list a probability mass or density function to memorise for either.
| Formula | Used for | Booklet? |
|---|---|---|
| \(E(X)=\sum xP(X=x)\) | Expected value of a discrete random variable | ✓ Yes |
| \(X\sim B(n,p)\), \(E(X)=np\) | Binomial distribution and its mean | ✓ Yes |
| \(X\sim N(\mu,\sigma^2)\) | Normal distribution notation | Not in booklet - notation, not a formula |
| 68% / 95% / 99.7% rule | Quick estimates within \(\mu\pm\sigma,\ \mu\pm2\sigma,\ \mu\pm3\sigma\) | Not in booklet - prior knowledge |
Binomial vs normal
Both are named distributions found using your GDC, but they model fundamentally different kinds of data.
| Feature | Binomial | Normal |
|---|---|---|
| Data type | Discrete (counts) | Continuous (measurements) |
| Notation | \(X\sim B(n,p)\) | \(X\sim N(\mu,\sigma^2)\) |
| Parameters | Number of trials \(n\), probability \(p\) | Mean \(\mu\), standard deviation \(\sigma\) |
| GDC function | binompdf / binomcdf | normalcdf / invNorm |
Discrete random variables
Probabilities sum to 1
\[\sum P(X=x)=1\]
A useful check, or a way to find a missing probability in a distribution table.
Not in the formula booklet - prior knowledgeFair game
\[E(X)=0\]
If \(X\) represents a player's net gain, \(E(X)=0\) means the game is fair on average.
Not in the formula booklet - key exam ideaBinomial and normal distributions
Conditions for binomial
Fixed number of trials \(n\); two outcomes per trial; constant probability \(p\); independent trials.
All four conditions must hold before you can write \(X\sim B(n,p)\).
Not in the formula booklet - key exam idea68-95-99.7 rule
About 68% of data lies within \(\mu\pm\sigma\), 95% within \(\mu\pm2\sigma\), 99.7% within \(\mu\pm3\sigma\).
A fast sanity check for normal probabilities without reaching for the GDC.
Not in the formula booklet - property to knowpdf vs cdf
pdf gives the probability of an exact value; cdf gives the probability of "at most" that value.
Combine cdf with the complement for "at least" or "more than" questions.
Not in the formula booklet - key exam ideaWorked examples
Two full exam-style questions, marked exactly like the real thing. Try each one yourself before checking the worked solution.
A fair coin is tossed 6 times.
(a) Find the probability of exactly 4 heads.
(b) Find the probability of at least one head.
Worked solution
(a) Exactly 4 heads. \(X\sim B(6,0.5):\) \(P(X=4)=\binom{6}{4}(0.5)^4(0.5)^2=15(0.5)^6=\dfrac{15}{64}\) M1
\(\approx 0.234.\) A1
(b) At least one head. Use the complement: \(P(X\ge 1)=1-P(X=0)=1-(0.5)^6=\dfrac{63}{64}\approx 0.984.\) M1 A1
Heights are \(X\sim N(170, 6^2)\) cm.
(a) Find \(P(164 < X < 176)\).
(b) Find the height exceeded by only 10% of people.
Worked solution
(b) Height exceeded by only 10%. Need \(x\) with \(P(X>x)=0.10,\) so \(z=1.282.\) M1
\(x=170+1.282(6)\approx 177.7\) cm. A1
Common mistakes
The four slip-ups that account for most of the marks lost on this topic - worth reading before you start practising, not just after you get one wrong.
- Using binompdf when the question asks for "at most" or "at least". binompdf gives an exact value only; "at most" needs binomcdf, and "at least" needs \(1-\)binomcdf of one below the threshold.
- Entering the wrong area into invNorm. invNorm always needs the area to the LEFT of the value you want - for a "top 10%" cutoff, that means entering 0.90, not 0.10.
- Not checking the four binomial conditions. A situation is only binomial if trials are fixed in number, independent, and have a constant probability of success - sampling without replacement from a small population breaks this.
- Confusing E(X) with a value X can actually take. \(E(X)\) is a long-run average and can easily be a number \(X\) never equals, like \(E(X)=1.6\) for a variable that only takes whole-number values.
Using your GDC
Every step below is a real button sequence, not a vague "use your calculator" hint - covering the TI-84 Plus, TI-Nspire, and Casio fx-9860/fx-CG50. Pick your model to filter down to just the steps that apply to you.
Used to check the mean and spread of a discrete distribution entered as a list, and as the underlying tool behind binomial mean/variance reasoning.
- Enter the data (or the values of \(X\) with their frequencies) into a list.
- STAT → Edit → type values into L1 (and probabilities/frequencies into L2). Then STAT → CALC → 1:1-Var Stats, choose L1, L2, Calculate.TI-84
- Add a Lists & Spreadsheet page, name two columns and enter data; then a Calculator page → menu → Statistics → Stat Calculations → One-Variable Statistics.Nspire
- Statistics menu → enter data in List 1, frequencies in List 2 → CALC (F2) → 1-VAR (set the frequency list).Casio
- Read \(\bar{x}\) (mean), \(S_x\) (sample sd) or \(\sigma_x\) (population sd).
Tip: Sx vs σx: use σx (population) for a complete data set, Sx (sample) for a sample. IB usually wants σx.
Useful for checking whether paired data on a random variable looks approximately linear or normal-shaped before committing to a model.
- Enter the x-values in one list and the y-values in another (equal lengths).
- 2nd → Y= (STAT PLOT) → Plot1 On → scatter icon → set Xlist and Ylist, then ZOOM → 9:ZoomStat.TI-84
- On a Data & Statistics page click the x- and y-axis labels to assign the two variables.Nspire
- In Statistics set GRPH → SET to Scatter, assign XList and YList, then DRAW.Casio
- Describe the shape (positive/negative, strong/weak, linear or not) from the plot.
Tip: The scatter shape tells you whether a linear, quadratic or exponential model is sensible for the underlying random variable.
See the full GDC guide for more calculator models and topics.
Ready to practise properly?
Random variable and distribution questions, marked instantly like the real exam.
Quick answers
The questions students on this topic ask most often.
What does E(X) actually mean?
E(X), the expected value, is the long-run average outcome of a random variable if the trial were repeated many, many times - not a value X will necessarily ever take. For a game, E(X)=0 means the game is fair on average, even though any single play wins or loses a specific amount.
How do I know if a situation is binomial?
Check for four things: a fixed number of trials n, each trial has only two outcomes (success/failure), the probability of success p is constant across trials, and the trials are independent of each other. If all four hold, X ~ B(n, p).
Why don't I need to standardize for AA SL normal probabilities?
The GDC's normal probability and inverse normal functions take the mean and standard deviation directly, so you don't need to convert to the standard normal variable Z to find a probability or an unknown value. Standardization (finding z) is only required separately at SL4.12, mainly to interpret how many standard deviations a value is from the mean.
What's the difference between binompdf and binomcdf?
binompdf(n, p, r) gives the probability of exactly r successes. binomcdf(n, p, r) gives the probability of at most r successes (a running total from 0 up to r). Use cdf whenever you see "at most", and combine it with the complement for "at least" or "more than". See the GDC guide for model-specific instructions.
Sub-topics
Random Variables & Distributions broken down into its individual skills, each with its own focused page.
Related topics
More Statistics & Probability topics from the same AA SL syllabus unit, in case you want to keep going.