Educerie · IB Diploma · Mathematics: analysis and approaches
Topic 4 Statistics and probability · 4.8 The binomial distribution
What you must be able to do
| You must be able to | Level | What it looks like in the exam |
|---|---|---|
| Recognise when the binomial distribution is an appropriate model, and state the conditions | SL, HL | "State two assumptions needed to model X with a binomial distribution" (2 marks) |
| Write the model in notation, X ~ B(n, p), with the variable defined in words | SL, HL | "Write down the distribution of X" (1 to 2 marks) |
| Find P(X = r) with technology | SL, HL | "Find the probability that exactly 7 seeds germinate" (2 marks), Paper 2 |
| Find cumulative probabilities: at most, fewer than, at least, more than, between | SL, HL | "Find the probability that at least 20 parcels arrive on time" (2 to 3 marks), Paper 2 |
| Use E(X) = np and Var(X) = np(1 − p) from the formula booklet | SL, HL | "Find the expected number…" (1 to 2 marks); "Given E(X) = 6 and Var(X) = 4.2, find n and p" (4 to 5 marks), Paper 1 |
| Solve for an unknown n, usually "the least number of trials so that…" | SL, HL | "Find the minimum number of cards that must be bought" (3 to 5 marks), Paper 2 |
| Use a binomial probability inside a longer, two-stage problem | SL, HL | Section B: find p from one model, then count successes with B(n, p) |
Before you start
You need the basic probability of 4.5 and 4.6: probabilities of independent events multiply, P(A′) = 1 − P(A), and the expected number of occurrences of an event in n trials is np. The binomial coefficient ⁿCᵣ from 1.9 (the binomial theorem) counts the routes through a tree, and it explains where the probabilities come from, although in the exam your GDC does the counting. Section 4.7 introduced discrete random variables and their expected value; the binomial is the most important named example of one.
1The idea in one paragraph
Many chance situations have the same skeleton: a trial with two outcomes, success or failure, repeated a fixed number of times, n, with the same probability of success, p, every time and no trial affecting another. Guess the answers to ten multiple-choice questions; plant twenty seeds; send thirty parcels. The number of successes, X, is then a random variable with a binomial distribution, written X ~ B(n, p). Its probabilities come from counting routes through a tree, and your GDC finds them in one line. Its mean is np, which is just the expected number of occurrences you already know, and its variance is np(1 − p). The skill that earns most marks is not the arithmetic. It is deciding whether the model fits, and turning phrases such as "at least 4" or "fewer than 4" into the right whole numbers.
2When the binomial model fits
A binomial random variable counts the number of successes in a set of trials. It is the right model only when all four of these hold, and Figure 1 lays them out as a checklist.
- A fixed number of trials, n, decided before you start.
- Two outcomes per trial, success or failure. "Success" is just the outcome you are counting; it can be a faulty bulb.
- The same probability of success, p, on every trial.
- The trials are independent: the result of one does not change the probability for another.
Practise on examples, because the exam asks you to judge.
- A student guesses every answer on a 10-question test where each question has 4 options, and X is the number correct. Fixed n = 10, correct or wrong, p = 1/4 each time, and one guess does not affect the next. Binomial, X ~ B(10, 0.25).
- You roll a die until you get a six, and X is the number of rolls. There is no fixed n: the count of trials is the thing that varies. Not binomial.
- You draw 5 cards from a pile of 20 in which 8 are red, without replacement, and X is the number of red cards. The first card is red with probability 8/20, but after a red has gone the next is red with probability 7/19. p changes, so the trials are not independent. Not binomial. With replacement it would be.
- X is the number of rainy days in a week in one town. There is a fixed n = 7 and two outcomes, but rain tends to come in runs (a wet day makes a wet tomorrow more likely), so independence is doubtful. A binomial model here is an assumption you would have to state, and a weak one.
Real samples are usually drawn without replacement, so the last two points matter. If you choose 10 people from a town of 40 000, removing one person changes the proportion of the rest by a negligible amount. The binomial is then an excellent approximation, and the exam treats it as the model. When the population is small, as with the 20 cards, it is not.
When a question asks you to state an assumption, name the condition that could fail in that context, and put it in the context's own words: "the probability that a seed germinates is the same for every seed, and seeds germinate independently of one another." A list of four textbook conditions with no context earns less.
X ~ B(n, p) needs: fixed n, two outcomes, constant p, independent trials. Define X in words every time: "X is the number of seeds, out of 20, that germinate."
3Where the probabilities come from
A student guesses three questions, each with four options, so the probability of guessing a question correctly is 1/4 and of getting it wrong is 3/4. Let X be the number correct. Figure 2 draws the tree.
There are 2³ = 8 routes. Look at the three routes with exactly two correct: CCW, CWC and WCC. Each has two C branches and one W branch, so each has probability (1/4)² × (3/4) = 3/64, whatever the order. There are three of them, so
That is the whole idea. Every route with r successes out of n has the same probability, pʳ(1 − p)ⁿ⁻ʳ, and the number of such routes is the number of ways to choose which r of the n trials are the successes, ⁿCᵣ. So
The four values for three guesses are 27/64, 27/64, 9/64 and 1/64 for X = 0, 1, 2, 3, and they add to 64/64 = 1, as a probability distribution must. The numbers of routes, 1, 3, 3, 1, are the row of the binomial coefficients you met in 1.9, and that is why this is called the binomial distribution: the probabilities are the terms of the expansion of (q + p)ⁿ, where q = 1 − p. The triangle of these coefficients is usually named after Pascal, who wrote about it in the seventeenth century, but the Chinese mathematician Yang Hui had published it in the thirteenth century. A name is not always a record of who found something first.
You need this formula to understand the distribution and to do the few cases Paper 1 can ask without a calculator: P(X = 0) = (1 − p)ⁿ, because every trial must fail; P(X = n) = pⁿ, because every trial must succeed; and a small case like the one above. The guide says binomial probabilities in examinations are found with technology, so on Paper 2 you use the GDC and the formula stays in your head.
4Using the GDC, and turning words into numbers
Every GDC has two binomial functions. Their names differ by model, but they do the same two jobs.
| What you want | Function | TI-84 Plus name | Casio name |
|---|---|---|---|
| P(X = r), one value | binomial pdf | binompdf(n, p, r) | Binomial PD |
| P(X ≤ r), everything up to and including r | binomial cdf | binomcdf(n, p, r) | Binomial CD |
Some models, such as the TI-Nspire and Casio, let the cumulative function take a lower and an upper bound; others take only the upper bound, and then you subtract. Either way, the cumulative function counts up from 0 and includes r itself. Every mistake in this section comes from forgetting one half of that sentence.
The exam's wording has to be turned into whole numbers first. Because X can only be 0, 1, 2, …, "fewer than 4" and "at most 3" are the same event, and "more than 4" is the same as "at least 5". Figure 3 rings the values for each phrase.
| Words | Values of X | On the GDC |
|---|---|---|
| exactly 4 | 4 | pdf at 4 |
| at most 4, no more than 4 | 0 to 4 | cdf at 4 |
| fewer than 4, less than 4 | 0 to 3 | cdf at 3 |
| at least 4, no fewer than 4 | 4 to n | 1 − cdf at 3 |
| more than 4 | 5 to n | 1 − cdf at 4 |
| between 2 and 6 inclusive | 2 to 6 | cdf at 6 − cdf at 1 |
The "at least" line uses the complement. P(X ≥ 4) is everything from 4 upwards, which is everything except 0, 1, 2 and 3, so P(X ≥ 4) = 1 − P(X ≤ 3). The number after "1 −" is always one less than the number in the question.
Worked example 1: ten guesses. A student guesses all 10 questions on a test where each has four options. Let X be the number correct, so X ~ B(10, 0.25). Figure 4 shows the whole distribution.
So a student who passes by getting at least half the questions right has less than an 8% chance of doing it by pure guessing.
Worked example 2: parcels. A courier delivers each parcel on time with probability 0.92, independently. A shop sends 25 parcels. Let Y be the number delivered on time, so Y ~ B(25, 0.92).
Write the probability statement, such as P(Y ≤ 21), as a line of working before the number. It shows the examiner the whole numbers you chose, and if the value is wrong because of a keying slip, that line still earns the method mark.
5Mean and variance
The formula booklet gives, for X ~ B(n, p):
E(X) = np and Var(X) = np(1 − p). The standard deviation is √(np(1 − p)).
The mean is the expected number of occurrences from 4.5 under a new name. If you guess 10 questions with p = 0.25, you expect 10 × 0.25 = 2.5 correct. You cannot get 2.5 correct in one test; E(X) is the long-run average number of successes over many repetitions, so it need not be a whole number and should not be rounded to one. Figure 4 marks it as the balance point of the bars.
The variance measures the spread. For ten guesses, Var(X) = 10 × 0.25 × 0.75 = 1.875, and the standard deviation is √1.875 = 1.37 (3 s.f.). For the parcels, E(Y) = 25 × 0.92 = 23 and Var(Y) = 25 × 0.92 × 0.08 = 1.84.
The factor p(1 − p) is largest when p = 0.5 and shrinks towards 0 as p approaches 0 or 1. If success is almost certain, or almost impossible, the results hardly vary. Figure 5 shows three distributions with n = 10. The peak sits at or next to np, the distribution is symmetric when p = 0.5, and it leans away from whichever end p is close to.
The guide does not ask you to prove these formulas, but it does ask you to use them, and one use is a Paper 1 favourite: finding n and p from a given mean and variance. Divide the variance by the mean and n cancels.
Worked example 3: find n and p. X ~ B(n, p) with E(X) = 6 and Var(X) = 4.2.
Check: 20 × 0.3 × 0.7 = 4.2. On Paper 2 the question might go on: find P(X = 6). With X ~ B(20, 0.3), that is 0.192 (3 s.f.).
6Finding n: "the least number of trials so that…"
Sometimes the unknown is n. The most common version asks for the smallest n that makes "at least one success" likely enough, and the complement makes it neat, because "at least one" is "not none".
Worked example 4: scratch cards. Each card in a scratch game wins a prize with probability 0.15, independently. Find the least number of cards a player must buy so that the probability of at least one win is greater than 0.9.
Let X be the number of winning cards out of n, so X ~ B(n, 0.15).
The logarithm step reverses the inequality because ln 0.85 is negative. If that step worries you, the GDC gives a safer route: tabulate y = 1 − 0.85ˣ and read down the table. At n = 14 the value is 0.897, just short; at n = 15 it is 0.913. Figure 6 plots those values. Either way, show both neighbours in your working: 14 fails and 15 works is what proves 15 is the least.
The same GDC table method handles harder versions where the complement is not a single term, such as "the least n so that P(X ≥ 3) > 0.95": tabulate 1 − binomcdf(n, p, 2) against n and find the first value that passes.
Two-stage problems. A binomial count often sits on top of another probability. A machine fills bottles and one bottle is underfilled with a probability you first work out from a normal model (4.9). Then the number of underfilled bottles in a crate of 24 is B(24, that probability). Keep the first probability to full calculator accuracy before you use it as p; rounding it to two figures first can change the third figure of the final answer.
7Where marks are lost
Not defining X. "B(20, 0.3)" floating in the working is not a model. Write "Let X be the number of … out of 20; X ~ B(20, 0.3)" once, at the start.
"Fewer than" treated as "at most". P(X < 4) is P(X ≤ 3). Using the cdf at 4 includes a value the question excluded.
The wrong number after "1 −". P(X ≥ 5) is 1 − P(X ≤ 4), not 1 − P(X ≤ 5). The complement of "5 or more" is "4 or fewer".
Subtracting the wrong cumulative value for a range. P(3 ≤ X ≤ 7) is P(X ≤ 7) − P(X ≤ 2). Subtracting P(X ≤ 3) throws away X = 3.
Using the binomial when a condition fails. Sampling without replacement from a small group, or counting trials until a first success, is not binomial. If a question asks whether the model is suitable, the answer is a reason in context, not "yes".
Rounding the mean. E(X) = 2.5 is the expected number, and it stays 2.5. Rounding it to 2 or 3 loses the accuracy mark.
Confusing variance with standard deviation. np(1 − p) is the variance. If the question asks for the standard deviation, take the square root.
Stopping at one side of the boundary. For "the least n", the answer needs evidence that n works and n − 1 does not, or a correctly handled inequality. A single value with no check is a guess.
8Work it right
- Define the variable in words and write X ~ B(n, p) with numbers.
- If asked for assumptions, name the conditions in the context of the question.
- Turn the words into whole numbers before touching the calculator: ring the values on a number line if unsure.
- Write the probability statement, P(X ≤ 21), as a line of working, then the value.
- For "at least" and "more than", use 1 − P(X ≤ …) with the number one below the lowest value you want.
- Give probabilities to 3 significant figures unless told otherwise, and keep full accuracy in anything you reuse.
- For mean and variance, quote E(X) = np and Var(X) = np(1 − p), substitute, and do not round a mean to a whole number.
- For "the least n", show the value that works and the one below it that does not.
9Try it
Marks in brackets. Q1 and Q2 are Paper 1 style, no calculator. Q3 to Q5 are Paper 2 style, with a GDC.
Q1. The random variable X ~ B(n, p) has mean 12 and variance 3.
(a) Find the value of p and the value of n. 4 marks
(b) Write down an expression for P(X = n), leaving your answer as a power. 2 marks
Q2. A biased coin lands heads with probability 2/3. It is tossed four times, and H is the number of heads.
(a) Find P(H = 3). 2 marks
(b) Find the probability of at least one tail. 2 marks
(c) Write down E(H). 1 mark
Q3. At a help desk, each call is resolved on first contact with probability 0.68. On one morning the desk takes 40 calls. Let R be the number resolved on first contact.
(a) State one assumption needed to model R with a binomial distribution. 1 mark
(b) Find P(R = 30). 2 marks
(c) Find the probability that at least 30 calls are resolved on first contact. 2 marks
(d) Find the probability that more than 25 but fewer than 30 calls are resolved on first contact. 2 marks
(e) Find the mean and the standard deviation of R. 2 marks
Q4. On any day, a bird-watcher at a river has a probability of 0.12 of seeing a kingfisher, independently of other days.
(a) Find the expected number of days, out of 30, on which she sees a kingfisher. 1 mark
(b) Find the least number of days she must visit so that the probability of seeing a kingfisher at least once is greater than 0.95. 4 marks
Q5. An archer hits the centre of the target with probability 0.35 on each arrow, independently. She shoots in sets of 8 arrows, and a set is called good if at least 4 of its arrows hit the centre.
(a) Find the probability that a set is good. 3 marks
(b) She shoots 5 sets. Find the probability that at least 2 of them are good. 3 marks
10In one breath
When the same trial is repeated a fixed number of times n, each trial either succeeds or fails, the probability of success p stays the same and the trials are independent, the number of successes is X ~ B(n, p). Each route with r successes has probability pʳ(1 − p)ⁿ⁻ʳ and there are ⁿCᵣ such routes, which is where the probabilities come from, but in the exam the GDC finds them: pdf for exactly r, cdf for r or fewer. Turn words into whole numbers first: fewer than 4 means at most 3, and at least 5 means 1 − P(X ≤ 4). The mean is np, the expected number of successes, and the variance is np(1 − p), both in the formula booklet; divide one by the other to find p. For the least n, use the complement, 1 − (1 − p)ⁿ, and show the value that works and the one that does not.
Answers
Q1. (a) np = 12 and np(1 − p) = 3. Dividing, 1 − p = 3/12 = 1/4, so p = 3/4. Then n = 12 ÷ (3/4) = 16. M1 for both equations np = 12 and np(1 − p) = 3, M1 for dividing to eliminate n, A1 for p = 3/4, A1 for n = 16.
(b) P(X = 16) = (3/4)¹⁶. M1 for recognising that every trial must succeed, so pⁿ, A1 for (3/4)¹⁶. Follow through from their n and p.
Q2. (a) H ~ B(4, 2/3). P(H = 3) = ⁴C₃ × (2/3)³ × (1/3) = 4 × 8/27 × 1/3 = 32/81. M1 for 4 × (2/3)³ × (1/3) or an equivalent count of the four routes, A1 for 32/81. Omitting the factor 4 gives 8/81 and scores M0.
(b) At least one tail is the complement of four heads: 1 − (2/3)⁴ = 1 − 16/81 = 65/81. M1 for 1 − P(H = 4), A1 for 65/81.
(c) E(H) = 4 × 2/3 = 8/3. A1. Rounding to 3 scores A0.
Q3. (a) For example: the probability that a call is resolved on first contact is the same, 0.68, for every call; or whether one call is resolved does not affect whether another is. R1 for one condition stated in context. "The trials are independent", with no reference to calls, is not enough.
(b) R ~ B(40, 0.68). P(R = 30) = 0.0902 (3 s.f.). M1 for recognising the binomial with n = 40 and p = 0.68, A1 for 0.0902.
(c) P(R ≥ 30) = 1 − P(R ≤ 29) = 0.220 (3 s.f.). M1 for 1 − P(R ≤ 29) or an equivalent, A1 for 0.220. Using 1 − P(R ≤ 30) gives 0.130 and scores M0 A0.
(d) More than 25 and fewer than 30 means 26, 27, 28 or 29: P(R ≤ 29) − P(R ≤ 25) = 0.502 (3 s.f.). M1 for the correct values 26 to 29, shown by the inequality or the subtraction, A1 for 0.502.
(e) E(R) = 40 × 0.68 = 27.2. Var(R) = 40 × 0.68 × 0.32 = 8.704, so the standard deviation is √8.704 = 2.95 (3 s.f.). A1 for 27.2, A1 for 2.95. Giving the variance 8.70 as the standard deviation scores A0.
Q4. (a) 30 × 0.12 = 3.6 days. A1. The expected number need not be a whole number.
(b) Let X ~ B(n, 0.12). P(X ≥ 1) = 1 − 0.88ⁿ > 0.95, so 0.88ⁿ < 0.05 and n > ln 0.05 ÷ ln 0.88 = 23.4… . The least number is 24 days. Check: n = 23 gives 0.947 and n = 24 gives 0.953. M1 for P(X ≥ 1) = 1 − P(X = 0), M1 for the inequality 1 − 0.88ⁿ > 0.95 or 0.88ⁿ < 0.05, A1 for 23.4 or for the two table values either side of 0.95, A1 for 24. An answer of 23 scores A0: rounding 23.4 down gives a probability below 0.95.
Q5. (a) Let A be the number of centre hits in a set, A ~ B(8, 0.35). P(good) = P(A ≥ 4) = 1 − P(A ≤ 3) = 0.294 (3 s.f.; 0.293600… on the GDC). M1 for B(8, 0.35), M1 for 1 − P(A ≤ 3), A1 for 0.294.
(b) Let G be the number of good sets out of 5, G ~ B(5, 0.293600…). P(G ≥ 2) = 1 − P(G ≤ 1) = 0.459 (3 s.f.). M1 for recognising a second binomial with n = 5 and their p from (a), M1 for 1 − P(G ≤ 1), A1 for 0.459. Follow through from their (a).
Educerie · written from the published IB Diploma Programme Mathematics: analysis and approaches guide, first assessment 2021, section 4.8 The binomial distribution. Original text, examples and questions. Diagrams drawn by Educerie. Last reviewed 25 September 2026.