Educerie · IB Diploma · Mathematics: analysis and approaches
Topic 4 Statistics and probability · 4.5 Sample spaces and probability
What you must be able to do
| You must be able to | Level | What it looks like in the exam |
|---|---|---|
| Use the words trial, outcome, equally likely, sample space U and event correctly | SL, HL | "Write down the sample space" or "List the outcomes in event A" (1 to 2 marks) |
| Represent a sample space as a list, a table or a grid | SL, HL | "Complete the sample space diagram" (2 marks), Paper 1 |
| Find a probability by counting: P(A) = n(A) ÷ n(U) | SL, HL | "Find the probability that the total is at least 10" (2 marks) |
| Use complementary events: P(A′) = 1 − P(A) | SL, HL | "Find the probability that at least one of the dice shows a six" (2 to 3 marks) |
| Estimate a probability from data as a relative frequency, and compare it with a theoretical probability | SL, HL | "Estimate the probability that…" then "Comment on whether the spinner is fair" (1 to 3 marks), Paper 2 |
| Find an expected number of occurrences, n × P(A) | SL, HL | "Find the expected number of days on which the bus is late" (2 marks) |
Before you start
You need fractions, decimals and percentages, and the ability to list things systematically so that nothing is missed or counted twice. Set notation from earlier years (curly brackets, n(A) for the number of elements) is used throughout. The formula booklet gives P(A) = n(A) ÷ n(U) and P(A) + P(A′) = 1. The expected-number rule is not in the booklet; it is one line, so learn it.
1The idea in one paragraph
Probability measures how likely an event is on a scale from 0 (impossible) to 1 (certain). There are two ways to get the number. If you can list every possible result and they are all equally likely, count: the probability is the number of results you want divided by the number there are. If you cannot, experiment: repeat the thing many times and use the fraction of times the event happened, its relative frequency, as your estimate. The more trials, the better the estimate. The probability that something does not happen is 1 minus the probability that it does. And once you have a probability, multiplying it by the number of trials tells you how many times to expect the event: not a promise, but a long-run average.
2The vocabulary: trial, outcome, sample space, event
Every probability question uses the same six words, and the marks depend on using them precisely. Figure 1 shows four of them for one roll of a die.
- A trial is one performance of an experiment whose result is uncertain: one roll of a die, one toss of a coin, one patient given a test, one day's weather.
- An outcome is one possible result of a trial. Rolling a die has six outcomes: 1, 2, 3, 4, 5, 6.
- The sample space, written U, is the set of all possible outcomes, each listed once. For the die, U = {1, 2, 3, 4, 5, 6}, so n(U) = 6.
- An event is any collection of outcomes, a subset of U. "An even score" is the event A = {2, 4, 6}. An event can be a single outcome ("a 5"), or even every outcome ("a score below 7", which is certain).
- Outcomes are equally likely when each has the same chance of happening. A fair die makes the six scores equally likely. A loaded die does not.
- The relative frequency of an event is the number of times it happened divided by the number of trials. Section 5 is about it.
Say the event in words and in set notation. "The score is prime" is B = {2, 3, 5}. Writing it out takes five seconds and removes arguments with yourself about whether 1 counts (it does not; 1 is not prime).
Representing a sample space. The guide expects you to use whatever form fits the trial.
- A list works for one simple trial: U = {H, T} for a coin, or the 20 numbers on 20 numbered counters.
- A table or grid, the sample space diagram, works for two trials together. Put the outcomes of the first along the top, the second down the side, and fill each cell with the combined result. Figure 2 does this for two dice, with the total in each cell. There are 6 × 6 = 36 cells, and because each die is fair and the dice do not affect each other, the 36 cells are equally likely.
A two-way list such as {HH, HT, TH, TT} for two coins is a sample space too. Notice HT and TH are different outcomes: the first coin shows heads and the second tails, or the other way round. Treating "one head and one tail" as a single outcome is the error that section 3 warns against.
3Probability by counting
When all the outcomes in U are equally likely, the probability of an event A is the fraction of those outcomes that belong to A.
P(A) = n(A) ÷ n(U), valid only when every outcome in U is equally likely. Every probability satisfies 0 ≤ P(A) ≤ 1.
A bag holds 20 counters numbered 1 to 20, and one is taken at random. "At random" means every counter is equally likely to be chosen, so the formula applies.
With two dice, read the answer off Figure 2. In panel (a) the total is 7 in six cells, so P(total is 7) = 6/36 = 1/6. For "total is at least 10", count the cells showing 10, 11 or 12: 3 + 2 + 1 = 6, so the probability is again 6/36 = 1/6. Only one cell shows 12, so P(total is 12) = 1/36.
The equally-likely trap. The possible totals of two dice are 2, 3, …, 12, which is eleven values. It is tempting to say P(total is 7) = 1/11. That is wrong, because the eleven totals are not equally likely: a total of 7 can happen six ways and a total of 2 only one way. The formula counts equally likely outcomes, so count cells in the grid, not values in a list of totals. Whenever you write n(A) ÷ n(U), ask whether the things you are counting really do have the same chance.
Figure 3 places these probabilities on the probability scale. An impossible event (a total of 13) has probability 0; a certain event (a total of at least 2) has probability 1. An answer outside 0 to 1 is a signal to find your mistake, not an answer.
On Paper 1 give probabilities as exact fractions, simplified if you can: 6/36 is correct, 1/6 is neater, and either usually earns the mark. On Paper 2 decimals to 3 significant figures are accepted: 0.167.
4Complementary events
The complement of A, written A′ (read "not A"), is the event that A does not happen: every outcome in U that is not in A. A and A′ cannot both happen, and one of them must happen, so their probabilities add to 1. Figure 4 draws this: A is a region inside the rectangle U, and A′ is everything else in the rectangle.
P(A) + P(A′) = 1, so P(A′) = 1 − P(A).
This is more than a tidy fact. It is often the fastest route to an answer, because the complement of an awkward event is frequently simple. The classic case is "at least one".
Two dice are rolled. Find the probability that at least one of them shows a six.
"At least one six" means one six or two sixes, which is several cases. Its complement, "no six at all", is one clean case: each die shows 1 to 5, so there are 5 × 5 = 25 cells with no six.
Check against panel (b) of Figure 2: 6 cells in the last column plus 6 in the last row, minus the (6, 6) cell counted twice, gives 11. The same answer both ways. The counting route works here because the grid is small; the complement route still works when the grid is far too big to draw, which is why it is worth learning now.
The complement is also how you check work. If a question asks for P(A) and P(A′) separately, they must add to exactly 1.
5Relative frequency: probability from experiment
Some probabilities cannot be found by counting, because the outcomes are not equally likely and there is no symmetry to lean on. What is the probability that a plastic bottle cap lands top-down when tossed? The two outcomes, top-down and top-up, are not equally likely; the shape decides. You have to toss it and see.
The relative frequency of an event is
relative frequency = number of times the event happened ÷ number of trials
and it is used as an estimate of the probability. This estimate is called an experimental probability. A probability found by counting equally likely outcomes, as in section 3, is a theoretical probability.
A student tosses a bottle cap 250 times and it lands top-down 165 times.
How good is the estimate? Figure 5 answers this with a simulation: a computer rolled a fair die 1000 times and, after each roll, recorded the fraction of rolls so far that were sixes.
Three things are worth reading off it. After 10 rolls there had been no sixes at all, so the relative frequency was 0, wildly wrong. After 100 rolls it was 0.14. After 1000 rolls it was 157/1000 = 0.157, close to the true 1/6 ≈ 0.167 but still not equal to it. So:
- A relative frequency from a few trials is unreliable. Its value depends heavily on luck.
- As the number of trials grows, the relative frequency tends to settle close to the true probability. More trials give a better estimate.
- It never becomes a guarantee. Even after many trials the relative frequency will usually differ a little from the probability.
This is how the two kinds of probability are linked. A theoretical probability predicts what the relative frequency will approach in the long run; a relative frequency from many trials estimates the probability when no theory is available. Insurers price policies this way, using the recorded frequency of past claims, and it is why a simulation, run thousands of times on a computer, can estimate a probability that nobody knows how to calculate.
Is it fair? Comparing the two tells you whether a theoretical model fits. If a spinner with five equal sectors is fair, each number has probability 1/5 = 0.2. Spun 500 times, a relative frequency of 0.236 for one number is higher than 0.2, but chance alone often produces a gap like that, so it is not proof of bias. A relative frequency of 0.4 after 500 spins would be strong evidence. At SL you are asked to comment, not to test formally: say what the model predicts, what the data show, and whether the gap is large compared with the number of trials.
6Expected number of occurrences
If an event has probability p and you carry out n independent trials, you expect it to happen about
expected number of occurrences = n × p
times. This follows from section 5 read backwards: in the long run the relative frequency is close to p, so the number of occurrences is close to n × p.
Roll two dice 180 times. P(total is 7) = 1/6, so you expect a total of 7 about 180 × 1/6 = 30 times.
Take the bottle cap. Using the estimate from section 5, in 600 more tosses you expect about 600 × 0.66 = 396 top-down landings. The answer inherits the uncertainty of the estimate: it is only as good as the 250 tosses behind it.
The answer need not be a whole number. A school bus is late on any given day with probability 0.15. Over a 22-day month, the expected number of late days is 22 × 0.15 = 3.3. You cannot have 3.3 late days in one month, and nobody claims you will. It means that over many months the average number of late days per month is 3.3. Do not round an expected number to a whole number unless the question tells you to; 3.3 is the answer.
An application. An insurer has 3500 policies of one kind, and past records suggest that each policy leads to a claim in a year with probability 0.02 (invented figures). The expected number of claims in a year is 3500 × 0.02 = 70. If an average claim costs €4000, the insurer should expect to pay about 70 × €4000 = €280,000 and must charge enough, spread over 3500 customers, to cover it: at least €80 each, before costs. Expected numbers are how insurance, staffing and stock levels get planned.
In 4.7 this idea becomes the expected value of a random variable, and in 4.8 it becomes the mean np of the binomial distribution. Same formula, same meaning: the long-run average.
7Where marks are lost
Using n(A) ÷ n(U) when the outcomes are not equally likely. The totals of two dice, the number of heads in three tosses, the colours on an unequal spinner: none of these lists are equally likely. Count equally likely outcomes, from a grid or a full list, instead.
Missing outcomes, or counting one twice. With two coins, HT and TH are two outcomes, not one. With two dice, (2, 5) and (5, 2) are two cells. Draw the grid rather than trust memory.
Forgetting the double-counted cell. "At least one six" is 11 cells, not 12: the (6, 6) cell sits in both the six-row and the six-column. The complement route avoids the trap entirely.
Treating a relative frequency as exact. 165/250 = 0.66 is an estimate of the probability, not the probability. Say "estimate" when the question does, and remember that more trials give a better estimate.
Rounding an expected number. An expected number of 3.3 late days is correct as 3.3. Writing 3 loses the accuracy mark.
Writing a probability greater than 1 or negative. It is always a mistake. Usually the numerator and denominator have been swapped or a complement has been taken the wrong way round.
Decimals on Paper 1 where fractions were intended. 11/36 is exact; 0.31 is not. Paper 1 answers should be exact unless the question asks otherwise.
8Work it right
- Define the sample space first: list it, or draw the grid, and write n(U).
- Name the event in words and as a set, then write n(A).
- Before dividing, check that the outcomes you counted are equally likely.
- For "at least one", "not" or "fewer than", ask whether the complement is quicker, and write P(A) = 1 − P(A′) as a line of working.
- For data, write relative frequency = frequency ÷ number of trials, and call the result an estimate.
- For an expected number, write n × p with both numbers, and leave the answer unrounded unless told otherwise.
- Check: every probability lies between 0 and 1, and an event and its complement add to 1.
9Try it
Marks in brackets. Q1, Q2 and Q4 are Paper 1 style, no calculator. Q3 and Q5 are Paper 2 style, with a GDC.
Q1. A card is chosen at random from 25 cards numbered 1 to 25.
(a) Find the probability that the number is a square number. 2 marks
(b) Find the probability that the number is not a multiple of 3. 3 marks
Q2. Two fair four-sided dice, each numbered 1, 2, 3, 4, are rolled. The score is the product of the two numbers.
(a) Draw a sample space diagram showing all the possible scores. 2 marks
(b) Find the probability that the score is odd. 2 marks
(c) Find the probability that the score is at least 6. 2 marks
(d) The two dice are rolled 80 times. Find the expected number of times the score is odd. 2 marks
Q3. A factory tests a random sample of 400 light bulbs and finds that 14 are faulty.
(a) Estimate the probability that a bulb from this factory is faulty. 1 mark
(b) The factory makes 12,000 bulbs a week. Find the expected number of faulty bulbs in a week. 2 marks
(c) State one way in which the estimate in (a) could be made more reliable. 1 mark
Q4. For an event A, P(A) = 3k and P(A′) = k + 0.2.
(a) Find the value of k. 3 marks
(b) The trial is carried out 45 times. Find the expected number of times A occurs. 2 marks
Q5. A spinner has five equal sectors numbered 1 to 5. It is spun 500 times, with these results.
| Number | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|
| Frequency | 92 | 108 | 97 | 118 | 85 |
(a) Find the relative frequency of a 4. 1 mark
(b) Use the data to estimate the probability that the spinner lands on an odd number. 2 marks
(c) Write down the number of 4s you would expect in 500 spins if the spinner were fair. 1 mark
(d) Comment on whether these results show that the spinner is biased. 2 marks
10In one breath
A trial is one go at an experiment, an outcome is one possible result, the sample space U lists every outcome once, and an event is any set of outcomes. When the outcomes are equally likely, P(A) = n(A) ÷ n(U), so draw the list or the grid and count, and never count totals that are not equally likely. Every probability lies between 0 and 1. The complement A′ is "not A", and P(A′) = 1 − P(A), which turns "at least one" into one minus "none". When nothing is equally likely, experiment: the relative frequency, frequency ÷ trials, estimates the probability, and it settles closer to the truth as the trials grow. Multiply a probability by the number of trials to get the expected number of occurrences, a long-run average that need not be a whole number.
Answers
Q1. (a) The square numbers are 1, 4, 9, 16, 25, so n(A) = 5 and P = 5/25 = 1/5. A1 for listing or counting 5 squares, A1 for 1/5 or 5/25. Missing 1 or 25 gives 3/25, which scores A0 A0. (b) The multiples of 3 are 3, 6, 9, 12, 15, 18, 21, 24, which is 8 numbers. P(multiple of 3) = 8/25, so P(not a multiple of 3) = 1 − 8/25 = 17/25. A1 for 8 multiples, M1 for using the complement (or for counting the 17 directly), A1 for 17/25.
Q2. (a) A 4 × 4 grid with the first die across and the second down. Rows: 1, 2, 3, 4 / 2, 4, 6, 8 / 3, 6, 9, 12 / 4, 8, 12, 16. A2 for all 16 products correct, A1 for at least 12 correct. (b) The odd products are 1, 3, 3, 9, which is 4 cells of 16, so P = 4/16 = 1/4. M1 for counting from their grid over 16, A1 for 1/4. (c) Products of at least 6: 6, 8, 6, 9, 12, 8, 12, 16, which is 8 cells, so P = 8/16 = 1/2. M1 for counting from their grid, A1 for 1/2. Counting the distinct values 6, 8, 9, 12, 16 as 5 out of the 9 distinct products scores M0: the products are not equally likely. (d) 80 × 1/4 = 20. M1 for 80 × their answer to (b), A1 for 20. Follow-through from (b).
Q3. (a) 14/400 = 0.035. A1. (b) 12,000 × 0.035 = 420 faulty bulbs. M1 for 12,000 × their (a), A1 for 420. (c) Test a larger random sample of bulbs. R1 for a larger number of trials; "test again" alone scores 0.
Q4. (a) P(A) + P(A′) = 1, so 3k + k + 0.2 = 1, giving 4k = 0.8 and k = 0.2. M1 for setting the sum equal to 1, A1 for 4k = 0.8, A1 for k = 0.2. (b) P(A) = 3 × 0.2 = 0.6, so the expected number is 45 × 0.6 = 27. M1 for 45 × their P(A), A1 for 27. Using 45 × k = 9 scores M0.
Q5. (a) 118/500 = 0.236. A1. (b) (92 + 97 + 85)/500 = 274/500 = 0.548. M1 for adding the frequencies of 1, 3 and 5, A1 for 0.548. (c) 500 × 1/5 = 100. A1. (d) If the spinner were fair each number would be expected about 100 times. The frequencies range from 85 to 118, and 118 is 18 more than expected, but with 500 spins gaps of this size are quite possible by chance, so the results do not show clearly that the spinner is biased. More spins would be needed to be confident. R1 for comparing the observed frequencies (or relative frequencies) with 100 (or 0.2), R1 for a conclusion that recognises chance variation. "Yes, because the numbers are not all 100" scores R1 R0 at most.
Educerie · written from the published IB Diploma Programme Mathematics: analysis and approaches guide, first assessment 2021, section 4.5 Sample spaces and probability. Original text, examples and questions. Diagrams drawn by Educerie. Last reviewed 25 September 2026.