Educerie
Level

Educerie · IB Diploma · Mathematics: analysis and approaches

Topic 4 Statistics and probability · 4.5 Sample spaces and probability

Level
SL and HL. Nothing here is HL only, so every section is examinable for both.
Themes (key concepts)
quantity, approximation, validity. A probability is a quantity between 0 and 1; a relative frequency from an experiment is an approximation to it; and the formula n(A) ÷ n(U) is only valid when the outcomes you are counting are equally likely.
The question this unit answers
how do you put a number on how likely something is, either by counting what could happen or by watching what did happen, and what does that number let you predict?
Where it is examined
Paper 1 and Paper 2, usually as the opening parts of a probability question: list or draw a sample space (1 to 2 marks), find a probability by counting (2 marks), use P(A′) = 1 − P(A) (1 to 2 marks), estimate a probability from experimental data (1 to 2 marks), and find an expected number of occurrences (2 marks). Paper 1 wants exact fractions such as 11/36; on Paper 2 a decimal to 3 significant figures is fine. Everything in 4.6, 4.7 and 4.8 is built on the words and rules on this page.

What you must be able to do

You must be able toLevelWhat it looks like in the exam
Use the words trial, outcome, equally likely, sample space U and event correctlySL, HL"Write down the sample space" or "List the outcomes in event A" (1 to 2 marks)
Represent a sample space as a list, a table or a gridSL, HL"Complete the sample space diagram" (2 marks), Paper 1
Find a probability by counting: P(A) = n(A) ÷ n(U)SL, HL"Find the probability that the total is at least 10" (2 marks)
Use complementary events: P(A′) = 1 − P(A)SL, HL"Find the probability that at least one of the dice shows a six" (2 to 3 marks)
Estimate a probability from data as a relative frequency, and compare it with a theoretical probabilitySL, HL"Estimate the probability that…" then "Comment on whether the spinner is fair" (1 to 3 marks), Paper 2
Find an expected number of occurrences, n × P(A)SL, HL"Find the expected number of days on which the bus is late" (2 marks)

Before you start

You need fractions, decimals and percentages, and the ability to list things systematically so that nothing is missed or counted twice. Set notation from earlier years (curly brackets, n(A) for the number of elements) is used throughout. The formula booklet gives P(A) = n(A) ÷ n(U) and P(A) + P(A′) = 1. The expected-number rule is not in the booklet; it is one line, so learn it.


1The idea in one paragraph

Probability measures how likely an event is on a scale from 0 (impossible) to 1 (certain). There are two ways to get the number. If you can list every possible result and they are all equally likely, count: the probability is the number of results you want divided by the number there are. If you cannot, experiment: repeat the thing many times and use the fraction of times the event happened, its relative frequency, as your estimate. The more trials, the better the estimate. The probability that something does not happen is 1 minus the probability that it does. And once you have a probability, multiplying it by the number of trials tells you how many times to expect the event: not a promise, but a long-run average.

2The vocabulary: trial, outcome, sample space, event

Every probability question uses the same six words, and the marks depend on using them precisely. Figure 1 shows four of them for one roll of a die.

Figure 1 · From a trial to an event Figure 1 · From a trial to an event trial roll one fair die sample space U, n(U) = 6 event A: an even score 1 3 5 2 4 6 each circle is one outcome A = {2, 4, 6}, n(A) = 3, so P(A) = 3/6 = 1/2 The sample space U lists every outcome once. An event is any collection of those outcomes.
Figure 1 · From a trial to an event
  • A trial is one performance of an experiment whose result is uncertain: one roll of a die, one toss of a coin, one patient given a test, one day's weather.
  • An outcome is one possible result of a trial. Rolling a die has six outcomes: 1, 2, 3, 4, 5, 6.
  • The sample space, written U, is the set of all possible outcomes, each listed once. For the die, U = {1, 2, 3, 4, 5, 6}, so n(U) = 6.
  • An event is any collection of outcomes, a subset of U. "An even score" is the event A = {2, 4, 6}. An event can be a single outcome ("a 5"), or even every outcome ("a score below 7", which is certain).
  • Outcomes are equally likely when each has the same chance of happening. A fair die makes the six scores equally likely. A loaded die does not.
  • The relative frequency of an event is the number of times it happened divided by the number of trials. Section 5 is about it.

Say the event in words and in set notation. "The score is prime" is B = {2, 3, 5}. Writing it out takes five seconds and removes arguments with yourself about whether 1 counts (it does not; 1 is not prime).

Representing a sample space. The guide expects you to use whatever form fits the trial.

  • A list works for one simple trial: U = {H, T} for a coin, or the 20 numbers on 20 numbered counters.
  • A table or grid, the sample space diagram, works for two trials together. Put the outcomes of the first along the top, the second down the side, and fill each cell with the combined result. Figure 2 does this for two dice, with the total in each cell. There are 6 × 6 = 36 cells, and because each die is fair and the dice do not affect each other, the 36 cells are equally likely.
Figure 2 · The sample space for two dice Figure 2 · The sample space for two dice (a) The event “total is 7” 6 cells, so P(total 7) = 6/36 = 1/6 1 1 2 2 3 3 4 4 5 5 6 6 2 3 4 5 6 7 3 4 5 6 7 8 4 5 6 7 8 9 5 6 7 8 9 10 6 7 8 9 10 11 7 8 9 10 11 12 first die second die (b) “At least one six” and its complement 11 amber cells; the other 25 have no six 1 1 2 2 3 3 4 4 5 5 6 6 2 3 4 5 6 7 3 4 5 6 7 8 4 5 6 7 8 9 5 6 7 8 9 10 6 7 8 9 10 11 7 8 9 10 11 12 first die second die 36 equally likely outcomes. Each cell shows the total of the two scores.
Figure 2 · The sample space for two dice

A two-way list such as {HH, HT, TH, TT} for two coins is a sample space too. Notice HT and TH are different outcomes: the first coin shows heads and the second tails, or the other way round. Treating "one head and one tail" as a single outcome is the error that section 3 warns against.

3Probability by counting

When all the outcomes in U are equally likely, the probability of an event A is the fraction of those outcomes that belong to A.

P(A) = n(A) ÷ n(U), valid only when every outcome in U is equally likely. Every probability satisfies 0 ≤ P(A) ≤ 1.

A bag holds 20 counters numbered 1 to 20, and one is taken at random. "At random" means every counter is equally likely to be chosen, so the formula applies.

U = {1, 2, …, 20}, n(U) = 20
A = multiple of 3 = {3, 6, 9, 12, 15, 18}, n(A) = 6
P(A) = 6/20 = 3/10
B = prime = {2, 3, 5, 7, 11, 13, 17, 19}, n(B) = 8
P(B) = 8/20 = 2/5

With two dice, read the answer off Figure 2. In panel (a) the total is 7 in six cells, so P(total is 7) = 6/36 = 1/6. For "total is at least 10", count the cells showing 10, 11 or 12: 3 + 2 + 1 = 6, so the probability is again 6/36 = 1/6. Only one cell shows 12, so P(total is 12) = 1/36.

The equally-likely trap. The possible totals of two dice are 2, 3, …, 12, which is eleven values. It is tempting to say P(total is 7) = 1/11. That is wrong, because the eleven totals are not equally likely: a total of 7 can happen six ways and a total of 2 only one way. The formula counts equally likely outcomes, so count cells in the grid, not values in a list of totals. Whenever you write n(A) ÷ n(U), ask whether the things you are counting really do have the same chance.

Figure 3 places these probabilities on the probability scale. An impossible event (a total of 13) has probability 0; a certain event (a total of at least 2) has probability 1. An answer outside 0 to 1 is a signal to find your mistake, not an answer.

Figure 3 · Every probability lies between 0 and 1 Figure 3 · Every probability lies between 0 and 1 0 0.25 0.5 0.75 1 impossible certain evens total 12: 1/36 total 7: 1/6 even score: 1/2 no six in two rolls: 25/36 at least one six: 11/36 Probabilities from this page, placed on the scale. Nothing can go below 0 or above 1.
Figure 3 · Every probability lies between 0 and 1

On Paper 1 give probabilities as exact fractions, simplified if you can: 6/36 is correct, 1/6 is neater, and either usually earns the mark. On Paper 2 decimals to 3 significant figures are accepted: 0.167.

4Complementary events

The complement of A, written A′ (read "not A"), is the event that A does not happen: every outcome in U that is not in A. A and A′ cannot both happen, and one of them must happen, so their probabilities add to 1. Figure 4 draws this: A is a region inside the rectangle U, and A′ is everything else in the rectangle.

Figure 4 · An event and its complement Figure 4 · An event and its complement U A at least one six 11 outcomes A′ no six 25 outcomes P(A) = 1 − P(A′) = 1 − 25/36 = 11/36 A and A′ do not overlap and together fill U, so P(A) + P(A′) = 1.
Figure 4 · An event and its complement

P(A) + P(A′) = 1, so P(A′) = 1 − P(A).

This is more than a tidy fact. It is often the fastest route to an answer, because the complement of an awkward event is frequently simple. The classic case is "at least one".

Two dice are rolled. Find the probability that at least one of them shows a six.

"At least one six" means one six or two sixes, which is several cases. Its complement, "no six at all", is one clean case: each die shows 1 to 5, so there are 5 × 5 = 25 cells with no six.

P(no six) = 25/36
P(at least one six) = 1 − 25/36 = 11/36

Check against panel (b) of Figure 2: 6 cells in the last column plus 6 in the last row, minus the (6, 6) cell counted twice, gives 11. The same answer both ways. The counting route works here because the grid is small; the complement route still works when the grid is far too big to draw, which is why it is worth learning now.

The complement is also how you check work. If a question asks for P(A) and P(A′) separately, they must add to exactly 1.

5Relative frequency: probability from experiment

Some probabilities cannot be found by counting, because the outcomes are not equally likely and there is no symmetry to lean on. What is the probability that a plastic bottle cap lands top-down when tossed? The two outcomes, top-down and top-up, are not equally likely; the shape decides. You have to toss it and see.

The relative frequency of an event is

relative frequency = number of times the event happened ÷ number of trials

and it is used as an estimate of the probability. This estimate is called an experimental probability. A probability found by counting equally likely outcomes, as in section 3, is a theoretical probability.

A student tosses a bottle cap 250 times and it lands top-down 165 times.

relative frequency = 165/250 = 0.66
estimate: P(top-down) ≈ 0.66
P(top-up) ≈ 1 − 0.66 = 0.34

How good is the estimate? Figure 5 answers this with a simulation: a computer rolled a fair die 1000 times and, after each roll, recorded the fraction of rolls so far that were sixes.

Figure 5 · Relative frequency settles as the trials pile up Figure 5 · Relative frequency settles as the trials pile up 0.05 0.1 0.15 0.2 0.25 0.3 0 200 400 600 800 1000 number of rolls, n relative frequency of a six theoretical P(six) = 1/6 ≈ 0.167 after 1000 rolls: 157/1000 = 0.157 in the first few dozen rolls it jumps about 1000 simulated rolls of a fair die. After each roll: sixes so far ÷ rolls so far.
Figure 5 · Relative frequency settles as the trials pile up

Three things are worth reading off it. After 10 rolls there had been no sixes at all, so the relative frequency was 0, wildly wrong. After 100 rolls it was 0.14. After 1000 rolls it was 157/1000 = 0.157, close to the true 1/6 ≈ 0.167 but still not equal to it. So:

  • A relative frequency from a few trials is unreliable. Its value depends heavily on luck.
  • As the number of trials grows, the relative frequency tends to settle close to the true probability. More trials give a better estimate.
  • It never becomes a guarantee. Even after many trials the relative frequency will usually differ a little from the probability.

This is how the two kinds of probability are linked. A theoretical probability predicts what the relative frequency will approach in the long run; a relative frequency from many trials estimates the probability when no theory is available. Insurers price policies this way, using the recorded frequency of past claims, and it is why a simulation, run thousands of times on a computer, can estimate a probability that nobody knows how to calculate.

Is it fair? Comparing the two tells you whether a theoretical model fits. If a spinner with five equal sectors is fair, each number has probability 1/5 = 0.2. Spun 500 times, a relative frequency of 0.236 for one number is higher than 0.2, but chance alone often produces a gap like that, so it is not proof of bias. A relative frequency of 0.4 after 500 spins would be strong evidence. At SL you are asked to comment, not to test formally: say what the model predicts, what the data show, and whether the gap is large compared with the number of trials.

6Expected number of occurrences

If an event has probability p and you carry out n independent trials, you expect it to happen about

expected number of occurrences = n × p

times. This follows from section 5 read backwards: in the long run the relative frequency is close to p, so the number of occurrences is close to n × p.

Roll two dice 180 times. P(total is 7) = 1/6, so you expect a total of 7 about 180 × 1/6 = 30 times.

Take the bottle cap. Using the estimate from section 5, in 600 more tosses you expect about 600 × 0.66 = 396 top-down landings. The answer inherits the uncertainty of the estimate: it is only as good as the 250 tosses behind it.

The answer need not be a whole number. A school bus is late on any given day with probability 0.15. Over a 22-day month, the expected number of late days is 22 × 0.15 = 3.3. You cannot have 3.3 late days in one month, and nobody claims you will. It means that over many months the average number of late days per month is 3.3. Do not round an expected number to a whole number unless the question tells you to; 3.3 is the answer.

An application. An insurer has 3500 policies of one kind, and past records suggest that each policy leads to a claim in a year with probability 0.02 (invented figures). The expected number of claims in a year is 3500 × 0.02 = 70. If an average claim costs €4000, the insurer should expect to pay about 70 × €4000 = €280,000 and must charge enough, spread over 3500 customers, to cover it: at least €80 each, before costs. Expected numbers are how insurance, staffing and stock levels get planned.

In 4.7 this idea becomes the expected value of a random variable, and in 4.8 it becomes the mean np of the binomial distribution. Same formula, same meaning: the long-run average.

7Where marks are lost

Using n(A) ÷ n(U) when the outcomes are not equally likely. The totals of two dice, the number of heads in three tosses, the colours on an unequal spinner: none of these lists are equally likely. Count equally likely outcomes, from a grid or a full list, instead.

Missing outcomes, or counting one twice. With two coins, HT and TH are two outcomes, not one. With two dice, (2, 5) and (5, 2) are two cells. Draw the grid rather than trust memory.

Forgetting the double-counted cell. "At least one six" is 11 cells, not 12: the (6, 6) cell sits in both the six-row and the six-column. The complement route avoids the trap entirely.

Treating a relative frequency as exact. 165/250 = 0.66 is an estimate of the probability, not the probability. Say "estimate" when the question does, and remember that more trials give a better estimate.

Rounding an expected number. An expected number of 3.3 late days is correct as 3.3. Writing 3 loses the accuracy mark.

Writing a probability greater than 1 or negative. It is always a mistake. Usually the numerator and denominator have been swapped or a complement has been taken the wrong way round.

Decimals on Paper 1 where fractions were intended. 11/36 is exact; 0.31 is not. Paper 1 answers should be exact unless the question asks otherwise.

8Work it right

  1. Define the sample space first: list it, or draw the grid, and write n(U).
  2. Name the event in words and as a set, then write n(A).
  3. Before dividing, check that the outcomes you counted are equally likely.
  4. For "at least one", "not" or "fewer than", ask whether the complement is quicker, and write P(A) = 1 − P(A′) as a line of working.
  5. For data, write relative frequency = frequency ÷ number of trials, and call the result an estimate.
  6. For an expected number, write n × p with both numbers, and leave the answer unrounded unless told otherwise.
  7. Check: every probability lies between 0 and 1, and an event and its complement add to 1.

9Try it

Marks in brackets. Q1, Q2 and Q4 are Paper 1 style, no calculator. Q3 and Q5 are Paper 2 style, with a GDC.

Q1. A card is chosen at random from 25 cards numbered 1 to 25.

(a) Find the probability that the number is a square number. 2 marks

(b) Find the probability that the number is not a multiple of 3. 3 marks

Q2. Two fair four-sided dice, each numbered 1, 2, 3, 4, are rolled. The score is the product of the two numbers.

(a) Draw a sample space diagram showing all the possible scores. 2 marks

(b) Find the probability that the score is odd. 2 marks

(c) Find the probability that the score is at least 6. 2 marks

(d) The two dice are rolled 80 times. Find the expected number of times the score is odd. 2 marks

Q3. A factory tests a random sample of 400 light bulbs and finds that 14 are faulty.

(a) Estimate the probability that a bulb from this factory is faulty. 1 mark

(b) The factory makes 12,000 bulbs a week. Find the expected number of faulty bulbs in a week. 2 marks

(c) State one way in which the estimate in (a) could be made more reliable. 1 mark

Q4. For an event A, P(A) = 3k and P(A′) = k + 0.2.

(a) Find the value of k. 3 marks

(b) The trial is carried out 45 times. Find the expected number of times A occurs. 2 marks

Q5. A spinner has five equal sectors numbered 1 to 5. It is spun 500 times, with these results.

Number12345
Frequency921089711885

(a) Find the relative frequency of a 4. 1 mark

(b) Use the data to estimate the probability that the spinner lands on an odd number. 2 marks

(c) Write down the number of 4s you would expect in 500 spins if the spinner were fair. 1 mark

(d) Comment on whether these results show that the spinner is biased. 2 marks

10In one breath

A trial is one go at an experiment, an outcome is one possible result, the sample space U lists every outcome once, and an event is any set of outcomes. When the outcomes are equally likely, P(A) = n(A) ÷ n(U), so draw the list or the grid and count, and never count totals that are not equally likely. Every probability lies between 0 and 1. The complement A′ is "not A", and P(A′) = 1 − P(A), which turns "at least one" into one minus "none". When nothing is equally likely, experiment: the relative frequency, frequency ÷ trials, estimates the probability, and it settles closer to the truth as the trials grow. Multiply a probability by the number of trials to get the expected number of occurrences, a long-run average that need not be a whole number.


Answers

Q1. (a) The square numbers are 1, 4, 9, 16, 25, so n(A) = 5 and P = 5/25 = 1/5. A1 for listing or counting 5 squares, A1 for 1/5 or 5/25. Missing 1 or 25 gives 3/25, which scores A0 A0. (b) The multiples of 3 are 3, 6, 9, 12, 15, 18, 21, 24, which is 8 numbers. P(multiple of 3) = 8/25, so P(not a multiple of 3) = 1 − 8/25 = 17/25. A1 for 8 multiples, M1 for using the complement (or for counting the 17 directly), A1 for 17/25.

Q2. (a) A 4 × 4 grid with the first die across and the second down. Rows: 1, 2, 3, 4 / 2, 4, 6, 8 / 3, 6, 9, 12 / 4, 8, 12, 16. A2 for all 16 products correct, A1 for at least 12 correct. (b) The odd products are 1, 3, 3, 9, which is 4 cells of 16, so P = 4/16 = 1/4. M1 for counting from their grid over 16, A1 for 1/4. (c) Products of at least 6: 6, 8, 6, 9, 12, 8, 12, 16, which is 8 cells, so P = 8/16 = 1/2. M1 for counting from their grid, A1 for 1/2. Counting the distinct values 6, 8, 9, 12, 16 as 5 out of the 9 distinct products scores M0: the products are not equally likely. (d) 80 × 1/4 = 20. M1 for 80 × their answer to (b), A1 for 20. Follow-through from (b).

Q3. (a) 14/400 = 0.035. A1. (b) 12,000 × 0.035 = 420 faulty bulbs. M1 for 12,000 × their (a), A1 for 420. (c) Test a larger random sample of bulbs. R1 for a larger number of trials; "test again" alone scores 0.

Q4. (a) P(A) + P(A′) = 1, so 3k + k + 0.2 = 1, giving 4k = 0.8 and k = 0.2. M1 for setting the sum equal to 1, A1 for 4k = 0.8, A1 for k = 0.2. (b) P(A) = 3 × 0.2 = 0.6, so the expected number is 45 × 0.6 = 27. M1 for 45 × their P(A), A1 for 27. Using 45 × k = 9 scores M0.

Q5. (a) 118/500 = 0.236. A1. (b) (92 + 97 + 85)/500 = 274/500 = 0.548. M1 for adding the frequencies of 1, 3 and 5, A1 for 0.548. (c) 500 × 1/5 = 100. A1. (d) If the spinner were fair each number would be expected about 100 times. The frequencies range from 85 to 118, and 118 is 18 more than expected, but with 500 spins gaps of this size are quite possible by chance, so the results do not show clearly that the spinner is biased. More spins would be needed to be confident. R1 for comparing the observed frequencies (or relative frequencies) with 100 (or 0.2), R1 for a conclusion that recognises chance variation. "Yes, because the numbers are not all 100" scores R1 R0 at most.


Educerie · written from the published IB Diploma Programme Mathematics: analysis and approaches guide, first assessment 2021, section 4.5 Sample spaces and probability. Original text, examples and questions. Diagrams drawn by Educerie. Last reviewed 25 September 2026.

Mocks: in the future, hold tight!