Educerie
Level

This whole subtopic is higher level. Nothing in it is on an SL paper.

Educerie · IB Diploma · Mathematics: analysis and approaches

Topic 4 Statistics and probability · 4.13 Bayes' theorem

Level
HL only, the whole subtopic. If you are SL, none of this is on your papers.
Themes (key concepts)
change, systems, validity. Bayes' theorem measures how a probability should change when new evidence arrives; it treats a medical test, a factory or a delivery network as a system whose parts feed into one observed outcome; and it tests the validity of the tempting conclusion that a positive result means you almost certainly have the condition.
The question this unit answers
when you know how likely the evidence is for each possible cause, how do you work out how likely each cause is, now that you have seen the evidence?
Where it is examined
HL Paper 1 and Paper 2, as a part worth 3 to 6 marks inside a longer probability question, usually after a tree diagram has been drawn or completed: "given that the parcel was damaged, find the probability that it passed through the south centre". Paper 1 uses fractions that come out exactly; Paper 2 uses decimals and contexts such as screening tests. It can also sit inside a Paper 3 problem, where a parameter is unknown and you must work backwards from a Bayes probability.

What you must be able to do

You must be able toLevelWhat it looks like in the exam
Derive Bayes' theorem for two events from the definition of conditional probabilityHL only"Show that P(B ∣ A) = …" (3 marks)
Use Bayes' theorem for B and B′: P(B ∣ A) = P(B)P(A ∣ B) ÷ [P(B)P(A ∣ B) + P(B′)P(A ∣ B′)]HL only"Given that the test is positive, find the probability that the person has the condition" (4 marks)
Use the law of total probability to find the denominator, P(A)HL only"Find the probability that a randomly chosen parcel is damaged" (2 to 3 marks)
Use Bayes' theorem with a partition into three events B₁, B₂, B₃HL only"Given that the item is faulty, find the probability that it came from machine 3" (4 to 5 marks)
Interpret a Bayes probability in context, including the effect of a rare conditionHL only"Comment on the usefulness of the test" (1 to 2 marks)
Apply Bayes' theorem twice to update a probability after a second piece of evidenceHL only"The person is tested again and is again positive. Find…" (3 to 4 marks), Paper 2
Relate Bayes' theorem to independence: if P(A ∣ B) = P(A ∣ B′), the evidence changes nothingHL only"Hence show that A and B are independent" (2 to 3 marks)
Work backwards from a Bayes probability to an unknownHL only"Given that P(D ∣ +) = 1/2, find x" (4 marks), Paper 1 or Paper 3

Before you start

Bayes' theorem is the SL content of 4.11 written as one formula. You need the definition P(A | B) = P(A ∩ B) ÷ P(B), the rearranged form P(A ∩ B) = P(B)P(A | B), tree diagrams whose second branches are conditional, and the "reversed tree" calculation of P(rain | late) from 4.11 section 3. You also need independent events from 4.6 and 4.11. The formula booklet (AHL section) prints Bayes' theorem for two events and for three, so the marks are for using it correctly, not for remembering it.


1The idea in one paragraph

Most real information comes the wrong way round. A test manufacturer tells you how often the test is positive for people who have the condition, P(positive | condition). A patient wants to know how likely they are to have the condition given that the test was positive, P(condition | positive). These are different numbers, sometimes wildly different. Bayes' theorem converts one into the other. It uses three ingredients: how common each possible cause is before any evidence, P(B); how likely the evidence is under each cause, P(A | B); and the total probability of the evidence, which comes from adding up every route to it. The answer is always the same shape: the route through the cause you care about, divided by all the routes that produce what you saw.

2From the definition to the theorem

Start from the definition of conditional probability, aimed the way you want to go:

P(B | A) = P(A ∩ B) / P(A)

Now replace the top and the bottom with things you know. The top comes from the alternate form in 4.11: P(A ∩ B) = P(B)P(A | B). The bottom is the total probability of A. A happens either with B or with B′, and those two routes cannot overlap, so

P(A) = P(A ∩ B) + P(A ∩ B′)
P(A) = P(B)P(A | B) + P(B′)P(A | B′)the law of total probability

Put both into the definition:

P(B | A) = P(B)P(A | B) ÷ [P(B)P(A | B) + P(B′)P(A | B′)] · Bayes' theorem, from the formula booklet.

Figure 1 shows where each piece sits on a tree. The tree is drawn with B first, so every quantity on the right-hand side of the formula is on a branch. The numerator is the product along the teal path, B then A. The denominator is the sum of the two paths that end in A, teal and amber. Bayes' theorem is nothing more than the "path over paths" recipe you used in 4.11, written in symbols.

Figure 1 · Bayes' theorem reads a tree backwards Figure 1 · Bayes' theorem reads a tree backwards P(B) B P(B′) B′ P(A | B) A P(B) P(A | B) P(A′ | B) A′ P(A | B′) A P(B′) P(A | B′) P(A′ | B′) A′ P(B | A) = P(B) P(A | B) P(B) P(A | B) + P(B′) P(A | B′) The tree gives P(A | B). Bayes' theorem gives P(B | A): the teal path over both paths that end in A.
Figure 1 · Bayes' theorem reads a tree backwards

So you always have a choice: quote the formula and substitute, or draw the tree and compute path over paths. Markers accept both. The formula is safer when the question gives you conditional probabilities in words and no tree; the tree is safer when you need several answers from the same situation.

3A screening test

Here is the context the guide names: medical studies. A screening test is used for a condition that affects 1% of the people screened. Among people who have the condition, the test is positive 95% of the time. Among people who do not, it is still positive 8% of the time (a false positive). A person tests positive. How likely is it that they have the condition?

Let D be "has the condition" and + be "tests positive". Then P(D) = 0.01, P(+ | D) = 0.95 and P(+ | D′) = 0.08.

P(+) = P(D)P(+ | D) + P(D′)P(+ | D′)
= 0.01 × 0.95 + 0.99 × 0.08
= 0.0095 + 0.0792 = 0.0887
P(D | +) = P(D)P(+ | D) / P(+)
= 0.0095 / 0.0887
= 0.107 (3 s.f.)

A positive result means about an 11% chance of having the condition. Most people find that shocking; a test that is "95% accurate" sounds as though a positive result should mean 95%. Figure 2 shows why it does not, using natural frequencies: imagine 10 000 people. About 100 have the condition, and 95 of them test positive. Of the 9 900 without it, 8% is 792, and every one of those tests positive too. So 887 people test positive, and only 95 of them have the condition. The false positives swamp the true ones because there are so many more healthy people for the 8% to act on.

Figure 2 · The screening test in 10 000 people Figure 2 · The screening test in 10 000 people 10 000 people have the condition 1% → 100 do not have it 99% → 9 900 test positive 95% → 95 test negative 5 test positive 8% → 792 test negative 9 108 All 887 positive results, to scale 95 true 792 false positives Of the 887 people who test positive, only 95 have the condition: 95 ÷ 887 ≈ 0.107.
Figure 2 · The screening test in 10 000 people

The 1% is called the prior probability, P(D) before any evidence; the answer 0.107 is the posterior probability, P(D | +) after it. You will not be asked to use these words, but they make the story easy to tell: the evidence raised the probability from 0.01 to 0.107, more than tenfold, yet not anywhere near certainty, because the prior was so small.

Figure 3 plots the same test used in groups where the condition is more or less common. The posterior rises steeply with the prior. In a group where 30% have the condition, a positive result means 0.836. The test did not change; the people did. That is why screening programmes aim tests at high-risk groups, and why a positive screening result is usually followed by a second, different test.

Figure 3 · The rarer the condition, the less a positive result means Figure 3 · The rarer the condition, the less a positive result means P(D | +) prevalence P(D), % 0 0.25 0.5 0.75 1 0 10 20 30 40 50 1%: P(D | +) = 0.107 10%: 0.569 30%: 0.836 Test with P(+ | condition) = 0.95 and P(+ | no condition) = 0.08, used in groups of different prevalence.
Figure 3 · The rarer the condition, the less a positive result means

A negative result is a different question. P(D | −) uses the paths that end in a negative:

P(−) = 1 − 0.0887 = 0.9113
P(D ∩ −) = 0.01 × 0.05 = 0.0005
P(D | −) = 0.0005 / 0.9113 = 0.000549 (3 s.f.)

A negative result is very reassuring: about 5 in 10 000. Questions often ask for both, then ask you to comment. A good comment names both numbers: "the test is good at ruling the condition out, since P(D | −) is tiny, but poor at ruling it in, since P(D | +) is only 0.107."

4Updating twice: a second positive test

Suppose the person who tested positive is tested again with an independent repeat of the same test, and is positive again. Now the right starting point is not 1%. It is 0.107, what you know after the first test. Apply the theorem again with that as the prior. The assumption, which the question will state, is that the two results are independent for a person of known status.

P(D) = 0.10710 (new prior, kept to 5 s.f.)
P(D | second +) = 0.10710 × 0.95 / (0.10710 × 0.95 + 0.89290 × 0.08)
= 0.10175 / (0.10175 + 0.07143)
= 0.10175 / 0.17318
= 0.588 (3 s.f.)

Figure 4 shows the three stages. One positive raised 0.01 to 0.107; a second raised it to 0.588. You get the same 0.588 in one step by treating "positive twice" as the evidence, with P(++ | D) = 0.95² = 0.9025 and P(++ | D′) = 0.08² = 0.0064: then P(D | ++) = 0.009025 ÷ (0.009025 + 0.99 × 0.0064) = 0.588. Two routes, one answer, which is a good check.

Figure 4 · Each positive result updates the probability Figure 4 · Each positive result updates the probability before testing 0.01 after one positive 0.107 after two positives 0.588 0 0.25 0.5 0.75 1 probability of having the condition The answer after the first test becomes the starting probability for the second.
Figure 4 · Each positive result updates the probability

5Bayes' theorem with three events

The guide asks for Bayes' theorem "for a maximum of three events". That means the possible causes are split into three: B₁, B₂ and B₃, which are mutually exclusive (only one can happen) and exhaustive (one of them must happen), so P(B₁) + P(B₂) + P(B₃) = 1. The theorem keeps its shape: one path on top, every path to A underneath.

P(Bᵢ | A) = P(Bᵢ)P(A | Bᵢ) ÷ [P(B₁)P(A | B₁) + P(B₂)P(A | B₂) + P(B₃)P(A | B₃)] · from the formula booklet.

A courier company sorts parcels at three centres: 45% go through the north centre (N), 35% through central (C) and 20% through south (S). The probability that a parcel is damaged (D) is 0.02 at N, 0.04 at C and 0.05 at S. A customer receives a damaged parcel. Which centre most probably handled it?

Figure 5 draws the tree with the three products.

Figure 5 · Three sorting centres: the tree for Bayes with three events Figure 5 · Three sorting centres: the tree for Bayes with three events 0.45 N 0.35 C 0.20 S 0.02 D 0.45 × 0.02 = 0.009 0.98 D′ 0.04 D 0.35 × 0.04 = 0.014 0.96 D′ 0.05 D 0.20 × 0.05 = 0.010 0.95 D′ N, C, S: north, central, south centre D: the parcel is damaged P(S | D) = 0.010 ÷ 0.033 = 10/33 ≈ 0.303 P(D) = 0.009 + 0.014 + 0.010 = 0.033. Given damage, each centre's share is its own path over 0.033.
Figure 5 · Three sorting centres: the tree for Bayes with three events
P(D) = 0.45 × 0.02 + 0.35 × 0.04 + 0.20 × 0.05
= 0.009 + 0.014 + 0.010 = 0.033
P(N | D) = 0.009 / 0.033 = 3/11 ≈ 0.273
P(C | D) = 0.014 / 0.033 = 14/33 ≈ 0.424
P(S | D) = 0.010 / 0.033 = 10/33 ≈ 0.303
check: 9/33 + 14/33 + 10/33 = 1

Central is the most probable source (0.424), even though south has the worst damage rate. South handles too few parcels to dominate. North handles the most parcels, but its low damage rate keeps its share down. Bayes' theorem weighs both things at once: how much traffic each cause has, and how likely each is to produce the evidence.

Always check that the posteriors add to 1. Given the evidence, one of the three causes happened, so the three answers must sum to 1. It costs ten seconds and catches a wrong denominator.

Conditioning on the complement. "Given that a parcel arrived undamaged, find the probability that it went through the north centre" uses the D′ branches: P(D′) = 1 − 0.033 = 0.967, and P(N | D′) = 0.45 × 0.98 ÷ 0.967 = 0.441 ÷ 0.967 = 0.456 (3 s.f.). An undamaged parcel is slightly more likely than average to have come through N (0.456 against 0.45), because N rarely damages anything.

6Bayes on Paper 1: fractions

Without a calculator the numbers are chosen to be fractions with a common denominator in reach. Noor gets to school by bus with probability 1/2, by bike with probability 1/3, and on foot with probability 1/6. She is late with probability 1/5 by bus, 1/10 by bike and 1/4 on foot. One morning she is late. Find the probability that she cycled.

P(L) = (1/2)(1/5) + (1/3)(1/10) + (1/6)(1/4)
= 1/10 + 1/30 + 1/24
= 12/120 + 4/120 + 5/120 = 21/120 = 7/40
P(bike | L) = (1/30) / (7/40)
= (1/30) × (40/7) = 40/210 = 4/21

The bus paths give P(bus | L) = (1/10) ÷ (7/40) = 4/7 and walking gives 5/21; with 4/21 for the bike, the three sum to 12/21 + 4/21 + 5/21 = 1. Find a common denominator for P(L) before dividing: dividing by a sum of three unlike fractions is where Paper 1 answers go wrong.

7When the evidence tells you nothing: the link to independence

The guide links this subtopic to independent events. Suppose the evidence A is equally likely whatever the cause: P(A | B) = P(A | B′) = q. Then

P(B | A) = P(B) q / (P(B) q + P(B′) q)
= P(B) q / (q (P(B) + P(B′)))
= P(B) q / q = P(B)since P(B) + P(B′) = 1

The posterior equals the prior: seeing A does not change the probability of B at all. That is exactly the definition of independence from 4.11, P(B | A) = P(B). A test that is positive 20% of the time whether or not you have the condition is useless, and Bayes' theorem says so. The further apart P(A | B) and P(A | B′) are, the more the evidence moves the probability.

Working backwards. Because the theorem is an equation, a question can give the posterior and hide something else. For the screening test with P(D) = 0.01 and P(+ | D) = 0.95, how low must the false-positive rate x = P(+ | D′) be for a positive result to mean at least a 50% chance of having the condition?

0.0095 / (0.0095 + 0.99x) ≥ 0.5
0.0095 ≥ 0.5(0.0095 + 0.99x)the denominator is positive, so multiply through
0.00475 ≥ 0.495x
x ≤ 0.00960 (3 s.f.)

The false-positive rate would need to fall from 8% to below about 0.96%. For a rare condition, a screening test needs a very low false-positive rate before a single positive result means much.

8Where marks are lost

Answering P(A | B) when P(B | A) was asked. "The test is positive 95% of the time for people with the condition" is P(+ | D). The question "given a positive result, how likely is the condition?" is P(D | +). Write both in symbols before you start.

Forgetting a path in the denominator. P(A) is the sum of every path that ends in A: two with B and B′, three with B₁, B₂, B₃. Using only P(B)P(A | B) in the denominator gives an answer of 1.

Using P(A | B′) = 1 − P(A | B). The complement rule works along the branches from one node: P(A′ | B) = 1 − P(A | B). It does not connect P(A | B) with P(A | B′), which come from different nodes.

Starting the update from the old prior. For a second test, the prior is the posterior from the first test, 0.107, not 0.01.

Rounding the prior before the second update. Carry at least five significant figures (0.10710) into the second step.

Posteriors that do not add to 1. With a partition into three, P(B₁ | A) + P(B₂ | A) + P(B₃ | A) = 1. If yours do not, a path is missing or a product is wrong.

Calling the rare-condition answer a "mistake". An answer of 0.107 is correct. The comment the examiner wants explains why it is low: the condition is rare, so false positives from the large healthy group outnumber the true positives.

9Work it right

  1. Define your events with letters and write what is given in symbols: P(D) = 0.01, P(+ | D) = 0.95, P(+ | D′) = 0.08.
  2. Write what you want in symbols, with the known event after the bar: P(D | +).
  3. Draw or sketch the tree with the cause first and the evidence second, so every given number sits on a branch.
  4. Find the total probability of the evidence as a sum of products, one per path. It is often a separate part worth marks.
  5. Put the one relevant path over that total. Quote the booklet formula as the first line if you are not using a tree.
  6. With three causes, find all three posteriors if asked for the "most likely" one, and check they sum to 1.
  7. For repeated evidence, use the previous posterior as the new prior, unrounded.
  8. Paper 1: common denominators before dividing. Paper 2: 3 significant figures at the end only.

10Try it

Marks in brackets. Q1, Q4 and Q5 are Paper 1 style, no calculator. Q2 and Q3 are Paper 2 style.

Q1. Mia's music app starts with playlist A with probability 2/3 and playlist B with probability 1/3. The probability that the first track is jazz is 1/4 for playlist A and 5/8 for playlist B.

(a) Find the probability that the first track is jazz. 2 marks

(b) Given that the first track is jazz, find the probability that the app started with playlist A. 3 marks

Q2. A condition affects 0.4% of a population. A test for it gives a positive result for 98% of people who have the condition and for 3% of people who do not.

(a) Find the probability that a randomly chosen person tests positive. 2 marks

(b) A person tests positive. Find the probability that they have the condition. 3 marks

(c) Comment on your answer to part (b). 1 mark

Q3. A shop buys stock from three suppliers. Supplier A provides 50% of deliveries, supplier B 30% and supplier C 20%. The probability that a delivery is late is 0.04 for A, 0.10 for B and 0.15 for C.

(a) Find the probability that a randomly chosen delivery is late. 2 marks

(b) Given that a delivery is late, find the probability that it came from supplier C. 2 marks

(c) Given that a delivery is on time, find the probability that it came from supplier A. 2 marks

Q4. For events A and B, P(B) = p where 0 < p < 1, and P(A | B) = P(A | B′) = q where 0 < q < 1. Use Bayes' theorem to show that P(B | A) = p, and state what this tells you about A and B. 3 marks

Q5. A test for an illness is used in a group where P(D) = 0.1. For people with the illness, P(+ | D) = 0.9. For people without it, P(+ | D′) = x. Given that P(D | +) = 1/2, find the value of x. 4 marks

11In one breath

Bayes' theorem reverses a conditional probability. From the definition, P(B | A) = P(A ∩ B) ÷ P(A); the top is P(B)P(A | B), and the bottom, by the law of total probability, is the sum of every path to A, so P(B | A) = P(B)P(A | B) ÷ [P(B)P(A | B) + P(B′)P(A | B′)], and with three causes the bottom has three terms. On a tree drawn cause first, it is the path you want over all the paths that end in what you saw. When the cause is rare, even a good test gives a low P(cause | positive), because false positives from the large unaffected group outnumber the true ones; a second positive result updates again, with the first answer as the new prior. Posteriors over a partition add to 1. If the evidence is equally likely under every cause, the posterior equals the prior, which is independence.


Answers

Q1. (a) P(J) = (2/3)(1/4) + (1/3)(5/8) = 1/6 + 5/24 = 4/24 + 5/24 = 3/8. M1 for the sum of two products, A1 for 3/8.

(b) P(A | J) = (1/6) ÷ (3/8) = (1/6) × (8/3) = 4/9. M1 for a conditional probability with their P(J) as denominator, A1 for 1/6 as numerator, A1 for 4/9. Using 1/4 as the numerator is P(J | A) and scores M1 A0 A0.

Q2. (a) P(+) = 0.004 × 0.98 + 0.996 × 0.03 = 0.00392 + 0.02988 = 0.0338. M1 for the sum of two products, A1 for 0.0338.

(b) P(D | +) = 0.00392 ÷ 0.0338 = 0.116 (3 s.f.). M1 for Bayes' theorem or path over paths, A1 for 0.00392 on top, A1 for 0.116.

(c) Only about 12% of people who test positive have the condition. Because the condition is rare, the 3% false positives from the large unaffected group outnumber the true positives, so a positive result should be followed by further testing. R1 for linking the low value to the rarity of the condition or the false positives. "The test is not accurate" alone scores 0.

Q3. (a) P(L) = 0.5 × 0.04 + 0.3 × 0.10 + 0.2 × 0.15 = 0.02 + 0.03 + 0.03 = 0.08. M1 for the sum of three products, A1 for 0.08.

(b) P(C | L) = 0.03 ÷ 0.08 = 0.375. M1 for their path for C over their P(L), A1 for 0.375.

(c) P(L′) = 0.92 and P(A ∩ L′) = 0.5 × 0.96 = 0.48, so P(A | L′) = 0.48 ÷ 0.92 = 0.522 (3 s.f.). M1 for using the on-time branches with P(L′) = 0.92 as denominator, A1 for 0.522.

Q4. P(B′) = 1 − p. By Bayes' theorem, P(B | A) = pq ÷ [pq + (1 − p)q] = pq ÷ [q(p + 1 − p)] = pq ÷ q = p. So P(B | A) = P(B): knowing A occurred does not change the probability of B, so A and B are independent. M1 for a correct substitution into Bayes' theorem, A1 for simplifying to p with the factor q shown, R1 for the conclusion of independence.

Q5. P(D | +) = 0.1 × 0.9 ÷ (0.1 × 0.9 + 0.9x) = 0.09 ÷ (0.09 + 0.9x). Setting this equal to 1/2: 0.18 = 0.09 + 0.9x, so 0.9x = 0.09 and x = 0.1. M1 for substituting into Bayes' theorem with x, A1 for a correct expression, M1 for setting it equal to 1/2 and solving, A1 for x = 0.1.


Educerie · written from the published IB Diploma Programme Mathematics: analysis and approaches guide, first assessment 2021, section 4.13 Bayes' theorem. Original text, examples and questions. Diagrams drawn by Educerie. Last reviewed 25 September 2026.

Mocks: in the future, hold tight!