Educerie · IB Diploma · Mathematics: analysis and approaches
Topic 4 Statistics and probability · 4.11 Conditional probability and independence
What you must be able to do
| You must be able to | Level | What it looks like in the exam |
|---|---|---|
| State and use the definition P(A ∣ B) = P(A ∩ B) ÷ P(B) | SL, HL | "Find P(A ∣ B)" from given probabilities (2 marks), Paper 1 |
| Read a conditional probability from a Venn diagram, a two-way table or a tree | SL, HL | "Given that a student studies Spanish, find the probability that…" (2 to 3 marks) |
| Use the rearranged form P(A ∩ B) = P(B)P(A ∣ B), including along the branches of a tree | SL, HL | "Find the probability that it rains and the bus is late" (2 marks) |
| Find a "reversed" conditional probability from a tree, such as P(rain ∣ late) | SL, HL | "Given that the bus was late, find the probability that it rained" (3 to 4 marks), Paper 2 |
| Use P(A ∣ B) = P(A) = P(A ∣ B′) and P(A ∩ B) = P(A)P(B) for independent events | SL, HL | "A and B are independent. Find P(A ∪ B)" (3 marks) |
| Test whether two events are independent, and justify the conclusion with numbers | SL, HL | "Determine whether A and B are independent. Justify your answer" (2 to 3 marks) |
| Find an unknown probability from a condition of independence | SL, HL | "Given that A and B are independent, find the value of p" (4 marks), Paper 1 |
| Tell independent events from mutually exclusive events | SL, HL | "Explain why A and B cannot be independent" (2 marks) |
| Interpret conditional probabilities in context, such as a risk factor in a medical study | SL, HL | "Comment on whether smoking and a persistent cough are independent" (1 to 2 marks) |
Before you start
This page formalises what 4.6 introduced. You need the rules for combined events there: P(A ∪ B) = P(A) + P(B) − P(A ∩ B), mutually exclusive events with P(A ∩ B) = 0, and the habit of reading a probability off a Venn diagram, a sample space diagram or a tree. The formula booklet gives P(A | B) = P(A ∩ B) ÷ P(B) and, for independent events, P(A ∩ B) = P(A)P(B). The extra statement P(A | B) = P(A) = P(A | B′) is not printed; it follows in one line from the definition, as section 4 shows.
1The idea in one paragraph
A probability always depends on what you already know. Before you look outside, the chance that your bus is late might be 0.19; once you see that it is raining, the chance is 0.4. The conditional probability P(A | B), read "the probability of A given B", is the probability of A once you know that B has happened. Knowing B throws away every outcome outside B, so B becomes the new sample space, and the probability of A is the share of B that A takes up: P(A ∩ B) ÷ P(B). Two events are independent when knowing one tells you nothing about the other: P(A | B) is the same as P(A), whether or not B happened. That turns into the multiplication rule P(A ∩ B) = P(A)P(B), which is the test you use in an exam.
2The definition: shrink the sample space to B
Here is a gym with its members as the sample space. Let A be the event "a member uses the pool" and B the event "a member attends a group class". Records show P(A) = 0.45, P(B) = 0.30 and P(A ∩ B) = 0.12. The Venn diagram in Figure 1(a) fills in every region: 0.12 in the overlap, 0.45 − 0.12 = 0.33 in A only, 0.30 − 0.12 = 0.18 in B only, and 1 − (0.33 + 0.12 + 0.18) = 0.37 outside both.
Now pick a member who, you are told, attends a class. What is the probability that this member uses the pool? Everyone outside B is now irrelevant, so Figure 1(b) keeps only the circle B, of total probability 0.30. Of that 0.30, the part inside A is 0.12. The answer is the fraction of B that also lies in A:
P(A | B) = P(A ∩ B) ÷ P(B), for P(B) > 0. From the formula booklet.
So 40% of class-goers use the pool, compared with 45% of members overall. Dividing by P(B) rescales B so that it has total probability 1, as U did.
The order matters. P(B | A) is a different question: of the pool users, what fraction attend a class?
Same overlap on top, a different event underneath, a different answer. The event after the bar is always the one you divide by, because it is the one you know has happened. Confusing P(A | B) with P(B | A) is the most expensive mistake in probability; 4.13 (HL) is built on the gap between the two.
Conditional probability from a two-way table. With counts you do not need the formula: restrict attention to the row or column you are given and read the fraction within it. The guide allows this route explicitly. A school of 240 Diploma students records who takes Biology HL and who takes Chemistry HL:
| Chemistry HL | Not Chemistry HL | Total | |
|---|---|---|---|
| Biology HL | 36 | 54 | 90 |
| Not Biology HL | 24 | 126 | 150 |
| Total | 60 | 180 | 240 |
"Given that a student takes Chemistry HL, find the probability that they take Biology HL." Look only at the Chemistry HL column, 60 students, of whom 36 take Biology HL: P(Bio | Chem) = 36/60 = 3/5. The formula gives the same thing, because (36/240) ÷ (60/240) = 36/60. The other way round, P(Chem | Bio) = 36/90 = 2/5.
3The multiplication form, and why trees work
Multiply the definition through by P(B) and you get the form the guide calls the alternate form:
P(A ∩ B) = P(B) P(A | B)
It says: for A and B both to happen, B must happen, and then A must happen given B. That is exactly what you do when you multiply along the branches of a tree diagram, and it is why the second set of branches on any tree carries conditional probabilities.
A commuter's morning bus is late more often in the rain. Let R be "it rains" and L "the bus is late". On a given morning P(R) = 0.3, P(L | R) = 0.4 and P(L | R′) = 0.1. Figure 2 draws the tree. The first branches are P(R) and P(R′); each second branch is a probability given the branch it grows from.
Reading the tree backwards. Now the harder question, which Paper 2 asks often: given that the bus was late, find the probability that it rained. The tree was drawn with rain first, so P(R | L) is not on any branch. Go back to the definition. You know L happened, so L is the new sample space, of probability 0.19; the part of it where it also rained is 0.12.
Knowing the bus was late raises the chance of rain from 0.3 to 0.632, because rain makes lateness four times as likely. The recipe is always the same: the path you want on top, the sum of all the paths consistent with what you know underneath. At HL this recipe, written as a formula, is Bayes' theorem (4.13); at SL you do it exactly as above, straight from the definition.
4Independent events: knowing B tells you nothing about A
Two events are independent when the occurrence of one does not change the probability of the other. In the language of conditional probability:
A and B are independent if P(A | B) = P(A) = P(A | B′).
Read the three parts as a story. Whether B happens (P(A | B)), or B does not happen (P(A | B′)), or you have no idea about B (P(A)), the probability of A is the same number. B carries no information about A.
Figure 3 makes this visible. Each square has area 1, split into a strip for B and a strip for B′, with width P(B) and P(B′). Inside each strip the shaded band is A, and its height is the probability of A within that strip, that is P(A | B) or P(A | B′). In panel (a) the band is level: A takes 30% of B and 30% of B′, so knowing which strip you are in tells you nothing. In panel (b) the band steps down: A takes 60% of B but only 20% of B′, so learning that B happened triples the chance of A.
From the definition to the multiplication rule. If P(A | B) = P(A), put that into the alternate form:
For independent events, P(A ∩ B) = P(A) P(B). From the formula booklet.
In panel (a) of Figure 3 the A-and-B rectangle has width 0.4 and height 0.3, so its area is 0.4 × 0.3 = 0.12 = P(A)P(B). In panel (b) it is 0.4 × 0.6 = 0.24, while P(A) = 0.24 + 0.6 × 0.2 = 0.36 and P(A)P(B) = 0.36 × 0.4 = 0.144. They differ, so those events are not independent.
Why P(A | B′) comes along for free. If P(A ∩ B) = P(A)P(B), then the part of A outside B is P(A ∩ B′) = P(A) − P(A)P(B) = P(A)(1 − P(B)) = P(A)P(B′). Divide by P(B′):
So the three statements P(A | B) = P(A), P(A | B′) = P(A) and P(A ∩ B) = P(A)P(B) stand or fall together, and the same argument shows that if A and B are independent, so are A and B′, A′ and B, and A′ and B′. Independence is a two-way street as well: if P(A | B) = P(A), then P(B | A) = P(A ∩ B) ÷ P(A) = P(A)P(B) ÷ P(A) = P(B).
Where independence comes from. Either the question gives it as a modelling assumption (two separate machines, two dice), and you multiply; or it hands you numbers and asks whether the events are independent, and you test. That is the next section.
5Testing for independence
To test, compute two things that must be equal if the events are independent, and compare them. Any one of these three comparisons is a valid test; use whichever the data make easiest.
| Compare | Independent if | Best when you have |
|---|---|---|
| P(A ∩ B) with P(A) × P(B) | they are equal | the three probabilities, or a Venn diagram |
| P(A ∣ B) with P(A) | they are equal | a conditional probability already worked out |
| P(A ∣ B) with P(A ∣ B′) | they are equal | a tree, or a two-way table split by B |
Example 1 · from probabilities. For the gym, P(A) = 0.45, P(B) = 0.30 and P(A ∩ B) = 0.12.
Equally, P(A | B) = 0.4 ≠ 0.45 = P(A). A full-marks answer shows the two numbers, the ≠ sign and the conclusion in words; "not independent" alone scores nothing.
Example 2 · from a two-way table, in a medical study. Medical researchers use exactly this test to decide whether something is a risk factor: if a disease is more likely among people with some habit than among people without it, the two are not independent. The data below are invented for this page. A clinic records, for 500 adults, whether each smokes (S) and whether each has a persistent cough (C).
| Cough | No cough | Total | |
|---|---|---|---|
| Smoker | 48 | 72 | 120 |
| Non-smoker | 57 | 323 | 380 |
| Total | 105 | 395 | 500 |
Figure 4 puts the three figures side by side. A smoker in this group is 0.40 ÷ 0.15 ≈ 2.7 times as likely to have a cough as a non-smoker, and in context that is the sentence the question is fishing for: "Smokers in this sample are more likely to have a persistent cough, so the events are not independent; smoking appears to be a risk factor."
Two cautions for the comment. First, the guide says that at SL a data set is treated as the population unless you are told otherwise, so you compare the numbers exactly: 0.40 ≠ 0.15 means not independent. (In real research a small difference in a sample could be chance, and statisticians use formal tests for that; you are not asked to.) Second, "not independent" does not prove that one event causes the other. Say "associated with" or "more likely among", not "causes".
Example 3 · using independence to find an unknown (Paper 1). Events A and B are independent, with P(A) = x, P(B) = 0.4 and P(A ∪ B) = 0.64. Find x.
Check: P(A ∩ B) = 0.16 and 0.4 + 0.4 − 0.16 = 0.64. The move that earns the method mark is replacing P(A ∩ B) with the product before anything else.
6Independent is not the same as mutually exclusive
Students muddle these two. They mean nearly opposite things.
- Mutually exclusive: A and B cannot happen together. P(A ∩ B) = 0. On a Venn diagram the circles do not overlap.
- Independent: A and B can happen together, and the chance of one is untouched by the other. P(A ∩ B) = P(A)P(B). On a Venn diagram the circles overlap by exactly that amount.
Figure 5 draws both with the same P(A) = 0.3 and P(B) = 0.5.
In panel (a), learning that B happened tells you that A certainly did not: P(A | B) = 0. That is the strongest possible influence, so the events are about as far from independent as events can be. Formally, if P(A) > 0 and P(B) > 0 then P(A)P(B) > 0, but P(A ∩ B) = 0, so the multiplication rule fails. Mutually exclusive events with non-zero probabilities are never independent. In panel (b) the overlap is 0.3 × 0.5 = 0.15, and P(A | B) = 0.15 ÷ 0.5 = 0.3, which is P(A) again.
7Where marks are lost
Dividing by the wrong event. P(A | B) divides by P(B), the event after the bar, the one you are told has happened. Writing P(A ∩ B) ÷ P(A) answers a different question.
Treating P(A | B) and P(B | A) as the same. 0.4 and 0.267 in the gym; 0.632 and 0.4 for the rain and the bus. They are only equal when P(A) = P(B).
Adding the second branch of a tree to the first instead of multiplying. Along a path you multiply; across separate paths you add. P(L) = 0.3 × 0.4 + 0.7 × 0.1, not 0.3 + 0.4.
Using the unconditional probability as a second branch. The branch after R is P(L | R), not P(L). If the question gives P(L) and P(R) and says nothing about independence, you cannot build the tree by multiplying them.
Assuming independence without being told. P(A ∩ B) = P(A)P(B) is a property to be tested or a given assumption, never a default. If the question does not say the events are independent, do not multiply.
Saying "independent" because the events are mutually exclusive. They are the opposite of independent, as section 6 shows.
Concluding without numbers. "Determine whether… justify" needs the two values compared, the ≠ or = sign, and the conclusion. Rounding before comparing can also hide a difference: compare exact fractions on Paper 1.
Writing "causes" in a context comment. A difference in conditional probabilities shows an association, not a cause.
8Work it right
- Write the definition as the first line: P(A | B) = P(A ∩ B) ÷ P(B). It is the method mark.
- Before dividing, say out loud which event is known. That one goes underneath.
- On a tree, label every second branch as a conditional, P(L | R), not just 0.4.
- For a "reversed" conditional, put the one path you want on top and the sum of all paths consistent with what you know underneath.
- To test independence, name the two quantities you are comparing, compute both, write = or ≠, and conclude in words, in context if there is one.
- When independence is given, substitute P(A ∩ B) = P(A)P(B) at once, then solve.
- Paper 1: keep fractions exact. Paper 2: carry full calculator values and round only the final answer to 3 significant figures.
9Try it
Marks in brackets. Q1, Q2 and Q5 are Paper 1 style, no calculator. Q3 and Q4 are Paper 2 style.
Q1. For two events A and B, P(A) = 0.5, P(B) = 0.4 and P(A ∪ B) = 0.7.
(a) Find P(A ∩ B). 2 marks
(b) Find P(A | B). 2 marks
(c) Show that A and B are independent. 2 marks
(d) Hence write down P(A | B′). 1 mark
Q2. Events A and B are independent, with P(A) = 2/5 and P(A ∪ B) = 13/20. Find P(B). 4 marks
Q3. A clinic records, for 600 patients, whether each takes regular exercise (E) and whether each has high blood pressure (H). Of the 240 patients who exercise regularly, 36 have high blood pressure. Of the 360 who do not, 108 have high blood pressure. A patient is chosen at random.
(a) Find P(H). 1 mark
(b) Find P(H | E). 2 marks
(c) Determine whether the events E and H are independent. Justify your answer, and interpret it in context. 3 marks
(d) Given that the patient has high blood pressure, find the probability that they exercise regularly. 2 marks
Q4. At a café, 65% of customers order a coffee. Of those who order a coffee, 40% also order a cake. Of those who do not order a coffee, 25% order a cake.
(a) Find the probability that a randomly chosen customer orders a cake. 3 marks
(b) Given that a customer orders a cake, find the probability that they also ordered a coffee. 3 marks
(c) State, with a reason, whether ordering a coffee and ordering a cake are independent events. 2 marks
Q5. (a) Events A and B are independent. Show that A and B′ are also independent. 3 marks
(b) Events C and D are mutually exclusive, with P(C) = 0.3 and P(D) = 0.5. Explain why C and D are not independent. 2 marks
10In one breath
P(A | B) is the probability of A once you know B happened: B becomes the whole sample space, so P(A | B) = P(A ∩ B) ÷ P(B), always dividing by the event after the bar, and P(A | B) is not P(B | A). Rearranged, P(A ∩ B) = P(B)P(A | B), which is why you multiply along a tree whose second branches are conditional. To reverse a tree, put the path you want over the sum of every path consistent with what you know. A and B are independent when P(A | B) = P(A) = P(A | B′), which is the same as P(A ∩ B) = P(A)P(B); to test it, compute two of these, compare them with = or ≠, and conclude in words. If you are told events are independent, replace P(A ∩ B) by the product straight away. Mutually exclusive events with non-zero probabilities are never independent, and a difference in conditional probabilities shows an association, not a cause.
Answers
Q1. (a) P(A ∩ B) = P(A) + P(B) − P(A ∪ B) = 0.5 + 0.4 − 0.7 = 0.2. M1 for rearranging the addition rule, A1 for 0.2.
(b) P(A | B) = 0.2 ÷ 0.4 = 0.5. M1 for their P(A ∩ B) divided by P(B), A1 for 0.5. Dividing by P(A) gives 0.4 and scores M0 A0.
(c) P(A | B) = 0.5 = P(A), so A and B are independent. Or: P(A)P(B) = 0.5 × 0.4 = 0.2 = P(A ∩ B). R1 for a correct comparison with both values shown, A1 for the conclusion. "Show that" means the numbers must appear; the word "independent" alone scores 0.
(d) 0.5, since for independent events P(A | B′) = P(A). A1. Check: P(A ∩ B′) ÷ P(B′) = 0.3 ÷ 0.6 = 0.5.
Q2. Let P(B) = p. Independent, so P(A ∩ B) = (2/5)p. Then 13/20 = 2/5 + p − (2/5)p, so (3/5)p = 13/20 − 8/20 = 1/4, giving p = 5/12. M1 for P(A ∩ B) = (2/5)p, M1 for substituting into the addition rule, A1 for (3/5)p = 1/4 or equivalent, A1 for 5/12. Check: 2/5 + 5/12 − 1/6 = 24/60 + 25/60 − 10/60 = 39/60 = 13/20.
Q3. (a) P(H) = (36 + 108) ÷ 600 = 144/600 = 0.24. A1.
(b) P(H | E) = 36/240 = 0.15. M1 for restricting to the 240 who exercise, A1 for 0.15.
(c) P(H | E) = 0.15 but P(H) = 0.24 (or P(H | E′) = 108/360 = 0.3). These are not equal, so E and H are not independent. Patients who exercise regularly are less likely to have high blood pressure than patients in general. M1 for a valid comparison with both values, A1 for "not independent", R1 for an interpretation in context. "Exercise prevents high blood pressure" is a causal claim and does not earn the R1.
(d) P(E | H) = 36/144 = 0.25. M1 for restricting to the 144 with high blood pressure, A1 for 0.25. The answer 0.15 is P(H | E) and scores 0.
Q4. (a) P(cake) = 0.65 × 0.40 + 0.35 × 0.25 = 0.26 + 0.0875 = 0.3475 = 0.348 (3 s.f.). M1 for multiplying along a branch, M1 for adding the two paths, A1 for 0.348 (0.3475 is also accepted as exact).
(b) P(coffee | cake) = 0.26 ÷ 0.3475 = 0.748 (3 s.f.). M1 for a conditional probability with P(cake) underneath, A1 for 0.26 on top, A1 for 0.748. Using the rounded 0.348 gives 0.747, which is also accepted.
(c) Not independent, because P(cake | coffee) = 0.40 but P(cake | no coffee) = 0.25; ordering a coffee changes the probability of ordering a cake. R1 for comparing two relevant probabilities, A1 for the conclusion. Comparing 0.40 with P(cake) = 0.3475 is equally valid.
Q5. (a) P(A ∩ B′) = P(A) − P(A ∩ B) = P(A) − P(A)P(B), using independence. So P(A ∩ B′) = P(A)(1 − P(B)) = P(A)P(B′), which is the condition for A and B′ to be independent. M1 for P(A ∩ B′) = P(A) − P(A ∩ B), M1 for substituting P(A)P(B) and factorising, R1 for the conclusion referring to the multiplication rule. Numerical examples are not a proof.
(b) Mutually exclusive, so P(C ∩ D) = 0. But P(C)P(D) = 0.3 × 0.5 = 0.15 ≠ 0, so the multiplication rule fails and C and D are not independent. (Equally: P(C | D) = 0 ≠ 0.3 = P(C).) R1 for stating P(C ∩ D) = 0, R1 for comparing it with P(C)P(D) = 0.15 and concluding.
Educerie · written from the published IB Diploma Programme Mathematics: analysis and approaches guide, first assessment 2021, section 4.11 Conditional probability and independence. Original text, examples and questions. Diagrams drawn by Educerie. Last reviewed 25 September 2026.