Educerie · IB Diploma · Mathematics: analysis and approaches
Topic 4 Statistics and probability · 4.12 Standardization of normal variables
What you must be able to do
| You must be able to | Level | What it looks like in the exam |
|---|---|---|
| Calculate a standardized value z = (x − μ) ÷ σ and say what it means | SL, HL | "Find the z-value for a height of 177 cm" (2 marks) |
| Use z-values to compare results from different normal distributions | SL, HL | "Determine in which test Ana performed better relative to the other candidates" (3 marks) |
| Know that Z ~ N(0, 1) and that P(X < x) = P(Z < z) | SL, HL | "Write down the distribution of Z" (1 mark); any probability found via z |
| Convert back from a z-value to a value of X with x = μ + zσ | SL, HL | "Find the height exceeded by 10% of students" (2 marks) |
| Find an unknown mean from a known σ and a probability | SL, HL | "Find the value of μ" (4 marks), Paper 2 |
| Find an unknown standard deviation from a known μ and a probability | SL, HL | "Find the value of σ" (4 marks), Paper 2 |
| Find both μ and σ from two probabilities by solving simultaneous equations | SL, HL | "Find the mean and standard deviation of the lengths" (6 marks), Paper 2 |
| Use symmetry, and z-values given in the question, without a calculator | SL, HL | Paper 1: "Given that P(Z < 1.2) = 0.885, find P(44 < X < 56)" (3 marks) |
Before you start
You need all of 4.9: the normal curve, its symmetry about μ, the notation X ~ N(μ, σ²) with the variance second, the 68–95–99.7 rule, and your GDC's normal cumulative distribution and inverse normal functions. In 4.9 the inverse normal always came with μ and σ given; the guide says that version "does not involve" standardizing. This page is what you do when μ or σ is the thing you are looking for. The formula booklet gives z = (x − μ) ÷ σ. Everything else on the page is that one line, rearranged.
1The idea in one paragraph
A raw measurement means little on its own. Is 177 cm tall? Is 74 marks good? It depends on the mean and the spread of the group it came from. The standardized value or z-value, z = (x − μ) ÷ σ, answers by counting how many standard deviations x lies above the mean (positive z) or below it (negative z). Once every value is measured in standard deviations, every normal distribution becomes the same curve, the standard normal distribution Z ~ N(0, 1), and the area to the left of x on the original curve equals the area to the left of z on the standard one. That is all standardization is. Its exam payoff is running it backwards: if a question tells you a percentage and hides μ or σ, the GDC turns the percentage into a z-value, and z = (x − μ) ÷ σ becomes an equation you can solve for the unknown.
2The z-value: how many standard deviations from the mean
Heights of the students in a large school are modelled by X ~ N(165, 8²), in centimetres. So μ = 165 and σ = 8. A student is 177 cm tall.
z = (x − μ) ÷ σ · the number of standard deviations x is from the mean. From the formula booklet.
A student of 150 cm has z = (150 − 165) ÷ 8 = −1.875: nearly two standard deviations below the mean. The sign says which side of μ; the size says how far. The mean itself always has z = 0, and one standard deviation above it, 173 cm, has z = 1.
Figure 1 draws both scales under one curve. The top scale is centimetres. The bottom scale is z, and each step on it is one σ = 8 cm. Standardizing is just reading the lower scale instead of the upper one.
Seen this way, the 68–95–99.7 rule from 4.9 is a statement about z: about 68% of values have z between −1 and 1, about 95% between −2 and 2, and about 99.7% between −3 and 3. A z-value beyond ±3 is rare in any normal distribution.
Comparing across distributions. Because z has no units and the same meaning for every normal distribution, it lets you compare results that were measured on different scales. Ana sits two tests. In Physics she scores 74 where the scores are modelled by N(62, 8²). In Mathematics she scores 70 where the scores are modelled by N(55, 12²). Which is the better performance relative to the other candidates?
Her Mathematics score beat the mean by more marks (15 against 12), but the Mathematics marks were more spread out, so 15 marks there is only 1.25 standard deviations. Figure 2 puts both results on one standard scale.
The answer that scores the reasoning mark says both things: which z is larger, and that a larger z means further above the mean in standard deviations.
3The standard normal distribution Z ~ N(0, 1)
If X ~ N(μ, σ²) and you standardize every value, the new variable
Z = (X − μ) ÷ σ has the standard normal distribution, Z ~ N(0, 1).
Subtracting μ slides the curve so its centre is at 0. Dividing by σ squeezes or stretches it so its standard deviation is 1. Neither step changes the proportion of the distribution lying to the left of a point, because every value moves together. So for any value x with z-value z,
Figure 3 shows the two shaded areas side by side. They are the same area, 0.933. Your GDC will give P(X < 177) directly with μ = 165 and σ = 8, so for a plain probability you do not need z at all. You standardize when z itself is asked for, when you are comparing, and, above all, when μ or σ is unknown.
From a probability back to a value. Rearrange the z formula for x:
x = μ + zσ · the value that sits z standard deviations from the mean.
Find the height exceeded by only 10% of students. That height has 90% below it. The inverse normal function on the GDC, applied to Z ~ N(0, 1) with area 0.9, gives z = 1.28155… Then:
In 4.9 you would get the same number by entering μ = 165 and σ = 8 into the inverse normal directly. The point of doing it through z is that the line x = μ + zσ still makes sense when μ or σ is a letter.
A GDC warning. The inverse normal function works with the area to the left of the value you want. "10% are taller than x" means an area of 0.9 to the left, not 0.1. Some calculators let you choose a tail; make sure you know which yours is using, and when in doubt, convert to a left-hand area yourself.
4Finding an unknown mean
A machine fills bags of flour. The masses are normally distributed with standard deviation 6 g, a property of the machine. The mean can be adjusted. The label says 500 g, and the company wants only 2% of bags to weigh less than that. At what mean should the machine be set?
Here μ is the unknown, so the GDC cannot give a probability or a value directly: it would need μ. Instead, turn the 2% into a z-value, which needs no μ at all.
Figure 4 shows what happened: the curve was slid to the right until the tail below 500 g held exactly 2%. The negative z-value is essential, because 500 is below the mean. Had you used +2.05375 you would get μ = 487.7 g, which puts more than half the bags under weight. Sketch the curve before you calculate: a quick drawing shows which side of μ the value is on, and so which sign z must have.
5Finding an unknown standard deviation
A commuter's journey time T is modelled by T ~ N(42, σ²) minutes. On 90% of days the journey takes less than 50 minutes. Find σ, then the probability that a journey takes more than 55 minutes.
The same structure as before: known value, known mean, z from the GDC, solve. The only extra care is algebraic, since σ is in the denominator; multiply both sides by σ first if that is easier.
Keep full accuracy. Store σ in the calculator and use the stored value for the follow-on probability. If you round σ to 6.2 and then find P(T > 55), you get a noticeably different answer (0.0180 instead of 0.0186), and the accuracy mark is lost. The IB rule of thumb is simple: never round an intermediate value to fewer than the 3 significant figures of the final answer, and preferably not at all.
6Finding both the mean and the standard deviation
If both are unknown you need two pieces of information, and each one becomes an equation.
The lengths of a species of fish in a lake are normally distributed. 15% of the fish are shorter than 20 cm, and 10% are longer than 32 cm. Find the mean and standard deviation of the lengths.
Sketch first, as Figure 5 does. 20 cm is in the lower tail, so its z-value is negative. 32 cm has 10% above it, so 90% below it, and its z-value is positive.
The subtraction step has a clean meaning, which Figure 5 marks: the 12 cm between the two cut-off points is 1.03643 + 1.28155 = 2.31798 standard deviations. On Paper 2 you may also solve the pair of linear equations with the GDC's simultaneous equation solver; write the two equations down first, because they carry the method marks.
Check it. With μ = 25.365 and σ = 5.1769, the GDC gives P(X < 20) = 0.150 and P(X > 32) = 0.100. A ten-second check like this catches the most common error in the whole topic, using an area of 0.10 instead of 0.90 for the upper value.
7Paper 1: symmetry and given z-values
Without a calculator you cannot find a z-value from a probability, so a Paper 1 question either gives you the z-values or relies on the symmetry of the curve. Figure 6 shows the symmetry fact you need: points the same distance either side of μ cut off equal tails.
For a standard normal Z and any a > 0:
Example 1. X ~ N(μ, σ²), with P(X < 14) = 0.3 and P(X > 26) = 0.3. Find μ. The two tails are equal, so 14 and 26 are the same distance from μ, and μ is halfway: μ = 20. No calculator and no z needed.
Example 2. X ~ N(50, 5²). Given that P(Z < 1.2) = 0.885, find P(44 < X < 56).
Example 3. X ~ N(30, σ²) and P(X < 36) = 0.8413. Given that P(Z < 1) = 0.8413, find σ. The two statements have the same area, so 36 standardizes to z = 1: (36 − 30) ÷ σ = 1, and σ = 6. The logic is identical to section 5; the calculator's job has been done by the question.
8Where marks are lost
Using σ² where σ belongs. In X ~ N(165, 64) the standard deviation is 8, not 64. The formula and the GDC both want σ.
Getting the sign of z wrong. A value below the mean has a negative z-value. For the bags, z = −2.054, not +2.054. A sketch settles it.
Feeding the inverse normal the wrong tail. "10% are longer than 32 cm" means an area of 0.90 to the left of 32. Entering 0.10 gives z = −1.282 and a nonsense answer.
Putting a probability into the z formula. (0.15 − μ) ÷ σ is meaningless. A probability goes into the inverse normal; what comes out, the z-value, goes into the formula.
Rounding z or σ early. Using z = 1.28 or σ = 6.2 in a follow-on part costs the final accuracy mark. Store full values.
Trying to use the GDC's inverse normal with the unknown parameter. If μ is unknown you cannot enter it. Standardize: use μ = 0 and σ = 1 to get z, then solve.
Comparing raw scores instead of z-values. 15 marks above the mean is not better than 12 if the spread is larger. Compare z.
Forgetting to answer in context and to 3 significant figures. "μ = 512.3224" with no units is not a finished answer; "the machine should be set to a mean of 512 g (3 s.f.)" is.
9Work it right
- Write down the distribution with the unknown as a letter: X ~ N(μ, 6²).
- Sketch the curve, mark the known value, shade the given area. Decide the sign of z from the sketch.
- Convert the given probability to a left-hand area if needed, then write "P(Z < z) = 0.9, so z = 1.28155" as a line of working.
- Write the standardized equation (x − μ) ÷ σ = z with the numbers in. This line is the main method mark.
- Solve. With two unknowns, write both equations, then subtract or use the GDC's solver.
- Keep full accuracy throughout; give the final answer to 3 significant figures with units.
- Check by putting your μ and σ back into the GDC and recovering the given percentage.
- When comparing, state both z-values and say which is further from the mean in standard deviations.
10Try it
Marks in brackets. Q1 to Q4 are Paper 2 style, with a GDC. Q5 is Paper 1 style, no calculator.
Q1. Lina scores 68 in a test whose marks are modelled by N(60, 5²). Omar scores 81 in a different test whose marks are modelled by N(70, 8²).
(a) Calculate the z-value of each score. 2 marks
(b) Determine who performed better relative to the other candidates in their test, giving a reason. 1 mark
Q2. The masses of apples from an orchard are normally distributed with standard deviation 18 g. 12% of the apples have a mass greater than 200 g. Find the mean mass. 4 marks
Q3. The time T taken to complete a puzzle is modelled by T ~ N(25, σ²) minutes. P(T < 20) = 0.2.
(a) Find σ. 3 marks
(b) Find P(T > 32). 2 marks
Q4. The breaking strength of a type of rope is normally distributed. 5% of ropes break under a load of less than 800 N, and 20% of ropes withstand a load of more than 1000 N. Find the mean and the standard deviation of the breaking strength. 6 marks
Q5. X ~ N(40, σ²) and P(X > 46) = 0.0228.
(a) Write down P(X < 34). 1 mark
(b) Find P(34 < X < 46). 2 marks
(c) Given that P(Z < 2) = 0.9772, find σ. 2 marks
11In one breath
The z-value z = (x − μ) ÷ σ counts how many standard deviations x lies from the mean: positive above, negative below, 0 at the mean. It has no units, so it compares values from different normal distributions, and standardizing turns every X ~ N(μ, σ²) into Z ~ N(0, 1) with P(X < x) = P(Z < z). Going the other way, x = μ + zσ. When μ or σ is unknown, the GDC cannot use it, so convert the given percentage to a left-hand area, get z from the inverse normal with μ = 0 and σ = 1, write (x − μ) ÷ σ = z, and solve; with both unknown, two percentages give two linear equations. Sketch first so the sign of z is right, use 0.9 not 0.1 for "10% above", keep full accuracy, and check by recovering the given percentages. On Paper 1, symmetry or given z-values replace the calculator.
Answers
Q1. (a) Lina: z = (68 − 60) ÷ 5 = 1.6. Omar: z = (81 − 70) ÷ 8 = 1.375 (1.38 to 3 s.f.). A1 for each z-value.
(b) Lina, because her score is 1.6 standard deviations above her test's mean, further than Omar's 1.375. R1 for comparing z-values. "Omar, because he scored 11 above the mean" compares raw marks and scores 0.
Q2. P(X > 200) = 0.12, so P(X < 200) = 0.88, giving z = 1.17499. Then (200 − μ) ÷ 18 = 1.17499, so μ = 200 − 21.1498 = 178.85 ≈ 179 g. M1 for converting to an area of 0.88 (or using an upper-tail setting correctly), A1 for z = 1.17499 (accept 1.17 or better), M1 for the standardized equation, A1 for 179 g. Using z = −1.175 gives 221 g and scores M1 A0 M1 A0.
Q3. (a) P(T < 20) = 0.2 gives z = −0.841621. Then (20 − 25) ÷ σ = −0.841621, so σ = 5 ÷ 0.841621 = 5.9409 ≈ 5.94 minutes. A1 for z = −0.8416, M1 for the standardized equation with the correct sign, A1 for 5.94.
(b) With σ = 5.9409, P(T > 32) = 0.119 (3 s.f.). M1 for a correct GDC set-up with their σ, A1 for 0.119. Follow-through from (a).
Q4. P(X < 800) = 0.05 gives z₁ = −1.64485. P(X > 1000) = 0.2 means P(X < 1000) = 0.8, giving z₂ = 0.841621. So μ − 1.64485σ = 800 and μ + 0.841621σ = 1000. Subtracting, 2.48647σ = 200, so σ = 80.435 ≈ 80.4 N, and μ = 800 + 1.64485 × 80.435 = 932.30 ≈ 932 N. A1 for z₁, A1 for z₂ from an area of 0.8, M1 for forming both standardized equations, M1 for solving them simultaneously, A1 for σ, A1 for μ. Using an area of 0.2 for z₂ gives z₂ = −0.8416, σ = 249 and μ = 1210, which a sketch shows is impossible (1000 N would be below the mean); that scores at most A1 A0 M1 M1 A0 A0.
Q5. (a) By symmetry about μ = 40, P(X < 34) = P(X > 46) = 0.0228. A1.
(b) P(34 < X < 46) = 1 − 2 × 0.0228 = 0.9544. M1 for 1 − 2 × their tail, A1 for 0.9544.
(c) P(X < 46) = 1 − 0.0228 = 0.9772 = P(Z < 2), so 46 standardizes to z = 2: (46 − 40) ÷ σ = 2, giving σ = 3. M1 for (46 − 40) ÷ σ = 2, A1 for σ = 3.
Educerie · written from the published IB Diploma Programme Mathematics: analysis and approaches guide, first assessment 2021, section 4.12 Standardization of normal variables. Original text, examples and questions. Diagrams drawn by Educerie. Last reviewed 25 September 2026.