Educerie · IB Diploma · Mathematics: analysis and approaches
Topic 4 Statistics and probability · 4.9 The normal distribution
What you must be able to do
| You must be able to | Level | What it looks like in the exam |
|---|---|---|
| Describe the normal curve and its properties: bell-shaped, symmetric about μ, mean = median = mode, total area 1 | SL, HL | "Write down the median of X" (1 mark); "Sketch the distribution" |
| Recognise where a normal model is and is not reasonable, and why so many measurements are roughly normal | SL, HL | "Comment on the suitability of the model" (1 to 2 marks) |
| Draw the curve with the region you want shaded, and use the diagram to reason | SL, HL | "On the diagram, shade the region representing P(X > 34)" (1 to 2 marks) |
| Use the rule that about 68%, 95% and 99.7% of values lie within 1, 2 and 3 standard deviations of the mean | SL, HL | "Estimate the percentage of plants taller than 195 cm" (2 marks), Paper 1 |
| Use symmetry to find one probability from another | SL, HL | "Given P(X > 34) = 0.15, find P(26 < X < 34)" (3 marks), Paper 1 |
| Find normal probabilities with technology | SL, HL | "Find the probability that a loaf weighs less than 780 g" (2 marks), Paper 2 |
| Find a value of X from a given probability (inverse normal), with μ and σ given | SL, HL | "Find the mass exceeded by 10% of loaves" (2 to 3 marks), Paper 2 |
| Use a normal probability to find an expected number, or as p in a binomial | SL, HL | Section B: "Find the probability that at least one of 6 loaves is underweight" |
Before you start
You need the mean and standard deviation from 4.3, and what the standard deviation measures: a typical distance from the mean. You need to read a histogram (4.2), and to know that a probability is a number from 0 to 1 and that P(A′) = 1 − P(A). The binomial distribution (4.8) is not needed to learn this page, but the exam often joins the two, so it is used in section 7. The formula booklet has nothing to give you here at SL; the calculator does the work, and your job is the setting-up and the sketch.
1The idea in one paragraph
Weigh 500 loaves from one bakery line and draw the histogram: most are close to the target, a few are a little heavy or light, very few are far off, and the bars make a symmetric bell. The normal distribution is the smooth curve that models that bell. It is fixed by two numbers, the mean μ, which says where the centre is, and the standard deviation σ, which says how spread out it is, and we write X ~ N(μ, σ²). For a quantity like mass, which can take any value in a range, probability is area under the curve: the chance a loaf weighs between 790 g and 820 g is the area above that interval. Your GDC finds those areas, and it also works backwards, from an area to the boundary that produces it. Almost every mark on this page comes from setting the question up correctly, and a quick sketch is how you do that.
2A curve for measured quantities
Figure 1 shows a histogram of 500 loaf masses with a normal curve over it. The data are real-looking and a little ragged; the curve is the tidy model. The bars are drawn as frequency densities, so the area of each bar is its number of loaves. Scale everything down so the total area is 1 instead of 500, and area becomes proportion, which is probability.
This is the key change from 4.8. A binomial variable is discrete: it can be 3 or 4, never 3.6, and each value has its own probability, drawn as a bar. The normal distribution models a continuous variable, one that can take any value in an interval: a mass, a time, a length. For a continuous variable:
- probability is area under the curve, and the whole area is 1;
- the probability of one exact value is 0, because a line has no width. P(X = 800) = 0, which says that no loaf weighs exactly 800.000… g;
- so P(X < 780) and P(X ≤ 780) are the same number. Including or excluding the end point changes nothing, which is the opposite of the binomial, where it changes everything.
Where the normal distribution occurs. Many measured quantities come out roughly normal: the masses of items filled or baked by a machine set to a target, the heights of adults of one sex in one population, the lengths of leaves from one species of tree, repeated measurements of the same quantity in a laboratory, where the random errors scatter symmetrically around the true value. What these have in common is that each value is the result of many small, independent influences, some pushing up and some pushing down. Most of the time they roughly cancel, which is why values bunch near the middle; now and then most push the same way, which makes the thin tails.
Many quantities are not normal. Incomes and house prices have a long tail to the right: a few very large values and none below zero. Waiting times are also skewed. Section 8 shows how to tell when a normal model is wrong.
The curve has a history worth knowing. In the 1730s Abraham de Moivre found it as an approximation to binomial probabilities when the number of trials is large; you can already see the bell forming in the binomial bar charts of 4.8 when p = 0.5. A century later the Belgian statistician Adolphe Quetelet fitted it to human measurements and built from it the idea of l'homme moyen, the "average man". His enthusiasm was also a warning: treating the average as the ideal, and every departure from it as an error, is a use of the model that the mathematics does not justify.
3The properties of the normal curve
Figure 2 labels the shape. For X ~ N(μ, σ²):
- It is bell-shaped, with a single peak at x = μ.
- It is symmetric about the vertical line x = μ, so the left half is the mirror image of the right. Exactly half the area lies on each side: P(X < μ) = P(X > μ) = 0.5.
- Because of the symmetry, the mean, the median and the mode are all equal to μ.
- The total area under the curve is 1.
- The tails get ever closer to the horizontal axis but never touch it, so in the model any value is possible, even though values far from μ are very unlikely.
- The curve changes from bending downwards to bending upwards (its points of inflection) at x = μ − σ and x = μ + σ. That is a good guide for sketching: the "shoulders" of the bell are one standard deviation from the centre.
Notation. In N(μ, σ²) the second number is the variance, σ². So X ~ N(800, 144) and X ~ N(800, 12²) are the same statement, and the standard deviation is 12. The calculator wants σ, not σ².
What μ and σ do. Figure 3 changes one at a time. Changing μ slides the whole curve left or right without changing its shape. Changing σ changes the spread: a larger σ makes the curve wider and, because the area must stay 1, lower. A small σ makes it tall and narrow.
4The 68–95–99.7 rule
For every normal distribution, whatever μ and σ are, the same proportions lie within the same number of standard deviations of the mean. You need to know these three:
About 68% of values lie within μ ± σ, about 95% within μ ± 2σ, and about 99.7% within μ ± 3σ.
Figure 4 draws them for the loaves, X ~ N(800, 12²). About 68% weigh between 788 g and 812 g, about 95% between 776 g and 824 g, and almost all, 99.7%, between 764 g and 836 g.
Combine the rule with symmetry and you can split the areas further, as the figure does. The central 68% splits into 34% each side of μ. The band between one and two standard deviations holds (95 − 68) ÷ 2 = 13.5% on each side. Beyond two standard deviations, (100 − 95) ÷ 2 = 2.5% on each side.
Worked example 1 (no calculator). For X ~ N(800, 12²), estimate:
These are approximations; the GDC gives 0.159, 0.0228 and 0.819. The rule is for estimates, for checking a calculator answer is sensible, and for Paper 1. It also gives a quick test of whether data look normal: if far fewer or far more than about two-thirds of a data set lie within one standard deviation of its mean, the normal model is doubtful.
5Symmetry: the Paper 1 method
On Paper 1 you have no calculator, so a normal question will hand you one probability and ask for another. Symmetry is the whole method, and a sketch makes it safe.
Worked example 2. X ~ N(800, 12²) and P(X < 785) = 0.106. Find P(785 < X < 815).
785 is 15 g below the mean and 815 is 15 g above it, so by symmetry the two tails are equal, as Figure 5 shows.
Two more facts from symmetry come up constantly. P(800 < X < 815) is half of the middle piece, 0.394. And P(X < 815) = 1 − 0.106 = 0.894. Before you use symmetry, check that the two values really are the same distance from μ; a question that gives P(X < 785) and asks about 820 cannot be done this way.
6Probabilities with technology
On Paper 2 the GDC finds any area. Every model has a normal cdf function that takes a lower bound, an upper bound, μ and σ, and returns the area between them. On a TI-84 Plus it is normalcdf(lower, upper, μ, σ); on a Casio it is Normal CD. For a tail that goes on for ever, use a very large number as the missing bound, such as 10⁹⁹ (typed 1E99 on a TI) or −10⁹⁹.
The routine that earns full marks is always the same: sketch, shade, write, calculate.
Worked example 3. Loaf masses X ~ N(800, 12²). Find the probability that a loaf chosen at random weighs (a) less than 780 g, (b) between 790 g and 820 g, (c) more than 825 g.
Figure 6 is the sketch for (b). It catches the two commonest errors: shading the wrong side, and an answer that could not possibly be right. An area that is clearly more than half the curve cannot be 0.25.
From a probability to a number of items. If 600 loaves are baked, the expected number under 780 g is 600 × P(X < 780) = 600 × 0.047790… = 28.7. As in 4.5, an expected number need not be a whole number, so the answer is 28.7, not 29.
7Inverse normal: from an area back to a value
Sometimes you know the area and want the boundary: "find the mass exceeded by only the heaviest 10% of loaves". That is an inverse normal calculation. The GDC's inverse normal function (invNorm on a TI-84, Inverse Normal on a Casio) takes an area, μ and σ, and returns the value of x with that much area to its left. At SL, μ and σ will always be given in an inverse normal question; finding an unknown μ or σ needs z-values and is 4.12.
The one step that needs thought is turning the question's area into an area to the left, and Figure 7 shows the two common cases.
Worked example 4. X ~ N(800, 12²).
(a) The heaviest 10% of loaves weigh more than k grams. Find k.
(b) Find the interval, symmetric about the mean, that contains the middle 90% of masses.
Check each answer against the sketch. The top 10% boundary must be above the mean, and it is. The two ends of the middle interval must be the same distance from 800, and they are: 19.7 g each side.
Joining the normal to the binomial. A Section B question often goes one step further. Suppose loaves are sold in packs of 6, chosen at random. What is the probability that at least one loaf in a pack weighs less than 780 g?
Carry the full calculator value of p into the second stage. Rounding it first to 0.05 would give 0.265, and the third figure would be wrong.
8When to trust the model
A normal model always gives an answer, so the thinking has to come from you. Three checks catch most misuse.
- Is the variable continuous and measured? Masses, lengths and times, yes. Counts of people or goals, not in their own right.
- Is the data roughly symmetric and bell-shaped? Look at a histogram, and compare the mean and median: if they are far apart, the data are skewed and the normal model is poor.
- Does the model give impossible values much probability? The normal curve never ends, so it always gives some probability to values below zero. When μ is many standard deviations above zero, that probability is negligible. When it is not, the model is wrong.
Figure 8 shows the third check failing. A survey of daily screen time found a mean of 3.1 hours and a standard deviation of 2.4 hours. A normal model N(3.1, 2.4²) gives P(X < 0) = 0.0982: nearly one person in ten would spend a negative number of hours on screens. The real data are skewed to the right, so a normal model does not suit them. A quick rule: if μ − 2σ is below a value that is impossible (here, 3.1 − 4.8 = −1.7 hours), be suspicious.
Be most careful in the tails. A model fitted to the middle of a data set says least about the extremes, yet extremes are often what matters, such as the strongest flood a river defence must hold. Far out in the tails there are few data, and real data often have heavier tails than the normal curve allows.
9Where marks are lost
Entering the variance as σ. For X ~ N(800, 144), the standard deviation is 12. Typing 144 into the calculator gives a completely different curve. Take the square root first.
No sketch, wrong side. "More than 825" means the area to the right. Without a sketch, P(X < 825) = 0.981 gets written down in place of 0.0186. The sketch catches it.
Giving the inverse normal the wrong area. For "the top 10%", the area to the left is 0.9, not 0.1. Using 0.1 gives 784.6 g, a value below the mean, which the sketch shows cannot be right.
Treating ≤ and < differently. For a continuous variable, P(X ≤ 780) = P(X < 780). Adjusting a boundary by 1, as you would for the binomial, is wrong here.
Writing only calculator syntax. "normalcdf(790, 820, 800, 12) = 0.750" with no probability statement risks the method mark. Write P(790 < X < 820) = 0.750.
Using symmetry about the wrong point. The curve is symmetric about μ, not about 0 and not about the value in the question. P(X < 785) = P(X > 815) only because both are 15 from 800.
Rounding a probability before reusing it. A value of p carried into a binomial or multiplied by 600 should be the full calculator value. Rounded early, the final answer can be wrong in the third figure, and the accuracy mark goes.
Accepting a model that gives impossible values. When asked to comment on a model, check what it says about values below zero or above a physical limit.
10Work it right
- Write the model: X ~ N(μ, σ²), with X defined in words and units.
- Sketch a bell, mark μ in the middle, mark the value or values in the question on the correct side, and shade the region you want.
- Write the probability statement, such as P(X > 825), before any number.
- For probabilities, use the normal cdf with a lower bound, an upper bound, μ and σ; use ±10⁹⁹ for an open tail.
- For inverse normal, convert the given area to the area to the left of the boundary, then use μ and σ.
- Check the answer against the sketch: an area bigger or smaller than half, a boundary above or below the mean.
- On Paper 1, use symmetry and the 68–95–99.7 rule, and say which you used.
- Keep full accuracy for anything you reuse; give final answers to 3 significant figures.
11Try it
Marks in brackets. Q1 and Q2 are Paper 1 style, no calculator. Q3 to Q5 are Paper 2 style, with a GDC.
Q1. The random variable X is normally distributed with mean 30. It is given that P(X > 34) = 0.15.
(a) Sketch the distribution, shading the region that represents P(X > 34). 2 marks
(b) Write down P(X < 26). 1 mark
(c) Find P(26 < X < 34). 2 marks
(d) Find P(30 < X < 34). 1 mark
Q2. The heights of a variety of sunflower are normally distributed with mean 180 cm and standard deviation 15 cm.
(a) Write down the interval, symmetric about the mean, that contains about 95% of heights. 2 marks
(b) Estimate the percentage of sunflowers taller than 195 cm. 2 marks
(c) A field contains 400 of these sunflowers. Estimate how many are between 165 cm and 210 cm tall. 3 marks
Q3. The time T minutes that Amara's commute takes is modelled by T ~ N(42, 6.5²).
(a) Find the probability that her commute takes more than 50 minutes. 2 marks
(b) Find P(35 < T < 45). 2 marks
(c) On only 5% of days her commute takes longer than t minutes. Find t. 2 marks
(d) Over 20 working days, find the probability that her commute takes more than 50 minutes on more than 3 days. Assume the days are independent. 3 marks
Q4. The marks in a large examination are modelled by a normal distribution with mean 61 and standard deviation 13. Treat marks as continuous.
(a) The top 15% of candidates are awarded a distinction. Find the least mark needed for a distinction. 2 marks
(b) Find the lower quartile and the upper quartile of the marks, and hence the interquartile range. 3 marks
Q5. The amount a customer spends in a café, in euros, has mean €6.20 and standard deviation €4.50. A manager proposes to model the amount with a normal distribution.
(a) Using the proposed model, find the probability that a customer spends less than €0. 2 marks
(b) Hence comment on whether a normal distribution is a suitable model. 1 mark
12In one breath
A continuous measured quantity that bunches symmetrically around a centre is modelled by X ~ N(μ, σ²), where μ is the mean and σ² the variance, so the calculator gets σ. The curve is bell-shaped and symmetric about μ, mean = median = mode = μ, the total area is 1, and probability is area, so P(X = a) = 0 and < and ≤ are the same. About 68%, 95% and 99.7% of values lie within one, two and three standard deviations of the mean, and with symmetry that is how Paper 1 questions are done. On Paper 2, sketch and shade first, write the probability statement, then use the normal cdf with a lower and an upper bound; for a value from an area, turn the area into the area to the left and use the inverse normal. Keep full accuracy when a normal probability becomes a binomial p. And check the model: if it gives real probability to impossible values, it is the wrong model.
Answers
Q1. (a) A bell-shaped curve, symmetric about a vertical line at 30, with 34 marked to the right of 30 and the area to the right of 34 shaded. A1 for a bell shape with its centre labelled 30, A1 for 34 marked right of the centre with the correct tail shaded.
(b) 26 and 34 are both 4 from the mean, so by symmetry P(X < 26) = 0.15. A1.
(c) P(26 < X < 34) = 1 − 2(0.15) = 0.7. M1 for 1 − 2 × 0.15 or 1 − 0.15 − 0.15, A1 for 0.7.
(d) Half of (c): 0.35. A1. Also accepted: 0.5 − 0.15.
Q2. (a) μ ± 2σ = 180 ± 30, so 150 cm to 210 cm. M1 for 180 ± 2 × 15, A1 for both ends.
(b) 195 is μ + σ. Above it lies 50% − 34% = 16% (approximately). M1 for recognising 195 as one standard deviation above the mean and using 68% or 34%, A1 for 16%.
(c) 165 is μ − σ and 210 is μ + 2σ. The proportion is 34% + 34% + 13.5% = 81.5%, and 0.815 × 400 = 326 sunflowers (approximately). M1 for identifying the interval as from one σ below to two σ above, A1 for 0.815 or 81.5%, A1 for 326. The GDC gives 327; on Paper 2, either is accepted with the method shown.
Q3. (a) P(T > 50) = 0.109 (3 s.f.; 0.10920… on the GDC). M1 for a correct probability statement or a sketch with the right tail shaded, A1 for 0.109.
(b) P(35 < T < 45) = 0.537 (3 s.f.). M1, A1.
(c) P(T > t) = 0.05, so P(T < t) = 0.95 and t = invNorm(0.95, 42, 6.5) = 52.7 minutes (3 s.f.). M1 for using an area of 0.95 to the left, A1 for 52.7. Using 0.05 gives 31.3, below the mean, and scores M0.
(d) Let D be the number of days, out of 20, with a commute over 50 minutes; D ~ B(20, 0.10920…). P(D > 3) = 1 − P(D ≤ 3) = 0.168 (3 s.f.). M1 for recognising a binomial with n = 20 and p from (a), M1 for 1 − P(D ≤ 3), A1 for 0.168. Using the rounded p = 0.109 gives 0.167; keep the full value so the last figure is secure.
Q4. (a) Let M be the mark, M ~ N(61, 13²). P(M > d) = 0.15, so P(M < d) = 0.85 and d = invNorm(0.85, 61, 13) = 74.5 (3 s.f.). M1 for the area 0.85 to the left, A1 for 74.5.
(b) Q₁ = invNorm(0.25, 61, 13) = 52.2, Q₃ = invNorm(0.75, 61, 13) = 69.8, so the IQR = 69.768… − 52.231… = 17.5 (3 s.f.). M1 for inverse normal at 0.25 and 0.75, A1 for both quartiles, A1 for 17.5. Symmetry check: both quartiles are 8.77 from 61.
Q5. (a) Let A ~ N(6.2, 4.5²). P(A < 0) = 0.0841 (3 s.f.). M1 for P(A < 0) with the correct parameters, A1 for 0.0841.
(b) The model gives a probability of about 8% to a customer spending a negative amount, which is impossible, so a normal distribution is not a suitable model; the amounts are likely to be skewed to the right. R1 for a conclusion that refers to the impossible negative values and their non-negligible probability. "Not suitable" with no reason scores 0.
Educerie · written from the published IB Diploma Programme Mathematics: analysis and approaches guide, first assessment 2021, section 4.9 The normal distribution. Original text, examples and questions. Diagrams drawn by Educerie. Last reviewed 25 September 2026.