Educerie
Level

Educerie · IB Diploma · Mathematics: analysis and approaches

Topic 4 Statistics and probability · 4.2 Presenting data

Level
SL and HL. Nothing here is HL only, so every section is examinable for both.
Themes (key concepts)
quantity, approximation, validity. A table or graph turns a list of quantities into a shape you can read, and a value read from a cumulative frequency graph is an approximation, valid only as far as the grouping allows.
The question this unit answers
a list of eighty numbers tells you almost nothing at a glance, so which pictures of it show the centre, the spread and the shape, and how do you read values back off them?
Where it is examined
both papers. Paper 1 gives you a table or a graph and asks you to complete a cumulative frequency column, read a median or quartile off a printed graph, or draw a box-and-whisker diagram from five given values (2 to 6 marks). Paper 2 asks for the same with a GDC doing the counting, and adds comparisons: "compare the two distributions" (2 to 4 marks). Box plots also appear as the first half of a question that goes on to the normal distribution (4.9), because the guide asks you to judge from one whether data could be normal.

What you must be able to do

You must be able toLevelWhat it looks like in the exam
Read and build a frequency table for discrete data, and for continuous data in class intervalsSL, HL"Complete the frequency table" (1 to 2 marks)
Draw and read a frequency histogram with equal class intervalsSL, HL"Draw a histogram to represent the data" (3 marks)
Complete a cumulative frequency table and draw the cumulative frequency graphSL, HL"Complete the cumulative frequency column"; "Draw the graph" (2 to 4 marks)
Read the median, quartiles, percentiles, range and IQR from a cumulative frequency graphSL, HL"Use the graph to estimate the 90th percentile" (2 marks)
Use the graph to count how many values lie above or below a valueSL, HL"Estimate the number of runners who took more than 38 minutes" (2 marks)
Draw a box-and-whisker diagram, with outliers marked by a crossSL, HL"Draw a box-and-whisker diagram for these data" (3 to 4 marks)
Compare two distributions from their box plotsSL, HL"Compare the two distributions" (2 to 4 marks), always in context
Judge from a box plot whether data could be normally distributedSL, HL"Explain why the data may be normally distributed" (1 to 2 marks)

Before you start

You need discrete and continuous data and the outlier rule from 4.1. Median and quartiles of a short list are prior learning; 4.3 covers how a GDC finds them. The formula booklet gives IQR = Q₃ − Q₁. Frequency density histograms, where bars have unequal widths, are not on this course: every histogram you meet has equal class widths.


1The idea in one paragraph

A frequency table counts how often each value, or each range of values, occurs. From it you draw a histogram, which shows the shape of the data: where it piles up, and whether it trails off to one side. Add the frequencies up as you go and you get cumulative frequency, the number of values up to each point; its graph lets you read off the median, quartiles and any percentile by going across and down. Those five numbers (minimum, lower quartile, median, upper quartile, maximum) give a box-and-whisker diagram, the fastest picture for comparing two data sets and for spotting skew, outliers, and data that might be normal.

2Frequency tables

Discrete data go in a table with one row per value. Thirty students were asked how many books they read last month (invented data):

Books read, x012345
Frequency, f379632

The frequencies add to 30, which is the first thing to check in any table. Figure 1 draws it: separate bars with gaps, because 2.5 books is not a possible answer.

Figure 1 · Books read last month by 30 students Figure 1 · Books read last month by 30 students 0 1 2 3 4 5 Number of books read 0 2 4 6 8 10 Frequency 3 7 9 6 3 2 Discrete data: one bar per value, with gaps, because nothing lies between 2 books and 3.
Figure 1 · Books read last month by 30 students

Continuous data are grouped into class intervals. The guide says these are written as inequalities without gaps, so every value belongs to exactly one class. Eighty runners' times at a park run (invented data):

Time t (minutes)Frequency
15 ≤ t < 204
20 ≤ t < 2511
25 ≤ t < 3022
30 ≤ t < 3519
35 ≤ t < 4013
40 ≤ t < 457
45 ≤ t < 504

A runner who took exactly 25.00 minutes goes in 25 ≤ t < 30, not in the class below, because the lower class has "< 25". The class width here is 5 minutes, and the lower and upper boundaries of the third class are 25 and 30. Grouping costs information: once a time is in the table you know only its class, not its value. That is why everything read from grouped data later on is an estimate.

3Histograms

A frequency histogram draws each class as a bar whose width spans the class and whose height is its frequency. Because the data are continuous, the bars touch: there is no gap between 24.99 minutes and 25.00. Figure 2 is the histogram of the park-run times.

Figure 2 · Frequency histogram of 80 park-run times Figure 2 · Frequency histogram of 80 park-run times 15 20 25 30 35 40 45 50 Time t (minutes) 0 4 8 12 16 20 24 28 Frequency 4 11 22 19 13 7 4 modal class 25 ≤ t < 30 Continuous data, equal class widths of 5 minutes: bars touch, and height is frequency.
Figure 2 · Frequency histogram of 80 park-run times

Read it for shape. The times pile up in 25 ≤ t < 30, the modal class, and trail off further to the right than to the left: a few runners walked it. A longer tail on the right is positive skew; a longer tail on the left is negative skew; a histogram that is roughly a mirror image either side of its peak is symmetric.

For the histogram to be marked correctly: frequency on the vertical axis with a scale, the variable and its units on the horizontal axis, bars starting and ending at the class boundaries, and equal widths.

4Cumulative frequency

The cumulative frequency at a value is the number of data items less than or equal to it. You build it by adding each frequency to the running total.

Time t (minutes)FrequencyCumulative frequency
t < 2044
t < 251115
t < 302237
t < 351956
t < 401369
t < 45776
t < 50480

The last entry must equal the total, 80. The cumulative frequencies belong to the upper boundaries: by 20 minutes, 4 runners had finished; by 25 minutes, 15 had. So the cumulative frequency graph plots (20, 4), (25, 15), (30, 37) and so on, and starts at (15, 0), because no one finished in under 15 minutes. Join the points with straight lines or a smooth curve; both are accepted, and readings will differ slightly between them. The graph always rises or stays level, because a running total can never go down. Figure 3 is the graph.

Plot cumulative frequency at the upper class boundary, and start the graph at the lower boundary of the first class with frequency 0.

Figure 3 · Cumulative frequency: median and quartiles Figure 3 · Cumulative frequency: median and quartiles 15 20 25 30 35 40 45 50 Time t (minutes) 0 10 20 30 40 50 60 70 80 Cumulative frequency 20 Q₁ ≈ 26.1 40 median ≈ 30.8 60 Q₃ ≈ 36.5 Go across from n/4, n/2 and 3n/4, then down. Q₁ ≈ 26.1, median ≈ 30.8, Q₃ ≈ 36.5 minutes.
Figure 3 · Cumulative frequency: median and quartiles

5Reading median, quartiles and percentiles

To read a value off a cumulative frequency graph, go across from the cumulative frequency you want to the graph, then down to the horizontal axis.

  • The median is the middle value: go across from n ÷ 2.
  • The lower quartile Q₁ has a quarter of the data below it: go across from n ÷ 4.
  • The upper quartile Q₃ has three quarters below it: go across from 3n ÷ 4.
  • The kth percentile has k% of the data below it: go across from (k ÷ 100) × n. The median is the 50th percentile, Q₁ the 25th and Q₃ the 75th.

For the 80 runners (Figure 3):

median: across from 80 / 2 = 40 → t ≈ 30.8 minutes
Q1: across from 80 / 4 = 20 → t ≈ 26.1 minutes
Q3: across from 3 × 80 / 4 = 60 → t ≈ 36.5 minutes
IQR = Q3 − Q1 ≈ 36.5 − 26.1 = 10.4 minutes

With a large n, IB questions use n ÷ 2 for the median rather than fussing over (n + 1) ÷ 2; the graph is not accurate enough for the difference to matter. Give readings to the accuracy the scale allows, and expect the mark scheme to accept a small range.

The range from a grouped table is less certain. The graph runs from 15 to 50, so the range is at most 50 − 15 = 35 minutes, but the fastest runner might have taken 16.2 minutes and the slowest 48.9. Say "at most" when that is all you can know.

The graph also answers counting questions. Figure 4 reads two things off the same curve.

Figure 4 · Reading a percentile and a count Figure 4 · Reading a percentile and a count 15 20 25 30 35 40 45 50 Time t (minutes) 0 10 20 30 40 50 60 70 80 Cumulative frequency 72 = 90% of 80 ≈ 42.1 ≈ 64 38 16 slower 90th percentile: across from 72, about 42.1 min. Slower than 38 min: 80 − 64 = 16 runners.
Figure 4 · Reading a percentile and a count
90th percentile: across from 0.9 × 80 = 72 → t ≈ 42.1 minutes
runners taking more than 38 minutes:
up from t = 38 → cumulative frequency ≈ 6464 took 38 minutes or less
80 − 64 = 16 runners

The trap in the second one is stopping at 64. The graph counts values below a point, so "more than" always needs the subtraction from the total.

6Box-and-whisker diagrams

A box-and-whisker diagram (box plot) draws five numbers on a scale: a box from Q₁ to Q₃ with a line at the median, and a whisker from each end of the box out to the most extreme value that is not an outlier. Any outlier is marked separately with a cross.

Nineteen apples from one orchard have these masses in grams (invented data), in order:

142, 148, 151, 153, 155, 156, 158, 160, 161, 163, 164, 166, 168, 170, 171, 174, 177, 181, 212

n = 19, so the median is the 10th value: 163
Q1 = median of the lower 9 values = 1555th value
Q3 = median of the upper 9 values = 17115th value
IQR = 171 − 155 = 16, and 1.5 × 16 = 24
fences: 155 − 24 = 131 and 171 + 24 = 195
212 > 195, so 212 is an outlier; nothing is below 131

So the left whisker runs to the minimum, 142, and the right whisker stops at 181, the largest value inside the fences. The 212 gets a cross. Figure 5 shows the finished diagram.

Figure 5 · Box-and-whisker diagram of 19 apple masses Figure 5 · Box-and-whisker diagram of 19 apple masses 140 150 160 170 180 190 200 210 220 Mass (grams) minimum 142 Q₁ = 155 median 163 Q₃ = 171 whisker end 181 212: outlier IQR = 16 Fences: 155 − 24 = 131 and 171 + 24 = 195. The whisker stops at 181, the largest value inside them.
Figure 5 · Box-and-whisker diagram of 19 apple masses

To draw one in an exam: a labelled scale with units first, then the box and median line, then the whiskers, then the crosses. A box plot drawn without a scale scores almost nothing, because its whole value is where things sit on the scale.

Half the data lie inside the box, a quarter along each whisker (counting any outliers on that side). A long whisker does not mean many values; it means the values in that quarter are spread out.

7Comparing two distributions

Box plots are built for comparison: put them on the same scale, one above the other. Orchard B's apples have minimum 138, Q₁ 150, median 157, Q₃ 162 and maximum 175, with no outliers. Figure 6 draws both orchards.

Figure 6 · Apple masses from two orchards Figure 6 · Apple masses from two orchards 140 150 160 170 180 190 200 210 220 Mass (grams) Orchard A Orchard B A's median is higher, so its apples are typically heavier; B's smaller IQR means more consistent masses.
Figure 6 · Apple masses from two orchards

A full comparison makes two points, each with numbers and each in context:

  • Centre. Orchard A's median (163 g) is higher than Orchard B's (157 g), so A's apples are typically heavier.
  • Spread. Orchard B's IQR (162 − 150 = 12 g) is smaller than A's (171 − 155 = 16 g), so the middle half of B's apples are more consistent in mass. The range tells the same story more dramatically (37 g against 70 g), but A's range is inflated by one outlier, which is why the IQR is the better measure to quote.

You may add a third point on symmetry: both boxes have the median near the centre, so neither orchard's middle half is strongly skewed. "A is bigger than B" scores nothing; "A has a higher median, 163 g compared with 157 g, so A's apples are typically heavier" scores the mark.

8Shape, and whether the data could be normal

The position of the median inside the box and the lengths of the whiskers tell you the shape, as Figure 7 shows.

Figure 7 · The shape of a distribution in both diagrams Figure 7 · The shape of a distribution in both diagrams Symmetric: could be normal Positively skewed Negatively skewed A long tail on one side stretches that whisker and pushes the median away from it inside the box.
Figure 7 · The shape of a distribution in both diagrams
  • Symmetric: the median is near the middle of the box and the two whiskers are about equal in length.
  • Positively skewed: the right whisker is longer and the median sits nearer the left of the box, because the values above the median are spread over a wider range.
  • Negatively skewed: the mirror image.

The normal distribution, which you meet in 4.9, is symmetric and bell-shaped. So the guide asks you to judge from a box plot whether data may be normally distributed, by looking for symmetry: median near the centre of the box, whiskers roughly equal. A normal distribution also has most of its values packed near the middle, so each whisker is typically longer than half the box. A skewed box plot rules normality out; a symmetric one only makes it plausible, since many non-normal distributions are symmetric too. Say "may be", not "is".

9Where marks are lost

Plotting cumulative frequency at the midpoint. The running total "by 25 minutes" belongs at 25, the upper boundary. Plotting at 22.5 shifts every reading left.

Forgetting the starting point. The graph starts at (lower boundary of the first class, 0). Without it, the first segment is missing and Q₁ is misread in small data sets.

Stopping at the reading in a "more than" question. The graph counts values below. More than 38 minutes is 80 − 64, not 64.

Whiskers drawn to an outlier. The whisker ends at the most extreme value that is not an outlier; the outlier gets its own cross.

Comparing without numbers or context. Every comparison quotes the two values it compares and says what the difference means for the apples, the runners, the students.

Using the range to compare when there is an outlier. One extreme value inflates the range. Compare with the IQR and say why.

Gaps between histogram bars for continuous data. Bars touch. Gaps are for discrete values, as in Figure 1.

"The data are normal." A symmetric box plot means the data may be normal. It is evidence, not proof.

10Work it right

  1. Check that the frequencies add to the stated total before anything else.
  2. Cumulative frequency: last value equals n; plot at upper boundaries; start at (first lower boundary, 0).
  3. Label both axes of every graph with the variable, the units and a scale.
  4. When reading off a graph, draw the across-and-down lines so the method mark is visible.
  5. Write which cumulative frequency you went across from: n ÷ 2, n ÷ 4, 3n ÷ 4, or (k ÷ 100) × n.
  6. For "more than" questions, subtract the reading from the total.
  7. For a box plot: calculate the fences, end each whisker at the last value inside them, and mark outliers with a cross.
  8. Compare with one statement on centre and one on spread, each with both values and a context sentence.

11Try it

Marks in brackets. Q1, Q2 and Q4 are Paper 1 style, no calculator. Q3 is Paper 2 style.

Q1. The lengths of 50 fish caught in a river are summarised below (invented data).

Length L (cm)20 ≤ L < 2525 ≤ L < 3030 ≤ L < 3535 ≤ L < 4040 ≤ L < 4545 ≤ L < 50
Frequency39171362

(a) Write down the cumulative frequencies. 2 marks

(b) Write down the class that contains the median length. 1 mark

(c) State the coordinates of the first two points you would plot on the cumulative frequency graph. 1 mark

Q2. A train was late on 11 days by these numbers of minutes, in order: 0, 1, 1, 2, 3, 3, 4, 5, 6, 8, 19.

(a) Find the median and the quartiles. 3 marks

(b) Show that 19 is an outlier. 2 marks

(c) Write down the values at which the two whiskers of the box-and-whisker diagram end. 1 mark

Q3. Figure 8 shows the cumulative frequency graph of the weekly homework time of 120 students (invented data).

Figure 8 · Homework time for 120 students Figure 8 · Homework time for 120 students 0 1 2 3 4 5 6 7 8 9 10 11 12 Homework time h (hours per week) 0 10 20 30 40 50 60 70 80 90 100 110 120 Cumulative frequency Cumulative frequency graph for Q3.
Figure 8 · Homework time for 120 students

(a) Use the graph to estimate the median and the interquartile range. 4 marks

(b) Estimate the 80th percentile. 2 marks

(c) Estimate the number of students who did more than 9 hours of homework. 2 marks

Q4. Two classes sat the same test, marked out of 100. The five-number summaries are: Class P: minimum 44, Q₁ 58, median 66, Q₃ 70, maximum 74. Class Q: minimum 40, Q₁ 51, median 58, Q₃ 65, maximum 76.

(a) Compare the two distributions. 4 marks

(b) One class's marks may be normally distributed. Identify which, and justify your answer. 2 marks

12In one breath

A frequency table counts values or classes; continuous classes are inequalities with no gaps, such as 25 ≤ t < 30. A histogram with equal widths shows shape: bars touching, frequency as height, the tallest bar the modal class, and a long tail to the right meaning positive skew. Cumulative frequency is a running total, plotted at upper class boundaries from a starting point of zero; go across from n ÷ 2 for the median, n ÷ 4 and 3n ÷ 4 for the quartiles, k% of n for the kth percentile, then down, and subtract from the total for "more than". A box plot draws minimum, Q₁, median, Q₃ and maximum on a scale, with whiskers stopping at the last non-outlier and outliers as crosses. Compare two box plots on centre and spread, with numbers and context, and read a symmetric box and equal whiskers as a sign the data may be normal.


Answers

Q1. (a) 3, 12, 29, 42, 48, 50. A2 for all correct, A1 for at least four correct. (b) The median is the 25th length; 12 lengths are below 30 cm and 29 below 35 cm, so it lies in 30 ≤ L < 35. A1. (c) (20, 0) and (25, 3). A1 for both. (22.5, 3) scores 0, because cumulative frequency is plotted at the upper boundary.

Q2. (a) n = 11, so the median is the 6th value, 3. Lower half 0, 1, 1, 2, 3 gives Q₁ = 1; upper half 4, 5, 6, 8, 19 gives Q₃ = 6. A1 for each. (b) IQR = 6 − 1 = 5, and 1.5 × 5 = 7.5. Upper fence = 6 + 7.5 = 13.5. Since 19 > 13.5, 19 is an outlier. M1 for 1.5 × IQR added to Q₃, A1 for 13.5 with the comparison 19 > 13.5. This is a "show that": the comparison must be written. (c) 0 and 8. A1. A right whisker ending at 19 scores 0.

Q3. Readings from the graph; answers within about 0.2 of these are accepted. (a) Median: across from 60, 6.6 hours. Q₁: across from 30, 4.7 hours. Q₃: across from 90, 8.4 hours. IQR ≈ 8.4 − 4.7 = 3.7 hours. A1 for the median, M1 for reading at 30 and 90, A1 for both quartiles, A1 for the IQR. (b) 0.8 × 120 = 96, and across from 96 gives 8.9 hours. M1 for 96, A1 for 8.9. (c) Up from 9 hours gives a cumulative frequency of 97, so 120 − 97 = 23 students. M1 for reading at 9 hours and subtracting from 120, A1 for 23. An answer of 97 scores M0 A0.

Q4. (a) Centre: Class P's median (66) is higher than Class Q's (58), so students in P typically scored more. Spread: P's IQR is 70 − 58 = 12, smaller than Q's 65 − 51 = 14, so the middle half of P's marks are more consistent, and P's range (30) is also smaller than Q's (36). A1 for a correct comparison of medians with values, R1 for its meaning in context, A1 for a correct comparison of IQRs (or ranges) with values, R1 for its meaning in context. (b) Class Q. Its median is in the centre of the box (58 − 51 = 7 and 65 − 58 = 7) and its whiskers are equal (11 each), so the distribution is symmetric, as a normal distribution is. Class P's left whisker (14) is far longer than its right (4) and its median sits in the right-hand part of the box, so its marks are negatively skewed. A1 for Q, R1 for a reason based on symmetry of the box and whiskers.


Educerie · written from the published IB Diploma Programme Mathematics: analysis and approaches guide, first assessment 2021, section 4.2 Presenting data. Original text, examples and questions. Diagrams drawn by Educerie. Last reviewed 25 September 2026.

Mocks: in the future, hold tight!