Educerie
Level

Educerie · IB Diploma · Mathematics: analysis and approaches

Topic 4 Statistics and probability · 4.3 Measures of centre and spread

Level
SL and HL. Nothing here is HL only, so every section is examinable for both.
Themes (key concepts)
quantity, approximation, generalization. One number for the centre and one for the spread stand in for a whole data set; from grouped data they are approximations, and the rules for constant changes generalize one calculation to every rescaled version of the data.
The question this unit answers
if you could keep only two numbers about a data set, one for where it sits and one for how spread out it is, which should they be, and how do you find them?
Where it is examined
both papers, in nearly every statistics question. Paper 1 (no calculator): a mean, median or mode from a short list or a frequency table, an unknown value or frequency from a given mean, and the effect of adding to or multiplying every value (3 to 6 marks). Paper 2: mean and standard deviation from raw, frequency or grouped data on the GDC, variance, and comparisons (2 to 6 marks). Standard deviation reappears in the normal distribution (4.9), so this page is groundwork for that one.

What you must be able to do

You must be able toLevelWhat it looks like in the exam
Find the mean, median and mode of a data setSL, HL"Find the mean number of goals" (2 marks), Paper 1 by hand or Paper 2 by GDC
Use x̄ = Σfx ÷ n for a frequency tableSL, HL"Given that the mean is 1.8, find the value of k" (3 to 4 marks)
Estimate the mean of grouped data using mid-interval valuesSL, HL"Write down the mid-interval value… Hence estimate the mean" (3 marks)
Write down the modal class (equal class intervals)SL, HL"Write down the modal class" (1 mark)
Find the standard deviation and variance with technologySL, HL"Find the standard deviation" (2 marks), Paper 2
Find the range and interquartile range; find quartiles of discrete data with technologySL, HL"Find the IQR" (2 marks)
Know why hand and technology quartiles can differSL, HLa 1-mark comment, or accepting your GDC's value
Find the new mean, standard deviation and variance after a constant change to the dataSL, HL"Each value is doubled and 5 is subtracted. Find the new mean and standard deviation" (3 to 4 marks), Paper 1

Before you start

You need the mean, median and mode of a short list from earlier years, the frequency tables and quartiles of 4.2, and the outlier rule from 4.1. The formula booklet gives the mean of a frequency table, x̄ = Σfx ÷ n with n = Σf, and IQR = Q₃ − Q₁. Remember the SL convention from 4.1: a data set in a question is the population unless it says otherwise.


1The idea in one paragraph

A measure of central tendency says where the data sit: the mean (add up, divide by how many), the median (the middle value in order) or the mode (the most common value). A measure of dispersion says how spread out they are: the range, the interquartile range, or the standard deviation, which is roughly the typical distance of a value from the mean. The mean and standard deviation use every value, so they are sensitive to outliers; the median and IQR are not. Your GDC produces all of them at once, and the marks are in entering the data correctly, choosing the right output, and saying what the numbers mean. Finally, add a constant to every value and the centre moves but the spread does not; multiply every value and both scale.

2Mean, median and mode

For five test scores 4, 7, 8, 9, 12:

  • the mean is x̄ = (4 + 7 + 8 + 9 + 12) ÷ 5 = 40 ÷ 5 = 8;
  • the median is the middle value once the data are in order: 8;
  • there is no mode, since no value repeats.

The mean has a physical meaning worth holding on to: it is the balance point. Put a weight at each value on a ruler, and the ruler balances on the mean. That is because the deviations from the mean, x − x̄, add up to zero: the values below pull left exactly as hard as the values above pull right. Figure 1 shows it.

Figure 1 · The mean is the balance point Figure 1 · The mean is the balance point 2 3 4 5 6 7 8 9 10 11 12 13 14 Value −4 −1 +1 +4 mean = 8 Deviations from the mean 8 are −4, −1, 0, +1, +4. They always add to zero.
Figure 1 · The mean is the balance point

For the median of n ordered values, count to position (n + 1) ÷ 2. With n odd that is a single value. With n even, it lands halfway between two values, and the median is their mean: for 3, 5, 6, 8, 9, 11 it is (6 + 8) ÷ 2 = 7.

Which average? The mean uses every value, which is its strength and its weakness. Nine employees of a small firm earn these monthly amounts in euros (invented data): 2100, 2200, 2300, 2300, 2400, 2600, 2800, 3100, 9500. Figure 2 plots them.

Figure 2 · One large salary pulls the mean, not the median Figure 2 · One large salary pulls the mean, not the median 2000 3000 4000 5000 6000 7000 8000 9000 10000 Monthly pay (€) mode median mean Monthly pay at a small firm (€, invented). Mode 2300, median 2400, mean about 3256.
Figure 2 · One large salary pulls the mean, not the median

The mean is 29 300 ÷ 9 ≈ €3256, higher than eight of the nine salaries. The median is €2400 and the mode €2300. Remove the 9500 and the mean drops to €2475 while the median barely moves, to €2350. So:

  • use the mean for data that are roughly symmetric with no outliers; it uses all the information;
  • use the median when the data are skewed or have outliers, because it describes a typical value;
  • use the mode for the most common value, which is the only average that makes sense for, say, shoe sizes a shop should stock.

In skewed data the mean is dragged towards the tail: positive skew puts the mean above the median.

3The mean of a frequency table

When values repeat, a frequency table saves writing them out. Each value x appears f times, so it contributes f × x to the total.

x̄ = Σfx ÷ n, where n = Σf · from the formula booklet.

Thirty students' books read last month (invented data, the same as in 4.2):

x (books)012345Total
f37963230
fx071818121065
x̄ = Σfx / Σf = 65 / 30 = 2.17 (3 s.f.)
median: position (30 + 1)/2 = 15.5, between the 15th and 16th values
running totals 3, 10, 19: both the 15th and 16th are 2, so median = 2
mode = 2the value with the highest frequency, 9

The formula also runs backwards, which is a Paper 1 favourite. The goals scored by a hockey team in its matches (invented) are 0 goals 5 times, 1 goal k times, 2 goals 8 times, 3 goals 4 times and 4 goals 3 times. The mean is 1.8. Find k.

Σfx = 0(5) + 1(k) + 2(8) + 3(4) + 4(3) = k + 40
Σf = 5 + k + 8 + 4 + 3 = k + 20
(k + 40) / (k + 20) = 1.8
k + 40 = 1.8k + 36
4 = 0.8k, so k = 5

The same idea finds a missing value: if the mean of 6 numbers is 15, their total is 6 × 15 = 90. Mean times count gives the total, and totals can be added and subtracted where means cannot.

4Grouped data: an estimate of the mean, and the modal class

Once data are grouped, the individual values are gone, so the mean can only be estimated. The guide's method is to treat every value in a class as if it sat at the mid-interval value, the midpoint of the class: (lower boundary + upper boundary) ÷ 2. Then use Σfx ÷ n with those midpoints as x. Figure 3 shows the idea for the park-run times from 4.2.

Figure 3 · Grouped data: every value treated as the class midpoint Figure 3 · Grouped data: every value treated as the class midpoint 15 20 25 30 35 40 45 50 Time t (minutes) 0 4 8 12 16 20 24 Frequency 17.5 22.5 27.5 32.5 37.5 42.5 47.5 estimated mean ≈ 31.4 Estimated mean = Σfx ÷ n = 2515 ÷ 80 ≈ 31.4 minutes, using x = 17.5, 22.5, …, 47.5.
Figure 3 · Grouped data: every value treated as the class midpoint
Time t (min)15 ≤ t < 2020 ≤ t < 2525 ≤ t < 3030 ≤ t < 3535 ≤ t < 4040 ≤ t < 4545 ≤ t < 50
Mid-interval x17.522.527.532.537.542.547.5
f41122191374
Σfx = 4(17.5) + 11(22.5) + 22(27.5) + 19(32.5) + 13(37.5) + 7(42.5) + 4(47.5)
= 70 + 247.5 + 605 + 617.5 + 487.5 + 297.5 + 190 = 2515
estimated mean = 2515 / 80 = 31.4 minutes (3 s.f.)

It is an estimate because the runners in each class are not really all at the midpoint; the hope is that those above and below it roughly cancel. On Paper 2 you enter the midpoints as a list and the frequencies as a second list, and the GDC does the sum.

The modal class is the class with the highest frequency, here 25 ≤ t < 30. The guide uses it only for equal class intervals, as here.

5Measures of dispersion

Two data sets can share a mean and still look nothing alike. In Figure 4, sets A and B both have mean 8, but B's values are far more spread out.

Figure 4 · Same mean, different standard deviation Figure 4 · Same mean, different standard deviation 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 Value A: σ ≈ 1.29 B: σ ≈ 3.96 mean 8 Both sets have mean 8. B's values sit further from 8 on average, so its σ is about three times A's.
Figure 4 · Same mean, different standard deviation
  • The range is largest − smallest. Quick, but it depends on the two most extreme values only, so one outlier can double it.
  • The interquartile range, IQR = Q₃ − Q₁, is the spread of the middle half. It ignores the extremes, so it pairs naturally with the median.
  • The standard deviation, σ, measures the typical distance of the values from the mean. It uses every value, so it pairs with the mean.
  • The variance is the square of the standard deviation, σ². It is in squared units, which is why the standard deviation is the one you interpret.

The guide says you find σ and σ² with technology, but one calculation by hand shows what the number is. For 4, 7, 8, 9, 12 with mean 8:

deviations x − μ: −4, −1, 0, 1, 4they add to 0, so square them
squared: 16, 1, 0, 1, 16, total 34
variance σ2 = 34 / 5 = 6.8mean of the squared deviations
standard deviation σ = √6.8 = 2.61 (3 s.f.)

So σ = √( Σ(x − μ)² ÷ n ): square each distance from the mean, average the squares, then take the square root to get back to the original units. Squaring stops the positive and negative deviations cancelling, and it makes values far from the mean count for more. For the sets in Figure 4, σ ≈ 1.29 for A and σ ≈ 3.96 for B.

On the GDC. Enter the values in a list (and frequencies in a second list if you have a table), then run one-variable statistics: on a TI-84 it is STAT, CALC, 1-Var Stats, with the list and, if needed, the frequency list; on a Casio it is the Statistics menu, CALC, 1-VAR, after setting the list and frequency in SET. The screen gives x̄, Σx, n, the quartiles and the median, and two standard deviations:

  • σx divides by n. This is the one to use in AA, because the data set is treated as the population.
  • sx divides by n − 1. It is used when a sample is meant to estimate a larger population's spread, and AA questions do not ask for it.

For 4, 7, 8, 9, 12 the GDC shows σx = 2.61 and sx = 2.92. Writing 2.92 loses the accuracy mark. For the books table σx = 1.34, and for the grouped park-run times, with midpoints and frequencies entered, σx ≈ 7.40 minutes, an estimate for the same reason the mean was.

Interpreting σ. Always in context: "the times are typically about 7.4 minutes either side of the mean of 31.4 minutes". A larger standard deviation means less consistency. When comparing two groups, compare means for centre and standard deviations for spread, with both numbers quoted.

6Constant changes to every value

If every value in a data set is changed in the same way, you do not need to recalculate from scratch. Think about what happens to the picture, as in Figure 5.

Figure 5 · Adding shifts the data; multiplying stretches it Figure 5 · Adding shifts the data; multiplying stretches it 0 2 4 6 8 10 12 14 16 18 20 22 24 26 Value x mean 8, σ ≈ 2.61 x + 3 mean 11, σ ≈ 2.61 2x mean 16, σ ≈ 5.22 Add 3: mean +3, σ unchanged. Double: mean × 2 and σ × 2, so the variance is × 4.
Figure 5 · Adding shifts the data; multiplying stretches it

Adding or subtracting a constant b slides every value the same distance. The whole data set moves, so every measure of centre moves by b. The gaps between values do not change, so every measure of spread stays the same.

Multiplying by a constant a (a > 0) stretches the data away from zero. Every value, every measure of centre, and every gap is multiplied by a, so the spread is multiplied by a too. The variance, a squared quantity, is multiplied by a².

Every value x becomesMean, median, mode, quartilesRange, IQR, standard deviationVariance
x + badd bunchangedunchanged
axmultiply by amultiply by amultiply by a²
ax + bmultiply by a, then add bmultiply by amultiply by a²

Adding moves the centre only. Multiplying scales both centre and spread; the variance scales by the square.

The guide's own two cases: subtract 3 from every value and the mean falls by 3 while the standard deviation is unchanged; double every value and both the mean and the standard deviation double.

Worked example. A week of noon temperatures in a town has mean 15 °C and standard deviation 4 °C. Convert to Fahrenheit with F = 1.8C + 32.

new mean = 1.8 × 15 + 32 = 59 °F
new standard deviation = 1.8 × 4 = 7.2 °Fthe + 32 moves nothing apart
new variance = 7.22 = 51.84 (°F)2or 1.82 × 16

A multiplier can be negative, for example if every value is subtracted from 100 (y = 100 − x). The mean becomes 100 − x̄, and the standard deviation is multiplied by |−1| = 1, because a spread cannot be negative.

7Quartiles of discrete data, and why methods disagree

For a short list, most GDCs find quartiles like this: put the data in order, find the median, split the data into a lower half and an upper half leaving the median out when n is odd, and take the median of each half. With n even there is no middle value to leave out, and the halves are simply the lower and upper n ÷ 2 values.

The guide asks you to know that different methods exist. Another common one keeps the median in both halves when n is odd, and spreadsheet functions use yet another rule that interpolates between values. Figure 6 applies the first two methods to nine values.

Figure 6 · Two methods, two sets of quartiles Figure 6 · Two methods, two sets of quartiles Median left out (most GDCs) 3 5 6 8 9 11 14 20 22 Q₁ = 5.5 Q₃ = 17 median 9 Median in both halves 3 5 6 8 9 11 14 20 22 Q₁ = 6 Q₃ = 14 median 9 Same nine values. Leaving the median out of the halves gives 5.5 and 17; putting it in gives 6 and 14.
Figure 6 · Two methods, two sets of quartiles
data: 3, 5, 6, 8, 9, 11, 14, 20, 22n = 9, median = 9
median left out: Q1 = (5 + 6)/2 = 5.5, Q3 = (14 + 20)/2 = 17, IQR = 11.5
median included: Q1 = 6, Q3 = 14, IQR = 8

Neither is wrong; they are different conventions, and the difference shrinks as the data set grows. In an exam, the GDC's answer is accepted, and a hand answer by a sensible method usually is too. Use one method consistently within a question.

8Where marks are lost

Writing sx instead of σx. In AA the data set is the population, so the standard deviation is σx. The sx value is larger and scores A0.

Dividing by the number of rows. For a frequency table the mean is Σfx ÷ Σf, not Σfx ÷ 6 because there are six values of x.

Forgetting that a grouped mean is an estimate. Use mid-interval values and say "estimate". Using the lower boundaries of each class gives a mean that is too small.

Adding a constant to the standard deviation. Adding b to every value leaves σ alone. Only the multiplier touches the spread.

Multiplying the variance by a instead of a². Doubling every value multiplies σ by 2 and σ² by 4.

Reporting the mean of skewed data as "typical". With outliers or strong skew, the median describes a typical value; say which average you chose and why.

Using the modal class as the mode. The modal class is an interval, 25 ≤ t < 30, not a single number.

Averaging means of groups of different sizes. The mean of a class of 10 with mean 60 and a class of 30 with mean 80 is not 70. Go through totals: (600 + 2400) ÷ 40 = 75.

9Work it right

  1. On Paper 2, write the list you entered in words ("midpoints 17.5, 22.5, … with frequencies…") so a slip in entry can still earn the method mark.
  2. Copy the needed outputs: x̄, σx, n, and check that n is the number of values you expected.
  3. For a frequency table by hand, add an fx row and show Σfx and Σf.
  4. For grouped data, write the midpoints first, then Σfx ÷ n, and call the answer an estimate.
  5. Give answers to 3 significant figures unless exact, with units.
  6. Variance is σ²: square the unrounded standard deviation, or use the GDC's value.
  7. For constant changes, state the rule you are using ("adding does not change σ") before the number.
  8. Interpret every measure in context in one sentence.

10Try it

Marks in brackets. Q1 to Q3 are Paper 1 style, no calculator. Q4 and Q5 are Paper 2 style.

Q1. The data set 5, 9, 4, 12, 9, 7, 10 is given.

(a) Find the mean, the median, the mode and the range. 4 marks

(b) Each value is increased by 3. Write down the new mean and the new range. 2 marks

Q2. The mean of 8 numbers is 12. A ninth number is added and the mean of the 9 numbers is 13. Find the ninth number. 3 marks

Q3. The number of goals a water-polo team scored in each quarter of its season is shown (invented data).

Goals01234
Quarters4k952

(a) The mean number of goals is 1.75. Find k. 3 marks

(b) Find the median number of goals. 2 marks

Q4. The daily screen time of 60 teenagers is summarised below (invented data).

Screen time h (hours)0 ≤ h < 22 ≤ h < 44 ≤ h < 66 ≤ h < 88 ≤ h < 10
Frequency51421137

(a) Write down the modal class. 1 mark

(b) Find an estimate of the mean screen time. 2 marks

(c) Find an estimate of the standard deviation. 2 marks

Q5. The prices of the items in an online shop have mean €24.50 and standard deviation €6.20. Every price is increased by 10%, and then a fixed packaging charge of €2 is added to each.

(a) Find the new mean price. 2 marks

(b) Find the new standard deviation and the new variance. 3 marks

11In one breath

The mean is the total divided by the count, x̄ = Σfx ÷ Σf for a table, and it is the balance point of the data; the median is the middle value in order; the mode is the most common. For grouped data, estimate the mean from mid-interval values and name the modal class. The range is largest minus smallest, the IQR is Q₃ − Q₁, and the standard deviation σ is the typical distance from the mean, found on the GDC as σx, not sx, with variance σ². Mean and σ use every value and suit symmetric data; median and IQR resist outliers and suit skewed data. Add b to every value and the centre moves by b while the spread stays; multiply by a and centre and σ scale by a while the variance scales by a². Quartiles by hand and by GDC can differ, because more than one method exists.


Answers

Q1. (a) Ordered: 4, 5, 7, 9, 9, 10, 12. Mean = 56 ÷ 7 = 8. Median = 4th value = 9. Mode = 9. Range = 12 − 4 = 8. A1 for each. (b) New mean = 11; new range = 8. A1 for each. Recalculating from the new data is accepted; the range must not change.

Q2. Total of 8 numbers = 8 × 12 = 96. Total of 9 numbers = 9 × 13 = 117. Ninth number = 117 − 96 = 21. M1 for finding either total, M1 for subtracting totals, A1 for 21.

Q3. (a) Σfx = 0(4) + 1(k) + 2(9) + 3(5) + 4(2) = k + 41 and Σf = k + 20, so (k + 41) ÷ (k + 20) = 1.75, giving k + 41 = 1.75k + 35, so 0.75k = 6 and k = 8. M1 for Σfx in terms of k, M1 for setting their Σfx ÷ Σf equal to 1.8, A1 for k = 5. (b) n = 28, so the median lies between the 14th and 15th values. Running totals are 4, 12, 21, so both are 2 and the median is 2 goals. M1 for identifying the 14th and 15th values, A1 for 2.

Q4. (a) 4 ≤ h < 6. A1. (b) Midpoints 1, 3, 5, 7, 9 with frequencies 5, 14, 21, 13, 7. Σfx = 306, so the estimated mean = 306 ÷ 60 = 5.1 hours. M1 for using mid-interval values with the frequencies, A1 for 5.1. (c) From the GDC with the same lists, σx = 2.23 hours (3 s.f.). M1 for evidence of midpoints and frequencies entered, A1 for 2.23. The sample value 2.25 scores M1 A0.

Q5. Each new price is y = 1.1x + 2. (a) New mean = 1.1 × 24.50 + 2 = €28.95. M1 for multiplying by 1.1 and adding 2, A1. (b) New standard deviation = 1.1 × 6.20 = €6.82; the €2 has no effect on spread. New variance = 6.82² = 46.5 (3 s.f.), that is 46.5124. M1 for multiplying σ by 1.1 only, A1 for 6.82, A1 for the variance. Adding 2 to the standard deviation scores M0.


Educerie · written from the published IB Diploma Programme Mathematics: analysis and approaches guide, first assessment 2021, section 4.3 Measures of centre and spread. Original text, examples and questions. Diagrams drawn by Educerie. Last reviewed 25 September 2026.

Mocks: in the future, hold tight!