Educerie
Level

Educerie · IB Diploma · Mathematics: analysis and approaches

Topic 4 Statistics and probability · 4.1 Populations, samples and sampling

Level
SL and HL. Nothing here is HL only, so every section is examinable for both.
Themes (key concepts)
validity, approximation, generalization, quantity. A sample is an approximation to a population, and whether a generalization from it is valid depends on how the sample was chosen and how carefully each quantity was recorded.
The question this unit answers
before you calculate anything from a set of data, how do you decide whether the data deserve to be trusted?
Where it is examined
Paper 1 and Paper 2, almost always as the opening parts of a longer statistics question: classify the data (1 mark), name or describe a sampling technique (1 to 2 marks), work out the numbers in a stratified sample (2 to 3 marks), test a value for being an outlier (3 to 4 marks) and say whether it should stay (1 to 2 marks). The outlier test with given quartiles is Paper 1 arithmetic; with raw data it is usually Paper 2, with the quartiles from your GDC. The same ideas decide whether the data in your exploration are fit to use.

What you must be able to do

You must be able toLevelWhat it looks like in the exam
Say what the population and the sample are in a study, and what makes a sample randomSL, HL"State the population in this investigation" (1 mark)
Classify data as discrete or continuousSL, HL"Write down whether the data are discrete or continuous" (1 mark)
Judge the reliability of a data source and identify bias in a sampleSL, HL"Explain why the results of this survey may be biased" (2 marks)
Deal with missing values and errors in the recording of dataSL, HL"One value was recorded incorrectly. Find the mean of the corrected data" (2 to 3 marks), Paper 2
Test whether a value is an outlier using 1.5 × IQRSL, HL"Determine whether 58 is an outlier. Justify your answer" (3 to 4 marks)
Interpret an outlier: a valid part of the data, or an errorSL, HL"Suggest whether this value should be removed" (1 to 2 marks)
Describe simple random, convenience, systematic, quota and stratified sampling, and judge how effective each isSL, HL"Identify the sampling technique" (1 mark); "Find how many juniors should be in the stratified sample" (2 marks)

Before you start

You need percentages and proportions, and the idea of a median and quartiles from earlier years. Quartiles are taught properly in 4.2 and 4.3; on this page you only use them. The formula booklet gives IQR = Q₃ − Q₁. The outlier rule itself is one line, and you should learn it by heart.


1The idea in one paragraph

Statistics answers questions about a group, the population, by measuring some of its members, the sample. That trade is only worth making if the sample behaves like the population, and it only does that if chance, not a person's convenience, chose it. So before any mean or graph, ask four things: who the data are about, how the members were chosen, what kind of values were recorded, and whether anything in the list is missing, impossible or extreme. Random methods (simple random, systematic, stratified) protect you against bias; non-random ones (quota, convenience) invite it. A value far from the rest, an outlier, is tested with a precise rule, then judged in context: it may be the most interesting fact in the data, or a typing slip.

2Populations and samples

The population is the whole set of people or things you want to draw a conclusion about. The question defines it, not convenience: if a head teacher asks how long her students sleep, the population is every student in the school, not the students in one tutor group.

Measuring every member is a census. Sometimes it is possible, but usually it is too slow, too expensive, or impossible because measuring destroys the item: a factory that tests every match by striking it has nothing left to sell. So we measure a sample, a subset of the population, and use what we find to estimate the same quantity for the whole population. Figure 1 shows the trade.

Figure 1 · A sample stands in for the population Figure 1 · A sample stands in for the population Population: every student in the school Sample: chosen at random select Statistic from the sample mean sleep 7.1 hours a night estimate for the population Measure the sample, then use what you find to estimate the same quantity for the population.
Figure 1 · A sample stands in for the population

A random sample is one chosen by a process in which chance alone decides who is included, so that every member of the population has the same chance of being chosen. Randomness matters because any human choice lets preferences leak in: a teacher asked to "pick a few typical students" will pick the ones she thinks of first, and those are not typical. A random sample does not guarantee a perfect match with the population; an unlucky draw can happen. What it guarantees is that the errors come from chance rather than from a built-in lean, and chance errors shrink as the sample grows.

One exam convention to know now: at SL, the data set in a question is treated as the population unless the question says otherwise. This matters in 4.3, when your GDC offers two standard deviations and you must pick the right one.

3Discrete and continuous data

Discrete data can take only particular, separate values, usually because they are counted: the number of siblings, goals in a match, cars in a car park. Shoe sizes (7, 7½, 8) are discrete too, even though they are not whole numbers, because nothing lies between size 7 and size 7½.

Continuous data are measured, and can take any value within a range: time, mass, length, temperature. The value you record is rounded to whatever your instrument can show, so a sprint time written as 13.47 s really means "somewhere between 13.465 s and 13.475 s". Figure 2 puts the two side by side.

Figure 2 · Discrete values are separate; continuous values fill a range Figure 2 · Discrete values are separate; continuous values fill a range 0 1 2 3 4 5 6 Discrete: number of siblings only these values are possible; 2.4 siblings is not 11 12 13 14 15 16 17 Continuous: time to run 100 m (seconds) 13.47 s, recorded to the nearest 0.01 s Between two continuous values there is always another. Between two discrete values there may be none.
Figure 2 · Discrete values are separate; continuous values fill a range

The quick test: between any two possible values, is there always another possible value? If yes, the data are continuous. Two cases need a sentence of judgement. Age is continuous (you grow older every second), but "age in completed years" is recorded like discrete data. Money is strictly discrete, since it moves in cents, but large amounts such as salaries are usually treated as continuous. If a question hinges on it, say which you are assuming and why.

The distinction decides how the data are presented. Continuous data are grouped into class intervals written as inequalities with no gaps, such as 20 ≤ t < 25, so every value has exactly one home. That is 4.2.

4Reliability of data sources, and bias in sampling

A data source is reliable when you have good reason to believe its figures describe what they claim to. Figure 3 gives the six questions to ask of any data set, in an exam or in a newspaper.

Figure 3 · Six questions to ask before you trust a data set Figure 3 · Six questions to ask before you trust a data set 1 Who collected it, and why? a reason to want one answer? 2 What is the population? who are the results about? 3 How was the sample chosen? random, or whoever was easy? 4 How big, and who replied? size and non-response 5 How was it measured or asked? instrument, units, wording 6 What is missing or odd? gaps, impossible values, outliers If any answer is "we do not know", say so, and say how it could bias the result.
Figure 3 · Six questions to ask before you trust a data set

Bias is a systematic tendency for the sample to differ from the population in one direction. It is not bad luck, and that is why a bigger sample does not cure it: a biased method applied to 10,000 people gives a very precise wrong answer. The common sources are these.

  • Selection bias. The list the sample is drawn from, the sampling frame, leaves some groups out. An online survey cannot reach people without internet access.
  • Non-response bias. The people who reply differ from those who do not. Busy people rarely fill in long questionnaires.
  • Self-selection. Volunteers choose themselves, so only people with strong views take part in a phone-in poll.
  • Response bias. The question or the setting pushes the answer: "Do you agree the canteen is overpriced?" invites yes, and people overstate how much exercise they take.

The guide points to a famous real case. Before the 1936 US presidential election, the magazine The Literary Digest mailed about ten million straw-poll ballots to names taken largely from telephone directories, car registrations and its own subscriber list. More than two million came back, and they pointed to a clear win for Alf Landon. Franklin Roosevelt won in a landslide. In 1936 telephones and cars were far more common in wealthier households, which leaned towards Landon, and the people who bothered to return a ballot were not typical either. George Gallup, with a far smaller but better-chosen sample, called the winner correctly. Two million responses could not rescue a biased frame.

A modern version: the city of Boston tried a smartphone app, Street Bump, that logged jolts as drivers went over potholes. Reports could only come from people who owned smartphones and installed the app, so there was a real risk of repair crews being sent disproportionately to better-off districts. The data were plentiful, and still not representative.

5Missing data and errors in recording

Real data arrive with gaps and mistakes. Two rules handle almost everything.

A missing value is left out, never written as 0. An impossible or clearly mistyped value is corrected from the original record if you can, removed if you cannot, and you say what you did.

A PE teacher records twelve students' 400 m times in seconds (invented data): 72.4, 68.1, 75.0, 81.3, 0, 70.6, 7.25, 77.9, 69.8, 74.2, 73.5, 79.5. The 0 is a student who was absent. The 7.25 is impossible for 400 m (a world-class time is over 40 s), and looks like 72.5 with the decimal point in the wrong place.

mean of all 12 as recorded = 749.55 / 12 = 62.5 sdragged down by 0 and 7.25
remove both: 742.3 / 10 = 74.2 sn is now 10, not 12
if the record confirms 72.5: 814.8 / 11 = 74.1 sn = 11

The raw mean of 62.5 s describes no student in the class. Either cleaned figure is honest, provided you state that n changed and why. Never "fix" a value by guessing without evidence: if the original sheet is lost, remove the 7.25 rather than invent a replacement.

6Outliers: the rule, then the judgement

An outlier is a data item more than 1.5 × IQR from the nearest quartile, where the interquartile range IQR = Q₃ − Q₁ is the width of the middle half of the data.

x is an outlier if x < Q₁ − 1.5 × IQR or x > Q₃ + 1.5 × IQR.

The two values Q₁ − 1.5 × IQR and Q₃ + 1.5 × IQR are called the fences. Everything between them is ordinary.

Worked example. Fifteen employees of a small firm record their commuting times in minutes (invented data), already in order:

12, 14, 15, 17, 18, 19, 20, 21, 22, 22, 24, 25, 27, 30, 58

With n = 15 the median is the 8th value, 21. The lower half is the seven values below it, and its median is Q₁; the upper half is the seven values above it, and its median is Q₃. This is the method most GDCs use.

Q1 = 17, Q3 = 254th and 12th values
IQR = 25 − 17 = 8
1.5 × IQR = 12
lower fence = 17 − 12 = 5
upper fence = 25 + 12 = 37
58 > 37, so 58 is an outlier; 12 > 5, so nothing is an outlier at the low end

Figure 4 draws the fences. Notice that they are measured from the quartiles, not from the median.

Figure 4 · The outlier fences for the commuting times Figure 4 · The outlier fences for the commuting times 0 5 10 15 20 25 30 35 40 45 50 55 60 65 Commuting time (minutes) lower fence 5 upper fence 37 Q₁ = 17 Q₃ = 25 1.5 × IQR = 12 12 58 > 37: outlier IQR = 25 − 17 = 8. The fences sit 1.5 × 8 = 12 beyond each quartile, at 5 and 37. Only 58 lies outside.
Figure 4 · The outlier fences for the commuting times

The rule tells you a value is unusual. It does not tell you what to do with it. That needs context.

  • A valid outlier is a genuine member of the population. The 58-minute commuter probably lives in the next town. Removing that person would make the firm's commuting look shorter than it is, so keep the value and mention it.
  • An error is a value that could not be true, or that has a clear recording explanation: 580 minutes, a height of 17.2 m, a mass typed in grams among masses in kilograms. Correct it from the source, or remove it and say so.

Outliers matter because they pull some measures much harder than others. With the 58 the mean commute is 22.9 minutes; without it, 20.4 minutes. The median moves only from 21 to 20.5. You will use this in 4.3 when choosing which average to report, and in 4.2, where a box-and-whisker diagram marks each outlier with a cross.

7Five sampling techniques

Take one invented setting for all five. A tennis club has 480 members: 144 juniors, 264 adults and 72 seniors. The committee wants the views of 40 members on new opening hours.

Simple random sampling. Number the members 1 to 480, then let a random number generator choose 40 different numbers, ignoring repeats. Every possible group of 40 is equally likely to be chosen. It needs a complete list of the population, and by bad luck it can under-represent a small group such as the seniors.

Systematic sampling. Take the list in some order and choose every kth member, where k = population size ÷ sample size.

k = 480 / 40 = 12
random start between 1 and 12, say 5
members 5, 17, 29, 41, …, 4735 + 39 × 12 = 473, the 40th member

It is quick and spreads the sample through the whole list. Its weakness is a list with a repeating pattern: if the club's list happened to run in blocks of 12 by court booking slot, every chosen member could come from the same slot. Only the start is random, so it is not a simple random sample: members 5 and 6 can never both be chosen.

Stratified sampling. Split the population into strata, non-overlapping groups that might answer differently (here, the age groups), then take a random sample from each stratum in proportion to its size.

juniors: (144 / 480) × 40 = 12
adults: (264 / 480) × 40 = 22
seniors: (72 / 480) × 40 = 6
check: 12 + 22 + 6 = 40

Within each stratum the members are chosen at random, for example by simple random sampling. When the proportions do not come out whole, round to the nearest whole number and check the total. A college with 680 first-year and 560 second-year students wanting a sample of 50 gets 27.4 and 22.6, so 27 and 23, which total 50. Stratified sampling is usually the most representative method, because every group appears in its true proportion. The cost is that you must know which stratum each member is in and how big each stratum is.

Quota sampling. The collector is given a quota for each group, 12 juniors, 22 adults and 6 seniors, and fills it with whoever is available until each quota is met. It looks like stratified sampling, and the difference is the whole point: in a quota sample, the people within each group are not chosen at random. The adults found at the club at 10 a.m. on a Tuesday are not typical adults. Quota sampling is quick and needs no list, but the collector's choices bring bias.

Convenience sampling. Ask whoever is easiest to reach, such as the first 40 members through the door on Saturday. It is fast and cheap, and it is almost always biased, because the people who are easy to reach share habits. Figures 5 and 6 show all five methods on the same group of 40.

Figure 5 · Three random methods, 10 chosen from 40 Figure 5 · Three random methods, 10 chosen from 40 Simple random any 10 equally likely Systematic every 4th, random start at 3 3 7 11 15 19 23 27 31 35 39 Stratified 6 juniors and 4 adults, each at random Teal: 24 juniors. Amber: 16 adults. Ringed: chosen. Chance does the choosing in all three.
Figure 5 · Three random methods, 10 chosen from 40
Figure 6 · Two non-random methods Figure 6 · Two non-random methods Quota 6 juniors, 4 adults, first found Convenience the 10 nearest the door door Quota fixes how many from each group but lets the collector pick. Convenience takes the nearest.
Figure 6 · Two non-random methods

How effective is each? A sampling method is effective when it is likely to give a representative sample at a cost you can bear.

MethodChosen by chance?Needs a full list?Groups in true proportion?Main risk
Simple randomyesyesonly on averagea small group missed by bad luck
Systematiconly the startyesonly on averagea pattern in the list
Stratifiedyes, within stratayes, with each member's groupyesneeds detailed information; more work
Quotanonoyes, by designthe collector's choices
Conveniencenononopeople who are easy to reach are not typical

When a question asks you to "comment on the effectiveness" or "suggest an improvement", name the flaw, say which way it could push the result, and name a random method that removes it.

8Where marks are lost

Calling haphazard "random". "I stood in the corridor and asked people at random" describes a convenience sample. Random means a chance process such as a random number generator, applied to a list.

Mixing up quota and stratified. Both fix the numbers from each group. Only stratified chooses at random inside each group. Say that explicitly when you distinguish them.

Believing a large sample cures bias. Size reduces chance error, not bias. The 1936 poll had more than two million replies and still got the winner wrong.

Judging outliers by eye. "58 looks much bigger" earns nothing. Calculate IQR, then both fences, then compare, and write the comparison: 58 > 37.

Measuring 1.5 × IQR from the median. The fences sit 1.5 × IQR beyond the quartiles: Q₁ − 1.5 × IQR and Q₃ + 1.5 × IQR.

Deleting every outlier automatically. Only errors are removed. A genuine extreme value is part of the population, and the question will usually reward you for saying so.

Writing a missing value as 0. A zero is a measurement. An absence is not, and a zero drags the mean down.

Stratified numbers that do not add up. After rounding, check that the strata total the sample size.

9Work it right

  1. Name the population in a full phrase: who, where, when.
  2. For discrete or continuous, give the reason: counted or measured.
  3. For a sampling method, name it and state the feature that identifies it (a random start and fixed interval; random within strata; a quota filled by choice).
  4. For stratified numbers, write (stratum size ÷ population size) × sample size for each stratum, round, and check the total.
  5. For an outlier test, show IQR, 1.5 × IQR, both fences, and the comparison, in that order.
  6. Finish an outlier question with a sentence in context: valid and kept, or an error and removed.
  7. When you clean data, state the new n and why it changed.
  8. For bias, say which group is over- or under-represented and which way that pushes the answer.

10Try it

Marks in brackets. Q1 to Q4 are Paper 1 style, no calculator. Q5 is Paper 2 style, with a GDC.

Q1. Write down whether each of these is discrete or continuous. 2 marks

(a) the number of emails a person receives in a day

(b) the mass of a letter

(c) the number of pages in a novel

(d) the time taken to read the novel

Q2. A college has 680 first-year students and 560 second-year students. The principal takes a stratified sample of 50 students, stratified by year.

(a) Find the number of students chosen from each year. 3 marks

(b) Give one reason why a stratified sample may be better than a simple random sample here. 1 mark

Q3. The masses of 40 parcels have lower quartile 1.8 kg and upper quartile 3.4 kg. The lightest parcel has a mass of 0.2 kg and the heaviest 6.1 kg.

(a) Find the interquartile range. 1 mark

(b) Determine whether either the lightest or the heaviest parcel is an outlier. 4 marks

(c) Comment on whether any outlier found should be removed from the data. 1 mark

Q4. A gym wants to estimate how many hours a week its members exercise. The manager hands a questionnaire to the first 30 members who arrive on Monday morning.

(a) Identify the sampling technique. 1 mark

(b) Explain why the results may be biased. 2 marks

(c) The gym has a list of its 900 members. Describe how a systematic sample of 30 members could be taken from it. 2 marks

Q5. A small museum records its visitors on 11 consecutive days (invented data): 212, 198, 225, 240, 0, 231, 219, 2210, 205, 236, 228. The museum was closed on the fifth day because of a burst pipe, and the 2210 is a typing error for 221.

(a) Find the mean of the 11 values as recorded. 1 mark

(b) Explain why the 0 should be removed rather than kept. 1 mark

(c) Remove the 0 and correct the 2210. For the cleaned data, find the mean and the quartiles, and determine whether any value is an outlier. 5 marks

11In one breath

The population is everyone the question is about; the sample is the part you measure, and a random sample is chosen by chance so that every member has the same chance of selection. Counted data are discrete; measured data are continuous, recorded to the precision of the instrument. Trust a source only after asking who collected it, from whom, how, and what is missing, because bias is a lean in one direction that no sample size can cure. Leave missing values out rather than writing 0, fix or remove impossible values, and say what you did. An outlier lies below Q₁ − 1.5 × IQR or above Q₃ + 1.5 × IQR; keep it if it is genuine, remove it only if it is an error. Simple random, systematic and stratified samples are chosen by chance, stratified in proportion to each group; quota and convenience samples are not, and that is where bias gets in.


Answers

Q1. (a) Discrete, since emails are counted. (b) Continuous, since mass is measured. (c) Discrete. (d) Continuous. A2 for all four correct, A1 for three correct.

Q2. (a) First year: (680 ÷ 1240) × 50 = 27.4…, so 27. Second year: (560 ÷ 1240) × 50 = 22.6 (3 s.f.), so 23. Check: 27 + 23 = 50. M1 for the proportion of the population times 50, A1 for 27, A1 for 23. Two answers that do not total 50 lose the final A1. (b) The two years may answer differently, and stratifying guarantees each year appears in its true proportion, which a simple random sample only achieves on average. R1 for a reason that refers to the groups being represented in proportion.

Q3. (a) IQR = 3.4 − 1.8 = 1.6 kg. A1. (b) 1.5 × IQR = 2.4. Lower fence = 1.8 − 2.4 = −0.6 kg; upper fence = 3.4 + 2.4 = 5.8 kg. 0.2 > −0.6, so the lightest parcel is not an outlier. 6.1 > 5.8, so the heaviest parcel is an outlier. M1 for 1.5 × their IQR, A1 for both fences, R1 for "0.2 is not an outlier" with its comparison, R1 for "6.1 is an outlier" with its comparison. "6.1 is much bigger than the rest" scores 0. (c) A parcel of 6.1 kg is perfectly possible, so it is probably a genuine value and should be kept. R1 for a judgement in context.

Q4. (a) Convenience sampling. A1. (b) Members who arrive first on a Monday morning are not typical: they are likely to be the keenest regulars, or people free on weekday mornings, so the sample would probably overstate how much members exercise. R1 for identifying a group that is over- or under-represented, R1 for the direction of the effect on the estimate. (c) k = 900 ÷ 30 = 30. Choose a random starting number between 1 and 30, say 11, then take members 11, 41, 71, … and every 30th member after that, until 30 have been chosen. A1 for the interval 30, A1 for a random start then every 30th member.

Q5. (a) 4204 ÷ 11 = 382 visitors (3 s.f.). A1. (b) The museum was closed, so 0 is not a count of visitors on an open day; it is a missing value, and keeping it would drag the mean down. R1. (c) Cleaned data, ordered: 198, 205, 212, 219, 221, 225, 228, 231, 236, 240, with n = 10. Mean = 2215 ÷ 10 = 221.5. Median = (221 + 225) ÷ 2 = 223. Q₁ = 212, Q₃ = 231, IQR = 19. 1.5 × 19 = 28.5, so the fences are 212 − 28.5 = 183.5 and 231 + 28.5 = 259.5. The smallest value 198 and the largest 240 both lie inside, so there are no outliers. A1 for mean 221.5, A1 for both quartiles, M1 for 1.5 × their IQR, A1 for both fences, R1 for the conclusion with a comparison. Follow-through from their quartiles.


Educerie · written from the published IB Diploma Programme Mathematics: analysis and approaches guide, first assessment 2021, section 4.1 Populations, samples and sampling. Original text, examples and questions. Diagrams drawn by Educerie. Last reviewed 25 September 2026.

Mocks: in the future, hold tight!