Educerie · IB Diploma · Mathematics: analysis and approaches
Topic 4 Statistics and probability · 4.1 Populations, samples and sampling
What you must be able to do
| You must be able to | Level | What it looks like in the exam |
|---|---|---|
| Say what the population and the sample are in a study, and what makes a sample random | SL, HL | "State the population in this investigation" (1 mark) |
| Classify data as discrete or continuous | SL, HL | "Write down whether the data are discrete or continuous" (1 mark) |
| Judge the reliability of a data source and identify bias in a sample | SL, HL | "Explain why the results of this survey may be biased" (2 marks) |
| Deal with missing values and errors in the recording of data | SL, HL | "One value was recorded incorrectly. Find the mean of the corrected data" (2 to 3 marks), Paper 2 |
| Test whether a value is an outlier using 1.5 × IQR | SL, HL | "Determine whether 58 is an outlier. Justify your answer" (3 to 4 marks) |
| Interpret an outlier: a valid part of the data, or an error | SL, HL | "Suggest whether this value should be removed" (1 to 2 marks) |
| Describe simple random, convenience, systematic, quota and stratified sampling, and judge how effective each is | SL, HL | "Identify the sampling technique" (1 mark); "Find how many juniors should be in the stratified sample" (2 marks) |
Before you start
You need percentages and proportions, and the idea of a median and quartiles from earlier years. Quartiles are taught properly in 4.2 and 4.3; on this page you only use them. The formula booklet gives IQR = Q₃ − Q₁. The outlier rule itself is one line, and you should learn it by heart.
1The idea in one paragraph
Statistics answers questions about a group, the population, by measuring some of its members, the sample. That trade is only worth making if the sample behaves like the population, and it only does that if chance, not a person's convenience, chose it. So before any mean or graph, ask four things: who the data are about, how the members were chosen, what kind of values were recorded, and whether anything in the list is missing, impossible or extreme. Random methods (simple random, systematic, stratified) protect you against bias; non-random ones (quota, convenience) invite it. A value far from the rest, an outlier, is tested with a precise rule, then judged in context: it may be the most interesting fact in the data, or a typing slip.
2Populations and samples
The population is the whole set of people or things you want to draw a conclusion about. The question defines it, not convenience: if a head teacher asks how long her students sleep, the population is every student in the school, not the students in one tutor group.
Measuring every member is a census. Sometimes it is possible, but usually it is too slow, too expensive, or impossible because measuring destroys the item: a factory that tests every match by striking it has nothing left to sell. So we measure a sample, a subset of the population, and use what we find to estimate the same quantity for the whole population. Figure 1 shows the trade.
A random sample is one chosen by a process in which chance alone decides who is included, so that every member of the population has the same chance of being chosen. Randomness matters because any human choice lets preferences leak in: a teacher asked to "pick a few typical students" will pick the ones she thinks of first, and those are not typical. A random sample does not guarantee a perfect match with the population; an unlucky draw can happen. What it guarantees is that the errors come from chance rather than from a built-in lean, and chance errors shrink as the sample grows.
One exam convention to know now: at SL, the data set in a question is treated as the population unless the question says otherwise. This matters in 4.3, when your GDC offers two standard deviations and you must pick the right one.
3Discrete and continuous data
Discrete data can take only particular, separate values, usually because they are counted: the number of siblings, goals in a match, cars in a car park. Shoe sizes (7, 7½, 8) are discrete too, even though they are not whole numbers, because nothing lies between size 7 and size 7½.
Continuous data are measured, and can take any value within a range: time, mass, length, temperature. The value you record is rounded to whatever your instrument can show, so a sprint time written as 13.47 s really means "somewhere between 13.465 s and 13.475 s". Figure 2 puts the two side by side.
The quick test: between any two possible values, is there always another possible value? If yes, the data are continuous. Two cases need a sentence of judgement. Age is continuous (you grow older every second), but "age in completed years" is recorded like discrete data. Money is strictly discrete, since it moves in cents, but large amounts such as salaries are usually treated as continuous. If a question hinges on it, say which you are assuming and why.
The distinction decides how the data are presented. Continuous data are grouped into class intervals written as inequalities with no gaps, such as 20 ≤ t < 25, so every value has exactly one home. That is 4.2.
4Reliability of data sources, and bias in sampling
A data source is reliable when you have good reason to believe its figures describe what they claim to. Figure 3 gives the six questions to ask of any data set, in an exam or in a newspaper.
Bias is a systematic tendency for the sample to differ from the population in one direction. It is not bad luck, and that is why a bigger sample does not cure it: a biased method applied to 10,000 people gives a very precise wrong answer. The common sources are these.
- Selection bias. The list the sample is drawn from, the sampling frame, leaves some groups out. An online survey cannot reach people without internet access.
- Non-response bias. The people who reply differ from those who do not. Busy people rarely fill in long questionnaires.
- Self-selection. Volunteers choose themselves, so only people with strong views take part in a phone-in poll.
- Response bias. The question or the setting pushes the answer: "Do you agree the canteen is overpriced?" invites yes, and people overstate how much exercise they take.
The guide points to a famous real case. Before the 1936 US presidential election, the magazine The Literary Digest mailed about ten million straw-poll ballots to names taken largely from telephone directories, car registrations and its own subscriber list. More than two million came back, and they pointed to a clear win for Alf Landon. Franklin Roosevelt won in a landslide. In 1936 telephones and cars were far more common in wealthier households, which leaned towards Landon, and the people who bothered to return a ballot were not typical either. George Gallup, with a far smaller but better-chosen sample, called the winner correctly. Two million responses could not rescue a biased frame.
A modern version: the city of Boston tried a smartphone app, Street Bump, that logged jolts as drivers went over potholes. Reports could only come from people who owned smartphones and installed the app, so there was a real risk of repair crews being sent disproportionately to better-off districts. The data were plentiful, and still not representative.
5Missing data and errors in recording
Real data arrive with gaps and mistakes. Two rules handle almost everything.
A missing value is left out, never written as 0. An impossible or clearly mistyped value is corrected from the original record if you can, removed if you cannot, and you say what you did.
A PE teacher records twelve students' 400 m times in seconds (invented data): 72.4, 68.1, 75.0, 81.3, 0, 70.6, 7.25, 77.9, 69.8, 74.2, 73.5, 79.5. The 0 is a student who was absent. The 7.25 is impossible for 400 m (a world-class time is over 40 s), and looks like 72.5 with the decimal point in the wrong place.
The raw mean of 62.5 s describes no student in the class. Either cleaned figure is honest, provided you state that n changed and why. Never "fix" a value by guessing without evidence: if the original sheet is lost, remove the 7.25 rather than invent a replacement.
6Outliers: the rule, then the judgement
An outlier is a data item more than 1.5 × IQR from the nearest quartile, where the interquartile range IQR = Q₃ − Q₁ is the width of the middle half of the data.
x is an outlier if x < Q₁ − 1.5 × IQR or x > Q₃ + 1.5 × IQR.
The two values Q₁ − 1.5 × IQR and Q₃ + 1.5 × IQR are called the fences. Everything between them is ordinary.
Worked example. Fifteen employees of a small firm record their commuting times in minutes (invented data), already in order:
12, 14, 15, 17, 18, 19, 20, 21, 22, 22, 24, 25, 27, 30, 58
With n = 15 the median is the 8th value, 21. The lower half is the seven values below it, and its median is Q₁; the upper half is the seven values above it, and its median is Q₃. This is the method most GDCs use.
Figure 4 draws the fences. Notice that they are measured from the quartiles, not from the median.
The rule tells you a value is unusual. It does not tell you what to do with it. That needs context.
- A valid outlier is a genuine member of the population. The 58-minute commuter probably lives in the next town. Removing that person would make the firm's commuting look shorter than it is, so keep the value and mention it.
- An error is a value that could not be true, or that has a clear recording explanation: 580 minutes, a height of 17.2 m, a mass typed in grams among masses in kilograms. Correct it from the source, or remove it and say so.
Outliers matter because they pull some measures much harder than others. With the 58 the mean commute is 22.9 minutes; without it, 20.4 minutes. The median moves only from 21 to 20.5. You will use this in 4.3 when choosing which average to report, and in 4.2, where a box-and-whisker diagram marks each outlier with a cross.
7Five sampling techniques
Take one invented setting for all five. A tennis club has 480 members: 144 juniors, 264 adults and 72 seniors. The committee wants the views of 40 members on new opening hours.
Simple random sampling. Number the members 1 to 480, then let a random number generator choose 40 different numbers, ignoring repeats. Every possible group of 40 is equally likely to be chosen. It needs a complete list of the population, and by bad luck it can under-represent a small group such as the seniors.
Systematic sampling. Take the list in some order and choose every kth member, where k = population size ÷ sample size.
It is quick and spreads the sample through the whole list. Its weakness is a list with a repeating pattern: if the club's list happened to run in blocks of 12 by court booking slot, every chosen member could come from the same slot. Only the start is random, so it is not a simple random sample: members 5 and 6 can never both be chosen.
Stratified sampling. Split the population into strata, non-overlapping groups that might answer differently (here, the age groups), then take a random sample from each stratum in proportion to its size.
Within each stratum the members are chosen at random, for example by simple random sampling. When the proportions do not come out whole, round to the nearest whole number and check the total. A college with 680 first-year and 560 second-year students wanting a sample of 50 gets 27.4 and 22.6, so 27 and 23, which total 50. Stratified sampling is usually the most representative method, because every group appears in its true proportion. The cost is that you must know which stratum each member is in and how big each stratum is.
Quota sampling. The collector is given a quota for each group, 12 juniors, 22 adults and 6 seniors, and fills it with whoever is available until each quota is met. It looks like stratified sampling, and the difference is the whole point: in a quota sample, the people within each group are not chosen at random. The adults found at the club at 10 a.m. on a Tuesday are not typical adults. Quota sampling is quick and needs no list, but the collector's choices bring bias.
Convenience sampling. Ask whoever is easiest to reach, such as the first 40 members through the door on Saturday. It is fast and cheap, and it is almost always biased, because the people who are easy to reach share habits. Figures 5 and 6 show all five methods on the same group of 40.
How effective is each? A sampling method is effective when it is likely to give a representative sample at a cost you can bear.
| Method | Chosen by chance? | Needs a full list? | Groups in true proportion? | Main risk |
|---|---|---|---|---|
| Simple random | yes | yes | only on average | a small group missed by bad luck |
| Systematic | only the start | yes | only on average | a pattern in the list |
| Stratified | yes, within strata | yes, with each member's group | yes | needs detailed information; more work |
| Quota | no | no | yes, by design | the collector's choices |
| Convenience | no | no | no | people who are easy to reach are not typical |
When a question asks you to "comment on the effectiveness" or "suggest an improvement", name the flaw, say which way it could push the result, and name a random method that removes it.
8Where marks are lost
Calling haphazard "random". "I stood in the corridor and asked people at random" describes a convenience sample. Random means a chance process such as a random number generator, applied to a list.
Mixing up quota and stratified. Both fix the numbers from each group. Only stratified chooses at random inside each group. Say that explicitly when you distinguish them.
Believing a large sample cures bias. Size reduces chance error, not bias. The 1936 poll had more than two million replies and still got the winner wrong.
Judging outliers by eye. "58 looks much bigger" earns nothing. Calculate IQR, then both fences, then compare, and write the comparison: 58 > 37.
Measuring 1.5 × IQR from the median. The fences sit 1.5 × IQR beyond the quartiles: Q₁ − 1.5 × IQR and Q₃ + 1.5 × IQR.
Deleting every outlier automatically. Only errors are removed. A genuine extreme value is part of the population, and the question will usually reward you for saying so.
Writing a missing value as 0. A zero is a measurement. An absence is not, and a zero drags the mean down.
Stratified numbers that do not add up. After rounding, check that the strata total the sample size.
9Work it right
- Name the population in a full phrase: who, where, when.
- For discrete or continuous, give the reason: counted or measured.
- For a sampling method, name it and state the feature that identifies it (a random start and fixed interval; random within strata; a quota filled by choice).
- For stratified numbers, write (stratum size ÷ population size) × sample size for each stratum, round, and check the total.
- For an outlier test, show IQR, 1.5 × IQR, both fences, and the comparison, in that order.
- Finish an outlier question with a sentence in context: valid and kept, or an error and removed.
- When you clean data, state the new n and why it changed.
- For bias, say which group is over- or under-represented and which way that pushes the answer.
10Try it
Marks in brackets. Q1 to Q4 are Paper 1 style, no calculator. Q5 is Paper 2 style, with a GDC.
Q1. Write down whether each of these is discrete or continuous. 2 marks
(a) the number of emails a person receives in a day
(b) the mass of a letter
(c) the number of pages in a novel
(d) the time taken to read the novel
Q2. A college has 680 first-year students and 560 second-year students. The principal takes a stratified sample of 50 students, stratified by year.
(a) Find the number of students chosen from each year. 3 marks
(b) Give one reason why a stratified sample may be better than a simple random sample here. 1 mark
Q3. The masses of 40 parcels have lower quartile 1.8 kg and upper quartile 3.4 kg. The lightest parcel has a mass of 0.2 kg and the heaviest 6.1 kg.
(a) Find the interquartile range. 1 mark
(b) Determine whether either the lightest or the heaviest parcel is an outlier. 4 marks
(c) Comment on whether any outlier found should be removed from the data. 1 mark
Q4. A gym wants to estimate how many hours a week its members exercise. The manager hands a questionnaire to the first 30 members who arrive on Monday morning.
(a) Identify the sampling technique. 1 mark
(b) Explain why the results may be biased. 2 marks
(c) The gym has a list of its 900 members. Describe how a systematic sample of 30 members could be taken from it. 2 marks
Q5. A small museum records its visitors on 11 consecutive days (invented data): 212, 198, 225, 240, 0, 231, 219, 2210, 205, 236, 228. The museum was closed on the fifth day because of a burst pipe, and the 2210 is a typing error for 221.
(a) Find the mean of the 11 values as recorded. 1 mark
(b) Explain why the 0 should be removed rather than kept. 1 mark
(c) Remove the 0 and correct the 2210. For the cleaned data, find the mean and the quartiles, and determine whether any value is an outlier. 5 marks
11In one breath
The population is everyone the question is about; the sample is the part you measure, and a random sample is chosen by chance so that every member has the same chance of selection. Counted data are discrete; measured data are continuous, recorded to the precision of the instrument. Trust a source only after asking who collected it, from whom, how, and what is missing, because bias is a lean in one direction that no sample size can cure. Leave missing values out rather than writing 0, fix or remove impossible values, and say what you did. An outlier lies below Q₁ − 1.5 × IQR or above Q₃ + 1.5 × IQR; keep it if it is genuine, remove it only if it is an error. Simple random, systematic and stratified samples are chosen by chance, stratified in proportion to each group; quota and convenience samples are not, and that is where bias gets in.
Answers
Q1. (a) Discrete, since emails are counted. (b) Continuous, since mass is measured. (c) Discrete. (d) Continuous. A2 for all four correct, A1 for three correct.
Q2. (a) First year: (680 ÷ 1240) × 50 = 27.4…, so 27. Second year: (560 ÷ 1240) × 50 = 22.6 (3 s.f.), so 23. Check: 27 + 23 = 50. M1 for the proportion of the population times 50, A1 for 27, A1 for 23. Two answers that do not total 50 lose the final A1. (b) The two years may answer differently, and stratifying guarantees each year appears in its true proportion, which a simple random sample only achieves on average. R1 for a reason that refers to the groups being represented in proportion.
Q3. (a) IQR = 3.4 − 1.8 = 1.6 kg. A1. (b) 1.5 × IQR = 2.4. Lower fence = 1.8 − 2.4 = −0.6 kg; upper fence = 3.4 + 2.4 = 5.8 kg. 0.2 > −0.6, so the lightest parcel is not an outlier. 6.1 > 5.8, so the heaviest parcel is an outlier. M1 for 1.5 × their IQR, A1 for both fences, R1 for "0.2 is not an outlier" with its comparison, R1 for "6.1 is an outlier" with its comparison. "6.1 is much bigger than the rest" scores 0. (c) A parcel of 6.1 kg is perfectly possible, so it is probably a genuine value and should be kept. R1 for a judgement in context.
Q4. (a) Convenience sampling. A1. (b) Members who arrive first on a Monday morning are not typical: they are likely to be the keenest regulars, or people free on weekday mornings, so the sample would probably overstate how much members exercise. R1 for identifying a group that is over- or under-represented, R1 for the direction of the effect on the estimate. (c) k = 900 ÷ 30 = 30. Choose a random starting number between 1 and 30, say 11, then take members 11, 41, 71, … and every 30th member after that, until 30 have been chosen. A1 for the interval 30, A1 for a random start then every 30th member.
Q5. (a) 4204 ÷ 11 = 382 visitors (3 s.f.). A1. (b) The museum was closed, so 0 is not a count of visitors on an open day; it is a missing value, and keeping it would drag the mean down. R1. (c) Cleaned data, ordered: 198, 205, 212, 219, 221, 225, 228, 231, 236, 240, with n = 10. Mean = 2215 ÷ 10 = 221.5. Median = (221 + 225) ÷ 2 = 223. Q₁ = 212, Q₃ = 231, IQR = 19. 1.5 × 19 = 28.5, so the fences are 212 − 28.5 = 183.5 and 231 + 28.5 = 259.5. The smallest value 198 and the largest 240 both lie inside, so there are no outliers. A1 for mean 221.5, A1 for both quartiles, M1 for 1.5 × their IQR, A1 for both fences, R1 for the conclusion with a comparison. Follow-through from their quartiles.
Educerie · written from the published IB Diploma Programme Mathematics: analysis and approaches guide, first assessment 2021, section 4.1 Populations, samples and sampling. Original text, examples and questions. Diagrams drawn by Educerie. Last reviewed 25 September 2026.