Educerie · IB Diploma · Mathematics: analysis and approaches
Topic 4 Statistics and probability · 4.4 Correlation and regression
What you must be able to do
| You must be able to | Level | What it looks like in the exam |
|---|---|---|
| Draw and describe a scatter diagram: positive, negative or no correlation; strong or weak | SL, HL | "Describe the correlation shown" (1 to 2 marks) |
| Find Pearson's product-moment correlation coefficient r with technology, and interpret it | SL, HL | "Find r" (2 marks); "Interpret your value of r" (1 mark), Paper 2 |
| Know that r measures only linear correlation | SL, HL | "Explain why r is not a suitable measure here" (1 mark) |
| Use a given critical value of r to decide whether correlation is significant | SL, HL | "Using the critical value 0.632, determine…" (2 marks) |
| Draw a line of best fit by eye through the mean point | SL, HL | "Plot the mean point and draw a line of best fit" (2 marks) |
| Find the regression line of y on x with technology and use it to predict | SL, HL | "Find the equation of the regression line of y on x" (2 marks); "Estimate y when x = 24" (2 marks) |
| Interpret a and b in y = ax + b in context | SL, HL | "Interpret the meaning of a" (1 mark) |
| Explain the dangers of extrapolation, and why a y-on-x line should not be used to predict x | SL, HL | "Comment on the reliability of this estimate" (1 mark) |
| Distinguish correlation from causation | SL, HL | "Does this show that x causes y? Explain" (1 to 2 marks) |
Before you start
You need straight lines from 2.1, especially the meaning of gradient and y-intercept, and the mean from 4.3. On your GDC you need to be able to enter two lists and run a linear regression. The IB writes the regression line as y = ax + b, so here a is the gradient and b the intercept; do not confuse that with y = mx + c from 2.1, which is the same line with different letters.
1The idea in one paragraph
Bivariate data are pairs of measurements taken on the same individuals: the temperature and the drinks sold on the same day. Plot them as a scatter diagram and look for a pattern. If the points cluster around a straight line, the variables show linear correlation, positive if they rise together and negative if one falls as the other rises. Pearson's product-moment correlation coefficient, r, puts a number on it: r = 1 is a perfect rising line, r = −1 a perfect falling line, r near 0 no linear pattern at all. The regression line of y on x is the straight line that best predicts y from x, and your GDC gives its equation. Use it to predict y inside the range of the data; outside that range, or backwards to predict x, or as proof that x causes y, it can mislead badly.
2Scatter diagrams and the language of correlation
In a scatter diagram, the independent variable (the one you think does the influencing, or the one you control) goes on the horizontal axis as x, and the dependent variable goes on the vertical axis as y. Figure 1 shows the six patterns you must be able to name.
A description has two parts:
- Direction. Positive correlation: as x increases, y tends to increase. Negative correlation: as x increases, y tends to decrease. No correlation: no linear trend either way.
- Strength. Strong: the points lie close to a straight line. Weak: there is a trend, but the points are widely scattered around it.
The word tends matters. In a correlation, individual points go against the trend; it is the overall pattern that rises or falls. The last panel of Figure 1 is the warning: the points follow a perfect arch, yet r ≈ 0, because r only detects straight-line patterns.
3Pearson's r
Here is one data set, used for the rest of the page. A drinks kiosk recorded the day's maximum temperature and the number of iced drinks it sold on ten days (invented data):
| Temperature x (°C) | 14 | 17 | 18 | 20 | 22 | 23 | 25 | 27 | 29 | 31 |
|---|---|---|---|---|---|---|---|---|---|---|
| Iced drinks sold y | 42 | 55 | 51 | 64 | 70 | 68 | 81 | 86 | 90 | 101 |
On the GDC. Put x in one list and y in another, then run a linear regression (TI-84: STAT, CALC, LinReg(ax+b), with diagnostics switched on so r appears; Casio: Statistics menu, CALC, REG, then the ax+b option). The output gives a, b and r together.
For the kiosk, r = 0.990 (3 s.f.): a strong positive linear correlation.
What r means. r always lies between −1 and 1.
| r | Description |
|---|---|
| r = 1 | perfect positive linear correlation: every point on a rising line |
| close to 1 | strong positive |
| a little above 0 | weak positive |
| about 0 | no linear correlation |
| a little below 0 | weak negative |
| close to −1 | strong negative |
| r = −1 | perfect negative linear correlation |
The guide sets no exact cut-offs. A common working guide is: |r| above about 0.75 strong, 0.5 to 0.75 moderate, 0.25 to 0.5 weak, below 0.25 very weak or none. Use words that match the size, and always say "linear".
Why r behaves like this: one calculation by hand. The guide says technology finds r, but seeing inside it helps. Draw lines through the mean point (x̄, ȳ) = (22.6, 70.8), splitting the diagram into four quadrants as in Figure 2. For each point, multiply its distance from x̄ by its distance from ȳ. A point above and to the right gives (+) × (+) = +; below and to the left gives (−) × (−) = +; the other two quadrants give negative products. Add them all up:
The sum is large and positive because almost every point sits in a positive quadrant. If the points were scattered evenly over all four quadrants, the positive and negative products would cancel and r would be near 0. The bottom line only rescales, so that r lands between −1 and 1 whatever the units. You are never asked to do this in an exam, but it explains three facts: r has no units; swapping x and y does not change r; and a single point far from the mean point, in the "wrong" quadrant, can pull r a long way, as Figure 6 later shows.
Critical values. With only a few points, a pattern can appear by chance. When a question gives a critical value for your number of pairs, compare: if |r| is greater than the critical value, the linear correlation is significant, meaning it is unlikely to be a fluke of this small sample. Suppose a question gives 0.632 as the critical value for n = 10. The kiosk's r = 0.990 > 0.632, so the correlation is significant. You do not need to know where critical values come from; they are always supplied.
4A line of best fit by eye
A line of best fit summarises the trend. Drawn by eye, it should pass through the mean point (x̄, ȳ) and have roughly as many points above it as below, spread along its length. Figure 3 draws one for the kiosk.
The method: calculate x̄ and ȳ, plot the mean point clearly (a circle or a cross), put a ruler on it, and turn the ruler about that point until the points are balanced either side. The line does not have to pass through the origin, or through the first and last points; those are the two commonest mistakes.
5The regression line of y on x
Different people draw different lines by eye. The regression line of y on x is the one line that every calculator agrees on. It is the line that makes the vertical distances from the points to the line as small as possible overall (it minimises the sum of their squares, which is why it is also called the least squares line). Figure 4 shows those vertical distances.
From the GDC, for the kiosk:
The regression line always passes through (x̄, ȳ). That fact is worth marks on Paper 1: given the gradient and the mean point, you can find b. If a regression line y = 2.5x + b passes through the mean point (12, 41), then 41 = 2.5 × 12 + b, so b = 11.
Interpreting a and b. Say what each means in the context, with units:
- a = 3.37: for each extra 1 °C of maximum temperature, the kiosk sells about 3.37 more iced drinks, on average.
- b = −5.44: the model's prediction of sales on a day with maximum temperature 0 °C. It is negative, which is impossible, and 0 °C is far outside the data (14 °C to 31 °C). So b has no sensible meaning here; it is only where the line crosses the axis. Saying so is itself worth the mark.
Keep full calculator precision when you use the line, and round only the answer.
6Predictions: interpolation, extrapolation, and which way round
Substitute an x value to predict y. On a 24 °C day:
24 °C lies inside the range of the data, so this is interpolation, and with r = 0.990 it is reliable. Figure 5 shades the safe region.
Extrapolation is predicting outside the range of the x values you have. The line says 40 °C would bring about 129 drinks, but nothing in the data tells you the straight-line pattern continues: the kiosk might run out of stock, or people might stay indoors. At 0 °C it predicts −5 drinks, which is plainly nonsense. Extrapolated predictions are unreliable, however large r is, and a comment on reliability should say exactly that: "x = 40 is outside the data range 14 ≤ x ≤ 31, so this is extrapolation and the prediction is unreliable."
Predicting x from y. The y-on-x line is built to predict y from x, because it minimises vertical distances. To estimate x from a known y you need a different line, the regression line of x on y, which minimises horizontal distances; that is subtopic 4.10. The two lines are the same only when r = ±1, and the weaker the correlation, the further apart they are. So if a question asks for the temperature on a day when 75 drinks were sold, do not rearrange y = 3.37x − 5.44; the guide says you cannot always make that prediction reliably.
Use the y-on-x line only to predict y from an x inside the data range. A strong r does not rescue an extrapolation.
How reliable is a prediction? Check three things and mention whichever apply: is |r| large (strong linear correlation, ideally significant)? Is the x value inside the data range? Is the sample big enough and from the right population? A prediction for a different kiosk, in a different country, may not follow this line at all.
Outliers and r. One unusual pair can transform r. In Figure 6 a single extra day, 30 °C with only 40 drinks sold because the kiosk ran out of ice, drops r from 0.990 to 0.623. Look at the scatter diagram before trusting r, and treat that point as 4.1 taught: an error or a special circumstance should be explained, and possibly removed, with a sentence saying so.
7Correlation is not causation
A strong correlation says two variables move together. It does not say that one causes the other. There are three common reasons for correlation without direct causation:
- A hidden (confounding) variable drives both. Over a summer, iced drink sales and sunburn cases are strongly correlated, but neither causes the other: hot, sunny weather causes both, as Figure 7 shows.
- Coincidence, especially with few data points or when many pairs of variables have been tried.
- Cause the other way round. Towns with more police officers may record more crime, partly because more officers find more crime, and partly because high-crime towns hire more officers.
The guide's example shows the other side. The link between smoking and lung cancer was first found as a correlation in the data, and it was the later scientific work on mechanisms, not the correlation alone, that established the cause. In an exam, "correlation does not imply causation" alone rarely earns the mark; name a plausible hidden variable in context.
8Where marks are lost
Describing r without "linear". r measures linear correlation only. "Strong positive linear correlation" is the full description.
Using r for a curved pattern. If the scatter diagram bends, r is misleading however small or large it is. Say that a linear model is not appropriate.
Reading r as the gradient. r = 0.99 does not mean y rises by 0.99 for each unit of x. That is a. A shallow line and a steep line can both have r = 0.99.
Swapping a and b. In the IB's y = ax + b, a is the gradient and b the intercept. Copy them from the GDC into the right places.
Line of best fit through the origin. A line by eye goes through the mean point, not (0, 0), unless the data happen to demand it.
Extrapolating without comment. Any prediction outside the data range needs the word "extrapolation" and the word "unreliable".
Rearranging the y-on-x line to find x. It is built to predict y. Predicting x reliably needs the x-on-y line.
"x causes y" from a large r. Correlation shows association. Name the hidden variable.
9Work it right
- Put the independent variable on the x-axis and label both axes with units.
- Write r to 3 significant figures, with its sign, and describe it in words: direction, strength, "linear".
- If a critical value is given, compare |r| with it and state "significant" or "not significant".
- Write the regression line with the letters of the context if given (for example d = 3.37t − 5.44), coefficients to 3 s.f., and keep full precision for predictions.
- Interpret a as "for each increase of 1 [unit of x], [y] increases by a [units], on average"; interpret b only if x = 0 is meaningful.
- For every prediction, say whether it is interpolation or extrapolation, and comment on reliability.
- For causation, name a plausible hidden variable in context.
10Try it
Marks in brackets. Q1 and Q2 are Paper 1 style, no calculator. Q3 and Q4 are Paper 2 style.
Q1. Four data sets have correlation coefficients r = −0.95, r = −0.31, r = 0.04 and r = 0.88.
(a) Write down which value of r describes: strong positive, weak negative, strong negative, and no linear correlation. 2 marks
(b) The scatter diagram of the data set with r = 0.04 shows the points lying close to an arch-shaped curve. Explain what this shows about r. 1 mark
Q2. For 12 students, the regression line of test score y on hours of practice x is y = 2.4x + b. The mean number of hours is 15 and the mean score is 50.
(a) Find the value of b. 2 marks
(b) Interpret, in context, the value 2.4. 1 mark
(c) Use the line to estimate the score of a student who practised for 20 hours. 1 mark
(d) Explain why the line should not be used to estimate the practice time of a student who scored 70. 1 mark
Q3. The ages x (years) and prices y (thousands of euros) of eight used cars of the same model are shown (invented data).
| Age x (years) | 1 | 2 | 3 | 3 | 4 | 5 | 6 | 8 |
|---|---|---|---|---|---|---|---|---|
| Price y (€ thousand) | 21.5 | 19.0 | 17.2 | 18.1 | 15.0 | 13.4 | 12.1 | 8.9 |
(a) Find the value of r. 2 marks
(b) Describe the correlation between age and price. 1 mark
(c) Find the equation of the regression line of y on x. 2 marks
(d) Interpret the gradient of the line in context. 1 mark
(e) Estimate the price of a 7-year-old car of this model. 2 marks
(f) Comment on using the line to estimate the price of a 15-year-old car. 1 mark
Q4. In a survey of 10 towns, the number of cafés and the number of bicycle shops have correlation coefficient r = 0.58. The critical value of r for 10 pairs at the level used is 0.632.
(a) Determine whether there is significant linear correlation. 2 marks
(b) A councillor claims that opening more cafés would bring more bicycle shops to a town. Comment on this claim. 2 marks
11In one breath
Bivariate data are pairs, plotted as a scatter diagram with the independent variable on x. Describe the pattern by direction (positive, negative, none) and strength (strong, weak), and measure it with Pearson's r from the GDC: between −1 and 1, sign for direction, size for strength, and meaningful only for straight-line patterns; if |r| beats a given critical value the correlation is significant. A line of best fit by eye passes through the mean point (x̄, ȳ). The regression line of y on x, y = ax + b from the GDC, minimises vertical distances and also passes through the mean point; a is the change in y per unit of x, and b is the value at x = 0, often meaningless. Predict y from x inside the data range; extrapolation is unreliable, predicting x needs the x-on-y line, one outlier can wreck r, and correlation never proves causation: look for the hidden variable.
Answers
Q1. (a) Strong positive: 0.88. Weak negative: −0.31. Strong negative: −0.95. No linear correlation: 0.04. A2 for all four, A1 for two or three. (b) There is a strong relationship, but it is not linear; r measures only linear correlation, so a value near 0 does not mean the variables are unrelated. R1 for "r only measures linear correlation" in some form.
Q2. (a) The regression line passes through (15, 50), so 50 = 2.4 × 15 + b = 36 + b, giving b = 14. M1 for substituting the mean point, A1 for 14. (b) For each extra hour of practice, the score increases by about 2.4 marks, on average. A1. "The gradient is 2.4" scores 0; it must be in context. (c) y = 2.4 × 20 + 14 = 62. A1, follow-through from their b. (d) This is the regression line of y on x, built to predict score from practice time; predicting x from y needs the regression line of x on y. R1.
Q3. (a) r = −0.993 (3 s.f.). A2. A1 if the sign is missing or the value is given to fewer than 3 s.f. (b) Strong negative linear correlation: older cars of this model tend to be cheaper. A1 for strong, negative (linear). (c) y = −1.79x + 22.8 (a = −1.79444…, b = 22.8277…). A1 for a, A1 for b. (d) For each extra year of age, the price falls by about €1790, on average. A1, must say "falls" (or a negative change) and give units. (e) y = −1.79444 × 7 + 22.8278 = 10.27, so about €10 300. M1 for substituting x = 7 into their line, A1 for 10.3 thousand euros. 7 is inside the data range, so this is interpolation. (f) The line gives −1.79444 × 15 + 22.8278 = −4.09, a negative price, which is impossible. x = 15 is well outside the data range (1 to 8 years), so this is extrapolation and the estimate is not valid. R1 for identifying extrapolation or the impossible negative value, with a conclusion.
Q4. (a) 0.58 < 0.632, so the linear correlation is not significant: a value of r this size could easily arise by chance with only 10 towns. M1 for comparing r with 0.632, A1 for the correct conclusion. (b) The correlation is not significant, and even a significant correlation would not show causation. A hidden variable, such as the town's population, is likely to raise both the number of cafés and the number of bicycle shops, so opening cafés would not by itself bring bicycle shops. R1 for "correlation does not imply causation" or the non-significance, R1 for a plausible hidden variable in context.
Educerie · written from the published IB Diploma Programme Mathematics: analysis and approaches guide, first assessment 2021, section 4.4 Correlation and regression. Original text, examples and questions. Diagrams drawn by Educerie. Last reviewed 25 September 2026.