Educerie
Level

Educerie · IB Diploma · Mathematics: analysis and approaches

Topic 4 Statistics and probability · 4.10 Regression line of x on y

Level
SL and HL. Nothing here is HL only, so every section is examinable for both.
Themes (key concepts)
validity, approximation, generalization. A regression line is an approximation built for one job, predicting one particular variable; using it for the other job, or outside the data, gives a number that is not valid, however precise it looks.
The question this unit answers
the regression line of y on x predicts y, so what do you use when the value you know is y and the value you want is x, and when can you trust the answer?
Where it is examined
Paper 2, as 3 to 6 marks at the end of a bivariate data question that began in 4.4: find r and the y on x line with your GDC, then find the x on y line, use it to estimate x, and say why the other line would be wrong or why the estimate is unreliable. Paper 1, without a calculator, where both lines are given as equations and you find the mean point by solving them together, or choose the right line for a prediction (3 to 5 marks).

What you must be able to do

You must be able toLevelWhat it looks like in the exam
Find the equation of the regression line of x on y with technology, in the form x = cy + dSL, HL"Write down the equation of the regression line of x on y" (2 marks), Paper 2
Use it to predict a value of x from a given value of ySL, HL"Use your equation to estimate the diameter of a tree 19 m tall" (2 marks)
Choose the correct line for a prediction, and explain why the other one is wrongSL, HL"Explain why the regression line of y on x should not be used to estimate x" (1 mark)
Know that you cannot reliably predict y from x with the x on y line (nor x from y with the y on x line)SL, HLPart of the explanation above, often worth R1
Use the fact that both lines pass through the mean point (x̄, ȳ)SL, HL"Find x̄ and ȳ" from two given line equations (3 marks), Paper 1
Judge whether a prediction is reliable: strength of correlation, interpolation or extrapolationSL, HL"Comment on the reliability of this estimate" (1 to 2 marks)
Interpret the gradient and intercept of x = cy + d in contextSL, HL"Interpret the value of c" (1 mark)

Before you start

This page finishes what 4.4 started, so you need all of it: scatter diagrams, Pearson's correlation coefficient r, the regression line of y on x, y = ax + b, found with technology, and its use for prediction, including the danger of extrapolation. You also need to rearrange a linear equation and solve two linear equations simultaneously (2.1). Neither regression line is in the formula booklet; the GDC finds both.


1The idea in one paragraph

The regression line of y on x is built to predict y. It is the line that makes the vertical distances from the points to the line as small as possible, because the vertical distance is exactly the error you make when you predict a y-value. If instead you know y and want x, the error that matters is the horizontal distance, and the line that makes those as small as possible is a different line: the regression line of x on y, written x = cy + d. The two lines are not rearrangements of each other. They cross at the mean point (x̄, ȳ), they are close together when the correlation is strong and far apart when it is weak, and each one is only for its own job. The rule to learn is short: use the line whose left-hand side is the variable you want.

2Why one line is not enough

In 4.4 the regression line of y on x was found by the least squares method. For every data point, measure the vertical distance to the line; that is the residual, the amount by which the line misses the point when it predicts y. Square those distances and add them. The regression line of y on x is the line that makes that total as small as it can be. Figure 1, left, shows the vertical misses for five points.

Figure 1 · Two lines, two ways of measuring the misses Figure 1 · Two lines, two ways of measuring the misses Regression line of y on x 0 1 2 3 4 5 6 x 0 1 2 3 4 5 6 7 y y = 0.83x + 1.77 Regression line of x on y 0 1 2 3 4 5 6 x 0 1 2 3 4 5 6 7 y x = 0.752y − 0.205 y on x makes the vertical misses small (it predicts y). x on y makes the horizontal misses small (it predicts x).
Figure 1 · Two lines, two ways of measuring the misses

Now turn the job round. Suppose you know a y-value and want to estimate x. Your error is now how far left or right you land, which is a horizontal distance. The line that makes the sum of squared horizontal distances smallest is the regression line of x on y, shown on the right of Figure 1. For those five points it is x = 0.752y − 0.205, while the y on x line is y = 0.83x + 1.77. Rearranging the first gives y = 1.33x + 0.273 (to 3 s.f.), a much steeper line than y = 0.83x + 1.77. They are genuinely different lines, because they answer different questions.

The names tell you which is which. "y on x" means y depends on x: y = ax + b, with y on the left. "x on y" means x depends on y: x = cy + d, with x on the left. The variable before "on" is the one the line predicts.

To predict y, use y on x. To predict x, use x on y. Never rearrange one line to get the other.

3Finding the x on y line with technology

The GDC has only one linear regression function, and it always fits "second list = gradient × first list + intercept". So to get x on y, give it the lists the other way round: the y-values as the first (independent) list and the x-values as the second.

  • On a TI-84 Plus: put x in L1 and y in L2, then run LinReg(ax+b) with Xlist set to L2 and Ylist set to L1.
  • On a Casio: in the regression settings, set XList to the list holding the y-data and YList to the list holding the x-data.
  • On a TI-Nspire: in the linear regression dialog, choose the y-data column as the X List and the x-data column as the Y List.

The calculator will still print its answer as "y = ax + b", because it does not know you swapped the lists. You must read it as x = (a)y + (b) and write it that way. Writing the output with the letters unchanged is the commonest way to lose the marks here.

Worked example 1: trees. A forester measures the height y (m) and the trunk diameter x (cm), at chest height, of eight trees of one species in a plantation.

Diameter x (cm)1822273136404552
Height y (m)13.011.817.514.220.116.422.821.0

From the GDC, as in 4.4:

r = 0.835strong positive linear correlation
y on x: y = 0.286x + 7.41lists in the usual order

Now swap the lists:

x on y: x = 2.44y − 7.82read the calculator's a as c and b as d

A drone survey can measure the height of a tree from above but not its diameter, so the forester wants to estimate the diameter of a tree 19 m tall. The known value is a height, y; the wanted value is a diameter, x. That is a job for x on y.

x = 2.4381 × 19 − 7.8165use the unrounded values from the GDC
= 38.5 cm (3 s.f.)

Figure 2 shows both lines with the data. Rearranging the y on x line instead, x = (19 − 7.41) ÷ 0.286, would give 40.5 cm, marked grey on the figure. That is a two-centimetre difference caused only by using the wrong line.

Figure 2 · Eight trees, both regression lines Figure 2 · Eight trees, both regression lines 10 20 30 40 50 60 x, trunk diameter (cm) 8 10 12 14 16 18 20 22 24 26 y, height (m) 40.5: rearranged y on x (wrong) 38.5 mean point (33.9, 17.1) x on y: x = 2.44y − 7.82 y on x: y = 0.286x + 7.41 Both lines pass through the mean point. For a tree 19 m tall, the x on y line predicts 38.5 cm.
Figure 2 · Eight trees, both regression lines

The reverse job has its own line. To estimate the height of a tree with diameter 30 cm, use y on x: y = 0.28615 × 30 + 7.4068 = 16.0 m. Rearranging the x on y line would give 15.5 m, which is again wrong.

Interpreting c and d. In x = 2.44y − 7.82, the gradient c = 2.44 means that for each extra metre of height, the predicted trunk diameter increases by 2.44 cm. The intercept d = −7.82 would be the predicted diameter of a tree of height 0 m, which is meaningless: no tree has height 0, and it is far outside the data. As in 4.4, an intercept often has no sensible interpretation, and saying so is the right answer.

4Why rearranging gives the wrong answer

The two lines disagree because the correlation is not perfect. Look at the tree that is 19 m tall. It is 1.9 m taller than the mean height of 17.1 m. The y on x line, rearranged, assumes that a tree's diameter is exactly as far above average as its height is, in the line's units, and gives 40.5 cm. But with r = 0.835, height does not fix diameter. Some tall trees are thin. The best estimate of the diameter of a tall tree is above average, but not by as much as that, so the x on y line pulls the estimate back towards the mean diameter, to 38.5 cm.

Each line pulls its predictions towards the mean of the variable it predicts. The y on x line pulls predicted heights towards ȳ; the x on y line pulls predicted diameters towards x̄. That is why each line is the right one only for its own variable.

Figure 3 shows how much the two lines differ as the strength of correlation changes. When r is weak, they open like scissors: one is nearly flat and the other nearly vertical, and predictions from the wrong one are wildly out. As |r| grows the scissors close, and when r = ±1 every point lies on one line and the two regression lines are the same line. That is the only case in which rearranging one gives the other.

Figure 3 · The stronger the correlation, the closer the two lines Figure 3 · The stronger the correlation, the closer the two lines r = 0.28 x y r = 0.80 x r = 0.99 x Weak correlation, the lines open like scissors. As |r| approaches 1 they close into one line.
Figure 3 · The stronger the correlation, the closer the two lines

A fact that is not on the syllabus but makes a good check: when both lines are written with their own variable on the left, y = ax + b and x = cy + d, the product of the gradients is a × c = r². For the trees, 0.28615 × 2.43810 = 0.698, and 0.835² = 0.698. If your two gradients multiply to more than 1, one of them is wrong.

5Both lines pass through the mean point

Every least squares regression line passes through the mean point (x̄, ȳ), the point whose coordinates are the mean of the x-values and the mean of the y-values. That was true of y on x in 4.4, and it is true of x on y too. So the two lines always cross at the mean point, as the cross in Figure 2 shows.

That gives Paper 1 a question that needs no data at all.

Worked example 2 (no calculator). For a set of bivariate data, the regression line of y on x is y = 0.8x + 2 and the regression line of x on y is x = 1.1y − 1.3. Find x̄ and ȳ.

The mean point lies on both lines, so solve them simultaneously. Figure 4 shows the crossing.

x = 1.1(0.8x + 2) − 1.3substitute y from the first line
x = 0.88x + 2.2 − 1.3
0.12x = 0.9
x = 7.5
y = 0.8(7.5) + 2 = 8check in line 2: 1.1(8) − 1.3 = 7.5 ✓

So x̄ = 7.5 and ȳ = 8.

Figure 4 · The lines cross at the mean point Figure 4 · The lines cross at the mean point 0 2 4 6 8 10 12 14 x 0 2 4 6 8 10 12 14 y (7.5, 8) y on x: y = 0.8x + 2 x on y: x = 1.1y − 1.3 Solve the two equations together: x̄ = 7.5 and ȳ = 8.
Figure 4 · The lines cross at the mean point

A shorter version gives one line and one mean. If the regression line of x on y is x = 2.5y − 4 and ȳ = 6, then the mean point lies on the line, so x̄ = 2.5 × 6 − 4 = 11. Notice that this works only because (x̄, ȳ) is on the line. You cannot find x̄ by putting some other value of y into the line.

The same Paper 1 question often goes on to ask for a prediction, with both lines printed. Choosing the right one is the whole mark: to estimate y when x = 10, use y = 0.8x + 2 and get 10; to estimate x when y = 10, use x = 1.1y − 1.3 and get 9.7.

6When a prediction can be trusted

Choosing the right line is necessary, but it does not make the prediction good. The same two checks from 4.4 apply to x on y, and Figure 5 puts the whole decision in one place.

Figure 5 · Choosing a line, then deciding whether to trust it Figure 5 · Choosing a line, then deciding whether to trust it What do you want to find? the unknown variable y from a given x: use y on x x from a given y: use x on y Is the prediction reliable? |r| strong (close to 1)? given value inside the data range? Both yes interpolation: a sensible estimate Either no weak r or extrapolation: unreliable Never rearrange one line to get the other. Pick the line by what you want to find. Then check strength and range before you believe the number.
Figure 5 · Choosing a line, then deciding whether to trust it
  1. Is the correlation strong? A line fitted to weakly correlated data predicts poorly in either direction. With r = 0.835 the tree predictions are reasonable estimates; with r = 0.3 they would be little better than guessing the mean.
  2. Is the given value inside the range of the data? Predicting inside the range of the observed values is interpolation, and is reasonable when r is strong. Predicting outside it is extrapolation, and is unreliable, because nothing tells you the linear pattern continues. For the x on y line, the range that matters is the range of the y-values, because y is the value you are given. The trees' heights run from 11.8 m to 22.8 m. A prediction for a 19 m tree is interpolation; a prediction for a 30 m tree is extrapolation, and the line would give 2.4381 × 30 − 7.8165 = 65.3 cm, a number with no evidence behind it.

A third point decides which line is sensible at all. The guide's warning, that you cannot always reliably predict y from x with an x on y line, is the same warning as 4.4's in reverse. A question may give both equations and ask you to explain your choice. The answer is always of this shape: "y is the known value and x the value to be estimated, so the regression line of x on y is used; the y on x line minimises vertical, not horizontal, distances, so it does not give the best estimate of x."

Correlation is still not causation. A strong x on y line lets you estimate a tree's diameter from its height; it does not say that height causes diameter. Both grow with the tree's age.

7Where marks are lost

Rearranging the y on x line to predict x. It gives a different, wrong answer unless r = ±1. Find the x on y line itself.

Leaving the calculator's letters unchanged. With the lists swapped the GDC still prints "y = ax + b". Writing y = 2.44x − 7.82 describes a completely different line. Write x = 2.44y − 7.82.

Swapping the lists and then predicting the wrong variable. Once you have x = cy + d, substitute the given y-value and get x. Substituting an x-value into it is meaningless.

Rounding the coefficients before predicting. Use the full values stored in the calculator, then round the answer to 3 significant figures. Rounded coefficients can change the third figure.

Checking the wrong range for extrapolation. For an x on y prediction, the given value is a y, so compare it with the range of the y-data.

Assuming a prediction is good because the right line was used. A weak r or an extrapolated value still makes the estimate unreliable, and the question often wants you to say so.

Using the mean point with the wrong values. Only (x̄, ȳ) is guaranteed to lie on both lines. A data point generally lies on neither.

8Work it right

  1. Decide first which variable is known and which is wanted; the wanted one goes on the left of the line you use.
  2. For x on y on the GDC, swap the lists, then write the result as x = cy + d with the letters corrected.
  3. Give coefficients to 3 significant figures in the equation, but predict with the unrounded values.
  4. Substitute the given value, and state the answer with units and to 3 significant figures.
  5. Check that the given value lies within the range of that variable in the data; if not, say the prediction is extrapolation and unreliable.
  6. Quote r and comment on its strength when asked whether an estimate is reliable.
  7. With two equations given and no data, find (x̄, ȳ) by solving them simultaneously.
  8. Interpret c as "for each increase of 1 unit in y, the predicted x changes by c units", in context.

9Try it

Marks in brackets. Q1 and Q2 are Paper 1 style, no calculator. Q3 and Q4 are Paper 2 style, with a GDC.

Q1. For a set of bivariate data, the regression line of y on x is y = 0.5x + 3 and the regression line of x on y is x = 1.6y − 2.

(a) Find x̄ and ȳ. 3 marks

(b) Estimate the value of y when x = 16. 2 marks

(c) Estimate the value of x when y = 12. 2 marks

Q2. For a set of data the regression line of x on y is x = −0.4y + 25, and ȳ = 30.

(a) Find x̄. 2 marks

(b) State whether the correlation between x and y is positive or negative, giving a reason. 1 mark

(c) Interpret the value −0.4 in the equation. 1 mark

Q3. A dealer records the age x (years) and the price y (thousand euros) of eight used cars of the same model.

Age x (years)234567810
Price y (€000)18.519.914.216.811.614.99.810.5

(a) Write down the value of Pearson's correlation coefficient r. 1 mark

(b) Write down the equation of the regression line of x on y. 2 marks

(c) Another car of this model is priced at €14 000. Use your equation to estimate its age. 2 marks

(d) Explain why the regression line of y on x should not be used to answer (c). 1 mark

(e) A buyer uses the equation in (b) to estimate the age of a car priced at €4000. Find the estimate and comment on its reliability. 2 marks

Q4. For twelve runners, x is the distance run in training each week (km) and y is their time for a 5 km race (minutes). The correlation coefficient is r = −0.79. The regression line of y on x is y = −0.12x + 27.4 and the regression line of x on y is x = −5.2y + 164. The training distances in the data ranged from 20 km to 80 km.

(a) Estimate the race time of a runner who trains 60 km a week. 2 marks

(b) Estimate the weekly training distance of a runner whose race time is 22 minutes. 2 marks

(c) Comment on using the y on x line to estimate the race time of a runner who trains 150 km a week. 1 mark

10In one breath

The regression line of y on x, y = ax + b, makes the vertical misses as small as possible and so predicts y; the regression line of x on y, x = cy + d, makes the horizontal misses as small as possible and so predicts x. They are different lines, so never rearrange one to get the other: use the line with the wanted variable on the left. On the GDC, swap the lists and then rewrite the answer as x = cy + d, correcting the calculator's letters. Both lines pass through the mean point (x̄, ȳ), so two given equations solved together give the means. The lines are far apart when correlation is weak and become one line only when r = ±1. A prediction is trustworthy only when r is strong and the given value lies inside the range of the data; outside it is extrapolation, and a correlation, however strong, is not a cause.


Answers

Q1. (a) Substitute y = 0.5x + 3 into x = 1.6y − 2: x = 0.8x + 4.8 − 2, so 0.2x = 2.8 and x̄ = 14. Then ȳ = 0.5(14) + 3 = 10. Check: 1.6(10) − 2 = 14. M1 for recognising that the lines meet at (x̄, ȳ) and attempting to solve simultaneously, A1 for x̄ = 14, A1 for ȳ = 10.

(b) x is known, so use y on x: y = 0.5(16) + 3 = 11. M1 for choosing the y on x line, A1 for 11. Using x = 1.6y − 2 gives 11.25 and scores M0.

(c) y is known, so use x on y: x = 1.6(12) − 2 = 17.2. M1 for choosing the x on y line, A1 for 17.2. Rearranging y = 0.5x + 3 gives 18 and scores M0.

Q2. (a) The mean point lies on the line, so x̄ = −0.4(30) + 25 = 13. M1 for substituting ȳ into the line, A1 for 13.

(b) Negative, because the gradient of the regression line is negative: as y increases, x tends to decrease. R1 for the conclusion with the reason.

(c) For each increase of 1 in y, the predicted value of x decreases by 0.4. A1. "It is the gradient" alone scores 0.

Q3. (a) r = −0.845 (3 s.f.). A1.

(b) With the lists swapped, x = −0.605y + 14.4 (coefficients to 3 s.f.). A1 for −0.605, A1 for 14.4, in an equation with x as the subject. y = −0.605x + 14.4 scores A1 A0 at most.

(c) x = −0.60523 × 14 + 14.416 = 5.94 years (3 s.f.). M1 for substituting y = 14 into their x on y line, A1 for 5.94. Substituting 14 000 in place of 14 scores M0: the prices are in thousands.

(d) The price is known and the age is to be estimated, so the line of x on y must be used; the y on x line minimises the vertical distances and is for predicting price from age, not age from price. R1 for a reason that identifies which variable is known and which is being estimated.

(e) x = −0.60523 × 4 + 14.416 = 12.0 years (3 s.f.). This is unreliable: €4000 is outside the range of prices in the data (€9800 to €19 900), so it is extrapolation. A1 for 12.0, R1 for identifying extrapolation with reference to the range of y. A comment only about r does not earn the R1.

Q4. (a) Training distance is known, so use y on x: y = −0.12(60) + 27.4 = 20.2 minutes. M1 for choosing the y on x line, A1 for 20.2.

(b) Race time is known, so use x on y: x = −5.2(22) + 164 = 49.6 km. M1 for choosing the x on y line, A1 for 49.6. Rearranging the y on x line gives 45 km and scores M0.

(c) 150 km is far outside the range of the data, 20 km to 80 km, so the estimate would be extrapolation and is not reliable. (The line would give 9.4 minutes for 5 km, which is not a realistic time.) R1 for extrapolation, with reference to the data range.


Educerie · written from the published IB Diploma Programme Mathematics: analysis and approaches guide, first assessment 2021, section 4.10 Regression line of x on y. Original text, examples and questions. Diagrams drawn by Educerie. Last reviewed 25 September 2026.

Mocks: in the future, hold tight!