Educerie

Educerie · SAT · Math

Problem-Solving and Data Analysis · PSDA.4 Two-variable data: models and scatterplots

Where it is examined
both modules, usually one question with a scatterplot or a table attached.
The question this unit answers
a cloud of points and a line drawn through it. What do the line's numbers mean, and what is it safe to say from them?
Before you start
linear functions. A line of best fit is a linear model, read in exactly the same way.

What you must be able to do

You must be able toWhat it looks like on the test
Read a scatterplotDirection, strength, and any point that sits apart
Interpret the slope of a line of best fitA rate, with units from the axes
Interpret its interceptThe predicted value when x is zero
Predict from the modelAnd say when a prediction is not safe
Tell a linear model from an exponential oneConstant difference or constant ratio

1The idea in one paragraph

A line of best fit is a summary, not a law. Its slope says how much the y variable changes for each unit of x on average across this data, and its intercept says what the model predicts at x = 0 — which is sometimes meaningless in the real situation. Reading those two numbers with the axis units attached answers most questions in this unit, and the rest are about not over-claiming.

Read the axis labels before the question. Slope is always "y units per x unit", and the whole answer is usually that phrase with the right numbers in it.

2Describing a scatterplot

Direction: positive if the cloud rises left to right, negative if it falls. Strength: tight to the line is strong, widely scattered is weak. Form: a straight cloud is linear; one that bends is not, and forcing a line through it is what the test wants you to notice. Outliers: a point far from the pattern. One outlier can tilt a line of best fit noticeably.

Correlation is not causation. A strong relationship in a scatterplot never establishes that one variable causes the other, however tight it looks.

3Slope and intercept in context

A line of best fit for a plot of maintenance cost (lira) against vehicle age (years) is C = 900 + 340a.

  • 340 is the predicted increase in annual maintenance cost for each additional year of age.
  • 900 is the predicted cost for a vehicle of age zero — a new one.

Both readings must carry units. Wrong options drop them, reverse the rate ("each 340 lira adds a year"), or turn the rate into a total.

4Predicting, and when not to

Putting a value into the model gives a prediction. Two cautions the test rewards:

Interpolation is safe; extrapolation is not. A model built on cars aged 1 to 10 years says nothing reliable about a 40-year-old car.

A prediction is not an observation. A question asking "how much did this car actually cost" is answered from the plotted point; "how much does the model predict" is answered from the line. When both appear as options, they are usually different numbers, and the difference is the question.

5Linear or exponential

From a table: constant differences → linear; constant ratios → exponential.

From a plot: a straight cloud → linear; a cloud that curves upward steeply → exponential growth; one that falls steeply and levels off → exponential decay.

In words: a quantity growing by a fixed amount each period is linear; growing by a fixed percentage is exponential.

6Lines of best fit are not exact

The line passes near the points, not through them. Asking "how many points lie above the line" is a counting question about the plot, and asking "by how much does the model overestimate this point" is a subtraction: observed minus predicted.


Where points are lost

  • Dropping the units in an interpretation.
  • Reversing the rate, reading x per y.
  • Extrapolating far beyond the data and treating the answer as reliable.
  • Reading a data point when the model was wanted, or the reverse.
  • Calling a strong correlation a cause.
  • Assuming linear because the x values are evenly spaced.

Work it right

  1. Read both axis labels, with units.
  2. Say the slope as "so many y per one x".
  3. Say the intercept as "the predicted y when x is zero", and check whether that makes sense here.
  4. For a prediction, substitute; for an observation, read the point.
  5. Check whether the question is inside the range of the data.

Try it

Q1. A line of best fit for weekly sales S, in thousands of lira, against advertising spend A, in thousands of lira, is S = 12 + 2.4A. What is the best interpretation of 2.4?

A) Sales are 2.4 thousand lira when nothing is spent on advertising. B) Each additional thousand lira of advertising is associated with 2.4 thousand lira more in sales. C) Advertising accounts for 2.4% of sales. D) Each additional thousand lira of sales requires 2.4 thousand lira of advertising.

Q2. Using the same model, what does the model predict for sales when 5 thousand lira is spent on advertising, in thousands of lira? (Type your answer.)

Q3. A scatterplot of 30 points shows a strong positive linear relationship between hours of sunlight and tomato yield. Which conclusion is supported?

A) More sunlight causes higher yield. B) Higher yield causes more sunlight. C) Plots with more sunlight tended to have higher yields. D) Sunlight is the only factor affecting yield.

Q4. A table of values shows 6, 12, 24, 48 for x = 0, 1, 2, 3. Which model fits?

A) Linear with slope 6 B) Linear with slope 12 C) Exponential with multiplier 2 D) Exponential with multiplier 6

Q5. For a data point where A = 4, the observed sales were 20 thousand lira. Using S = 12 + 2.4A, by how much does the model overestimate the observation, in thousands of lira? (Type your answer.)

In one breath

A line of best fit is a summary: its slope is y units per x unit and its intercept is what the model predicts at zero, and both readings need the axis units attached. Predictions inside the range of the data are reasonable and predictions far outside it are not, an observed point is not the same thing as a modelled one, and no scatterplot however tight establishes that one variable causes the other.

Answers

Q1. B. slope is y units per x unit, and it is an association A describes the intercept, which is 12. C turns a rate into a percentage. D reverses the variables.

Q2. 24. substitute

S = 12 + 2.4A
S = 12 + 2.4(5)
S = 24thousand lira

Q3. C. describe the association, claim no cause A and B both claim causation from a scatterplot. D claims sunlight is the only factor, which nothing here supports.

Q4. C. constant ratio, not constant difference

differences: 6, 12, 24not constant
ratios: 12/6 = 2, 24/12 = 2, 48/24 = 2
y = 6 × 2x

A takes the first difference as a slope.

Q5. 1.6. predicted minus observed

predicted = 12 + 2.4(4) = 21.6
observed = 20
21.6 − 20 = 1.6

Answering 21.6 gives the prediction rather than the gap — the difference between reading the line and reading the point.


Educerie · written from the published College Board* Assessment Framework for the Digital SAT Suite *(Math, Problem-Solving and Data Analysis, skill/knowledge testing point "Two-variable data: models and scatterplots"). All questions and explanations are original Educerie text. Last reviewed 12 September 2026.

Check your understanding

The main ideas of this note. Tick each one you could do now, in an exam, without looking back up. Anything you cannot tick yet is the part to read again.

Practise this skillMocks: in the future, hold tight!