FiveWay Premium

Unlock every question

1,292 more practice questions, 134 Killer problems, timed mock exams in the 2027 format, and sets built from the skills you miss.

Current accessGuestYou are browsing without an account. Sign in to keep your record.

Free

Always, no account needed to read

  • A 7-question diagnostic and the first 8 practice questions in every topic
  • A worked solution and a note on each wrong choice for those questions
  • Every concept note, in every chapter
  • Your record and My page

Premium

Everything in Free, plus

  • The other 1,292 questions — every chapter, to the end
  • 134 Killer problems, pitched above the exam ceiling
  • Timed mock exams in the 2027 format — 42 questions, 105 minutes
  • Sets built from the skills you keep missing, refilled weekly
  • Analytics — accuracy per skill, time per question, your weak chapters
Unlock early access

Early access is free while we test. No card, no timer. Compare plans

Sign in to FiveWay

Sign in to keep your answers and see which skills to fix.

  • Free to start
  • No password
  • Progress saved on every device

We store your answers and the skills they belong to. Your first name and last initial appear only on the leaderboard and in Community.

Topic 2.6

Competing Function Model Validation

Nova just launched a new video series. Its total views (in thousands) at the end of each week look like they're curving upward, but many functions curve upward. A linear model, a quadratic model and an exponential model could all be fitted to the same data, and a calculator will happily produce all three. Which one should we trust?

9 MIN READ7 IDEAS33 PROBLEMS6 flashcards

Read this first

30 sec

  1. 01

    Appropriate model → residuals show no pattern. Pattern → wrong type of model.

Remember from 2.2 and 2.5B: linear change is additive, exponential change is proportional, and a regression gives the best curve of a chosen type. This note is about choosing the TYPE: using the data's pattern, the context, and a new tool called residuals.

01

Three Kinds of Change

modelwhat the data does over equal x-stepstypical context
lineardifferences are roughly constanta fixed amount added per unit (a plan that charges $4 per GB)
quadraticdifferences change at a roughly constant rate (2nd differences roughly constant)area from a length (pizza area from its diameter), height of a thrown ball
exponentialratios are roughly constanta fixed percent per unit (growth, decay, doubling, half-life)

Remember from Unit 1: a quadratic is the function whose rate of change changes at a constant rate. That is why its second differences are constant, just as a linear function's first differences are.

Here is what “second differences” means with a simple example. For y=x2y = x^{2} at x=1x = 1, 2, 3, 4, 5, the outputs are 1, 4, 9, 16, 25. The first differences (subtract neighbors) are 3, 5, 7, 9: not constant, so not linear. Now subtract THOSE: 2, 2, 2. Constant second differences are the fingerprint of a quadratic.

x12345
y = x2x^{2}1491625
1st differences—3579
2nd differences——222
02

Nova's New Series

w (week)012345678
views (thousands)12.416.324.132.745.565.090.7125.8177.5

Worked example

Use the data's patterns to decide which kind of model is most appropriate.

  1. 01

    Differences: 3.9, 7.8, 8.6, 12.8, 19.5, 25.7, 35.1, 51.7. They keep growing, so linear is out.

  2. 02

    Second differences: 3.9, 0.8, 4.2, 6.7, 6.2, 9.4, 16.6. They are not roughly constant; they grow too. That's a warning sign for quadratic.

  3. 03

    Ratios: 1.31, 1.48, 1.36, 1.39, 1.43, 1.40, 1.39, 1.41. After the first week they stay close to 1.4. The outputs change roughly proportionally, so an exponential model is the best choice.

    Here are all three regressions drawn on the data:

    Figure

    The line clearly misses. The quadratic and the exponential both look close. Looking at the graph alone isn't enough to decide between them.

03

Residuals

CONCEPT

What a Residual Is

For each data point, the model makes a prediction. The residual is how far the actual value is from that prediction.

Positive residual: the actual value is ABOVE the model, so the model underestimated. Negative residual: the actual value is BELOW the model, so the model overestimated.

residual=actual value−predicted value\text{residual} = \text{actual value} - \text{predicted value}

REAL-LIFE EXAMPLE

Ridge's Real Numbers

In 2.1 we modeled Ridge with R(t)=1200+300tR(t) = 1200 + 300t. Its actual counts for months 0 to 6 were 1180, 1530, 1790, 2120, 2350, 2710, 2990.

At month 4 the model predicts 1200+300(4)=24001200 + 300(4) = 2400, but the actual count was 2350. Residual = 2350−2400=−502350 - 2400 = -50: the model overestimated by 50 subscribers.

All seven residuals: −20, 30, −10, 20, −50, 10, −10. They bounce above and below 0 with no trend, which is exactly what a good model's residuals look like.

COMMON MISTAKE

“Residual = predicted − actual.”

The order is actual − predicted. Getting it backwards flips every sign, and then “underestimate” and “overestimate” get swapped. A quick check: if the point is above the curve, the residual must be positive.

Worked example

Example (by hand). A linear model y=2.1x+3y = 2.1x + 3 was fitted to the data (1,5.4)(1, 5.4), (2,7.0)(2, 7.0), (3,9.6)(3, 9.6), (4,11.1)(4, 11.1). Find each residual and say whether the model over- or underestimated.

  1. 01

    Predicted values: 2.1(1)+3=5.12.1(1) + 3 = 5.1, 2.1(2)+3=7.22.1(2) + 3 = 7.2, 2.1(3)+3=9.32.1(3) + 3 = 9.3, 2.1(4)+3=11.42.1(4) + 3 = 11.4.

  2. 02

    Residuals = actual − predicted: 5.4−5.1=0.35.4 - 5.1 = 0.3, 7.0−7.2=−0.27.0 - 7.2 = -0.2, 9.6−9.3=0.39.6 - 9.3 = 0.3, 11.1−11.4=−0.311.1 - 11.4 = -0.3.

  3. 03

    Signs: + (underestimate at x=1x = 1), − (overestimate at x=2x = 2), + (under at x=3x = 3), − (over at x=4x = 4). They alternate and stay small, with no pattern, so the linear model is appropriate for this data.

04

Residual Plots: Look for a Pattern

A residual plot graphs each residual against its input. The rule is simple: if the model is the right type, the residuals are scattered randomly above and below 0. If they form a pattern, the model is systematically missing something.

Figure

Residuals of the three models for Nova's series. Notice the different vertical scales: the linear residuals reach 35, the quadratic 6.5, the exponential only 1.0.

  1. 01

    Linear: positive at both ends, negative in the middle, a U-shape. The data curves and the line can't. Not appropriate.

  1. 02

    Quadratic: negative, then positive (weeks 1 to 4), then negative (weeks 5 to 7), then positive. The residuals change sign in long runs, a wave. The quadratic bends at the wrong rate. Not appropriate.

  1. 03

    Exponential: the signs jump around (+−+−−++−+)(+ - + - - + + - +) with no shape. Appropriate, and it matches the conclusion from the ratios.

KEY RULE

Appropriate model → residuals show NO pattern. Pattern → wrong type of model.

COMMON MISTAKE

“The linear model has r≈0.938r \approx 0.938, which is close to 1, so it's a good model.”

A high rr only says the points roughly go up together; it does not say the right shape was chosen. The linear residuals form a clear UU, so the linear model is not appropriate despite its rr.

“The quadratic residuals are small, so the quadratic is fine.”

Small is not the test; patternless is. A pattern of small residuals still means the model is the wrong shape, and the errors will grow once you predict beyond the data: at week 10 the exponential model predicts about 345.1 thousand views but the quadratic only about 270.6.

Quick check

A residual plot has points above and below zero, but they trace a clear U shape. Is the model appropriate?

05

Overestimate or Underestimate?

Sometimes the context tells you which kind of error is worse.

REAL-LIFE EXAMPLE

Ordering Milk for a Café

A café's model predicted it would need 42 liters of milk on Saturday. It actually used 45 liters. Residual = 45−42=345 - 42 = 3 liters: the model underestimated, and the café ran out before closing.

For the café, an overestimate (a little extra milk) is a smaller problem than an underestimate (turning customers away). In a situation like this, a model that tends to overestimate slightly is preferable.

Worked example

Try it yourself. Which type of model (linear, quadratic, or exponential) fits each situation best? Explain in one phrase. (a) the cost of a phone plan: $15 per month plus $4 per GB (b) the area of a circular pizza as a function of its diameter (c) a rumor where each person tells two new people every hour

Answers: (a) Linear: a constant $4 is added per GB. (b) Quadratic: area depends on the diameter squared. (c) Exponential: the number of people is multiplied by the same factor every hour.

Quick check

A model predicts 52.452.4 at x=6x = 6; the actual value was 49.149.1. What is the residual, and did the model over- or underestimate?

06

Practice

Worked example

P1. Using the model y=2.1x+3y = 2.1x + 3, predict yy at x=6x = 6. The actual value was 14.8. Find the residual and interpret it.

Answer: Predicted 2.1(6)+3=15.62.1(6) + 3 = 15.6. Residual = 14.8−15.6=−0.814.8 - 15.6 = -0.8: the model overestimated by 0.8.

Worked example

P2. An exponential regression's residual plot shows residuals positive at both ends and negative in the middle. Is the exponential model appropriate?

Answer: No. The residuals form a U-shaped pattern, so the model is systematically off. Try a different type of model.

Worked example

P3. The area of a square garden is recorded for side lengths 1, 2, 3, 4, 5 m. Which type of model should fit, and what should its second differences look like?

Answer: Quadratic (area = side²). The second differences should be constant (2, 2, 2).

Common slips

  • “Residual = predicted − actual.”

    The order is actual − predicted. Getting it backwards flips every sign, and then “underestimate” and “overestimate” get swapped. A quick check: if the point is above the curve, the residual must be positive.

  • “The linear model has r≈0.938r \approx 0.938, which is close to 1, so it's a good model.”

    A high rr only says the points roughly go up together; it does not say the right shape was chosen. The linear residuals form a clear UU, so the linear model is not appropriate despite its rr.

    “The quadratic residuals are small, so the quadratic is fine.”

    Small is not the test; patternless is. A pattern of small residuals still means the model is the wrong shape, and the errors will grow once you predict beyond the data: at week 10 the exponential model predicts about 345.1 thousand views but the quadratic only about 270.6.

Lock it in

Try the flashcards

6 cards · Semi-log plots and residuals

Start

Recap card

6 lines to re-read the night before.

  1. 01

    Linear: roughly constant differences. Quadratic: roughly constant second differences. Exponential: roughly constant ratios. Context often decides too.

  2. 02

    Residual = actual − predicted. Positive → the model underestimated; negative → it overestimated.

  3. 03

    An appropriate model has a residual plot with no pattern. A U-shape, a wave, or any trend means the wrong type of model.

  4. 04

    A high rr or small residuals do not prove the model type is right; the pattern of the residuals does.

  5. 05

    In context, decide whether overestimating or underestimating is the worse mistake.

  6. 06

    Next, in 2.7: putting functions inside other functions, composition.

Premium feature

Unlock Premium

Every question, sets from your misses, mock exams and more.

Current accessFree

  • Every practice question
  • Sets from your misses
  • Mock exams

Free during early access — no card. Your progress stays exactly where it is. Compare plans