Scaling Laws and the Compute Budget · seed 1 · A4, ink-friendly. The answer key prints on its own page for grown-ups.

Reading the scaling curve

Computing · Machine Learning · ages 22-23
Name ______________________   Date ____________
  1. How do you predict loss at a larger budget from the curve?

    • Guess the smallest number possible
    • Extend the straight log-axes line
    • Copy the last measured value
  2. What falls as a smooth power of parameters, tokens and compute?

    • Held-out loss
    • The number of researchers
    • The office rent
  3. A scaling curve tells you whether new abilities will appear at larger scale.

    Circle one:   True   False

  4. A huge model trains on tiny data and lands above its expected line. What is wrong?

    • The axes were labelled wrong
    • Loss curves never apply to large models
    • The budget was wasted on size the data cannot feed
  5. Loss is 2.5 at one budget, and each tenfold compute step multiplies loss by 0.8. What is the loss after two such steps?

    Answer: ______________

  6. The size curve has flattened but the data curve still falls. Where does budget go?

    • More size at any cost
    • Toward more tokens
    • Toward a new office
  7. Two teams share one budget. Team A doubles size on fixed data, team B grows both. Who defends better?

    • Team B, balance follows both curves
    • Team A, size is all that matters
    • Neither, budgets decide nothing
  8. A plan extrapolates far past all measurements with a new data mix. What is the flaw?

    • Log axes forbid long extensions
    • Large budgets always break curves
    • Extrapolation assumes the same mix and recipe hold
LightMySky · lightmysky.comW1-mt_tzl0sQblIM-s1

Answer key

For grown-ups. Fold this page away before handing over the rest.

Reading the scaling curve W1-mt_tzl0sQblIM-s1

  1. Extend the straight log-axes line · The straight line is the prediction tool.
  2. Held-out loss · That falling loss is what the scaling curve plots.
  3. False · It covers loss only, not emergent abilities or inference cost.
  4. The budget was wasted on size the data cannot feed · Oversized models on starved data sit above the line.
  5. 1.6 · Two ratio steps of 0.8 bring 2.5 down to 1.6.
  6. Toward more tokens · Fund the side that still improves.
  7. Team B, balance follows both curves · Balanced scaling tracks both curves instead of flattening one.
  8. Extrapolation assumes the same mix and recipe hold · Prediction assumes continuity of mix, recipe and smoothness.
Worksheet · LightMySky