How do you predict loss at a larger budget from the curve?
What falls as a smooth power of parameters, tokens and compute?
A scaling curve tells you whether new abilities will appear at larger scale.
Circle one: True False
A huge model trains on tiny data and lands above its expected line. What is wrong?
Loss is 2.5 at one budget, and each tenfold compute step multiplies loss by 0.8. What is the loss after two such steps?
Answer: ______________
The size curve has flattened but the data curve still falls. Where does budget go?
Two teams share one budget. Team A doubles size on fixed data, team B grows both. Who defends better?
A plan extrapolates far past all measurements with a new data mix. What is the flaw?