Fitting a Model to Data: Least Squares, Chi-Square and Goodness of Fit · seed 1 · A4, ink-friendly. The answer key prints on its own page for grown-ups.

Fitting lines and judging the fit

Science · Scientific Inquiry · ages 22-23
Name ______________________   Date ____________
  1. A scientist has some data points and a model with free parameters. What does the method of least squares actually do to choose the best parameters?

    • It makes the sum of the squared residuals as small as possible
    • It makes the line pass through every data point
    • It makes the sum of the residuals exactly zero
    • It picks the parameters that look best on a graph
  2. For one data point, the observed value is y and the model predicts y-hat. What is the residual?

    • y minus y-hat
    • y-hat minus x
    • The largest value in the data set
  3. A reduced chi-square near 1 means the model describes the data about as well as the size of the uncertainties says it should.

    Circle one:   True   False

  4. A scientist has data points and a model with free parameters. What does least squares do to choose them?

    • It makes the line pass through every data point
    • It picks the parameters that look nicest on a graph
    • It makes the sum of the squared residuals as small as possible
  5. You can get meaningful uncertainties on fitted parameters even if the data points have no reported measurement uncertainties.

    Circle one:   True   False

  6. After fitting a constant to three points, the residuals are 1, -2, and 1, and every point has uncertainty 1. What is the chi-square statistic for this fit?

    Answer: ______________

  7. You fit a model to 20 points with 2 free parameters and get a reduced chi-square very close to 1. What does that tell you?

    • The model describes the data about as well as the size of the uncertainties says it should
    • The model is definitely the true law of nature
    • The uncertainties were probably estimated too large
    • The fit failed and the model must be rejected
  8. In weighted least squares, each squared residual is divided by the square of that point's uncertainty before being added up. Why do we do that?

    • So that points measured more precisely pull the fit harder than sloppy points
    • So that all the residuals come out positive
    • So that the fit ignores the points with large values
    • So that the maths works without a computer
  9. A fit returns a reduced chi-square of 0.05 and residuals tiny beside the error bars. What is the most likely explanation?

    • The model is a perfect law of nature
    • The data were measured with zero error
    • The uncertainties were overestimated, or the model has too many parameters for the data
  10. You measure one quantity three times with equal uncertainties and get 5.0, 7.0, and 9.0. Your model is a constant y equals a, fitted by least squares. Type the best value of a.

    Answer: ______________

LightMySky · lightmysky.comW1-mt_NN0WYlp0md-s1

Answer key

For grown-ups. Fold this page away before handing over the rest.

Fitting lines and judging the fit W1-mt_NN0WYlp0md-s1

  1. It makes the sum of the squared residuals as small as possible · Least squares means exactly what it says: you adjust the parameters until the sum of the squared residuals is as small as it can be. Squaring makes positive and negative misses both count, and big misses count extra.
  2. y minus y-hat · The residual is the observed value minus the predicted value.
  3. True · The typical miss is about one sigma, which is exactly what honest uncertainties predict.
  4. It makes the sum of the squared residuals as small as possible · Least squares means exactly that: adjust the parameters until the squared misses bottom out.
  5. False · Parameter error bars come from propagating the point uncertainties through the fit. Without knowing how trustworthy each point is, you cannot say how trustworthy the slope or intercept is, and chi-square cannot even be computed.
  6. 6 · Chi-square is the sum of (residual / uncertainty) squared. Here that is 1^2 + (-2)^2 + 1^2 = 6.
  7. The model describes the data about as well as the size of the uncertainties says it should · Reduced chi-square is chi-square per degree of freedom. A value near 1 means the typical residual is about one sigma, exactly what you expect when the model fits and the uncertainties are honest.
  8. So that points measured more precisely pull the fit harder than sloppy points · A point with a small uncertainty is trustworthy, so missing it should cost a lot. Dividing by the uncertainty squared makes precise points weigh more in the sum, and without uncertainties on the points the fit means very little.
  9. The uncertainties were overestimated, or the model has too many parameters for the data · Misses far smaller than the bars mean the bars were set too wide or the model just follows noise.
  10. 7 · With equal uncertainties the best constant is the middle of the readings, 7.
Worksheet · LightMySky