A model trained through 2021 is asked about 2024 events. What holds?
- It updates itself nightly
- It knows nothing past its cut-off
- Dates never matter
Which three questions fit a dataset before trusting its model?
- How heavy is it and how blue is the cover
- Which font and which printer
- Who is in it, who is missing, when collected
Whose permission matters when gathering training records?
- The printer
- Nobody at all
- The people the records describe
Blurry photos of some tanks and sharp photos of others taught photo quality, not tanks. What is the lesson?
- Blurry photos always help
- Data quirks become the model's rules
- Tanks cannot be photographed
Why keep a separate test set?
- To train twice as fast
- To delete the training data
- To catch tricks like the tank story before real use
Before trusting a model that was trained on a dataset, which set of questions is most useful to ask about that dataset?
- Who is in the data, who is missing, and when was it collected
- How big the files are, how fast the computer was, and what the model is named
- How many colors are in the charts, who made the charts, and what font was used
- How much the computer cost, where it was stored, and who owns the building
A fitness model trained only on males is used on everyone. What is most likely?
- Equal results for all
- Worse predictions for the missing group
- Better results for the missing group
A group of researchers collected all of their data from males in an athletic association, then used that data to build a model that predicts fitness test results. Later the model is used on everyone, including women and girls. What is the most likely result?
- The model works just as well for everyone, because more math always fixes missing data
- The model tends to make worse predictions for the people who were missing from the training data
- The model refuses to make any prediction for women and girls
- The model works better for women and girls, because it had fewer examples to memorize
Two models read job applications. One trained on all groups with consent, one on a narrow slice. Which earns trust for decisions about people?
- The one trained on all groups with consent
- The narrow one, since small data is cleaner
- Neither needs any data questions
A cut-off date only affects spelling, never facts about events.
Circle one: True False