What does the network predict at each step?
What does the forward process need training for?
Generation starts from pure noise and denoises stepwise into a sample.
Circle one: True False
Why is stepwise denoising easier to learn than one-shot image making?
Samples show rough textures and odd speckles. What is the cheapest fix?
Generation is too slow for a demo. What do you trade away by cutting steps?
Two configs give equal quality but one uses half the steps. Which ships?
A student trains the network to output clean images directly from pure noise. What is the conceptual error?