Comment by programjames
11 hours ago
You are referencing this from the article:
> There were some technical findings along the way—for example, Kavish found that DDPM, while theoretically ideal for denoising in diffusion models, diverged, and that flow matching performed much better.
When they say "theoretically ideal" they mean a specific thing: the DDPM loss is information theoretically the number of error correction bits needed to recover an image. At very low noises, this quantity goes singular. This is not actually a problem—you should bound the log-SNR of your schedule by the image quantization level (usually 9 bits) anyway, and a trained bound lands within 10% of this.
Flow matching does not diverge because it removes the log-SNR rate from the loss. Since the rate can span several orders of magnitude over your schedule, you should importance sample during training. My guess is Kavish forgot to do this, leading to divergence.
No comments yet
Contribute on Hacker News ↗