← Back to context

Comment by visarga

3 years ago

Modern generative image models are trained on curated data, not raw internet data. Sometimes the captions are regenerated to fit the image better. Only high quality images with high quality descriptions.

I wouldn't call what Stable Diffusion et al are trained on "high quality". You need only look through the likes of LAION to see the kind of captions and images they get trained on.

It's not random but it's not particularly curated either. Most of the time, any curation is done afterwards.