Comment by joefourier

3 hours ago

VAEs don't have to inhibit fine detail, you can trivially increase the channel count like the Flux VAE did, but if you're not careful, that will make training on that latent space much more difficult.

Older models like the SD1.5 work like you mentioned, in large part due to relying on PatchGan which just ensures that individual patches are plausible in isolation. There's nothing preventing a VAE from rendering tiny text with the correct architecture and training. If pushed to an extreme, it could essentially start guessing the letters themselves and write different, but still legible words instead of nonsense scribbles.