← Back to context

Comment by barrkel

7 hours ago

If you think about it there's overlap in what the VAE encodes and main model encodes. Objects at a distance resemble texture and textures zoomed in gain structure. The VAE makes textures more efficiently representable at the cost of reducing the representable space of pixels. So things like tiny text become nonsense scribbles. Working in pixel space, especially with something with recursive or cascaded structure, opens the possibility of using the learnt structure of real writing at a higher level to perfect tiny details that actually cannot be approximated without being obviously wrong.

Time is "just" another dimension. There's temporal continuity between frames, a video VAE would be learning and representing those temporal shifts, but there's nothing to say that e.g. a recursively applied generative model at the pixel level also doesn't learn and represent those things.

As ever, figuring out how to train the thing is the hard bit I expect.

(Handwaving over "textures" here, VAEs encode more like somewhat macro blocks of image whose content is also conditioned on surrounding blocks, rather than tiny patches of patterned pixels.)

(And yes I'm a total imposter layman here, I just see VAEs as seeming to be a crutch that reduce data size - super super helpful of course - but being strictly speaking redundant and inhibiting correct fine detail.)

VAEs don't have to inhibit fine detail, you can trivially increase the channel count like the Flux VAE did, but if you're not careful, that will make training on that latent space much more difficult.

Older models like the SD1.5 work like you mentioned, in large part due to relying on PatchGan which just ensures that individual patches are plausible in isolation. There's nothing preventing a VAE from rendering tiny text with the correct architecture and training. If pushed to an extreme, it could essentially start guessing the letters themselves and write different, but still legible words instead of nonsense scribbles.