← Back to context

Comment by joefourier

4 hours ago

I wonder how much of that is due to the use of a suboptimal VAE? The same optimisations could be applied to it, and my intuition tells me that your total compute spend would be even more optimal if you retrained the VAE yourself with a better approach (esp. ensuring translation and rotation invariance, + ability to rescale/blur the latents).

You could also have the same advantages of a draft in low resolution latent space, with an easier to learn data distribution that's more robust to perturbation, but instead of 512x512, you could get a 4096x4096 output for the same compute (assuming a 8x VAE).

There's also the advantage of being able to use a high number of diffusion/flow matching steps for the main LDM while the VAE can be single step and much smaller, since it does not have to handle language or significant scene understanding, just perceptual compression. This sounds especially important for a video model where I would be extremely hesitant to train a generative model without relying on interframe compression.

If you think about it there's overlap in what the VAE encodes and main model encodes. Objects at a distance resemble texture and textures zoomed in gain structure. The VAE makes textures more efficiently representable at the cost of reducing the representable space of pixels. So things like tiny text become nonsense scribbles. Working in pixel space, especially with something with recursive or cascaded structure, opens the possibility of using the learnt structure of real writing at a higher level to perfect tiny details that actually cannot be approximated without being obviously wrong.

Time is "just" another dimension. There's temporal continuity between frames, a video VAE would be learning and representing those temporal shifts, but there's nothing to say that e.g. a recursively applied generative model at the pixel level also doesn't learn and represent those things.

As ever, figuring out how to train the thing is the hard bit I expect.

(Handwaving over "textures" here, VAEs encode more like somewhat macro blocks of image whose content is also conditioned on surrounding blocks, rather than tiny patches of patterned pixels.)

(And yes I'm a total imposter layman here, I just see VAEs as seeming to be a crutch that reduce data size - super super helpful of course - but being strictly speaking redundant and inhibiting correct fine detail.)

This is a totally fair point and definitely worth exploring!

The jumping point for this no-VAE work was three-fold:

1) Our goal here is to get 32x32 token reduction to make video training and inference downstream cheaper. To date, the best open-weight Image VAEs like Flux-2 seem to cap out at 16x16 token reduction (8x8 VAE + 2x2 linear patchification). Others like H3 have pushed to 32x32 reduction but requires them swapping out the small VAE decoder with a 2B parameter decoder. So, this is a foray to get 32x32 compression without compromising quality.

2) We believe that end-to-end trained networks will tend to perform better than modularly trained networks (e.g. VAE + DiT). This hypothesis comes from work like REPA-E, where authors are able to get much better results by backpropagating through the pretrained VAE. The latent space for perception / reconstruction seems to have a different "optimal" configuration than a latent space specifically built for generation. That's why we liked the idea of trained E2E here.

3) There's work with VAEs that show that providing additional modalities (e.g. text captions) can help the VAEs improve as well. That's natural to this construction, so we thought language might actually help with the compression, not hurt. To be honest, this this is the most hypothetical of three ideas; and definitely warrants specific ablations.