← Back to context

Comment by schopra909

5 hours ago

This is a totally fair point and definitely worth exploring!

The jumping point for this no-VAE work was three-fold:

1) Our goal here is to get 32x32 token reduction to make video training and inference downstream cheaper. To date, the best open-weight Image VAEs like Flux-2 seem to cap out at 16x16 token reduction (8x8 VAE + 2x2 linear patchification). Others like H3 have pushed to 32x32 reduction but requires them swapping out the small VAE decoder with a 2B parameter decoder. So, this is a foray to get 32x32 compression without compromising quality.

2) We believe that end-to-end trained networks will tend to perform better than modularly trained networks (e.g. VAE + DiT). This hypothesis comes from work like REPA-E, where authors are able to get much better results by backpropagating through the pretrained VAE. The latent space for perception / reconstruction seems to have a different "optimal" configuration than a latent space specifically built for generation. That's why we liked the idea of trained E2E here.

3) There's work with VAEs that show that providing additional modalities (e.g. text captions) can help the VAEs improve as well. That's natural to this construction, so we thought language might actually help with the compression, not hurt. To be honest, this this is the most hypothetical of three ideas; and definitely warrants specific ablations.

Understandable, (V)AEs are their own rabbit hole (I've spent way too much time working with them). I think they deserve more attention, there's been a few good papers recently about how to ensure that the latent space is optimal for generation without sacrificing too much reconstruction detail, but it's no easy task.

But how do you plan on scaling to 1080p24 (or higher resolution) video with end to end pixel space training? For a 10 second long video, that's 1920x more tokens than your 512x512 DiT experiments, and I'd be concerned about scaling linear patchification temporally and spatially past a certain point.