← Back to context

Comment by joefourier

1 hour ago

Understandable, (V)AEs are their own rabbit hole (I've spent way too much time working with them). I think they deserve more attention, there's been a few good papers recently about how to ensure that the latent space is optimal for generation without sacrificing too much reconstruction detail, but it's no easy task.

But how do you plan on scaling to 1080p24 (or higher resolution) video with end to end pixel space training? For a 10 second long video, that's 1920x more tokens than your 512x512 DiT experiments, and I'd be concerned about scaling linear patchification temporally and spatially past a certain point.