Comment by armcat
4 hours ago
Sure and that's what I do, but a video can be seen as a causal generation on discreet sequence of images, each image conditioned on the one before it. It can also be seen as a series of image editing tasks. It would be cool to get this working in imagegen because of the amount of control you would get. Right now with video generation you can at best specify start and end frame and hope for the best.
Image edit models can probably do a grid, but the temporal accuracy / coherence will never match what a video model, which is really a world model, can do.
Regarding your world model statement. This is completely FALSE. Learning the visual statistics of a physical world is NOT the same thing as learning its causal dynamics. The difference is observational likelihood versus intervention-dependent dynamics. There have been great studies disproving video models as world models, like this ICML paper: https://proceedings.mlr.press/v267/kang25g.html. Unfortunately lot of people treat them as world models, mostly because of their ability to reproduce increasingly convincing physical behaviour without ever discovering the underlying physical laws. This is due to many things that I could write an essay about, but better conditioning, latent space represtnation, scaling etc, all make them look awesome.
I can still get absolutely insane results with MiniMax H3 - insane in the sense that it would not make sense at all and would make your head spin.
They are proto world models (lots written about this - flux being an example of a video model whose weights also power world-action-engines used in robots) in that they attempt to model causality in time, the thing that is required for what OP is asking for and which image models will never do because it is out of domain.