← Back to context

Comment by ricardobeat

18 hours ago

To do this you need video/motion understanding, the intent cannot be judged from still images or state descriptions.

We’ve had the tool to do this since mid 2025, V-JEPA2 [1], Yann Lecun’s last work at Meta.

It runs at several FPS on a macbook and can even be trained locally. Chaining it with Jev for decision-making would probably work great!

[1] https://ai.meta.com/research/vjepa/

The technique i mentioned in the blogpost works with video too. I tested succesfully with Qwen/Qwen3-VL-8B-Instruct. Effecient caching is a little trickier though, but very doable.