← Back to context

Comment by famouswaffles

4 days ago

What did they do to Astra so cracked at vision (and computer use). That ARC 3 score turned out to be no joke/fluke. That huge gap between Astra and Fable (in this case) is basically every hard vison/spatial benchmark i've seen including non-benchmarks like playing games (Portal, Factorio, RimWorld).

SpatialBench - https://x.com/spicey_lemonade/status/2096365630190698516

ZeroBench - https://zerobench.github.io/

Robot Arms - https://openai.robocurve.org/gpt-6-astra/

I think Opus 5.5 is at same level now. I have seen too many videos made by Opus 5.5 today on twitter.

https://x.com/victormustar/status/2102707412704919910 horse galloping pixel art

https://x.com/LexnLin/status/2102133072585965759 moving train pixel art animation

https://x.com/jkeatn/status/2102441348075057539 painting with code

https://x.com/LCSlates/status/2102503027340988559 video, very detailed prompt though

https://x.com/aj_dev_smith/status/2102504509637587339 generated song/music with code

https://x.com/aj_dev_smith/status/2102575577563570450 another song

  • These are amazing but the parent comment is referring to vision comprehension, not generation.

    • Yup and I think these examples demonstrate just that. From my experience, both Claude and ChatGPT iterate over what they can see to build things like these. I don't think these examples are made without vision.

    • Aren't these the same thing, to some extent? I remember my music teacher told me, that as long as you can hear your false tones, she can teach you how to sing, no matter how bad you are at it. But if you can't hear it, she cannot help you.

      AI is the same - as long as it can see well, it can tell the difference between what it outputs and what its supposed to. If you subtract the two, you have an error, and you can hill-climb on that.

Well, they have the best in class image generator so that probably has something to do with it