Comment by tintor
12 hours ago
Pure software benchmarks might be getting saturated, but physical ones aren't.
Let LLM control a physical robot to perform tasks that average human can do.
12 hours ago
Pure software benchmarks might be getting saturated, but physical ones aren't.
Let LLM control a physical robot to perform tasks that average human can do.
What makes you think this task won't fall quickly too?
It might, but if there's one thing that hasn't changed since 2022, it's that the models tend to ace the tasks where you have gobs of training data and where verification loops are fast and cheap... and they are not nearly as amazing elsewhere. If it's close enough, they can generalize, e.g. translate one programming language to another. But there's a pretty steep cliff past a certain distance.
Case in point: you had hundreds of millions of JPEGs to vacuum up and bitmap image generation is amazing. But if you ask them to recreate the same scene as vector art, they will struggle to generate a decent SVG. Like, kindergarten-style pelicans on bicycles are the state of the art. It should generalize seamlessly, but somehow, doesn't?
I think it will happen, just like self-driving cars are happening, but it will probably be a slow process.