Comment by qayxc

12 hours ago

When put to the test in real-world environment, the capabilities don't look as impressive as benchmarks and synthetic tests might indicate. So doubts about actual spatial reasoning capabilities remain.

I see. Gary Marcus said that AI won’t be able to make a coffee in any arbitrary home kitchen.

I think it’s a good test and I think LLMs will reach it in 3 years. Current benchmarks maybe slightly incorrect.

I’m happy to make a 4:1 bet in my favour that I’m correct about the kitchen bet.

  • I cant make coffee in an arbitary house kitchen. People tend to put stuff anywhere but where at look for them...

    • Your prompter need merely to say “keep going” each time you report that you haven’t found the grounds yet.