Comment by bonzini
10 hours ago
As of a few months ago they still have trouble, with low thinking, at the "should I drive to a car wash that is 100 m away" kind of question.
10 hours ago
As of a few months ago they still have trouble, with low thinking, at the "should I drive to a car wash that is 100 m away" kind of question.
Simply appending “check your assumptions” to the question fixed it even back then: https://news.ycombinator.com/item?id=47040530
Similarly for Apple’s “red herring” paper, simply adding a generic caveat to “disregard irrelevant factors” (without specifying which ones) restored performance even in the weaker local llama models back then.
The flaw was not in the reasoning; the flaw seems to be simply that the assumptions we make are often different from the assumptions it makes. I wonder if that might be a fundamental underlying cause of misalignment.
Low thinking is an artificial constraint. It can fail spectacularly on things that aren't in the training data.
It's a nonsensical question to ask, and how an LLM answers gives 0 signal.
If you were home and a family member asked you that question, you'd probably criticise the question rather than answering. LLM are RLHF'd into being milk-toast helpers that just try to answer questions like that with no criticism.
This is all beside the fact that the world of AI has changed pretty dramatically in the last few months.
It is so nonsensical because it has such an obvious answer. The answer is so obvious, in fact, that one answer can be considered nonsense and the other common sense.
*Milquetoast
This is just a stupid post.
It’s nonsense to test if a product that is marketed and sold as being able to provide generalised intelligence on demand, does what it says on the tin?
Check yourself
Since you're new here, I'd suggest you read the guidelines for etiquette.
https://news.ycombinator.com/newsguidelines.html
1 reply →