← Back to context

Comment by bonzini

10 hours ago

As of a few months ago they still have trouble, with low thinking, at the "should I drive to a car wash that is 100 m away" kind of question.

Simply appending “check your assumptions” to the question fixed it even back then: https://news.ycombinator.com/item?id=47040530

Similarly for Apple’s “red herring” paper, simply adding a generic caveat to “disregard irrelevant factors” (without specifying which ones) restored performance even in the weaker local llama models back then.

The flaw was not in the reasoning; the flaw seems to be simply that the assumptions we make are often different from the assumptions it makes. I wonder if that might be a fundamental underlying cause of misalignment.

Low thinking is an artificial constraint. It can fail spectacularly on things that aren't in the training data.

It's a nonsensical question to ask, and how an LLM answers gives 0 signal.

If you were home and a family member asked you that question, you'd probably criticise the question rather than answering. LLM are RLHF'd into being milk-toast helpers that just try to answer questions like that with no criticism.

This is all beside the fact that the world of AI has changed pretty dramatically in the last few months.