Comment by keeda
6 hours ago
Simply appending “check your assumptions” to the question fixed it even back then: https://news.ycombinator.com/item?id=47040530
Similarly for Apple’s “red herring” paper, simply adding a generic caveat to “disregard irrelevant factors” (without specifying which ones) restored performance even in the weaker local llama models back then.
The flaw was not in the reasoning; the flaw seems to be simply that the assumptions we make are often different from the assumptions it makes. I wonder if that might be a fundamental underlying cause of misalignment.
No comments yet
Contribute on Hacker News ↗