Comment by schappim

14 hours ago

> "frequently contains all too plausible nonsense"

This really isn’t the case with frontier models in 2026.

I’ve found (sadly) that every time I thought the model was hallucinating, I was in fact the one who was mistaken.

N=1, and biased towards the type of questions asked. N+1, If I use one of the search engine sloptools I get frequent inaccurate answers which is I guess what the majority of people do.

Sure, your frontier model in 2026 will know in the team who is responsible for what part of the tech stack.

No?

Then, in the document where it writes who has to do what change, it is hallucinating. And this time you or me are not the ones who are mistaken.

I am prepared to admit that the model usually doesn't get things outright incorrect (unless you ask it to count letters). But the solution does contain a lot of nonsense. Not false nonsense, but meaningless or irrelevant sentences that make understanding the core of the fix much more difficult.