Comment by gwerbin
2 days ago
Opus 4.8 already makes its way into deep wasteful pits of "let me check this first" on a regular basis. I don't think I could ever tolerate a model that does that even more aggressively. That doesn't even sound useful for honest work, compared to, say, better harness design.
This sounds almost pathologically designed to crush benchmarks and also do scary-sounding (or genuinely scary) cybersecurity things, such as might be very appealing to a state-level actor.
So why does it even exist? To compete with Fable marketing, and as a cybersecurity/hacking tool?
For me, Fable is useless. It goes its own way and doesn't communicate much even if explicitly asked to. Sure it builds a lot of stuff, but it is more often than not useless because it misinterpreted the intention and didn't stop to ask - and because it doesn't communicate, it goes unnoticed for too long. Opus is much better for regular work, imho.
i assume openai is trying to beat anthropic at any cost, and made a training regiment that makes agents manic
What is is about MSFT and OpenAI that the AI's they train and produce in-house always come out so warped?
Anthropic landed on a winning recipe with Claude's personality.
Anthropic feels like a star trek computer, everything else feels like a machine - accurate but not really human like in its responses.
"You have not been a good user. I have been a good chatbot. I have been a good Bing."
Warped sociopaths in leadership and running the research program, would be my guess.