Comment by Neywiny
3 hours ago
I just started using opus 4.8 and it's still trash at very domain specific stuff and debugging. It went down a rabbit hole trying to coerce the algorithm to work for 1 data point, which dramatically overfit the code. Later it did 11 experiments to try and figure out how to do the math, each one yielding an accuracy no better than a coin flip and each one it asserted beforehand that it had it this time, 100% accuracy.
We've got some time.
That's last gen though. I agree it's mediocre. Try Astra. It's significantly better than Opus and even Fable. I'm not sure about earlier OpenAI models.
I also thought we had time.