Comment by aitchnyu
9 hours ago
Load-bearing (whoops) quirks were noticeable months back, but haven't most flagship models become predictable and reliable?
9 hours ago
Load-bearing (whoops) quirks were noticeable months back, but haven't most flagship models become predictable and reliable?
it does seem to be moving in that direction. There were really specific things (large, complex json outputs) that gemini-2.5 flash was basically the only model that seemed capable of reliably for a long period. gpt-5+ has covered the usecase for us now pretty well but still evals slightly below what 2.5 could do