← Back to context

Comment by aitchnyu

9 hours ago

Load-bearing (whoops) quirks were noticeable months back, but haven't most flagship models become predictable and reliable?

it does seem to be moving in that direction. There were really specific things (large, complex json outputs) that gemini-2.5 flash was basically the only model that seemed capable of reliably for a long period. gpt-5+ has covered the usecase for us now pretty well but still evals slightly below what 2.5 could do