← Back to context

Comment by 345f506fdc718

1 day ago

> 15 months ago, the top model on the Artificial Analysis index was GPT-o3. It scores 30 on the Artificial Analysis index.

There must be something really of with those benchmarks. Yes, hallucinations gotten better, but I don't see that the big frontier models got so much better in the last 12-18 Months. They just put out bigger wall of texts and feel smarter. But they still make way too many stupid errors

12 months ago "way too many stupid errors" was constant news. Today, you rarely hear about those anymore.

Sure, the novelty of the errors has worn off a bit and thus the reporting. Nevertheless the quality has improved immensely in this regard.

Also, AI video generation is now so good and accessible that it is very, very regularly used for memes, disinformation and proper (short) movie projects. AI image generation even more so (Mitch McConnell anyone?).

Pretending progress hasn't been mindboggling is insane.

  • Maybe it got a lot less and I just got used to it. True.

    Still feels too much for me. Breaks my workflow for no reason. Too much overhead for me, if I can't trust the output