Comment by colinhb
3 days ago
Opus 5 hasn't been available for that long - long enough for benchmarks, but not really use and develop a subjective view on
3 days ago
Opus 5 hasn't been available for that long - long enough for benchmarks, but not really use and develop a subjective view on
I suspect most of those comments on llms like the parents are generated by anthropic and openai to shape the discussion/mindset
They always give off the same astroturfing vibes that reddit became infested with after the early 2010s (just look at it's comment history)
Ofc unprovable for users. Ycombinatior could try to, but it'd just become a cat/mouse game which they'd likely lose because of the incentives
> just look at it's comment history
I checked one. Old account, nuanced takes, and shitting on everyone equally. The perfect HN user!
Ah, I really walked into that one. Yeah, the phrasing + placement of the remark implied that all comments are artificial.
That was not my actual intent, it was poorly expressed by me. I was specifically talking about the account which created the comment colinhb responded to. That'd make it the... Grand Grand Grand grandparent now I think?
That's not necessary. Sometimes we make choices and we feel the need to justify them in vivid flamewars regardless of how arbitrary. Vi vs emacs or amd vs intel, now anthropic vs openai.
We love to take sides, to belong to a small community of peers, no?
I have spent about 6 hours with it now and it is absurd to say it worse than 4.8. It is wonderful.
To me, there is this strange critique that seems to always happen now with a new model. Like shitting on the model for entertainment purposes seems more interesting to many than actually using the model.
It reminds me of looking up a new music album on youtube that has very few views with a review above it by The Needle Drop shitting on the album will have a few hundred thousands views.
I would rather just listen to the album and judge for myself.
I don’t know that individuals can really be expected to use a new model enough to develop a proper opinion though. If I try a new model, and it doesn’t seem as good as the one I’m using, I’m just going to stop using it. That’s not enough data to give anyone else a useful view on it, but it’s enough for me to make my mind up.
Especially because I’m probably trying it at work, and I can’t really justify using the company’s enterprise plan to develop my understanding of a model that I don’t think is going to be the one I use for my work.