Comment by arjie

2 days ago

Most of us aren’t deploying these in general purpose use-cases. E.g. I use Qwen mostly for vision in my personal assistant. I have an eval set for that. Pretty much each of my use cases has a pre-computed problem set.

Likewise I have a tool-use eval set and a browser-use eval set. I use frontier models for most interactive tasks that are not home assistant or background agents.

I can imagine someone could build evals for that but I have never done so.

Deepseek v4 Pro or GLM 5.3 for software architecture I feel are only deployed for general, rather doubtful those are for narrow stuff, if we are sticking with this threads mentioned models. For small models, sure, narrowly targeted sets which can be self evaluated are amazingly valuable, but I feel beyond 500B we are in a different dimension. Rating any model the size of GLM or V4 Pro in hours I doubt is done beyond pure vibes.