Comment by benjaminsky2
6 hours ago
I’ve validated Jev’s confidence score. Accuracy scales linearly with confidence for the 3 use cases I tested. >.9 it matched a human labeler. I immediately discovered a user behavior I didn’t expect for ~$3. I can now mitigate in real-time due to low cost and latency. This may have a major positive financial impact for all our customers.
Not sure why anyone would feel the need to dunk on this thing without showing a real failure example.
If you’ve validated the Jev scores, then it means you’ve built and run the evals. The exact thing the author is arguing needs to be done if you’re using Jev.
So they’re not dunking on that, they’re dunking on the idea that you can just trust Jev confidence scores without actually doing any validation.
TBF OP makes a good point: skipping out on the whole fine-tune sing and dance and diving directly into a good general classifier is a barrier to developing actual, deep intuition for the problem space (a pretty prevalant anti-AI argument).