Comment by baobabKoodaa
6 hours ago
Everything you said here is false. No, I didn't enter the AI space during the vibe coding era. I was training custom ML models back in 2017. And no, there haven't been models comparable to Jev before Jev was published.
Jev is:
- accurate
- general purpose
- fast and cheap
Models we had before Jev had at most 2/3 of above qualities, but none of them were 3/3.
Maybe you are a good person to ask my question then. I have not looked into Jev much, but is it much different from using a regular LLM and constraining its token output to the action space? (e.g. like using llama.cpp's GBNF grammars). Is it just that Jev's "confidence scores" are significantly better than the softmaxed logits? Or is there something else I am missing?
What you're missing is: cost and speed. Otherwise, it's very much like running an LLM and constraining the output.
I don't know how good Jev's "confidence scores" are, but I would be surprised if they were in any sense better than logits from some good LLM. One advantage of Jev here is that the confidence scores are easy to access. Most LLM API providers don't provide an easy/convenient way to access the logits. But that's a minor point, you could of course build something like this with LLMs (and many people have).
llama.cpp is already extremely fast for single-token responses (<5 ms). I can't see Jev being faster when taking network latency into account, except maybe for multimodal inputs.
1 reply →
> Models we had before Jev had at most 2/3 of above qualities, but none of them were 3/3.
You're the one being deceptive here. Jev is trading accuracy, speed, and cost for generality. It's less accurate, slower and more expensive than trained classifiers. So it's still 2 out of 3, but with decimals. Maybe 2.2 out of 3 if I'm being charitable.
And the reason we didn't have that before is because nobody thought it's a good tradeoff.
When you say "trained classifiers", you are referring to models which are trained (or fine tuned) to work on one specific problem, right? That is the opposite of "general purpose".
Would Jev be more accurate in a specific task if it had been developed only for that task, as opposed to general purpose? Of course it would. So, sure, Jev is trading accuracy for generality. According to you "nobody thought it's a good tradeoff", which again is false, there was huge demand for a cheap and accurate general purpose classifier.
> work on one specific problem, right? That is the opposite of "general purpose".
A business doesn't need Jev for the sake of Jev. Most business are solving specific problems.
And fine-tuning got a lot cheaper these days - I've seen claims here on HN that ~500 examples is enough to beat Jev.
> "nobody thought it's a good tradeoff", which again is false, there was huge demand for a cheap and accurate general purpose classifier
There wasn't. The hope is that there was a latent demand, but we've yet to see if it's truly latent or just manufactured.
Noone is saying "hell yeah, finally we got a general purpose classifier, my business needed it so much". The typical message is "this seems cool, let me see where I can apply it".
The fact that name itself is a play on Jevons Paradox illustrates that there was no demand until Jev was released.
5 replies →
>>I was training custom ML models back in 2017
but you truly do sound like an angry 19 year old from your arguments.
- accurate - on what? on trust me bro benchmarks?
- zero-shot model are fundamentally general purpose.
- fast and cheap ; models on hf are FREE and fast enough.
"Models on hf" (unspecified) are "FREE"? Like "free to download"? Sure, but nobody was talking about that. They cost money to run inference on. Unless you are talking about some tiny toy models that are useless for any non-toy problems. You clearly don't have any idea what you're talking about. Just stop, man.