Comment by dkersten
5 hours ago
Most of them appear to be small LLM’s fine tuned for the role.
That’s a different set of properties in terms of size, cost, and latency. Jev (apparently, not like I’ve seen its insides) is extremely cheap, extremely fast, doesn’t cost any output tokens as it speaks the output natively, can’t get the output wrong because it speaks the format natively, and (presumably based on the docs), the context is separate from the question, meaning it should be immune (or at least highly resistant) to prompt injection attacks.
It’s not just about the accuracy of the result, it’s a collection of all the properties that make Jev interesting.
Jev took years to develop, I strongly doubt that a copycat that was put together within days after Jev’s release will be able to match it on a sun of its properties. Even if fine tuned LLMs can outperform it on raw accuracy.
Jev is something your favorite LLM could zero-shot months ago, if you pointed it to the right arXiv paper (some of which are linked in this thread).
That probably explains why there were so many competitors around withing days of the Jev announcement. They are not starting with a moat, and there doesn't seem to be any moat in sight. Just buzzword recognition because everything is comparing to "jev".
There will always only be a very small percentage of people who want to discover and build things, even with llms, its a very small subset of people though, and most people want off the shelf solutions.
Also Jevs purpose isnt to become its own thing. It will get aquired in 18 months by one of Andressen Horowitz's incestuous circle of companies and everyone will make money, and the person who buys it wont necessarily care if Jev itself makes them a ton of money. They're just passing chips around the table.
Jev did not take years to develop. What it does was published in arxiv back in 2025. TypeSafe just marketed it.
What specific arxiv paper are you referencing here?
https://arxiv.org/abs/2503.23303
and https://arxiv.org/abs/2510.01237
6 replies →
> I strongly doubt that a copycat that was put together within days after Jev’s release will be able to match it on a sun of its properties
But why? If the simplest way to achieve Jev's capabilities (accuracy, cost, latency) is by fine tuning a small model, what makes you think that this isn't exactly what Typesafe did?
And even if they did something different - what makes you think it was a good idea in the first place, given how easy their results were replicated without any "secret sauce"?
My point is that it’s not replicated. You replicate the accuracy, but not the other properties. The Jev-competitors only proved that you can get or beat the accuracy, nothing about the other properties. Especially the “zero hallucination” output and the (if it works how the documentation make it sound) prompt injection resistant architecture. You can’t get that with a fine tuned LLM.
> zero hallucination
Plenty has been said about this claim. If you're still falling for this, I feel sorry for you.
If you remove wheels from your car, your car will get a "no speeding ticket" property, and yet there's nothing exciting about it.
> You can’t get that with a fine tuned LLM
Of course you can. All these claims are nothing but marketing.
I understand most major model providers support passing a JSON schema that is strictly followed in the output, accomplishing the same 'zero hallucination' and prompt injection resistance.
The only difference I am aware of is that probabilities are better calibrated with these decision models compared to regular LLMs which can output hallucinated numbers where your schema allows a number.