← Back to context

Comment by gpugreg

3 hours ago

    > I was training custom ML models back in 2017.

Maybe you are a good person to ask my question then. I have not looked into Jev much, but is it much different from using a regular LLM and constraining its token output to the action space? (e.g. like using llama.cpp's GBNF grammars). Is it just that Jev's "confidence scores" are significantly better than the softmaxed logits? Or is there something else I am missing?

What you're missing is: cost and speed. Otherwise, it's very much like running an LLM and constraining the output.

I don't know how good Jev's "confidence scores" are, but I would be surprised if they were in any sense better than logits from some good LLM. One advantage of Jev here is that the confidence scores are easy to access. Most LLM API providers don't provide an easy/convenient way to access the logits. But that's a minor point, you could of course build something like this with LLMs (and many people have).

  • llama.cpp is already extremely fast for single-token responses (<5 ms). I can't see Jev being faster when taking network latency into account, except maybe for multimodal inputs.

    • Sounds like you are running a tiny toy model if you can get generations in under 5 ms? Typical response times from LLMs for typical "jev-like" queries from OpenAI and Anthropic are 2s-10s. Not milliseconds. Seconds. Same queries from Jev are like 0.2s. and the cost is 1000x.