Comment by andy12_
13 hours ago
All the people that are just writing an Jev-like API on top of a normal LLM are missing the point. What makes Jev special is the training data; it's how it's trained. The architecture is probably nothing special. Just a text encoder with parallel prediction branches.
I have tried many of these open-source Jev-like models on some linguistic tasks and they are so bad compared to Jev.
It won't be long until people produce a decent training data set generation pipeline.
The number of people working on this is crazy. Something will coalesce.
I hope so. And I would really like to try an actual Jev open source model. But it will make it more difficult to market it when someone releases something like that because of so many of these "open source Jev-like model".
I'm just sitting back for a few weeks / a couple months to let it shake out, let others put in all the work, and then see if people are still interested and finding use cases that this access model fits better than the usual chat completions endpoint people are used to.
1 reply →