I saw Jev and without understanding it said to myself, “I can build that!” And brainstormed some weird Alex trebek jeopardy generator I called trebek, a bun ran typescript app that you give it an input, it decides if it’s a category, question or answer then generates what’s missing. Trained a SQLite-vec database on the English language for a few days with a small qwen embedding model to add vec embeds to the db then ripped the cord on the embeddings before I started feeding my llm the proper specs for Jev and now i have a cool jeopardy generator that now doubles as a Jev clone classifier with a custom v1 endpoint for system 0 or whatever it is. Level understanding here, fun experiment though ^>^ and produces useful outputs
That sounds like a lot of fun, love the creativity. Do you have a repo you can share? It would be nice to run it, and if you wanted, maybe I could add it to the demos on https://playground.jeffyclassify.com
Ha. I did something similar but made a Jeopardy judge for my family to play jeo-paw-dy hosted by the family dog. It uses our phones as buzzers and speech to text to let people just speak the answer
As an MLE who has been failing to get anyone interested in classifiers for many years, the hype around Jev makes me scream internally.
Yes, I get that a zero-shot classifier is more convenient than the traditional kind, it's very cool. Kind of. But then again plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now, so again I'm back to my internal screaming.
I think part of what made Jev catch on is that the API is like an if(...) or switch statement
People are just so used to the chat style APIs that they didn't even consider doing things like sending a bunch of emojis to a chat model and then asking for the optimal one in this context etc. Also chat models are pricier for the same behavior and can also output something random like a refusal
But yeah ironically I think in the initial breakthrough LLM paper on GPT-3 in 2020 some of the multiple choice questions were answered by comparing token probabilities of specific continuations rather than fill in the blank
Maybe I should have spent less time pitching to PMs and more time pitching ground up to devs, who have the right foundation to intuitively understand the usefulness.
> But then again plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now, so again I'm back to my internal screaming.
It's the scale of "perfectly cheap", jev (specifically) is so dirt cheap and fast that you can throw it at things that should not be justifiable in the past and you barely have to do any work other then quick testing.
I will also say, not having to get a team to build this for you, having to get business justification from your team for that teams hours, then spending time revising and testing that out and then you have to "prove" that it's worth having in your feature as a AI cost
Versus
"Let's put jev here and see how it works, if it works, then fantastic let's build a business case"
on the bright side, maybe this zero-shot classifier wave could be the back bone for more specialized classifiers (with more mindful selection of data and training)
> But then again plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now
I'm no prompt engineer, but I've found them to be too slow any time I wanted to use them that way. I have never tried Jev, but apparently it is supposed to be fast, so it seems like, according to the marketing, it could become usable where LLMs haven't been.
I also found inconsistent results (across time and similar inputs) which made it useless for me. I reverted to using LLMs for coding machine learning. I get deterministic outputs that make sense to me, but don’t have to know the ins and outs of how it works (given how it’s applied). I may try Jev in the future for a new problem but unlikely to refactor something that works.
Maybe instead of screaming you can take this as a chance to level up your engineering.
Good engineers don't treat approach as A == B or even A like B, when extremely integral parts of their applications differ.
Zero-shot isn't just "more convenient", in a low data regime: it's the only workable solution, and 100x so if your plan involves the acornym "BERT" (because even the largest of those models has the world knowledge of a fart to draw priors from)
Better ergonomics while being faster and cheaper as the existing things really is enough to justify callling what you've done a new thing, in a world of finite resources and time. It's actually making me scream how many people don't get that.
> in a low data regime: it's the only workable solution
Low data regimes no longer exist in the age of LLMs, one can trivially generate a training and eval set and distill a good classifier on any domain within a day.
But I do acknowledge zero-shot is more convenient. Personally I don't think Jev has any moat so I won't bother with their model specifically, but yes I do anticipate using this type of thing more in the future.
This is a great way to start but the perf of LLM models like Qwen are not ideal for local execution. I have adapted Laya (pure decision model) to run in a browser and I am able to get responses under 200ms. Give it a try: https://wexare-ai.github.io/browser-laya/
You can easily get under 100ms per decision in browser with a Qwen based decison model: https://alxnahas.github.io/strands-decider-web/?backend=engi...
There is nothing about Qwen or Laya that make one or the other better at running in the browser. It is really just: how large is your model and how does KV cache expand as input length increases
Laya is just ModernBert fine-tuned. It's still a language model (just a bidirectional one). Let's call it what it is instead of this weird Decision Model mysticism.
I am also not a big fan of everybody calling them system one models. AFAIK, system one also refers to activities like driving a car, but I don't want to have those decision models driving cars.
So I think I understand what is meant, but I don't like the comparison.
If you don't really care about knowing the specific probability (I think this is actually a huge info hazard), you can achieve this right now with a trivial tool calling arrangement. The benefit still stands. You are reducing the size of the action space from something that might be Turing complete (shell, code) into a multiple choice question.
To me the Jev moment isn't about Jev / Typesafe. It's articles like these. It's the letting a million Jevs bloom.
The bubble will come from a realization that a lot of people can train these models. AI Researchers will become more diffuse, work at more companies, and building a Jev or fine-tuning an LLM isn't some trillion dollar frontier lab exercise, but increasingly just something developers do.
decades ago you would train a classifier over a fixed set of classes. decision models are flexible regarding the possible outputs they can produce. also, you get text and images in input
But they’re only flexible because they can hallucinate an answer for anything. It’s like if you asked Gary Oldman to make all your decisions for you. He could do it, but he might have to make up answers when you get to more difficult topics
A coin flip can make infinite decisions on all possible input taxonomies.
Zero-shot is indeed very cool. And it's super useful/useful for the developer masses that don't know, care, or work pressures don't allow, for proper evaluation and calibration.
But it would be good that we don't over-hype these things.
This is very cool. If you are looking for something similar but more lightweight, that you can run (and train) on CPU, try out Jeffy: https://jeffyclassify.com/
I think Jev-like models are amazing for exploration and finding the right workflows, but the moment you have fixed classification tasks, it’s often more efficient to use an adhoc classifier, which you can quickly and easily train on CPU with not that much data (you can get an email classifier to 95% accuracy/f1 with 50-100 emails)
Edit: Would love to somehow mix both approaches automatically and have a general model which can take novel tasks, but then switch to a classifier after it gets enough data for training an adhoc model
Saw the question "Where would you most likely find a bat?", it occurs to me there's an innate tension, do we want the llm to be factually correct or do we want it to be more average human like? As an average human being not a sme on the subject my first instinct answer would be cave as well. I think it's reasonable to expect trainning on the aggregate of the internet means it would arrive at the same answer.
Edit: The context of the question does indeed make it sound more like the animal bat. The other answers sound more like gotchas to me.
The point is there is no right answer until we want it to be accurate in the domain we're working in. Like a sporting goods service would definitely want the baseball answer.
Does someone have examples of interesting stuff that has been built utilizing Jev/decision models? The way this is hyped up surely there must be some good stuff?
This started as a fun experiment, but I now use Jev on my mac as an advanced auto-correct. Much better than any other option I tried. https://levmiseri.com/nospace (the no space being more of a gimmick, but the autocorrection is good)
I am working on a game https://imperiaquiz.com/en that needs over 20,000 questions. I use jev to classify 'has statement' and 'can be a question' paragraphs/chunks taken from cli script to reduce token usage. Meaning, an LLM only starts work once I've chunked text and marked it as 'to be reviewed' instead of parsing the full content. This reduces tokens usage at least 10x on average
I see Jev as a major step towards commoditizing current LLM paradigm. One thing would be to further optimize this particular route to work purely on CPU. This will grant an option to embed this feature into any application, from MS Office to games. The other is integrating this into agent workflow to vastly minimize token consumption.
It would be interesting to see if a coding LLM trained for tool calling like GLM 5.3 would work well as a Jev model, or maybe even a flow where a model generates options (eg for a plan) and then uses a Jev to refine/optimize the path. Or similarly where else in a harness they’d help.
Not in a serious manner but I created a testing harness for a Nintendo 3DS game I'm making that uses the OpenAI Decisions API. The main advantage is the speed (~2-300ms per input) which I really need for this purpose.
I had the exact same thought "I can build that!" one month ago. So I built JobFit, a CV/job-post fit-scoring typed-decision model, training pipeline, and web app that runs in your browser: https://github.com/gw0/jobfit-model
if it works it works!
as long as the test set is reasonably large and diverse its better than nothing. You could characterize how robust it is by throwing dozens of different types of work at it and see how much the confidence varies
I think there's a huge problem with people getting into the machine learning field with the AI boom.
In prior settings, there used to be a clear separation of training, development/validation, testing partitions of any given task benchmark. The reason for this is so that you can tune hyperparameters: during training (e.g. learning rate), or after a training run (e.g. calibration), and then once you evaluate your system (could include the model and other pre/ post processing), that was it. The test set performance is the number you report.
There is a rationale behind this workflow, because when demonstrating a method, if you're adjusting ANY part of your system's performance against the result you finally report, you're overfitting to the test set.
Suppose you report a 90% performance on the test set, someone reading that would reasonably assume that the system works more often than it doesn't. But if you've overfit any part of your system (the prompt, the calibration, etc.), you could be tuning a 10% performance to 90%, shrugging and saying "Hey if it works it works!" and then happily reporting that number. Applying that same system to some other data that doesn't have the same quirks of this test set will fail.
How Jev manages to claim calibrated probabilities is beyond me. Calibrated to what?
Just for the author: On any of my iOS 26 browsers (orion, brave, safari), a page reload occurs when the model download completes, which resets the state, so I never get to interact with the model.
Nice read! I also gave a shot building one on Gemma3 and Gemma4. It was a fun exercise. I’m sharing it if anyone interested, Gemma4 based one is on a branch: https://github.com/onatm/gev
Take ModernBert, add a fully connected layer with 255 outputs, take a bunch of classification datasets from Huggingface, write the code to have the datasets fit the jev format on these 255 outputs, do supervised fine tuning on the datasets with that format. then use a confidence loss of some kind
Asking as a curious bystander: would that be sufficient?
I'm not familiar with ModernBert (my understanding stops around the original Bert), but it feels like this is asking it to do lots of heavy lifting. Can it do that much?
This is really cool to see. Being able to play the token generation was awesome. Amazing job with breaking down how to think about these models. This made the idea of Jev/decision models really easy to grasp for me. The idea of calibrating the model was helpful. I thought this was a great overview.
These are neat - and the source of the many Jev clones we've seen. I think their recent funding round is in part because of their algorithms/data. It remains to be seen if that's a big enough edge to be worth 1 billion+ dollars
Hey I have one question Is there any way we can also make a model, that can take a first decision about voice policing let's say I have speech to text tool like a whisper flow and if we want to make a voice policing decision model that can take first decision so that the voice policing works fast do it that can work out can you guys answer is.
this is just constrained generation? I thought the latest crop of decision models (inspired by Jev) do something fundamentally different in the architecture; they're not simply off-the-shelf models with a token mask
Indeed. Token masking just limits the costs associated with using an LLM. It therefore also limits the accuracy by limiting the amount of compute available.
But LLMs can mimick decision models, and I wouldn't be surprized if some labs are doing it this way at the moment.
I saw Jev and without understanding it said to myself, “I can build that!” And brainstormed some weird Alex trebek jeopardy generator I called trebek, a bun ran typescript app that you give it an input, it decides if it’s a category, question or answer then generates what’s missing. Trained a SQLite-vec database on the English language for a few days with a small qwen embedding model to add vec embeds to the db then ripped the cord on the embeddings before I started feeding my llm the proper specs for Jev and now i have a cool jeopardy generator that now doubles as a Jev clone classifier with a custom v1 endpoint for system 0 or whatever it is. Level understanding here, fun experiment though ^>^ and produces useful outputs
That sounds like a lot of fun, love the creativity. Do you have a repo you can share? It would be nice to run it, and if you wanted, maybe I could add it to the demos on https://playground.jeffyclassify.com
I'm 90% sure, but I keep getting surprised when I actually ask... /s?
2 replies →
Ha. I did something similar but made a Jeopardy judge for my family to play jeo-paw-dy hosted by the family dog. It uses our phones as buzzers and speech to text to let people just speak the answer
ha i made one to play the Oregon trail
As an MLE who has been failing to get anyone interested in classifiers for many years, the hype around Jev makes me scream internally.
Yes, I get that a zero-shot classifier is more convenient than the traditional kind, it's very cool. Kind of. But then again plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now, so again I'm back to my internal screaming.
I think part of what made Jev catch on is that the API is like an if(...) or switch statement
People are just so used to the chat style APIs that they didn't even consider doing things like sending a bunch of emojis to a chat model and then asking for the optimal one in this context etc. Also chat models are pricier for the same behavior and can also output something random like a refusal
But yeah ironically I think in the initial breakthrough LLM paper on GPT-3 in 2020 some of the multiple choice questions were answered by comparing token probabilities of specific continuations rather than fill in the blank
You know, it's a good point.
Maybe I should have spent less time pitching to PMs and more time pitching ground up to devs, who have the right foundation to intuitively understand the usefulness.
> But then again plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now, so again I'm back to my internal screaming.
It's the scale of "perfectly cheap", jev (specifically) is so dirt cheap and fast that you can throw it at things that should not be justifiable in the past and you barely have to do any work other then quick testing.
I will also say, not having to get a team to build this for you, having to get business justification from your team for that teams hours, then spending time revising and testing that out and then you have to "prove" that it's worth having in your feature as a AI cost
Versus
"Let's put jev here and see how it works, if it works, then fantastic let's build a business case"
1 reply →
on the bright side, maybe this zero-shot classifier wave could be the back bone for more specialized classifiers (with more mindful selection of data and training)
What are some pre-Jev classifier and ranker approaches? And how do you train them?
Jev is classification for normies. Look ma, no training.
Supermarkets are just agriculture for normies. Look ma, no subsistence farming.
1 reply →
I want to share the rage. Can you expound on what makes you scream?
For me, so much about building a traditional classifier goes into measuring and improving its performance on the data you’re making decisions about.
With Jev, we seem to have just ignored all that. There seems to be some magical thinking that, because it’s AI, its decisions must be accurate.
Which I don’t think is really justified given the narrowness of the benchmarks and breadth of tasks people want to use it for.
Idk, imagine you worked on Skype's B2B sales team for years and then COVID happens and Zoom blows up.
Is it a better thing? Yeah. Does it affect me in any tangible way? No.
But come on, really people? All you needed was like one tiny bell & whistle to take this from nothing to the hottest thing of all time?
2 replies →
> But then again plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now
I'm no prompt engineer, but I've found them to be too slow any time I wanted to use them that way. I have never tried Jev, but apparently it is supposed to be fast, so it seems like, according to the marketing, it could become usable where LLMs haven't been.
I also found inconsistent results (across time and similar inputs) which made it useless for me. I reverted to using LLMs for coding machine learning. I get deterministic outputs that make sense to me, but don’t have to know the ins and outs of how it works (given how it’s applied). I may try Jev in the future for a new problem but unlikely to refactor something that works.
Maybe instead of screaming you can take this as a chance to level up your engineering.
Good engineers don't treat approach as A == B or even A like B, when extremely integral parts of their applications differ.
Zero-shot isn't just "more convenient", in a low data regime: it's the only workable solution, and 100x so if your plan involves the acornym "BERT" (because even the largest of those models has the world knowledge of a fart to draw priors from)
Better ergonomics while being faster and cheaper as the existing things really is enough to justify callling what you've done a new thing, in a world of finite resources and time. It's actually making me scream how many people don't get that.
> Maybe instead of screaming you can take this as a chance to level up your engineering.
Probably also sales. System One thinking and other buzzwords will help.
> in a low data regime: it's the only workable solution
Low data regimes no longer exist in the age of LLMs, one can trivially generate a training and eval set and distill a good classifier on any domain within a day.
But I do acknowledge zero-shot is more convenient. Personally I don't think Jev has any moat so I won't bother with their model specifically, but yes I do anticipate using this type of thing more in the future.
3 replies →
> As an MLE
Yea but those models you were building were lame and inaccessible to play with for common devs.
Just because they have the same api doesnt mean you were building the same thing.
This is a great way to start but the perf of LLM models like Qwen are not ideal for local execution. I have adapted Laya (pure decision model) to run in a browser and I am able to get responses under 200ms. Give it a try: https://wexare-ai.github.io/browser-laya/
You can easily get under 100ms per decision in browser with a Qwen based decison model: https://alxnahas.github.io/strands-decider-web/?backend=engi... There is nothing about Qwen or Laya that make one or the other better at running in the browser. It is really just: how large is your model and how does KV cache expand as input length increases
I guess it depends on your application, but 100ms is pretty slow. Plus there is the model footprint (memory + cpu/gpu).
Why would a Qwen 1.7b be too slow? You should get well under 200ms there.
Laya is just ModernBert fine-tuned. It's still a language model (just a bidirectional one). Let's call it what it is instead of this weird Decision Model mysticism.
I am also not a big fan of everybody calling them system one models. AFAIK, system one also refers to activities like driving a car, but I don't want to have those decision models driving cars.
So I think I understand what is meant, but I don't like the comparison.
1 reply →
Is this mysticism what TypeSafe AI is valued at $7.5b for?
Maybe the people talking about an AI bubble do have a point. I am getting the impression here that investors are throwing money at everything.
Personally I believe AI labs have a solid business model that could soon be very profitable, but this here has me doubting now.
1 reply →
If you don't really care about knowing the specific probability (I think this is actually a huge info hazard), you can achieve this right now with a trivial tool calling arrangement. The benefit still stands. You are reducing the size of the action space from something that might be Turing complete (shell, code) into a multiple choice question.
To me the Jev moment isn't about Jev / Typesafe. It's articles like these. It's the letting a million Jevs bloom.
The bubble will come from a realization that a lot of people can train these models. AI Researchers will become more diffuse, work at more companies, and building a Jev or fine-tuning an LLM isn't some trillion dollar frontier lab exercise, but increasingly just something developers do.
So "decision model" is the new cool term for classifiers? You know, that thing that already existed decades ago
decades ago you would train a classifier over a fixed set of classes. decision models are flexible regarding the possible outputs they can produce. also, you get text and images in input
But they’re only flexible because they can hallucinate an answer for anything. It’s like if you asked Gary Oldman to make all your decisions for you. He could do it, but he might have to make up answers when you get to more difficult topics
3 replies →
If you think “bullet” is the cool new name for “arrow” then sure.
Decision models can perform zero-shot classification over almost unlimited text-input domains. This was science fiction decades ago.
A coin flip can make infinite decisions on all possible input taxonomies.
Zero-shot is indeed very cool. And it's super useful/useful for the developer masses that don't know, care, or work pressures don't allow, for proper evaluation and calibration.
But it would be good that we don't over-hype these things.
1 reply →
The difference is in the training, or more accurately what data the model is trained on.
This is very cool. If you are looking for something similar but more lightweight, that you can run (and train) on CPU, try out Jeffy: https://jeffyclassify.com/
On GitHub: https://github.com/nicobrenner/jeffy
Cool project too. If you are looking for a 500MB instead of gigabytes, with evals on Jevbench that you can run fast on CPU check out gutsy.
https://github.com/kouhxp/gutsy
Very cool, thank you for sharing
The banking77 numbers called my attention. Using a local classifier you can get 94%+ accuracy: https://playground.jeffyclassify.com/#model/banking77
I think Jev-like models are amazing for exploration and finding the right workflows, but the moment you have fixed classification tasks, it’s often more efficient to use an adhoc classifier, which you can quickly and easily train on CPU with not that much data (you can get an email classifier to 95% accuracy/f1 with 50-100 emails)
Edit: Would love to somehow mix both approaches automatically and have a general model which can take novel tasks, but then switch to a classifier after it gets enough data for training an adhoc model
1 reply →
Saw the question "Where would you most likely find a bat?", it occurs to me there's an innate tension, do we want the llm to be factually correct or do we want it to be more average human like? As an average human being not a sme on the subject my first instinct answer would be cave as well. I think it's reasonable to expect trainning on the aggregate of the internet means it would arrive at the same answer.
Edit: The context of the question does indeed make it sound more like the animal bat. The other answers sound more like gotchas to me.
The point is there is no right answer until we want it to be accurate in the domain we're working in. Like a sporting goods service would definitely want the baseball answer.
Does someone have examples of interesting stuff that has been built utilizing Jev/decision models? The way this is hyped up surely there must be some good stuff?
This started as a fun experiment, but I now use Jev on my mac as an advanced auto-correct. Much better than any other option I tried. https://levmiseri.com/nospace (the no space being more of a gimmick, but the autocorrection is good)
I am working on a game https://imperiaquiz.com/en that needs over 20,000 questions. I use jev to classify 'has statement' and 'can be a question' paragraphs/chunks taken from cli script to reduce token usage. Meaning, an LLM only starts work once I've chunked text and marked it as 'to be reviewed' instead of parsing the full content. This reduces tokens usage at least 10x on average
I see Jev as a major step towards commoditizing current LLM paradigm. One thing would be to further optimize this particular route to work purely on CPU. This will grant an option to embed this feature into any application, from MS Office to games. The other is integrating this into agent workflow to vastly minimize token consumption.
So, what exactly is the
> interesting stuff that has been built utilizing Jev/decision models
in that case?
It would be interesting to see if a coding LLM trained for tool calling like GLM 5.3 would work well as a Jev model, or maybe even a flow where a model generates options (eg for a plan) and then uses a Jev to refine/optimize the path. Or similarly where else in a harness they’d help.
Not in a serious manner but I created a testing harness for a Nintendo 3DS game I'm making that uses the OpenAI Decisions API. The main advantage is the speed (~2-300ms per input) which I really need for this purpose.
Not with Jev, but you can use classifiers for a lot of use cases, here’s a few: https://playground.jeffyclassify.com/
Also like to see a "layman harness" like LM Studio integrate an open decision model in its workflow.
use it to classify a bunch of research papers (by feeding it section by section, or summaries of sections if too long)
works like a charm
(not a product, so not much to share)
I use Jev in a Claude code hooks to detect dangerous commands.
I had the exact same thought "I can build that!" one month ago. So I built JobFit, a CV/job-post fit-scoring typed-decision model, training pipeline, and web app that runs in your browser: https://github.com/gw0/jobfit-model
Is simply changing the temperature so that the model appears calibrated over a particular benchmark after the fact “allowed”? Feels p-hacking esque.
if it works it works! as long as the test set is reasonably large and diverse its better than nothing. You could characterize how robust it is by throwing dozens of different types of work at it and see how much the confidence varies
I think there's a huge problem with people getting into the machine learning field with the AI boom.
In prior settings, there used to be a clear separation of training, development/validation, testing partitions of any given task benchmark. The reason for this is so that you can tune hyperparameters: during training (e.g. learning rate), or after a training run (e.g. calibration), and then once you evaluate your system (could include the model and other pre/ post processing), that was it. The test set performance is the number you report.
There is a rationale behind this workflow, because when demonstrating a method, if you're adjusting ANY part of your system's performance against the result you finally report, you're overfitting to the test set.
Suppose you report a 90% performance on the test set, someone reading that would reasonably assume that the system works more often than it doesn't. But if you've overfit any part of your system (the prompt, the calibration, etc.), you could be tuning a 10% performance to 90%, shrugging and saying "Hey if it works it works!" and then happily reporting that number. Applying that same system to some other data that doesn't have the same quirks of this test set will fail.
How Jev manages to claim calibrated probabilities is beyond me. Calibrated to what?
What is the intuition behind the temperature fitting for calibration step?
Flattening the output probability distribution curve makes some sense, but playing with the graph doesn't seem to show the 3.797 figure as the best.
And how do you curve fit for this single example?
https://en.wikipedia.org/wiki/Platt_scaling
Just for the author: On any of my iOS 26 browsers (orion, brave, safari), a page reload occurs when the model download completes, which resets the state, so I never get to interact with the model.
The memory limit for Safari on iOS is less than 400 MB so it's probably because of that
Nice read! I also gave a shot building one on Gemma3 and Gemma4. It was a fun exercise. I’m sharing it if anyone interested, Gemma4 based one is on a branch: https://github.com/onatm/gev
I'm looking forward to a "Build a Decision Model (From Scratch)" book.
Take ModernBert, add a fully connected layer with 255 outputs, take a bunch of classification datasets from Huggingface, write the code to have the datasets fit the jev format on these 255 outputs, do supervised fine tuning on the datasets with that format. then use a confidence loss of some kind
Asking as a curious bystander: would that be sufficient?
I'm not familiar with ModernBert (my understanding stops around the original Bert), but it feels like this is asking it to do lots of heavy lifting. Can it do that much?
Great guide! It works, made a 500MB model for CPU based off qwen 3.5 0.8B. It plays maze games like Pac man https://github.com/kouhxp/gutsy
This is really cool to see. Being able to play the token generation was awesome. Amazing job with breaking down how to think about these models. This made the idea of Jev/decision models really easy to grasp for me. The idea of calibrating the model was helpful. I thought this was a great overview.
These are neat - and the source of the many Jev clones we've seen. I think their recent funding round is in part because of their algorithms/data. It remains to be seen if that's a big enough edge to be worth 1 billion+ dollars
A lot of LLM "thinking" involve choosing between alternatives.
Offload the decision-making parts of LLM reasoning to a jev like model.
Hey I have one question Is there any way we can also make a model, that can take a first decision about voice policing let's say I have speech to text tool like a whisper flow and if we want to make a voice policing decision model that can take first decision so that the voice policing works fast do it that can work out can you guys answer is.
this is just constrained generation? I thought the latest crop of decision models (inspired by Jev) do something fundamentally different in the architecture; they're not simply off-the-shelf models with a token mask
Indeed. Token masking just limits the costs associated with using an LLM. It therefore also limits the accuracy by limiting the amount of compute available.
But LLMs can mimick decision models, and I wouldn't be surprized if some labs are doing it this way at the moment.
If a lab doesn't know how to put a new output head on a transformer, they shouldn't be considered a lab.
Great article, thanks!
Is this only about getting fix json output?
it doesnt generate text though, only probabilities
- i have to keep scrolling down on your home page https://nishtahir.com/ to see what posts you have
- could you kindly put all that in a /blog page with pagination and not infinite scroll?
[flagged]
[dead]
[dead]
[flagged]
[flagged]
[dead]