← Back to context

Comment by howunfortunate

1 day ago

As an MLE who has been failing to get anyone interested in classifiers for many years, the hype around Jev makes me scream internally.

Yes, I get that a zero-shot classifier is more convenient than the traditional kind, it's very cool. Kind of. But then again plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now, so again I'm back to my internal screaming.

I think part of what made Jev catch on is that the API is like an if(...) or switch statement

People are just so used to the chat style APIs that they didn't even consider doing things like sending a bunch of emojis to a chat model and then asking for the optimal one in this context etc. Also chat models are pricier for the same behavior and can also output something random like a refusal

But yeah ironically I think in the initial breakthrough LLM paper on GPT-3 in 2020 some of the multiple choice questions were answered by comparing token probabilities of specific continuations rather than fill in the blank

  • You know, it's a good point.

    Maybe I should have spent less time pitching to PMs and more time pitching ground up to devs, who have the right foundation to intuitively understand the usefulness.

> But then again plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now, so again I'm back to my internal screaming.

It's the scale of "perfectly cheap", jev (specifically) is so dirt cheap and fast that you can throw it at things that should not be justifiable in the past and you barely have to do any work other then quick testing.

  • I will also say, not having to get a team to build this for you, having to get business justification from your team for that teams hours, then spending time revising and testing that out and then you have to "prove" that it's worth having in your feature as a AI cost

    Versus

    "Let's put jev here and see how it works, if it works, then fantastic let's build a business case"

Around 8 years ago I built https://www.taggit.io/ and concluded that the market just wasn't there... Turns out it just wasn't there yet. I know the frustration you're talking about, pitching to PMs. Maybe the iron is hot again, and SotA perf/problem has definitely changed. Would love to chat and commiserate, hit me up if you're interested.

on the bright side, maybe this zero-shot classifier wave could be the back bone for more specialized classifiers (with more mindful selection of data and training)

I want to share the rage. Can you expound on what makes you scream?

  • For me, so much about building a traditional classifier goes into measuring and improving its performance on the data you’re making decisions about.

    With Jev, we seem to have just ignored all that. There seems to be some magical thinking that, because it’s AI, its decisions must be accurate.

    Which I don’t think is really justified given the narrowness of the benchmarks and breadth of tasks people want to use it for.

  • Idk, imagine you worked on Skype's B2B sales team for years and then COVID happens and Zoom blows up.

    Is it a better thing? Yeah. Does it affect me in any tangible way? No.

    But come on, really people? All you needed was like one tiny bell & whistle to take this from nothing to the hottest thing of all time?

    • I bet the timing was the key, people in the last year have been furiously building things that use LLMs as an API and as we work on this stuff we have systems with N LLM steps and M of them are frustrating because you want a specific choice picked or list of things ranked and sometimes the llm will just output something else entirely!

      So along comes this thing you can graft in that is more reliable, faster, and cheaper for that, and I get it immediately.

      A year ago I'd be like "kinda cool but what it for?"

      2 replies →

> But then again plain LLMs have been perfectly cheap and serviceable as zero-shot classifiers for quite some time now

I'm no prompt engineer, but I've found them to be too slow any time I wanted to use them that way. I have never tried Jev, but apparently it is supposed to be fast, so it seems like, according to the marketing, it could become usable where LLMs haven't been.

  • I also found inconsistent results (across time and similar inputs) which made it useless for me. I reverted to using LLMs for coding machine learning. I get deterministic outputs that make sense to me, but don’t have to know the ins and outs of how it works (given how it’s applied). I may try Jev in the future for a new problem but unlikely to refactor something that works.

Maybe instead of screaming you can take this as a chance to level up your engineering.

Good engineers don't treat approach as A == B or even A like B, when extremely integral parts of their applications differ.

Zero-shot isn't just "more convenient", in a low data regime: it's the only workable solution, and 100x so if your plan involves the acornym "BERT" (because even the largest of those models has the world knowledge of a fart to draw priors from)

Better ergonomics while being faster and cheaper as the existing things really is enough to justify callling what you've done a new thing, in a world of finite resources and time. It's actually making me scream how many people don't get that.

  • > in a low data regime: it's the only workable solution

    Low data regimes no longer exist in the age of LLMs, one can trivially generate a training and eval set and distill a good classifier on any domain within a day.

    But I do acknowledge zero-shot is more convenient. Personally I don't think Jev has any moat so I won't bother with their model specifically, but yes I do anticipate using this type of thing more in the future.

    • > one can trivially generate a training and eval set and distill a good classifier on any domain within a day.

      This might be possible, but it’s not ‘trivial’. Even getting my favorite agent to do it would involve a lot choices about what to tell the agent to do, and a bunch of testing.

      With Jev, I can just write my questions and get an answer - or get an agent to write code that generates questions and gets answers. The answers are pretty good! And the whole thing is pretty trivial to understand (where ‘trivial’ here is an order of magnitude more ‘trivial’ than in your comment).

    • I don't think jev has any moat either, but interestingly the distillation you described sounds like something imo someone would pay for. Convenience again. Personally I wouldn't. I suppose that just makes it a commodity.

    • > Low data regimes no longer exist in the age of LLMs, one can trivially generate a training and eval set and distill a good classifier on any domain within a day.

      They do exist and you can't easily generate them. Distilling is also pointless.

  • > Maybe instead of screaming you can take this as a chance to level up your engineering.

    Probably also sales. System One thinking and other buzzwords will help.

> As an MLE

Yea but those models you were building were lame and inaccessible to play with for common devs.

Just because they have the same api doesnt mean you were building the same thing.