← Back to context

Comment by hbrn

9 hours ago

Yeah, but I wouldn't be surprised OpenAI's decision API is a post-trained Luna with confidence calibration.

Typesafe claims that Jev is calibrated, but there are plenty of examples where it completely fails (predicting die roll being the most obvious one).

Unfortunately calibration is hard to benchmark.

The dice roll prediction is about the way the prompt is setup misunderstanding how Jev works (they treat the confidence score as a probability score, which it isn't).

If you instead give it a list of probability for each number and ask it whats the probability of each number, the result will be accurate.

  • > give it a list of probability for each number and ask it whats the probability of each number

    Did i hear that correctly? In order for Jev to be accurate you have to give it the answer before asking for the answer?

    (btw this is exactly how Jev is playing games).

    • That’s circular, sure. But if we think about potential real-world tasks where someone might, say, use it to classify on “does this advice correspond to our policy docs”, you’d absolutely push the answer (the policy docs) into its context..

      What am I missing?

      As in, neither LLMs nor Jev are truth engines. Truth comes from the provided context.. Plus weights.. kind of fuzzy, sorry I’m thinking out “loud”