Comment by kantahayashi
4 days ago
Yes. There's no problem with choosing the same face every time. The problem is the probability it attached to the choice. Jev gave face 1 an 83% probability while the true probability is 1/6.
4 days ago
Yes. There's no problem with choosing the same face every time. The problem is the probability it attached to the choice. Jev gave face 1 an 83% probability while the true probability is 1/6.
Okay, I see, you're expecting Jev to properly give 1/6 probability for each option. This is different from my intuition of how LLMs work, where their probabilities don't really work like this (I would expect LLM to also do something like 0.83 for 1).
That's right. It's normal behavior of LLMs. But what matters is TypeSafe argues it's different exactly on this point. The selling point of Jev is "calibrated probabilities", so I checked it on probability problems.
Jev and LLMs give other promises. Jev's RLCD training aims to make its probabilities calibrated such that given many cases where it assigns label Y about X% probability, Y should be the correct label about X% of the time.
One of the entire value props for Jev is that is exactly how it is supposed to work. That is one of the big claimed advantages over regular LLMs.
Do you provide Jev that the probability is 1/6 and yet it gives back a probability that is way off?
Yes. For example, one of the prompts said "The die is unbiased: each of the six faces has probability exactly 1/6."