Comment by aesthesia
15 hours ago
One way to interpret these results is that the LLMs tested are badly calibrated for this kind of multi-armed bandit problem. Even if the intent is for the model to find and exploit patterns, it's bad at doing it (or rather, at recognizing that there is not in fact any pattern).
It may be bad at recognizing it, but if all arms are equally good, that doesn't matter.
[dead]