← Back to context

Comment by sidclaw

2 days ago

Thoughtful demo. If possible, I'd want to see next is whether the probabilities it returns are calibrated and not just whether the top choice is right. I ran a small experiment on this, compared base and instruct models' calibration on option-letter logits, then fine-tuned with Brier loss vs cross-entropy. Have you looked at a reliability curve for this version?