Comment by catlifeonmars
1 hour ago
I think that’s a bad example, as classifiers tend to output a floating point number and you use some threshold/activation function to collapse the classifier into a particular state. In that sense there is nothing qualitatively different when the classifier outputs 0.85 vs 0.87.
LLMs also output a distribution over tokens at the output. I don't see the point you're making.