← Back to context

Comment by catlifeonmars

1 hour ago

I think that’s a bad example, as classifiers tend to output a floating point number and you use some threshold/activation function to collapse the classifier into a particular state. In that sense there is nothing qualitatively different when the classifier outputs 0.85 vs 0.87.

LLMs also output a distribution over tokens at the output. I don't see the point you're making.