← Back to context

Comment by dennisy

13 hours ago

Are you able to share how it works in that case?

I'll tell you this. Output isn't too cheap to meter, there is no decoder.

  • So an encoder-only model with a classifier trained on the heads or something? DeepSeek recently switched to an encoder-decoder architecture in an attempt to get the best of both worlds (fast prefill while preserving generation capability), I wonder if that might be the future?