Comment by redman25
17 hours ago
Many older models are still better at "creative" tasks because new models have been benchmarking for code and reasoning. Pre-training is what gives a model its creativity and layering SFT and RL on top tends to remove some of it in order to have instruction following.
No comments yet
Contribute on Hacker News ↗