Comment by aka-rider
2 days ago
Smaller models are overconfident and have a hard time to self-correct.
If it’s stuck, usually that’s it.
Bigger models “understand” better, both the prompt and the contents. If you will try to read a paper together with a smaller model, the difference is immediately obvious.
Bigger models will “forget” and drift much less.
No comments yet
Contribute on Hacker News ↗