Comment by Forgeties79
2 hours ago
I’m a little more novice than a lot of the people on this site so take my response with a grain of salt.
The wall I keep hitting is I can run models like I described (Q3-4 usually), but it’s very sensitive to context. Once I start getting past 7.5k or so it can really fall apart. Sometimes before that. It just really depends.
If I run smaller ones that offload less to ram, they stay somewhat coherent but don’t quite do what I want them to do.
Your token speeds are not that much slower than mine. I imagine part of it is I’m not fine-tuning it very well. On a good day I’ll get like…15tps.
No comments yet
Contribute on Hacker News ↗