Comment by zenoprax
9 hours ago
9070XT operator here: I'm using llama.cpp with the same model and quant and I'm getting 87,000 for my context limit. I tried the Unsloth models but they lowered it to around 30-40K so I went back to upstream.
I'm on Linux and using some sort of unholy mess of ROCM libraries that I don't understand.
No comments yet
Contribute on Hacker News ↗