Comment by nilsherzig
2 hours ago
Fyi, if you're trying to run this under AMD/HIP:
PTQ1_0 has no optimized MMQ-Path in their llama-cpp fork, try running PTQ2_0 (needs a bit more vram, but is about 2x faster on my 6700 XT)
https://gist.github.com/nilsherzig/b8266d001c5c01bdb3d81d209...
No comments yet
Contribute on Hacker News ↗