Comment by gitpusher42
2 hours ago
Thank you!
afaik ollama relies on llama.cpp and mmap. mmap loads pages on demand and doesn't use the same explicit cache or parallel reads like my engine. Most likely ollama/llama.cpp will be way slower in this case
No comments yet
Contribute on Hacker News ↗