Comment by prabhanjana_c
1 day ago
On my RTX3060 - 12GB VRAM + 24GB RAM , with below command
ollama run qwen3.8:27b --verbose "explain mmap”,
I got 2.41 Tokens/s, Not sure if that can be improved considering VRAM doesn’t fit the entire, model.
Additional details:
total duration: 8m18.2870918s load duration: 612.105ms prompt eval count: 12 token(s) prompt eval duration: 2.900965s prompt eval rate: 4.14 tokens/s eval count: 1193 token(s) eval duration: 8m14.660618s eval rate: 2.41 tokens/s
System spec: NVIDIA GeForce RTX3060 AMD Ryzen 5 1600 Six-Core B450 AORUS M Mother board. NVIDIA-SMI 620.02 Driver:620.02, CUDA Version: 13.2
No comments yet
Contribute on Hacker News ↗