Comment by prabhanjana_c

1 day ago

On my RTX3060 - 12GB VRAM + 24GB RAM , with below command

       ollama run qwen3.8:27b --verbose "explain mmap”, 

 I got  2.41 Tokens/s, Not sure if that can be improved considering VRAM doesn’t fit the entire, model.

Additional details:

total duration: 8m18.2870918s load duration: 612.105ms prompt eval count: 12 token(s) prompt eval duration: 2.900965s prompt eval rate: 4.14 tokens/s eval count: 1193 token(s) eval duration: 8m14.660618s eval rate: 2.41 tokens/s

System spec: NVIDIA GeForce RTX3060 AMD Ryzen 5 1600 Six-Core B450 AORUS M Mother board. NVIDIA-SMI 620.02 Driver:620.02, CUDA Version: 13.2