Comment by quietFalcon
10 hours ago
Nice, though generation speed is the easy half for MoE offload, what's your prompt processing look like at say 16k context?
10 hours ago
Nice, though generation speed is the easy half for MoE offload, what's your prompt processing look like at say 16k context?
They have some community benchmarks published https://github.com/Niko1221/Strata/tree/main/bench/results
Q2_0 does 33 tok/s decode and ~600t/s prompt processing at 128k context on RTX2060 8GB VRAM.
ISTA IQ3_XXS does ~21 tok/s decode and ~240t/s prompt processing