Comment by swalsh
2 days ago
WOW, my first try running on my 2 3090's, it was a bit slow... but it FEELS like opus 4.5, i gave it an image and a broad overview of what I wanted it to build, and it built the whole thing from beginning to end.
2 days ago
WOW, my first try running on my 2 3090's, it was a bit slow... but it FEELS like opus 4.5, i gave it an image and a broad overview of what I wanted it to build, and it built the whole thing from beginning to end.
For your setup, do you have both 3090's in parallel for the inference of the model?
Yes they run in parallel via LMStudio (250k context)
Why slow? I see ~50tps on a single 3090
Yeah make sure you're using MTP and potentially tensor parallelism.
Whats your setup? I have a single 3090 and am struggling to get it purring
5900x, 3090 24gb (slightly undervolted), 128gb ddr4, running via Ollama.
I am benchmarking it now locally, will put the results and speed/tps on aibenchy.com
2 replies →
yeah, i consider that slow.
Oh, ok, that's like the average tps for most AI providers