Comment by bilekas

2 days ago

I haven't tried this either but I'm guessing if you could pool the GPU memory over whatever the kids are using these days, I think it was SLI back in my day. The GPU memory should still be faster than the RAM?

Unfortunately I checked, SLI doesnt work for this situation. Because the program loading the LLM uses CUDA library, which doesnt account for/takes advantage of SLI for this purpose at least.