Comment by bityard
10 hours ago
A small model, yes! But not necessarily a good model.
With the additional caveat that I don't know whether that specific card is supported by modern drivers.
You'd be looking at one in the 6B or 7B parameters range at FP8. Or smaller. It's been quite some time since a recognizable company in the AI space released a model that small. You can try larger model that has been quantized down to that size, but they don't always fare well with that.
Modern text-to-speech and speech-to-text models also fit well into modest amounts of VRAM.
You're arguing for a very specific range of weights but many slightly smaller and slightly larger models have been released including QAT and MoE versions.
An old nVidia brand card with 8GB is more than enough to see those models running at usable speeds and accuracy.