Comment by 3836293648

17 days ago

But this also means tiny context windows. You can't fit gpt-oss:20b + more than a tiny file + instructions into 24GB

2 comments

3836293648

Gpt-oss is natively 4-bit, so you kinda can

3836293648 14 days ago

You can fit the weights + a tiny context window into 24GB, absolutely. But you can't fit anything of any reasonable size. Or Ollama's implementation is broken, but it needs to be restricted beyond usability for it not to freeze up the entire machine when I last tried to use it.