Comment by saejox

7 hours ago

Most of those people will be dissapointed when they experience Q4 variants of those models getting stuck in loops.

I would wait till the ram crisis is over to fetch a future 64gb ram gpu to run Q8 models. Cloud inference until than.