Comment by usef-
3 hours ago
It's impressive for its total size if you need local inference.
Though it's significantly slower in Token/s and also thinks a lot more without matching the same intelligence (xhigh qwen27b scores lower than haiku's medium setting, and haiku-med is $0.05 per task compared to Qwen27B's $1.01 on AA's comparison)
It seems like more a backup if you need to work offline, imo, unless time doesn't matter and/or your electricity is free. Or you want independence from the labs (fair enough).
No comments yet
Contribute on Hacker News ↗