I think they mean the new Haiku, which is mildly above Luna now . If you have a plan written by a smarter model (so the hard parts are solved) they can be great at implementation.
It's impressive for its total size if you need local inference.
Though it's significantly slower in Token/s and also thinks a lot more without matching the same intelligence (xhigh qwen27b scores lower than haiku's medium setting, and haiku-med is $0.05 per task compared to Qwen27B's $1.01 on AA's comparison)
It seems like more a backup if you need to work offline, imo, unless time doesn't matter and/or your electricity is free. Or you want independence from the labs (fair enough).
I think they mean the new Haiku, which is mildly above Luna now . If you have a plan written by a smarter model (so the hard parts are solved) they can be great at implementation.
I use Qwen3.8-27B as a daily driver, and for things I know will be quite hard I tend to get ChatGPT to do the planning. Works very well.
It's impressive for its total size if you need local inference.
Though it's significantly slower in Token/s and also thinks a lot more without matching the same intelligence (xhigh qwen27b scores lower than haiku's medium setting, and haiku-med is $0.05 per task compared to Qwen27B's $1.01 on AA's comparison)
It seems like more a backup if you need to work offline, imo, unless time doesn't matter and/or your electricity is free. Or you want independence from the labs (fair enough).