Comment by switchbak

1 hour ago

I have a subscription at work, and still I find myself wishing I could use a fast Chinese model. Something wired up to really fast inference - that rapidity of feedback is a feature in itself.

4.1 Flash seems to be in that sweet spot of very decent, really fast and really cheap. Even omitting the cost, it’s still compelling for staying in flow.