Comment by XCSme

3 hours ago

My concern is that reasoning could involve some sequential steps that instant models don't.

Not sure if modern models "think" only by outputting <thinking> blocks, or there is a more complex mechanism at play.

It's not really "instant", i.e. the text is still generated token-by-token, it's just super fast. Reasoning would work with this model without any changes to the chip but it's disabled for speed.