← Back to context

Comment by 2bitencryption

16 hours ago

this is just constrained generation? I thought the latest crop of decision models (inspired by Jev) do something fundamentally different in the architecture; they're not simply off-the-shelf models with a token mask

Indeed. Token masking just limits the costs associated with using an LLM. It therefore also limits the accuracy by limiting the amount of compute available.

But LLMs can mimick decision models, and I wouldn't be surprized if some labs are doing it this way at the moment.

  • If a lab doesn't know how to put a new output head on a transformer, they shouldn't be considered a lab.