Comment by 2bitencryption
17 hours ago
this is just constrained generation? I thought the latest crop of decision models (inspired by Jev) do something fundamentally different in the architecture; they're not simply off-the-shelf models with a token mask
Indeed. Token masking just limits the costs associated with using an LLM. It therefore also limits the accuracy by limiting the amount of compute available.
But LLMs can mimick decision models, and I wouldn't be surprized if some labs are doing it this way at the moment.
If a lab doesn't know how to put a new output head on a transformer, they shouldn't be considered a lab.