Comment by kombinatori
16 hours ago
Indeed. Token masking just limits the costs associated with using an LLM. It therefore also limits the accuracy by limiting the amount of compute available.
But LLMs can mimick decision models, and I wouldn't be surprized if some labs are doing it this way at the moment.
If a lab doesn't know how to put a new output head on a transformer, they shouldn't be considered a lab.