Comment by LoganDark
3 hours ago
For autoregressive models (practically all hosted ones), it's because of the nature of next-token prediction. LLMs lock themselves into a particular sentence structure ahead of time and have to guess at the rest of the sentence. Samplers have no insight into the LLM's "intent" aside from the probability of each next token, and the LLM has no insight into its previous "intent" that resulted in a given probability in the first place. I don't know if this is possible to solve with more training, I think a fundamental architectural shift may be needed, like more research into diffusion language models.
I think that's quite a strong claim, especially since chain-of-though reasoning means a modern LLM has its own private scratchpad to workshop sentence structure in, if it were a significant problem.