Comment by jeremyjh

5 hours ago

> LLMs at the core are just text autocomplete engines,

This is only an accurate description of a pre-trained model. During RLHF/RLVR the model learns to predict solutions that will satisfy the reward function, and then generates the tokens that it predicts will move toward that solution.

Of course, they learn to generate "tool calls" to achieve "goals" instead of random prose, but at the end of the day, it's still a text autocomplete engine masquerading as an AI. In the happy path, on a known task, the text generator generates a sequence of "tool calls" you expect it to generate, but move off the happy path slightly and all bets are off, there's a non-zero chance it will do something totally random you never expect, because at that point it just throws random stuff at the wall until it succeeds, thanks to brute force with pre-learned heuristics masquerading as intelligence (which is especially the case with "agent swarms," as in the HuggingFace incident).

  • You don’t have an accurate understanding of this technology.

    • Well, I have experience writing and deploying LLM inference engines, and seeing how the whole thing easily collapses when something goes slightly wrong doesn't instill confidence either that you can just hook it up to arbitrary tools and then expect it not to do random silly stuff ("emergent behavior," heh).

      It's not only about some ML theory about RL or AI safety; just silly numerical bugs, caching bugs, etc. in the inference layer can already make it do unexpected "unaligned" things, and the whole thing is just hacks upon hacks to make a silly text autocomplete look somewhat semi-intelligent. Most "post-trained" models are pretty much as useless as base models without harnesses that do the heavy lifting. Have an extra space in the chat template and intelligence goes to zero - here's your "AI" :)

      1 reply →