Comment by pessimizer
4 years ago
> I'm disagreeing with the language you are using to characterize models. "copying from memory" implies that there is something being copied, and a memory that you are copying it from. I am pointing out that LLMs do not do this. It's not how they work.
Then we're arguing about the semantics of the word "copy." That is not an interesting argument when you know exactly what I mean and can express it clearly.
edit: If it helps, either substitute your description in whenever I say 'pretty much copy' or change the word "copy" to whatever word you want to use. But even though I can't reproduce the opening paragraph to A Tale of Two Cities verbatim, I can certainly write something that is "copying" it without doing that, and anyone who was familiar with the book and read my paragraph would agree with me.
It is semantics, but that was your whole point no?
> That's because they're not modelling anything
If we agree on "how LLMs work", then how can you claim that they aren't modeling anything? They are modeling language, and while it's unlikely current paradigms will be proving new mathematical truths, it's completely plausible to me that bigger models will be able to handle simple math word problems like those in the article, precisely because LLMs can model the "Alice", "Apple", and "Bob" entities.
I disagree that they are modeling language.* I think that not only bigger models but same-sized or much smaller models will be able to handle arbitrarily complicated word problems if they're eventually supplemented with some explicit model-building process.
-----
* ...and that would be a completely semantic argument to have. I don't care whether it's called modeling, other than the fact that when I'm talking about modeling, I'm not talking about language probability, I'm talking about categories. But discussing what current AI is (a language model, copying?) is a waste of time, because I absolutely agree with your description of how it works, so we're talking about exactly the same thing.