Comment by WarmWash

19 hours ago

Just because something is in the training data, doesn't mean it is the root of an LLMs output.

Turn off web search and ask a model what a random redditor said about a random topic in 2015. You will only get hallucinations at best, even though that comment is definitely in the training set.

Sure. But it's possible to say: if the document isn't in the training data, it isn't the cause of the output. If it is in the training data, the question gets more complicated.