Comment by ndriscoll
14 hours ago
Attributing training data seems pointless for trustworthiness. The way you trust a model is the same way you trust a human; you ask it to:
1. Provide a chain of reasoning from agreed premises. These days LLMs can even do this airtight with proof assistants.
2. Cite data sources for non-agreed premises. I don't care where the model learned a fact. It might not have ever read a document directly from the primary source. I want it to link directly to either widely agreed facts (e.g. standard textbooks, and if necessary school syllabi demonstrating that the text is standard) or primary sources (e.g. datasets).
Training provenance is irrelevant. It's neither necessary nor sufficient to deal with truth.
No comments yet
Contribute on Hacker News ↗