← Back to context

Comment by dpark

7 hours ago

I don’t trust or expect AI companies to serve books at all. That’s not what they are scanning them for.

Not in a traditional sense, but obviously on the surface, they're using the info to regurgitate in some fashion and serve back.

The point is that even under the best intentions, hallucinations occur. Then there's the fact that most models have an ideological bias programmed into them.

The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order?

  • > they're using the info to regurgitate in some fashion and serve back.

    Sure, in the same sense that they regurgitate any other text they consume. LLMs by definition do not have the full training dataset available, though. It’s far larger than the resulting model. So they can’t reliably reproduce full text without an external source (or if it’s in the training data repeatedly). ChatGPT actually refused to give me a bible quote the other day, presumably because I ran into some general “book regurgitation” safety net.

    > The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order?

    Honestly, yeah. The idea that people or corporations should hold onto books forever because of a cultural “ick” about throwing out books is a bit ridiculous. Most books end up in landfills.

    They aren’t feeding Da Vinci manuscripts into this pipeline. They are feeding still-in-copyright books.

    • > It’s far larger than the resulting model.

      Is it? How many different books are we talking about, and how much information is that, after conversion to text and lossless compression? Images, maybe, but text?

      4 replies →

    • > LLMs by definition do not have the full training dataset available, though.

      That makes it even worse, then. This proves the original point.

      > The idea that people or corporations should hold onto books forever because of a cultural “ick” about throwing out books is a bit ridiculous.

      If we're building black and white straw man arguments, then sure, let's not archive anything.

      1 reply →