← Back to context

Comment by dpark

7 hours ago

> they're using the info to regurgitate in some fashion and serve back.

Sure, in the same sense that they regurgitate any other text they consume. LLMs by definition do not have the full training dataset available, though. It’s far larger than the resulting model. So they can’t reliably reproduce full text without an external source (or if it’s in the training data repeatedly). ChatGPT actually refused to give me a bible quote the other day, presumably because I ran into some general “book regurgitation” safety net.

> The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order?

Honestly, yeah. The idea that people or corporations should hold onto books forever because of a cultural “ick” about throwing out books is a bit ridiculous. Most books end up in landfills.

They aren’t feeding Da Vinci manuscripts into this pipeline. They are feeding still-in-copyright books.

> It’s far larger than the resulting model.

Is it? How many different books are we talking about, and how much information is that, after conversion to text and lossless compression? Images, maybe, but text?

  • These models are trained on way more than just books. GPT-3 was trained on about half a terabyte of filtered plaintext and the training corpuses have grown significantly by then by all accounts.

> LLMs by definition do not have the full training dataset available, though.

That makes it even worse, then. This proves the original point.

> The idea that people or corporations should hold onto books forever because of a cultural “ick” about throwing out books is a bit ridiculous.

If we're building black and white straw man arguments, then sure, let's not archive anything.

  • > That makes it even worse, then. This proves the original point.

    I don’t know what the “original point” is here, but these AI companies are not providing “book excerpt services” and do not claim to. ChatGPT at least will refuse to provide detailed book excerpts (I hit a week or two ago myself).

    > If we're building black and white straw man arguments, then sure, let's not archive anything.

    It seems like you are the one creating the straw man. Do you have evidence that these companies are shredding actually rare books? The only cited concrete examples (in this thread anyway) are all rather boring. I seriously doubt they are shredding 200 year old books because why would they?