Comment by lbrito

1 year ago

>Why would you assume anything?

Because they already used data without permission on a much larger scale, so it's a perfectly logical assumption that they would continue doing so with their users?

I don't think that logically makes sense.

Training on everything you can publicly scrape from the internet is a very different thing from training on data that your users submit directly to your service.

  • >Training on everything you can publicly scrape from the internet is a very different thing from training on data that your users submit directly to your service.

    Yes. It's way easier and cheaper when the data comes to you instead of having to scrape everything elsewhere.

  • OpenAI, Meta and X all train from user submitted data, in Meta and X’s case data that had been submitted long before the advent of LLMs.

    It’s not a leap to assume Anthropic does the same.

    • By X do you mean tweets? Can you not see how different that is from training on your private conversations with an LLM?

      What if you ask it for medical advice, or legal things? What if you turn on Gmail integration? Should I now be able to generate your conversations with the right prompt?

      1 reply →