← Back to context

Comment by pas

2 years ago

it's even worse. it's not theirs (it's the users'), they are merely hosting it and using it (ToS gives them a fancy irrevocable license I guess).

so they can do whatever they want with it and the actual owners/authors have no chance to really influence Reddit at all to make it crawlable. (the GDPR-like data takeout is nice, but ... completely useless in these cases where the value is in the composition and aggregation with other users' content.)

actually owners/authors like me would not want our stuff crawlable because that gives up our ownership.

When I am answering some random dude on reddit with a problem I want that dude to read my solution. I don't want this to be crawled and forever stored (probably deanonymized) or enshrined in a dozen commercial LLMs. There is substack for that stuff.

  • I'll be honest... I don't care about this thing. I view public posts as something that belongs to the public domain and if that is searchable, all the better, other people can reach my post. And if that is what LLM's are trained on, also, great, I hope it will be useful to some people. What I _do_ care about is an even playing fields and access to _all_ llms to be trained in that data.

    In the end, I view potential AGI's as a common consciousness brainchild.

    • OK. I see "AGI" as a buzzword and current ML trend as exploitation of creative works that profits big tech. They will always have bigger computers and get more out of what we create increasing wealth gap, unless we stop it

      7 replies →

  • have you heard about DMs?

    • That's why more and more exchanges are moving to DMs and closed communities. Because we don't plan on stuff we intend for participants in the discussion to be harvested but it more and more is. If reddit breaks that trend I only welcome it.

> the GDPR like data takeout is nice

Is there a way to export my history? How?