Comment by altcognito

19 hours ago

I'm pro LLM generally speaking, but LLMs are the most extreme version of corporatism you can imagine. Taking literal the entirety of human thought and existence, distilling it, and reselling it to you piecemeal with whatever god forsaken manipulation that corporations are likely to apply to it.

Look at the political landscape (in America) today, and why does government fail to serve its people? Because corporations by and large want to privatise everything because it gives them ever increasing power over people's lives. (so they agree to destroy institutions that have served people well)

Obviously Europe (yes, not a monoculture) and China are not immune, but it manifests itself in different ways that I am not really qualified to get into.

> Taking literal the entirety of human thought and existence, distilling it, and reselling it to you piecemeal with whatever god forsaken manipulation that corporations are likely to apply to it.

From earlier this month, "Aaron Swartz was prosecuted for scraping, while Meta does it without consequence":

* https://news.ycombinator.com/item?id=49379550 1700 points by speckx 9 days ago | flag | hide | past | favorite | 403 comments

LLMs wouldn't work without learning on massive amounts of text, so using the entirety of human thought and existence is a very practical step. I wouldn't worry about the reselling part, because AI is a race towards the best price/performance. Open weights models are driving the price down a huge amount, making this tech more practical for everyone.

  • > LLMs wouldn't work without learning on massive amounts of text, so using the entirety of human thought and existence is a very practical step.

    Practical or not, it's abhorrent.

    But it wasn't trained on the entirety of human thought. It's mostly trained with what is available on the internet.

What if you use the open weight models?

  • Open weights are one thing, but honestly, the holy grail would be the original data (and some reasonably cost model by which to train something equivalent).

    We've kind of reached the age of true industrial sized data. We've been acclimated to being able to eventually get to supercomputer levels of computation -- it will be a decade before we get there (if we ever do). It will be critical to retain the data that we have before it is locked off from the world into private vaults.

    I think the other sibling response to this comment is reasonable: corporations violate the law and they pay some minimal penalty. people violate the law and effectively lose their life.

    • But if the original data is the whole internet (and then some), what would that even look like? Would it be useful to anyone? Are individuals going to re-run the training?

      1 reply →