Comment by jazzyjackson
19 hours ago
They’ve settled some lawsuits and have a few licensing deals, IMHO they are not free from the accusations of pirating.
And look, I’ve pirated material in a past life, I was all about information wants to be free, but I’ve learned something about consent since then and try not to ignore the contract that creators offer when they publish something: you buy my book, and do whatever you want with it on the second hand market. Buy my book second hand that’s fine. But don’t go downloading every book that’s ever been scanned to create a service that destroys writers’ ability to make a living and act like you’re doing us all a favor.
The point is, the big improvements we’re seeing nowadays are coming from RL, not from scraping the internet.
Where are you getting your information from? From all I've seen the RL gives an incremental improvement, most of the capability increase comes from new model architectures (eg the jump from opus to fable is greater than the jump from opus 4.5 to 4.8)
the point isn’t scraping it’s taking your data and enterprises data
https://trustedrouter.com/blog/they-are-still-training-on-yo...