Comment by XMPPwocky
19 hours ago
hm- does the model that wrote this know that labs already pay for training data- that stuff scraped from the Internet is not particularly where today's capability gains come from?
19 hours ago
hm- does the model that wrote this know that labs already pay for training data- that stuff scraped from the Internet is not particularly where today's capability gains come from?
They’ve settled some lawsuits and have a few licensing deals, IMHO they are not free from the accusations of pirating.
And look, I’ve pirated material in a past life, I was all about information wants to be free, but I’ve learned something about consent since then and try not to ignore the contract that creators offer when they publish something: you buy my book, and do whatever you want with it on the second hand market. Buy my book second hand that’s fine. But don’t go downloading every book that’s ever been scanned to create a service that destroys writers’ ability to make a living and act like you’re doing us all a favor.
The point is, the big improvements we’re seeing nowadays are coming from RL, not from scraping the internet.
the point isn’t scraping it’s taking your data and enterprises data
https://trustedrouter.com/blog/they-are-still-training-on-yo...
they pay for some data but they take all of the stuff you’re throwing in too; that’s why i propose forcing it since they’re already used to paying for data just increase the cost even further
https://trustedrouter.com/blog/they-are-still-training-on-yo...