Comment by tremon

2 years ago

How do you "discriminate" data gathering at web-scale, though? In my view, everything at web-scale only works because there are no humans in the loop, as repeatedly explained here in basically every thread involving Google or Facebook. Yes, since it's a scientific paper they should have defined their usage of the word, but I see nothing wrong with the basic premise that automation at large-scale implies indiscrimate use of content.

You can use LLMs to vet the relevancy of the content, so you only select the most useful data. I believe most labs are doing this today.