Comment by fragmede
2 months ago
Yes but it's getting bot owners to use it is the problem. There's already the common crawl repository to start with but it isn't being used.
2 months ago
Yes but it's getting bot owners to use it is the problem. There's already the common crawl repository to start with but it isn't being used.
Common Crawl's dataset was downloaded in full 100 times in 2025.
We agree that it would be great if it was even more widely used.
as well as the bot owners could would never believe that the torrent has been kept up to date. the only way to do that would compare to the actual site, so why not just scrape the actual site and be done with it?
Common Crawl's archive has metadata that says when each record (html file) was crawled.
But who stores the metadata for the last date the site updated so you know if it needs to be refetched or not.
2 replies →