Comment by onlyrealcuzzo
2 years ago
This is an interesting development.
How many other sites might have leverage to charge to be indexed?
I don't want to live in a world where you have to use X search engine to get answers from Y site - but this seems like the beginning of that world.
From an efficiency perspective - it's obviously better for websites to just lease their data to search engines then both sides paying tons of bandwidth and compute to get that data onto search engines.
Realistically, there are only 2 search engines now.
This seems very bad for Kagi - but possibly could lead the old, cool, hobbiest & un-monetized web being reinvented?
Kagi uses at least Google and Mojeek
edit:
> Realistically, there are only 2 search engines now.
https://seirdy.one/posts/2021/03/10/search-engines-with-own-...
> Realistically, there are only 2 search engines now.
From the article:
This seems to assert that ~0 other search providers do any crawling at all. Ever. Are we sure that's the case?
It's a very long article so understandable that you did not read on and learn about other search engines crawling beyond GBY. Still there are indeed very few that are crawling at web scale, and internationally. We are at 8 billion pages and totally independent [0], hence expressing our concerns to 404 media after being blanked by Reddit.
[0] https://www.mojeek.com/about/why-mojeek
5 replies →
I believe Brave Search is also starting their own index. There are some tiny independent indexes too:
https://www.crawlson.com/ https://search.marginalia.nu/ https://wiby.me/ https://searchmysite.net/
Bing provides far fewer verbatim results for pretty much all search queries that I've tested.
And Yandex isn't much better for non cyrillic search, Baidu is only for the Chinese web effectively.
And all other search engines either don't even attempt to do full web crawls anymore and/or buy from one of the four above.
So realistically there's just one search engine for the full web that actually does the work.
10 replies →
I believe Kagi has its own crawler as well and it merges all the results and does whatever Kagi does behind the scenes to show the mix
Aside: Does anyone know how the GBY term became a thing and why it includes Yandex but not Baidu?
Doesn't it list three major ones, Google, Bing, and Yandex, plus Mojeek and a few other small ones? That's a bit more than two.
That seems like the business model for streaming. You subscribe to X provider to watch Y series. So, as for streaming, I suppose a pirate bay search engine will come up
Pirate Bay is probably not the most optimal analogy, more like Anna's Archive imho [1], individually offered by web property scrape runs compressed into a package, maybe served by torrents like this Academic Torrents site example [2].
Scraper engine->validation/processing/cleanup->object storage->index + torrent serving is rough pipeline sketch.
[1] https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu... ("HN Search: annas archive")
[2] https://academictorrents.com/details/9c263fc85366c1ef8f5bb9d... ("AcademicTorrents: Reddit comments/submissions 2005-06 to 2023-12 [2.52TB]")
> but this seems like the beginning of that world.
It's not the beginning, it's mere continuation.
Walled gardens have existed since the AOL days. They deteriorate over time but it doesn't prevent companies from trying (each time, in bigger attempts).
> but possibly could lead the old, cool, hobbiest & un-monetized web being reinvented?
It still exists. It just isn't that popular.
idk man i bet you five bucks and a handshake it's just going to play out like the existing startup grift.
There's an established player with institutional protections, then a scrappy upstart takes a bunch of VC money, converts it into runway, gives away the product for free, gradually replaces and becomes the standard, then puts out an s-1 document saying "we don't make money and we never have, want to invest?" and then they start to enjoy all the institutional protections. Or they don't. Either way you pay yourself handsomely from the runway money so who cares.
The upstart gets indexed and has an API, the established player doesn't.
The upstart is more easily found and modular but the institutional player can refuse to be indexed to own their data and they can block their API to prevent ai slop from getting in and dominating their content.