Comment by simonw

8 hours ago

My number one question about search APIs is always if they allow you to store and resyndicate results you get from them.

If I'm running an agent system but I'm not allowed to store the responses - or provide a "share transcript" button - that's a pretty significant limitation.

The answer to that question is inevitably buried deep in the terms. Here's the relevant section I found for Ceramic, in their list of things you can't do:

> (n) collect, aggregate, store, or compile Output, including search results, relevance scores, or rankings, for the purpose of creating or contributing to any database, dataset, index, or corpus, whether or not such database, dataset, index, or corpus is used for a purpose that competes with Ceramic; (o) resell, syndicate, or otherwise make Output available to any third party on a standalone basis or as a separately accessible component of another product or service; provided that you may display Output to your authorized end users within your own application so long as such Output is integrated into your application's functionality, is incident to the end user’s real-time query, and is not independently accessible, extractable, or downloadable by end users or third parties; or (p) retain, cache, or store Output beyond what is reasonably necessary to display such Output to your authorized end users in the ordinary and real-time course of use, unless expressly permitted in an applicable Order Form.

https://www.ceramic.ai/terms-of-service

Am I alone in caring about this?

It seemed like this part gives you the exception you wanted:

> provided that you may display Output to your authorized end users within your own application so long as such Output is integrated into your application's functionality, is incident to the end user’s real-time query ...

but it continues:

> ... and is not independently accessible, extractable, or downloadable by end users or third parties

How can you prevent end users from extracting it if its visible? Why even have the exception if you just throw it out with an impossible to meet restriction like this?

  • So they crawled the web, stole the information to populate their own database and then pretend it was "illegal" for others to steal it back?

    The weird attitude in the Internet Tech company scene is akin to Gold Rush scenarios.

    Who are the native people?

    • > So they crawled the web, stole the information to populate their own database and then pretend it was "illegal" for others to steal it back?

      These days, this seems to be "modus operandi". The bet is who can get closer to the administration to suddenly enforce the un-enforce-able. For your own safety. You wouldn't steal a car now, would you?

      4 replies →

    • Don't think it's unreasonable. They crawled it and made it available in an easy to digest form. You can't take their copy and start distributing it infinitely while paying them for one-time read. Write your crawler, and crawl the web on your own if you want... and give it away!

      4 replies →

  • Not to mention: "retain, cache, or store Output beyond what is reasonably necessary to display such Output to your authorized end users in the ordinary and real-time course of use" which would seem to preclude storing it in a long lived session.

My general stance on things like this is to think about the intent -- why does the company have that in their TOS. Use that as a proxy for assessing the likelihood of the company enforcing the terms against you.

  • This is only valid up to the level of risk you can tolerate for them pulling the rug out from under you.

    • Which is why, as much as I love Cloudflare, I wouldn't route things like this through their billing. The risk that a company's entire infrastructure goes down, perhaps even by a fraud/risk flag by an incorrectly-configured AI (including, say, if the search API providers are back-sharing their own potentially-broken abuse flagging metadata with Cloudflare), is far too great.

It’s a shit tier web scraping startup, just violate their terms, who cares.

  • This is the right way to think about it. If there's any fear of getting caught, just use a reputable VPN or one of the hundreds of residential proxy providers.

    • The whole point of this product from Cloudflare is to let LLMs "top up" their corpus of knowledge with up-to-the-minute search results after they have been trained on the contents of the entire Internet. The idea that courts would enforce intellectual property rights on little upstarts trying to make LLM wrappers without enforcing any TOS affecting the massive training scrapers is ridiculous. Probably true, but logically unjustifiable.