← Back to context

Comment by Aurornis

16 hours ago

> it seems really far off from the kind of experience even a basic $20/month subscription gets me.

The $20/month subs are much stronger than the local models you can run, even with how far local models have advanced lately.

The appeal of local models is that the data never leaves your network so you can feel safer putting sensitive content into it. It also feels “free” to use when you’ve already paid for the hardware.

But it doesn’t perform better and if you do the math you’re probably not saving money either. It’s helpful for things that you can’t or don’t want to outsource to a 3rd party.

There are a few use cases that are (somewhat) surprisingly unsuited for cloud providers:

- translations: cloud providers can bowdlerize (censor) bad words/content; also, if you want to do a translation for personal use of copyrighted materials, cloud providers may block it

- image generation: generating drawings with a style that even just resembles a copyrighted one (ie. Disney) may be blocked by cloud providers - for example, generating old cartoons style with GPT may not be possible.

  • What about a light but bulky AI job, like batch processing 50GB of files? I'm currently doing it on my used Macbook M1 Max 64gb, and it's chugging through it for the cost of electricity (free with my solar).

    • 64gb unified memory is incredibly powerful for a local machine. Most of us have 16GB which is almost useless compared to the API models.

  • Cloud does not mean censored. You can rent gpu time and run whatever model you want, with your data kept as private as any other cloud instance you personally run. Cloud is location, with (for some work) wayyyy cheaper access.

    • Cloud is still not "your computer" so it's probably wise to take that into account and act in accordance with your own threat model.

glm5.3 is matching fable in lots of places and beating it after more than 1 pass in many. now of course when i say this, folks would claim that it's not local, but it can be. if you happen to own a mac studio 512gb, you could run it.

I don't think it is really about "sensitive", but basically about any content you put in. Why would you give corporations your reasoning (data on how you interact with AI, how you "talk" etc.).

All of this is private, but not necessarily sensitive. You never know what is happening with this data. They might say they don't log it or don't sell it, then few years later you'll find it all online or read a book that has a story eerily similar to what you chatted about with GPT a year ago.

  • Although in this use case, it's likely because GPT guided you to write the same story as somebody else. Talking with an LLM about an idea is a great way to make it more predictable and homogenized. If you're fixing a bike or writing software, this is usually a good thing.

    • Gemini not long ago, when you said something "useful" said thank you, I will use it to help other users with similar problem. When asked "why would you do that, I thought our chat is private?" it would respond "Apologies. My mistake, of course this chat is private and your information will not be used." Funny.

> much stronger than the local models you can run

but depending on what you're doing, you may not need the "bleeding edge" performance

Banks, Biglaw, and the Pentagon all do it in the cloud. What could an individual be working on that is so secretive?

  • > Banks, Biglaw, and the Pentagon all do it in the cloud.

    In _a_ cloud: their own virtual private cloud. They also have enough power to negotiate contracts with strong privacy provisions.

    • > They also have enough power to negotiate contracts with strong privacy provisions.

      What privacy provisions would you want to add to AWS? Most of the reasonable strong privacy provisions you'd want are already there and/or available if you want to sign up for it, even including US Govt Top Secret data if you meet some approval.

      2 replies →

  • I like to buy specific brand of soap. I don't want them to know that, it is my right and so is running local LLM "wasting" money on local inference to keep track of my stack of soap.

  • Those companies also have data sharing/use agreements that they can get from Cloud AI providers due to their size and spend. The secrecy and data protection is largely what they are paying for. Those types of agreements just aren’t available to individual customers. It’s only when you’re spending $$$ that it becomes worth it for the provider.

It also takes some load off the AI data centers.

IDK if that might be a concern for Apple or their AI partners.

  • It worsens the supply crunch, no? A unit you use sparingly vs that memory going into a GPU that serves many more people.

    • Those will use different wafers, so unless that memory is allocated for unified memory vs gpu HBM it won't make a difference.

    • What supply crunch? Tons of RAM available for purchase. It's just expensive. That there is a "supply crunch" is made up to benefit from Trump administration not giving a shit how big corps operate

      There's cloud hosts out there with unused compute. Wasted cycles are all over businesses running unused cloud apps and subscribed to services they don't use.

      Still need a local computer to access the cloud; so a barely used gadget still exists. And this creates duplication of effort; we built RAM for servers AND the edge devices.

      Seems redundant when tech nerds and corporations are really the only people that care.

      And all that data in the web is meaningless yet we create a supply crunch storing it in servers.

      This an out of touch nickel and dime perspective given the big picture to say nothing of the mess of strip mining and manufacturing pipelines that go into every screw, wire, and such