Comment by vishvananda

5 hours ago

The reason people aren’t freaking out is because most people are using heavily subsidized subscriptions.

I tried the cheapest provider on openrouter and burned through $50 in a few days. Quality was ok, seems slightly above Luna quality perhaps? But that $50 is 1/4 of my codex subscription where I could have burned that many tokens or more using Astra within my weekly reset.

This won’t last forever but as long as the frontier labs are subsidizing this heavily the open models won’t matter.

Yea if i use opus 5.5 in api through openrouter and pi agent harness I will easily burn 50-100$ a day (and with fable 5.1 i could burn 200$ easily). Whereas i have now been using a claude code subscription for 2 weeks using 5.5 at all times and have never hit a limit. I often run 6+ agent sessions at once.

I do think its important long term to not be reliant on these companies as you don't have control over the system prompts, the thinking tokens, and once the subsidization stops or the company is public they will be required to start making money and thus raise prices.

But models may get more intelligent and cheaper once that time comes so it may be a non issue.

  • I dont think itll be an issue. Opus 5.5 now is way more than enough for me and open weight models will reach that level by the time subsidization stops

    • It's enough for you now, but I feel like part of the mythology of our future is that we'll be continued to be employed because we'll be working on more complex problems, with smarter LLMs at our side.

  • 100%

    I use a personal Claude account for personal projects and can let rabl run for an hour and barely make a dent into my usage

    On the enterprise I have to be a lot more careful or I can burn through 2k in a week

    The excuse they give is the guarantees you get with enterprise plans that they won’t look at your data

I am curious how you managed to spend that much on Deepseek via OpenRouter. I loaded $100 back in July while using v4-flash or whatever the cheap good model was at the time, and have upgraded as the new ones came out from Deepseek. I still have $16 and some of that spend also goes towards the AI usage from my customers (the context they need to load in is quite large too).

And I am using the Claude Code harness with DS as the endpoint. And I use it ~5-8hrs a day to do my coding.

  • It's wild how different usage patterns are between users.

    I have seen the Cursor leaderboard on my company and the vibe coders consume about 5x more tokens than the developers. They and other office workers also have Claude and their limits are often over around Wednesday.

    People are using millions of tokens to do very simple HTML reports. I have seen someone asking the LLM to download the entire data into the context and asking it to sort.

    Those usage patterns don't correlate to output.

    • Now that Microsoft allows you to do /cost for individual tasks (or whatever they call them today). So I tasked Sol, Astra and Fable in cowork with exactly the same vibe coding task on the exact same zip file containing a code project I needed an update for. Astra used 20x and Fable used 15x of what Sol did.

      Fable changed a lot of things I had explicitly told it not to change. Arguably a lot of them would've been correct if you didn't work in a place where abstractions are directly against the core principles, but what it produced was basically unusable. I'm not sure if Sol or Astra did best, they produced rather similar code outputs. Astra's was better, but Sol didn't do so bad. It forgot to clean up a few places after it's refactor and it made two bugs I had to correct but other than that it was fine. Astra on the flip-side might have produced code that didn't need changes but it also rewrote every piece of documentation so that it became horrible.

      As far as the "experiment" goes, it just shows you that the credit consumption is basically pure magic. You'd think that the Microsoft AI admin tools and the Agent365 FOMO DLC license they sell might give you some sort of reporting, but it doesn't. What you can see is how many tokens a user consumes and the total number of tasks they've initiated as well as whatever running agents they have. You can't see what models they use or which tasks are expensive, which makes it very hard to help them. Early on we had an employee who hit their limit in an hour, and it turned out they had basically uploaded a lot of information and run it in a single long task that kept going over it again and again. We told them it might be a good idea to only give it what it needed and to create more tasks, and even though it's been three months, they have yet to consume as many credits as they did that first hour.

      But that's how you support and track it. You see a user spend a lot, then you go to their computer and now that you can actually do the /cost thing, you go through their tasks and try and figure out where they're spending money...

      It's obviously improving. A month ago /cost wasn't there and they just released a new dashboard for cowork, but it's still black magic that is impossible to govern.

    • It's because this is how AI has been sold to everyone - just ask, and it will do it.

      The better pattern is to let it code the app and then you can use the app to target your data. So you only pay for it once, plus it's deterministic. But yeah, it requires setting up an environment, etc. It becomes "maintenance".

  • Likely the user doesn't know what they're doing or has extermely bad workflows. They're prob not managing their cache, and dont use compaction.. Letting context get to 500k and invalidating their cache every 10 tool calls because they have no providor fallback settings.

    I was running deepseek v4.1 pretty much non stop during work hours, with heavy tool/mcp usage and finding it very difficult to spend more than $75 in a month.

    Also the cheapest providers on Openroutrr can often have terrible cache hit %, short TTLs resulting in their effective price being much more expensive than people realize. 75% cache pretty much destroys any savings from a super cheap token perspective.

Yes. You have to find the provider with pricing that suits your usage.

I am having 98% my input in cache, so using Coralbricks makes sense due to them giving cache reads for free — you only pay for writes. I spend maybe 5-10 dollars a day and my agents basically work day and night implementing things for me.

If your tasks are write-heavy, find a provider with cheaper output.

If you build a customer-facing app, pay a bit extra for 400+ tok/s e.g. on Lithos.

  • I tried fireworks.ai, drawn in by their supposed blazing speed. Yawn. Mostly worse than vanilla Deepseek.

    Is Lithos actually fast for common usage?

  • What are your agents implementing day and night…sheesh. Built anything useful for anyone yet?

  • You have all the struggles for the price of Anthropic / cursor subscription. I use the first one I code large chunks some PR are 50k LOC and I have at least 2-3 like this a week . It’s a greenfield project .

    I still have quotas left I use it for home things build 3d model of my renovation projects, alerts for shopping list etc . And yeah I use cutting edge of cutting edge of models that saves me time and money , only discount monitor saved me ~$2k on my renovation project

    • Why the hell would you put 50k lines into a single PR?

      I mean, why even pretend you’re going to “review” something that large? Just build everything on main.

      2 replies →

  • I agree. I code a lot, a lot! And maybe my code is shitty, but yeah, I burn a lot of tokens, and I couldn’t do it without Chinese models. I’m just a random dev in the middle of nowhere. And I don’t feel like I’m missing out on anything with my setup at all.

I got insane amounts of Anthropic and OpenAI credits given to me for free for my startup, and I have not touched them.

I get privacy, freedom, and no rate limits with the GPUs I racked locally, and those are features I would never give up even if the surveillance capitalism labs paid -me- to use their models.

How many consumers are there like me? Probably not many, but once local inference hardware is plug and play, I bet the tides shift pretty quick. Also weights-on-silicon will serve the needs of most consumers locally with more speed than any GPU could deliver for a fraction of the cost.

Most people will be doing inference in their pocket or a wearable in 5 years and the giant datacenters will be like AWS, sold to only big organizations that need to auto-scale capacity of custom models on demand.

The industry surely knows this and the subsidized inference is just marketing to generate so much buzz and demand such that the tiny fraction of the market they will be able to keep in the end is big enough that they do not collapse under all the debt.

OpenAI and Anthropic will be Dell and IBM in 10 years if they survive at all.

This. At this point I don't really care about other models because max subscription are super cheap (relatively speaking) and I don't hit my limits. Even if the frontier models are only 5% better I might as well just use the best thing available if the price is reasonable.

Once the subsidization ends and cost becomes significant I will take a serious look around for the best value models and switch off the expensive providers, but that time hasn't come yet.

  • I have a subscription at work, and still I find myself wishing I could use a fast Chinese model. Something wired up to really fast inference - that rapidity of feedback is a feature in itself.

    4.1 Flash seems to be in that sweet spot of very decent, really fast and really cheap. Even omitting the cost, it’s still compelling for staying in flow.

  • > Once the subsidization ends and cost becomes significant I will take a serious look around for the best value models and switch off the expensive providers, but that time hasn't come yet.

    There's a reason the labs in the US frontier oligopoly are using “safety” to lobby for antitrust exemptions for mutual coordination as well as anticompetitive regulation.

  • I'm actually shocked by how many people seem to have max subscriptions.

    Something about renting that much compute doesn't sit right with me so I stick with the $20 subs.

How's the caching? I have 99.5% cache hit rate with deepseek when using their own API, it's dirt cheap.

But aren't you developing bad habits and learning patterns that won't work long term? Or do you think things will get cheap enough that you will be able to keep going with your current patterns post-subsidies?

  • Long term, we will see what happens and adapt. At worst we all go back coding by hand. Meanwhile what can I do, tell my customers that I'm raising my fee because I have to pay for token? The Claude Pro $20 plan is good enough for me and even in auto mode I never had to wait for the 5 hours reset.

  • > But aren't you developing bad habits and learning patterns that won't work long term?

    2 reasons - there's an advantage now, use it. 2nd the frontier providers, this is the "early cheap days" like when uber was initially cheap to compete vs standard cabs. they want you to become hooked and boy are we hooked.

  • I expect by that point we'll have local models that can do a decent job, I would guess give it a decade and we'll be running custom accelerators that are smarter than current frontier models.

    In the same way that only supercomputers used to have multiple processors and caches but it's now standard.

  • I use the frontier openai/anthropic models at work but exclusively open weight models (on cloud/hosted inference) for personal stuff and I think about it like this; 1) I don't see any reason GLM and DeepSeek won't eventually be as good as Claude, it's just a matter of time and 2) the open models are well and truly capable enough for most of what I'd want to do. I don't need nor want an LLM chewing away on a horrible enterprise spaghetti codebase, my employers can pay for that privilege.

  • Compared to what a lot of companies spend on software for chip and electronics design (we're talking about $10k-200k/seat per year), AI coding assistants have a long way to go in cost before companies won't be willing to pay for them. Companies pay a fortune for software when it enables their engineers to be productive.

    For my company, I'd honestly pay $4-8k/month for Claude if I had to (it would be painful, and I'd try to get cheaper options to work first). I know some enterprise Claude users are paying that much now since they have to pay for API tokens. I am certain it's at least a 2X productivity booster for our work. Compared to the cost of hiring another developer, it's well worth it.

    If they stop subsidising Claude Code for the pro/max users, there will be a lot of people priced out of it, especially the casual developer. But I don't see it going away for commercial use, even with a large price increase.

    • Nah dude I think it’s worse than that

      Old coding is done, as a workflow in teams. It’s the top down executive pressure of being non competitive as a company, and the bottom up pressure of human laziness

      Show me people handwriting code à la NASA

      And I mean we as coders have been trying to do this workflow for a while, I personally would refuse to code without IntelliJ magic complete

      For this workflow, there’s no going back. What’s hard to imagine is AI taking over the other workflows we predict it will; Customer service AI sucks ass for me as a customer, et cetera

I use $20 codex subscription, it's practically useless for anything other than luna. Deepseek v4.1 flash goes a LONG way for $20.

I freak out since months for Z.ai lite subscription, I use glm-5.3-flash every day for a ludicrous 8.5USD/month and it's as good as DS 4.1 flash, if not better.

  • Almost exact same experience here, but I'm using Opencode's $10/month sub. It's perma set to DS 4.1 flash and I have anywhere from 3-5 agents going at a time. Never once hit a cap of any sort. I have absolutely no idea why people would be paying $200/mo when you can get perfectly good AI for $10 from multiple places

OpenRouter is complete garbage.

Buy directly from DeepSeek's API.

You can literally get overcharged 100x on DeepSeek on OpenRouter (or more).

  • I'm all in for saving money and _can_ move to using DS directly from them, but maybe I am missing something here:

    OpenRouter Pricing:

    $0.02/M input tokens $0.60/M output tokens

    DeepSeek Pricing (cache miss, off-peak):

    $0.15/M Input $0.60/m output

    • When 98.5% of my requests are cache hits (according to Pi for the last week), the cache miss price isn’t that important to me, and $0.003-0.006 per 1M input tokens is shockingly cheap.

      It’s also the major difference between using DeepSeek directly vs other providers also serving it, though I have not looked lately: it’s possible other providers have matched its cache hit pricing better?

      2 replies →

    • I've heard that certain inference providers may have different quality of caching implementations, so even if the listed numbers are as you say, the practical cache hit % you get might be significantly different/incur significantly different costs.

    • There’s a big difference in speed & quality between using DeepSeek API directly with DSH vs. DeepSeek in Opencode Go with Opencode CLI. Can’t tell if it’s the provider or the harness - but worth to give it a try.

  • Just pin your config to a single provider, or several providers with the params `order` and `allow_fallbacks: false`. I regularly get ~98-99% cache hit rates with OpenCode. And some providers are much faster than DeepSeek; I was getting 200-300 tokens/second the other day with Together as my provider.

    It's regrettable that OpenRouter doesn't even try to pin you to a single provider per session, but once you know about it, it's a problem that's easily solved.

  • Zero Data Retention and not having company source code leak to "CHINA!" (said in Trumps annoying voice) would be two reasons not to

opencode DeepSeek v4.1 Flash isn’t us/eu hosted as of recently, so not sure how this impacts privacy / model training

Also DeepSeek usage is subsidized as well, it’s a power hungry model.

You should basically never pay API prices, they are always several times higher than subscriptions.

There are several open weight subscription providers. OpenCode Go used to be good but now it's complete shit. Charm Hyper is really great and the best value. Other subscriptions have a more limited model selection or provide less value but are still decent.