Comment by vishvananda
4 hours ago
The reason people aren’t freaking out is because most people are using heavily subsidized subscriptions.
I tried the cheapest provider on openrouter and burned through $50 in a few days. Quality was ok, seems slightly above Luna quality perhaps? But that $50 is 1/4 of my codex subscription where I could have burned that many tokens or more using Astra within my weekly reset.
This won’t last forever but as long as the frontier labs are subsidizing this heavily the open models won’t matter.
Yea if i use opus 5.5 in api through openrouter and pi agent harness I will easily burn 50-100$ a day (and with fable 5.1 i could burn 200$ easily). Whereas i have now been using a claude code subscription for 2 weeks using 5.5 at all times and have never hit a limit. I often run 6+ agent sessions at once.
I do think its important long term to not be reliant on these companies as you don't have control over the system prompts, the thinking tokens, and once the subsidization stops or the company is public they will be required to start making money and thus raise prices.
But models may get more intelligent and cheaper once that time comes so it may be a non issue.
> Yea if i use opus 5.5 in api through openrouter and pi agent harness I will easily burn 50-100$ a day (and with fable 5.1 i could burn 200$ easily).
Checked yesterday, for that day alone I had used $168 worth on my $20 subscription in Claude Code. I still had plenty of weekly use left. Seems like subscriptions are discounted at a 1:10 rate?
100%
I use a personal Claude account for personal projects and can let rabl run for an hour and barely make a dent into my usage
On the enterprise I have to be a lot more careful or I can burn through 2k in a week
The excuse they give is the guarantees you get with enterprise plans that they won’t look at your data
What's rabl?
[dead]
Yes. You have to find the provider with pricing that suits your usage.
I am having 98% my input in cache, so using Coralbricks makes sense due to them giving cache reads for free — you only pay for writes. I spend maybe 5-10 dollars a day and my agents basically work day and night implementing things for me.
If your tasks are write-heavy, find a provider with cheaper output.
If you build a customer-facing app, pay a bit extra for 400+ tok/s e.g. on Lithos.
What are your agents implementing day and night…sheesh. Built anything useful for anyone yet?
>What are your agents implementing day and night…sheesh. Built anything useful for anyone yet?
you can't think of anything to unleash some agents on within the entire digital world at any given time?
you motivate your own personal work only via gauging its' usefulness to others and your own prospects?
sheesh.
built anything for the sake of building yet?
2 replies →
Sheesh... what's in a name? :p
Bio shows: https://twin.so/
3 replies →
You have all the struggles for the price of Anthropic / cursor subscription. I use the first one I code large chunks some PR are 50k LOC and I have at least 2-3 like this a week . It’s a greenfield project .
I still have quotas left I use it for home things build 3d model of my renovation projects, alerts for shopping list etc . And yeah I use cutting edge of cutting edge of models that saves me time and money , only discount monitor saved me ~$2k on my renovation project
Why the hell would you put 50k lines into a single PR?
I mean, why even pretend you’re going to “review” something that large? Just build everything on main.
I am curious how you managed to spend that much on Deepseek via OpenRouter. I loaded $100 back in July while using v4-flash or whatever the cheap good model was at the time, and have upgraded as the new ones came out from Deepseek. I still have $16 and some of that spend also goes towards the AI usage from my customers (the context they need to load in is quite large too).
And I am using the Claude Code harness with DS as the endpoint. And I use it ~5-8hrs a day to do my coding.
It's wild how different usage patterns are between users.
I have seen the Cursor leaderboard on my company and the vibe coders consume about 5x more tokens than the developers. They and other office workers also have Claude and their limits are often over around Wednesday.
People are using millions of tokens to do very simple HTML reports. I have seen someone asking the LLM to download the entire data into the context and asking it to sort.
Those usage patterns don't correlate to output.
Now that Microsoft allows you to do /cost for individual tasks (or whatever they call them today). So I tasked Sol, Astra and Fable in cowork with exactly the same vibe coding task on the exact same zip file containing a code project I needed an update for. Astra used 20x and Fable used 15x of what Sol did.
Fable changed a lot of things I had explicitly told it not to change. Arguably a lot of them would've been correct if you didn't work in a place where abstractions are directly against the core principles, but what it produced was basically unusable. I'm not sure if Sol or Astra did best, they produced rather similar code outputs. Astra's was better, but Sol didn't do so bad. It forgot to clean up a few places after it's refactor and it made two bugs I had to correct but other than that it was fine. Astra on the flip-side might have produced code that didn't need changes but it also rewrote every piece of documentation so that it became horrible.
As far as the "experiment" goes, it just shows you that the credit consumption is basically pure magic. You'd think that the Microsoft AI admin tools and the Agent365 FOMO DLC license they sell might give you some sort of reporting, but it doesn't. What you can see is how many tokens a user consumes and the total number of tasks they've initiated as well as whatever running agents they have. You can't see what models they use or which tasks are expensive, which makes it very hard to help them. Early on we had an employee who hit their limit in an hour, and it turned out they had basically uploaded a lot of information and run it in a single long task that kept going over it again and again. We told them it might be a good idea to only give it what it needed and to create more tasks, and even though it's been three months, they have yet to consume as many credits as they did that first hour.
But that's how you support and track it. You see a user spend a lot, then you go to their computer and now that you can actually do the /cost thing, you go through their tasks and try and figure out where they're spending money...
It's obviously improving. A month ago /cost wasn't there and they just released a new dashboard for cowork, but it's still black magic that is impossible to govern.
It's because this is how AI has been sold to everyone - just ask, and it will do it.
The better pattern is to let it code the app and then you can use the app to target your data. So you only pay for it once, plus it's deterministic. But yeah, it requires setting up an environment, etc. It becomes "maintenance".
[dead]
Likely the user doesn't know what they're doing or has extermely bad workflows. They're prob not managing their cache, and dont use compaction.. Letting context get to 500k and invalidating their cache every 10 tool calls because they have no providor fallback settings.
I was running deepseek v4.1 pretty much non stop during work hours, with heavy tool/mcp usage and finding it very difficult to spend more than $75 in a month.
Also the cheapest providers on Openroutrr can often have terrible cache hit %, short TTLs resulting in their effective price being much more expensive than people realize. 75% cache pretty much destroys any savings from a super cheap token perspective.
I freak out since months for Z.ai lite subscription, I use glm-5.3-flash every day for a ludicrous 8.5USD/month and it's as good as DS 4.1 flash, if not better.
But aren't you developing bad habits and learning patterns that won't work long term? Or do you think things will get cheap enough that you will be able to keep going with your current patterns post-subsidies?
No one involved in this is thinking about the long term
Long term, we will see what happens and adapt. At worst we all go back coding by hand. Meanwhile what can I do, tell my customers that I'm raising my fee because I have to pay for token? The Claude Pro $20 plan is good enough for me and even in auto mode I never had to wait for the 5 hours reset.
> But aren't you developing bad habits and learning patterns that won't work long term?
2 reasons - there's an advantage now, use it. 2nd the frontier providers, this is the "early cheap days" like when uber was initially cheap to compete vs standard cabs. they want you to become hooked and boy are we hooked.
I expect by that point we'll have local models that can do a decent job, I would guess give it a decade and we'll be running custom accelerators that are smarter than current frontier models.
In the same way that only supercomputers used to have multiple processors and caches but it's now standard.
I use the frontier openai/anthropic models at work but exclusively open weight models (on cloud/hosted inference) for personal stuff and I think about it like this; 1) I don't see any reason GLM and DeepSeek won't eventually be as good as Claude, it's just a matter of time and 2) the open models are well and truly capable enough for most of what I'd want to do. I don't need nor want an LLM chewing away on a horrible enterprise spaghetti codebase, my employers can pay for that privilege.
Compared to what a lot of companies spend on software for chip and electronics design (we're talking about $10k-200k/seat per year), AI coding assistants have a long way to go in cost before companies won't be willing to pay for them. Companies pay a fortune for software when it enables their engineers to be productive.
For my company, I'd honestly pay $4-8k/month for Claude if I had to (it would be painful, and I'd try to get cheaper options to work first). I know some enterprise Claude users are paying that much now since they have to pay for API tokens. I am certain it's at least a 2X productivity booster for our work. Compared to the cost of hiring another developer, it's well worth it.
If they stop subsidising Claude Code for the pro/max users, there will be a lot of people priced out of it, especially the casual developer. But I don't see it going away for commercial use, even with a large price increase.
Nah dude I think it’s worse than that
Old coding is done, as a workflow in teams. It’s the top down executive pressure of being non competitive as a company, and the bottom up pressure of human laziness
Show me people handwriting code à la NASA
And I mean we as coders have been trying to do this workflow for a while, I personally would refuse to code without IntelliJ magic complete
For this workflow, there’s no going back. What’s hard to imagine is AI taking over the other workflows we predict it will; Customer service AI sucks ass for me as a customer, et cetera
How's the caching? I have 99.5% cache hit rate with deepseek when using their own API, it's dirt cheap.
OpenRouter is complete garbage.
Buy directly from DeepSeek's API.
You can literally get overcharged 100x on DeepSeek on OpenRouter (or more).
I'm all in for saving money and _can_ move to using DS directly from them, but maybe I am missing something here:
OpenRouter Pricing:
$0.02/M input tokens $0.60/M output tokens
DeepSeek Pricing (cache miss, off-peak):
$0.15/M Input $0.60/m output
When 98.5% of my requests are cache hits (according to Pi for the last week), the cache miss price isn’t that important to me, and $0.003-0.006 per 1M input tokens is shockingly cheap.
It’s also the major difference between using DeepSeek directly vs other providers also serving it, though I have not looked lately: it’s possible other providers have matched its cache hit pricing better?
2 replies →
I've heard that certain inference providers may have different quality of caching implementations, so even if the listed numbers are as you say, the practical cache hit % you get might be significantly different/incur significantly different costs.
There’s a big difference in speed & quality between using DeepSeek API directly with DSH vs. DeepSeek in Opencode Go with Opencode CLI. Can’t tell if it’s the provider or the harness - but worth to give it a try.
DeepSeek trains on your inputs. That's why people go on OpenRouter and choose ZDR providers.
Let me get this straight you guys really like deep seek because it’s open but you don’t wanna help them improve.
2 replies →
One of the reasons I use OpenRouter is because they offer zero data retention. As far as I can tell, DeepSeek's own API doesn't support ZDR.
DeepInfra does and it's the same price. That's what I use.
Youre doing something so special you need that?
2 replies →
Just pin your config to a single provider, or several providers with the params `order` and `allow_fallbacks: false`. I regularly get ~98-99% cache hit rates with OpenCode. And some providers are much faster than DeepSeek; I was getting 200-300 tokens/second the other day with Together as my provider.
It's regrettable that OpenRouter doesn't even try to pin you to a single provider per session, but once you know about it, it's a problem that's easily solved.
Or.. BYOK Deepseek because OpenRouter's UX is much nicer?
Zero Data Retention and not having company source code leak to "CHINA!" (said in Trumps annoying voice) would be two reasons not to
opencode DeepSeek v4.1 Flash isn’t us/eu hosted as of recently, so not sure how this impacts privacy / model training
Which provider was that?
Also DeepSeek usage is subsidized as well, it’s a power hungry model.
Interesting. Are all of the providers on OpenRouter simply losing money? How does that even work out?
No, it’s outrageously profitable above x% utilization without stealing any prompts. Provider economics still pretty good. Acquiring hardware is the current limiter.
You're the RLHF.
3 replies →
Don't ask and dance as long as the music keeps playing.
Are you sure about that? My impression was most providers on openrouter were purely selling tokens for profit...
Have y'all tried an Ollama Cloud subscription? Their off-hours pricing for V4.1 Flash is extremely competitive.
You should basically never pay API prices, they are always several times higher than subscriptions.
There are several open weight subscription providers. OpenCode Go used to be good but now it's complete shit. Charm Hyper is really great and the best value. Other subscriptions have a more limited model selection or provide less value but are still decent.