← Back to context

Comment by MisterMunchkin

4 hours ago

My company took away my Claude because it’s too expensive. I feel like there is a reckoning coming. The accountants are finally realising the cost of token maxing.

That's pretty stupid. Most people who are incurring significant costs are just tokenmaxxing rather than being efficient with usage. You can get 99% of jobs and work done with Haiku/Luna in a collaberating working enviroment.

I feel like people who are later to the AI game just like to "oneshot" and sink a bunch of usage into generating garbage

  • There really is a skill to using it effectively. I've tried coaching some of the devs on my team. Some get it, some don't.

    Our company has been tracking token usage and models used vs output (tickets, story points, PRs, deploys, etc...). A dev got chewed out, even after I warned him, because he spent over $2k in a single month almost exclusively on Opus while his actual productivity in terms of what he delivered was abysmal.

    • I get it but it goes against the grain for me. Isn't it ironic that we have to waste our precious and expensive human brain cycles to think about how to use AI cheaply so that it is not more expensive than us?

      In other words I want to spend 100% of my mental capacity in the problem domain for the things AI cannot do for me, like steering, grounding, verification and not for things AI could do.

      3 replies →

    • So you are going to blame this one that dev?

      Define productivity, and while at it, quality, maintainability , modularity and so forth.

    • rookie numbers. in one of the top companies, i know someone who tokenmaxed so hard they ended up spending $50000

    • It's funny, the thing that makes effective prompt also makes effective documentation/communication.

      It's bizzare to see people that made clown issues (not enough detail etc.) suddenly start writing detailed prompts just because it is AI that will do the task and not the human on the other side.

      1 reply →

    • there's a manifold to what "effective" means. The problem is once you get into the vibe flow, it's really difficult to eject yourself into the other realms of vscode or IDE or whatever it is you normal do because the vibing provides no anchor to what you're doing.

      Even if these models are smart enough to reorient themselves, they get entirely stuck in a desert and now you're asking someone to just pull up stakes and digg them out even thought they only watched them get there and the UI provides so much speed that no human can comprehend how they got there in the first place.

      It's like asking a pilot to take over in an emergency situation when they're not tasked with any of the every day requirements of the job. The orgs are relying on borrowed time of experienced professionals, and that's going to erode away and what replaces it is mostly people who understand how to navigate context but not use any of the _classic_ tools.

      It's a real conundrum and won't be easily surfaced but for a decade.

      1 reply →

  • That's what happens when token usage becomes a performance metric. As has been done at my company.

  • I feel like it's only within the past few months that opus got to the point where guiding the model is faster than doing things myself. I tried out sonnet recently and it was not a net positive to my work. I feel like anything that I'd trust haiku to handle isn't worth doing in the first place.

    For context, I'm doing a range of tasks, everything from one-shotting adhoc scripts to having 4 hour 10M+ token conversations debugging things.

    • That is always amazing to me.

      There is no way I can beat even local models at generating complex Python scripts fast.

      hn is filled with uber geniuses.

  • in my company there are a few who keep sharing screenshots of reaching limits on 3 separate subscriptions, 2 of them their personal on top of the company subscription

    • Wonder when subscription-hopping attacks will become more often (jumping from a personal model to injecting instructions into the business account and exfiling data)

  • What on earth do you even do with these models?

    Or does a "collaborating work environment" mean that everything is basically spoonfed to them? Or do you only ever use ghost suggestions?

    I genuinely cannot even fathom. Just how do you even get into a state where tasks are so clear and cookie cutter? These things are abhorrent. Not only are they not useful, it's an outright form of psychological torture to try and use them. They almost fight you.

    Luna doesn't even respond to steers properly! You try steering it and it immediately gets distracted and then just stops.

    I can imagine coercing Sonnet into doing some of my tasks okay, but Haiku? Especially 4.5? Really?

    • I think you might be overestimating the sort of projects most of us have worked on throughout our careers -- we haven't been doing much groundbreaking work. LLMs can easily and successfully write most code.

      5 replies →

  • Sorry, but that is nonsense. Compared to opus haiku doesn't cut it most of the time.

    • I think they mean the new Haiku, which is mildly above Luna now . If you have a plan written by a smarter model (so the hard parts are solved) they can be great at implementation.

  • > You can get 99% of jobs and work done with Haiku/Luna in a collaberating working enviroment.

    Optimally? Opus will pay for itself if you save just 10% of your time

That's funny. In a meeting with my manager recently, they specifically called me out for not spending enough money on tokens. It's not like I didn't use it, just apparently not enough, and apparently on too low a setting.

Since then I've had fable cranked up to 11 for even the most trivial of tasks.

I know plenty of people working at very well known and very large companies who's CEOs were boasting about "not hiring anyone anymore", "all the code will be generated in 3 months", etc. they all went from "unlimited budget per dev" + public dashboard with ranking to flex how much credits everyone was burning to hard caps at $500-$1000/month/employee real fast. Some are even not allowing their devs to use the more expensive models

You realise the subject here is Meta, which is all in on this stuff? Of course they are going to use Muse Spark over Claude.

>Great Depression style collapse and all the current AI companies go bankrupt.

Oh this is just a 33 day old doomer account.

Do you know the details of the Claude Code plan you and your company are using (if not part of some enterprise deal)? Does your individual capacity out run something like Claude Max 20x ($200/mo)?

$200 a month is too expensive yet they employ human developers?

  • It's unclear how much OP's company was spending. The article gives a figure of $100k/month per employee.

    But even $200/month is worth shaving if it doesn't generate value.

  • Corporations pay API rates.

    • So what, if they're cutting it, it has propagated to the last bean counter that the roi isn't there.

      This takes some doing and now is the time where it's dawning on the finance departments.

      1 reply →

    • I use Opus 5.5 heavily but only spend around $800/week at API rates. I mean, I say "only"... That's a lot in absolute terms, but trivial compared to my salary and EASY worth it.

    • And someday I hope to understand why they do that. CEO: "Let's see, I can pay $200/month for Bob's tokens, or I can pay $2000/month or so, and then hope he doesn't screw up and rack up a seven-figure bill. The service is the same either way. Hmm."

      3 replies →

  • Besides that the article states quite high numbers, budget is budget in these companies.

    You had budget for your normal salaries, for externals and now suddenly you have a few millions additional.

    What do you do? You compensate.

    Business people doing business things.

My company (me, I'm self employed) did so as well.

Since this summer coding on Opencode Go + Codex for a total 28$/month gives me more intelligence and token than 400$ did in may.

Also, SOTA models are increasingly useless for anything even barely tangential to security work.

taking away sounds insane, we had basically unlimited tokens (inference bought from aws) and they moved us back to the $100 sub to save money

The reckoning started years ago when we did the equivalent to token maxxing hiring coders for everything to crank LOC

Software is inherently a physics problem not all the job titles and specializations made up the last 20 years as dev job salaries kept attracting people

That was all illusory social construct to prop up jobs

Still a whole lot of that in tech but it's all at the top of the org now. Leadership sensory experience and thus innate habit to forecast future been programmed by years of yes men they refuse to accept the jig is up for them too

Sensory memory of being a useless figurehead fosters a lot of existential dread in priests, politicians, and the like. Completely aware their day to day effort is insufficient to sustain them they know how co-dependent they are. They'll dig in harder.

See Chris Matthews flame out shrieking about socialist execution squads. Dude seriously thought everyone wants to hang him from a lamp post. The reality is people just want a sense of control back and not have their perception dragged along by Chris Matthews.

We're just getting put on a budget.

Our velocity is twice as high as it was before Claude, so I doubt that we'll ever go back, but I could see efficiency being a priority.

  • Does all this velocity translate to increased income to the business though. At some point if we are releasing 20x more features, more games, more music, who is actually buying it all?

    • At this point we're solidifying a lot of stuff for the product: disaster recovery, getting workflows to scale so we can actually sell it to more people, paying off the Fort Knox of tech debt, so, at least in my case, yeah. YMMV.

  • > Our velocity

    Is this the new buzzword for the quarter? Last quarter was "granularity", I didn't get the memo yet

    • Is that not a widely-known word in project management circles for scrum and other methodologies? I've heard it for years. Basically how much work you get done through the lens of how long it took to get done via a pointing system.

      1 reply →

My thoughts were the reckoning would come when Infra teams started offloading AWS usage to LLMs and ended up token maxing and deploy maxing.

Good luck to the accountant that tries to tell leadership to slash AI usage.

I'm sure investors will love it.

Now we're starting to see real impact from AI, people are learning how to use it, and OpenAI cut prices by no less than 60% like a week ago.

You think now is the time they're going to cut the spend?

which is quite sad because opus 5.5 is really good. i say this as an anthropic hater. i wish I could move away to other models like 6.1 sol or deepseek or whatever, but they just all lack something. i _trust_ opus 5.5

i hope other labs catch up, especially chinese labs.

I mean, we’re not far from a situation where instead of how many story points you completed per sprint the metric to optimize is going to be what was your efficiency? How many story points did you complete while minimizing your token usage. In fact, that’s a pretty good idea. I’m going try and implement it at work with some sort of complexity normalization function