← Back to context

Comment by rajeevk

6 hours ago

What Chinese models/providers are you using for this? I'm hitting Claude's weekly limits much sooner than I used to with roughly the same workload, so I'm interested in trying alternatives, especially ones with strong coding/agentic performance.

> hitting Claude's weekly limits much sooner than I used to with roughly the same workload,

Anthropic had a +50% weekly tokens promotion since April (!) which just ran out last weekend after getting multiple extensions.

I've been feeling that too,and I suspect that's the true reason why they released opus 5.5 at a discount

Get yourself an OpenCode Go subscription and give DeepSeek Flash 4.1 a shot.

A common tactic is to used a big brain model like Opus for planning and reviewing, and a cheaper model for execution.

  • While it's a common tactic, and I'd vouch for it if you don't really know what you want to code in fact, but if you know what you want to get out of it, I haven't found anything I'd need opus 5.5 for instead of deepseek-flash (flash v4.1 hosted via platform.deepseek.com)

  • In my experience that tactic works well if the codebase is limited in size, or well maintained and separated. Otherwise I do notice a difference also letting fable do the execution, not just the planning for complex tasks.

    • In my experience it never works well on any real work. In fact, I'd go the opposite, plan with the dumb model and execute with the smart model because at least the model writing the code and solving the emergent problems is capable.

      In my experience (and I've been trying this a bunch): smart planner + dumb executor produces worse code with higher spend than simply using the smart planner to do both.

      It's easy to understand why:

      - If the planner has truly thought the issue through, properly designed the solution, solved all of the emergent problems, then the final "write" of the code is just a few more output tokens.

      - If the planner has NOT truly planned the issue completely, then you're letting a substantially dumber and less capable model make significant decisions, and trusting its problem solving, without having a better model check it.

      If you're highly cost conscious (paying for your own tokens and not making any money) then you have no choice but to trade your time and effort for tricks like this to save money by lowering the quality of your output.

      But if your employer is paying for tokens: just use the smarter model. You save your time preventing re-work and reducing code review, you save your employer money (primarily from the cost of your own labor and reduced rework), and you get a better output every time (Opus 5.5 mogs Deepseek 4.1 flash in every single way except cost).

      8 replies →

  • Been using DeepSeek Flash 4.0 and 4.1 for some random sideprojects via OC GO, its a great deal and for non-corporate work it's really great!

I have tried GLM on a subscription, and also DeepSeek and MiMo using API directly. MiMo in particular is extremely cheap.

For regular software development they have been pretty great.