Comment by gregwebs
2 hours ago
I have been using DeepSeek 4.1 flash intensively for over a month. If I run it all day long it costs $1-2. Its fast. Previously I was always quickly running up to my Claude/Codex 5 hour window (on the $20/month plan). The cost savings of DeepSeek is real as shown in this article and I am using subsidized plans.
DeepSeek is horrible at grilling sessions (the /grill* skills to make technical decisions). It doesn't know how to explain things. Maybe the skill could be adjusted. It also doesn't come up with as good solutions as Opus/Sol.
What I use it for is
* the orchestator of my coding workflows
* the tester/verifier of code changes
* the sub agent that explores code or does web searches
* putting together code base research reports
Previously I planned with Opus/Sol/Astra and then I used DeepSeek for coding, and then reviewed with Opus/Sol/Astra. With the cost improvements to Opus/Sol I am trying to use them for coding instead now so there will be less back and forth review needed.
They are all working together in Pi using the extension @tintinweb/pi-subagents where my workflow skill is calling different subagents that use different models.
Luna is cost competitive, but doesn't score as well on intelligence. I do need the intelligence for most of what I use it for, so I am not motivated to use Luna. Haiku also doesn't seem like a competitive price/performance mix.
Luna is 1 point being on AA's index at 1/4 the cost, yes it "doesn't score as well" but paying 4x for 1 point is crazy if you're going off benchmarks.
AA has Haiku 5.5 as cheaper than 4.1 Flash (both on Max, which isn't ideal but what can ya do) and a 4 point intelligence gap.
Why do people like to think open models are more competitive than they are?
It is super bad on a bit more complex workflows and starts repeating same errors with the same tool until the cycle breaker hits.
6 is worse than 5.6 here.
But it is amazing on generating a report on content generated by better agentic models such as DeepSeek or GLM, which both do a mediocre/bad job on reports.
DeepSeek's own paper advises against using Max, showing that it normally doesn't perform that much better. I am not using it on Max, so that's not a useful benchmark for me. I have seen other benchmarks where Flash does significantly (30%) better than Luna.