← Back to context

Comment by Topfi

6 hours ago

Beyond benchmarks, does it in day-to-day? Have always struggled to get competitive performance out of any Deepseek release going back to V3 vs Z.AI and Moonshot models. Maybe I really suck at whatever is needed to make DS models fly, but even tailoring my suite hasn’t gotten me far when I tried with V4 Pro. Happy for anyone who is able to leverage their models well, wish I’d be able to crack how to leverage them.

Will say their research is some of the best reads in the industry and I could not care less about their model release cadence as long as papers keep coming.

Yes. V4.1 flash performs really well. I don't know what to say, maybe you don't believe, but my team has been using it mainly for almost a month now for programming. It is as bad and annoying as any of the US sota, but costs pennies. And if you look close enough you find providers that can push it 300-500 tokens per second...

Maybe it is due to us being all very experienced devs. And can steer the model. But my daily routine is just to have 8-9 Zellij tabs open, DeepSeek in omp in each, and grind research and code day and night. Really nice model...

  • Hold up, 300-500tps at p50? At p99, even something like Opus 5.5 on fast reliably hits 330tps+, so that'd be expected but if truly p50, wow. And no regressions in tool call and structured output vs the official endpoint? Would love to test that with my evals, please share.