← Back to context

Comment by euazOn

3 days ago

The multimodal abilities are great, but if you deal with text only, what is the benefit of using this over DS V4 Flash/Pro? 13-26x cheaper with comparable intelligence, and available across many different inference providers.

I fail to see the usecase where DS V4 Pro is not enough, but Flash 3.7 is - except multimodal.

Luna is similar, and also 8x cheaper. Source: artificialanalysis

The only benefit I can see is the speed, that looks to be outstanding, probably thanks to their TPUs.

> 13-26x cheaper with comparable intelligence, and available across many different inference providers.

Well, compared to 2 months ago, it's no longer 100x more expensive for similar levels of quality...

If they continue monthly-ish releases by 3.9 - by Halloween - they should be close to the best in terms of what you get for what you pay for.

In 2 months, they've gone from basically the bottom of the pack to at least being somewhat usable and competitive.

OpenAI and Anthropic release in a month, and change things. OpenAI is claiming to be close to an Astra release - but that seems like a Fable type release - where they're just releasing a better more expensive model, not more cost effective models.

  • Surely Anthropic will release a new Haiku at some point soon?

    It's terribly outdated and way overpriced now and they do need something to compete at that "fast, cheap and ok" level.

for non-coding applications, i think speed is a real differentiator. Im building an app that uses LLMs for some functionality that the user would not have any reason to expect is using AI and therefore having then wait seconds or minutes is just not feasible. latency is a huge upside for me

  • Anything interacting with the real world seems like latency would be hugely important. Something more asynchronous friendly (like coding) is for obvious reasons over represented here

did you try to ingest 1M documents per hour with any provider except GCP with Flash? None work at scale. Deepseek, Luna, Mistral all fail. 1 in 3 requests is a fail. I stopped trying.

The only thing that works at scale is gemini flash.

> over DS V4 Flash/Pro? 13-26x cheaper with comparable intelligence, and available across many different inference providers.

That's why DS4 already had a huge price hike announcement.

  • The inference providers did not raise the prices no?

    Deepseek as a company can just increase prices for the crazily cheap cache they have, that's their only lever.

  • I guess the demand is just too high... But even after the price hike, ds is still much cheaper?

I guess the question then becomes "are you sure you'll do text only?"

I could probably do text only for my workflow (feature development/debugging for web microservices) but sometimes it is easier to just toss a screenshot into the Claude prompt, so that gives it an edge.

If your workflow is 100%, certifiably never ever going to involve an image, then yeah, this isn't going to be huge.

DSV4 Flash is in a tier of its own, until at least the price change arrives.