← Back to context

Comment by rvz

18 hours ago

The speed means absolutely nothing when it is finishing almost dead last when compared to the frontier AI companies.

It means something, because it an iterative workflow. If you're willing to burn tokens, it's possible for weaker models to implement tasks by incrementally improving drafts.

Not if your use case needs speed. For one of my products I can't use an LLM that has a p99 of >700ms for TTFT.

  • If it could output 1k tokens per second but needed 4 seconds to produce the first batch of 4k, would that not be viable?