Comment by jychang
10 hours ago
If only LLM benchmarks could benchmark it in the first day!
Still no Artificial Analysis benchmark yet. Or benchmark for Laguna S 2.1 or Meituan models or lots of other models.
10 hours ago
If only LLM benchmarks could benchmark it in the first day!
Still no Artificial Analysis benchmark yet. Or benchmark for Laguna S 2.1 or Meituan models or lots of other models.
According to a few tasks from my little personal coding benchmark it's very good at coding and kinda bad at web design. (Also excellent at "draw me a picture" one-shot prompts, for whatever that's worth)
On a sneaky one that involved parsing MIME headers and dealing with character encodings it did better than Kimi K3 at Max and for 38% lower cost.
Interestingly it seems noticeably better than the qwen3.8-max-preview model they offered just a few weeks ago.