Comment by scirob
2 days ago
We maintain German Langauge index as no one publishes or reruns these sepeartly.
Qwen 3.8 27B is a small improvement with some regressions in our benchmarks not a huge jump like benchmarks listed.
https://dach.peerbench.ai/compare?models=qwen%2Fqwen3.8-27b,...
German language has never been a big focus for asian models but they still outperform Gemma models https://dach.peerbench.ai/compare?models=openai%2FQwen%2FQwe...
So in production we have been using Gemini Flash Lite as primary and fall back to Qwen when gemini servers are overloaded or just giving us 429
Did you already test TranslateGemma? I use this model for my Android Studio Translation Plugin (https://plugins.jetbrains.com/plugin/30265-localizepipe) and so far it produces great results for its size.
If there are other models (of similar size) out there, that are better at this, please let me know.
From your benchmark, Qwen3.8 is nearer than Opus 4.8 than Qwen3.6. 0.1pp but still.
Also, a lot of people don't really care about german language capacity, maybe people programming in DDP idk.
PS: You benchmark seems saturated. Most values sit @>75% in a benchmark generally indicate that it's no longer as useful as a <70% one. I mean, Qwen3.8 is 77.5% and Fable5 80%, the poll of values is from 65% to 90%.