Comment by derin-picment

2 hours ago

Skimmed the first sections — the most interesting part to me isn't just the 78.1B total / 3.46B active MoE numbers, but the data story: 24T tokens with >20% German, including 2T+ German tokens curated/generated themselves.

That explains why they're framing it around sovereign deployment for public administration / aerospace rather than chasing general English benchmarks. The Pareto-frontier claim on throughput vs quality (Figure 1, 8xB200 evals) is also refreshingly honest — serving cost matters a lot for regulated on-prem use.

Would love to see more detail on how the synthetic German data was validated for quality, and how MergeMix data mixing affected German vs English trade-offs. Apache 2.0 open weights is a big plus here.