Comment by q8zd3

1 day ago

The article's premise is that USA based LLM providers are losing the AI (cold war) battle because it will not be as adopted as open-weight models, comparing it to closed vs open sourced software. I do not think this is the case because:

* The comparison is weird because open-weight is not the same as open-source software to begin with;

* People based in the USA are at an advantaged position since they have access to both american and chinese models;

* Isn't Running your own model training infrastructure more expansive?

* One can still leverage both, in different phases or use-cases. I do not see how this is an "one or the other" situation.

The Chinese models are usually not only open weight AND open-source but they also often publish their methodology in detailed scholarly publications that are themselves open-access. DeepSeek most famously

  • Training data, training methodology. All NOT OPEN.

    Until we know what a model is trained on, and how it is trained in high detail, I hesitate to call them "Open Source" in any way. They are free. But, we don't know what their priorities are etc. Witness the censorship we see in all models in one form or another. I'm not absolving any side of this.

    Just saying: Don't be blind.

    • Did you even read my comment? They explicitly DO share their training methodology in depth in Technical Reports on arXiv.

      DeepSeek completely revolutionized LLMs and every western LLM today uses or is inspired by the their innovations including Group Relative Policy Optimization and Multi-head Latent Attention.

      2 replies →