Comment by zelphirkalt
1 day ago
And also a "proprietary data set" hahaha... Probably just means they don't want to show it, and it is data, that either they shouldn't have, or that there is nothing special about their training data and it is just meant to sound like there is some secret ingredient, while there is none.
Not sharing the data is pretty standard because 1) it tends to get the lawyers involved and 2) good data is critical for getting good results.
Imo you can get better results with great data and generic modeling techniques than with incredible modeling techniques and crappy data. Because if you have crappy data, you won’t even know if your model is good because your evals will also be bad.
This is why Anthropic is throwing a fit about the Chinese distillation “attacks”. Clean reasoning traces are gold.
This isn't true.
Companies pay lots of money for proprietary agentic trajectories which are used during RL. These are things like "Task: summarize stock levels for months end accounting" which then traces the task though using SAP to look at different SKU stock levels, exporting them and generating summary Excel spreadsheets.
This is very different to the "scrape the internet" datasets that a table stakes for training a LLM.
Xiaomi released a fairly developer-centric dataset like this here: https://huggingface.co/datasets/XiaomiMiMo/MiMo-V2.6-RL-oss
SpreadsheetRL is another fairly specialized dataset: https://spreadsheet-rl.github.io/
Data has copyright issues, so one can't share it generally without getting permissions from all of the copyright holders. The data is not theirs to share, anyways. The derived (learned) weights are a different matter.
True for the pre-training data scraped from diverse sources. Less so for the later stage data for RL which is by all account more of a differentiator. In most cases the labs themselves produced the data so they are the copyright holders (or they are borrowing it from other labs via distillation). A lot of it is synthetic data, and since you can't copyright AI output, it becomes less about copyright and more about trade secrets.
If I train a model on a song's lyrics, and the user asks the model to recite the song, and it does, so in RLHF I downvote the generation because it refused-to-refuse to regurgitate the song's lyrics... is that RLHF session copyright-encumbered?
this is very normal for frontier lab companies. you need good data either synthetic or labelled (all the chinese open source models have their own armies of data labelers)