Comment by wg0
1 day ago
Suppose I inherited a data center spanning several hundred acres full of GPUs and free electricity.
Where do I get the data?
I mean, this many models. They have to start somewhere.
1 day ago
Suppose I inherited a data center spanning several hundred acres full of GPUs and free electricity.
Where do I get the data?
I mean, this many models. They have to start somewhere.
I guess public datasets on HuggingFace and some shadow libraries content is enough to start.
e.g. fineweb dataset is 50TB https://huggingface.co/datasets/HuggingFaceFW/fineweb
There are a lot of open-research on pre-training, post-training and RL data mixtures and sourcing.
I recommend checking papers from Datalogy, Nvidia Nemotron, Ai2 (Ollmo, Tulu, ...) and the recent model from Aleph Alpha if you want to learn more.
All of those are good references. Other folks in the thread are missing distinctions between pre training (~the internet + curated sources) and post training (~instructions and RL)
I heard you should ask Claude about this. Preferably with thousands of accounts, routed through residential proxies
If you ask a model, they will generally tell you where to get data. Modern frontier models have the large advantage of having tens if not hundreds of millions of users providing use cases to train against to improve their responses.
https://scale.com/data-engine - you just buy it.
any prices anywhere for anything specific?
Millions of dollars
You can also hire teams to create data for you for higher quality.
Get data from Claude. That's what the Chinese (allegedly) do.
Note that this sort of distillation is NOT for pre-training data (which is tens of trillions of tokens). I think the allegations against Chinese companies by Anthropic is more so that they distill SFT data (which is good for post-training, but you still need a strong base model)
forget the data....sell it and go live your life!