← Back to context

Comment by computerdork

21 hours ago

Interesting. Have heard about the need to create synthetic data for LLM's, but didn't realize it was such an important factor. Although, how big a factor synthetic data is the next question, but guess there is no way to truly verify how much difference it makes with these closed models.

This is literally what the labs have been doing over the past year. The whole idea of emergence is a lie, there is some interpolation and superhuman long evaluations the models can do, but almost all the gains are from synthetic data. They hire thousands of professionals and pay them as much as 200$/hour to create many tasks that they want the model to perform and use these to teach the model on how to do it with RL. OAI had 30,000 contractors from Mercor for Sol 5.6. When you see Opus suddenly becoming great at blender or some other 3D graphics, that's because they hired professionals and had them do similar tasks that people are looking for. They keep having better professionals at each iteration and so the quality improves. There is no emergence or "General" intelligence. The model doesn't learn to become better at a task because of scaling laws or emergence or whatever they might wanna say, it is literally RL on tasks that they want the model to perform well on.

  • Also interesting. Do you happen to know if once they've used those thousands of contractors to teach the model something like Blender (or some similar app) using RL, when they training a new model, do they need to use same contractors again to teach that new model the same behaviors?

    The reason I ask is even though it can take hundreds or thousands of contractors to teach a model a certain behavior, wonder if they really only need to do it once for each desired behavior? (of course, future models might expand and refine this previous training) Because if that is the case, then wow, then future models can really expand their capabilities really very fast.

    ...Am wondering if they somehow record a training session so they can play it back whenever they need to train a new model with the same info? Or maybe the new models can just use distillation from the old model to relearn the old behaviors?

    • The contractors don't directly teach the model. They create datasets. Mostly they create tasks within an environment that the current generation of models wouldn't be able to do, they then write maybe a solution, a set of rules for evaluating the response, and whatever is needed for the RL. These tasks form a dataset. You can see examples of a task in the Mimo dataset that was open-sourced recently. The dataset would then be used by engineers for post-training of whatever model. Some of the model iterations that are released every month tend to be just a further post-training of the same base model that was pre-trained months ago. OAI recently has been doing a lot more pre-training, but for a long time they had the same pre-trained base model. This is why you see so many releases done so fast by the labs, they just need to post-train the same base model on whatever task they think would be better suited, I would speculate based on what users want and what the benchmarks test for.

      Regarding your second question, I think if they want to further post-train a model using new data, they wouldn't feel the need to re-train it again on the data that it has already being trained on. But you never know. If the model is a completely new pre-trained base model, then they could either train the model using all the data and/or use a previous model to teach it. There is definitely a bunch of tricks they do to evaluate the models and check the performance or whatever their recipe is. It's really up to what the engineers would prefer. But you get the core idea, the models are not suddenly coming up with how to use the Blender on their own, they are explicitly being trained on a dataset curated by a professional that teaches the model how to use Blender. Surely there is another aspect that if the model gets better at coding, then it also helps it become better at Blender, and you have that transfer learning. However, there is no emergence or a deity popping up. But you get people who were evaluating theses models on blender use and suddenly seeing the model ace their tasks and they think they are dealing with a super-intelligence. They then undergo an AI psychosis once they try and extrapolate the (super)-exponential improvement in that one task over the next few months and across all other domains.

      Regarding your third question, I already answered at it. But, when it comes to training, they definitely freeze the weights after each run just in case an issue arises and they need to address it (a GPU not working or the loss value blowing up).

  • What you're talking about is not synthetic data. Synthetic data is data generated by an LLM. If the data is generated by a human, it's not synthetic.

    You're talking about real data created and curated by humans to help in training LLMs.

    • This is true and thank you for correcting me. I make a mistake of using synthetic data for what is human-generated data. There are augmentation tricks they could use on the human data but still human-generated data shouldn't be mistaken with synthetic data. I believe my core arguments still stand. Specifically that these models are learning their capabilities through specialized data that human professionals curate. There is no emergence or "general" intelligence. This is really what the open-weight models lack.