← Back to context Comment by speedgoose 14 hours ago I didn't know 2 thirds of the training data would be source code. 4 comments speedgoose Reply jerrygenser 14 hours ago that is the the "data used to improve the model" when signing up for the subscription plans leothetechguy 13 hours ago this is the rl run, not the pretraining run ahmadyan 13 hours ago even in pre-training, usually 30%-50% is code these days. leothetechguy 1 hour ago That would be far too high in my opinion. But happy if anybody can give insights from their own experience with pretraining runs.
jerrygenser 14 hours ago that is the the "data used to improve the model" when signing up for the subscription plans
leothetechguy 13 hours ago this is the rl run, not the pretraining run ahmadyan 13 hours ago even in pre-training, usually 30%-50% is code these days. leothetechguy 1 hour ago That would be far too high in my opinion. But happy if anybody can give insights from their own experience with pretraining runs.
ahmadyan 13 hours ago even in pre-training, usually 30%-50% is code these days. leothetechguy 1 hour ago That would be far too high in my opinion. But happy if anybody can give insights from their own experience with pretraining runs.
leothetechguy 1 hour ago That would be far too high in my opinion. But happy if anybody can give insights from their own experience with pretraining runs.
that is the the "data used to improve the model" when signing up for the subscription plans
this is the rl run, not the pretraining run
even in pre-training, usually 30%-50% is code these days.
That would be far too high in my opinion. But happy if anybody can give insights from their own experience with pretraining runs.