← Back to context Comment by danielmarkbruce 10 hours ago I don't think you've ever done either of these training steps. You are just handwaving. 2 comments danielmarkbruce Reply verdverm 10 hours ago you know what they say about making assumptions, yea?and then you are going to ignore all the research and results that clearly show otherwise? why?what might we infer about the importance of data from a learning algorithm like decision trees? danielmarkbruce 8 hours ago Read the paper. They train RLCR on existing big math problems. They subtract a brier score penalty from the correctness reward. No new confidence labels are needed.Existing datasets, different reward function.
verdverm 10 hours ago you know what they say about making assumptions, yea?and then you are going to ignore all the research and results that clearly show otherwise? why?what might we infer about the importance of data from a learning algorithm like decision trees? danielmarkbruce 8 hours ago Read the paper. They train RLCR on existing big math problems. They subtract a brier score penalty from the correctness reward. No new confidence labels are needed.Existing datasets, different reward function.
danielmarkbruce 8 hours ago Read the paper. They train RLCR on existing big math problems. They subtract a brier score penalty from the correctness reward. No new confidence labels are needed.Existing datasets, different reward function.
you know what they say about making assumptions, yea?
and then you are going to ignore all the research and results that clearly show otherwise? why?
what might we infer about the importance of data from a learning algorithm like decision trees?
Read the paper. They train RLCR on existing big math problems. They subtract a brier score penalty from the correctness reward. No new confidence labels are needed.
Existing datasets, different reward function.