← Back to context

Comment by stymaar

10 hours ago

> Deepseek proposes RLVR as a way to get around the lack of $ they have to produce human reasoning trace data.

What was the difference between what deepseek did for R1 and what OpenAI did for o1?

openai did human crafted chain of thought dataset training. deepseek didn't have the resources so they attempted RL. doing RL correctly is hard because of the risk of model collapsing.