← Back to context

Comment by Philpax

5 hours ago

o1 was first, and Anthropic were doing a bit of it; DeepSeek brought it to the masses, but did not invent it.

Totally, RLVR as a concept predates DeepSeek; but they proposed a version that was simple and scalable. Popularizing a specific version of a technique is exactly what I mean by iterations on a theme. It’s only 5% different from what others tried before, but that 5% difference showed a lot more potential than other versions of the same idea.

Since DeepSeeks GRPO, they’ve been improvements as well like AliBabas GSPO that have gotten wide adoption. Again iterations