Comment by NitpickLawyer
3 months ago
> But models did not become good at coding just because coding is replayable. It’s because there are countless repos, issues, Stack Overflow threads, and Reddit posts/comments/questions where a solution is clearly marked as “solved” or “that helped,” and AI can learn from that feedback.
That's at least 2yo take. Today's gains for SotA (either closed or open models) come from RLVR 100%. The model unrolls many iterations, those iterations get verified w/ tests/known tests/rubrics and the model learns from that (grpo or similar).
And what's cool about this (and why scale really matters now) is that you can mostly get this process automated (i.e. take a known good repo, ask one agent to remove one feature, keep the tests, ask another model to add that feature back, verify that old tests work on new implementation, repeat). This is why top labs are pulling away in the breadth of their capabilities, compared to open models. It's scale, pure and simple. And the better their models become, the larger the gap due to automating better cases.
No comments yet
Contribute on Hacker News ↗