← Back to context

Comment by qlte

1 day ago

I believe we are several generations past peak-RLHF at this point. Now it's much more RLVR (Reinforcement Learning with Verifiable Rewards), with a goal/evaluator loop.

Which, conveniently, fits neatly into the benchmaxxing arms race/agentic coding market fit, since you can basically train "directly" on a specific problem space for a benchmark/agentic goal (fudged sufficiently to avoid excess overfitting on public problems/bechmaxxing accusations if real world performance falls short).

The language evolution could be explained by reliance on ever increasing layers of a model judging a model, using a model developed eval, based on synthetic data from a model, etc. And by the time a human evaluator sees it both A/B choices already converged into weird Claude pseudo English as that was baked in much earlier in training.

This, 100%. I don’t think the industry knows how to scale LLMs’ general intelligence much further. The training paradigm is about maximizing very specific behaviors / very specific tasks, but doing lots and lots of them. Which can create the illusion of general intelligence if your tasks are similar to the ones the models were fitted for.

  • If you have watched The Substance, the transformation feels a bit like when things start falling apart in that one.

  • I tend to agree. We will see this demonstrated in novel research done by agents, or, more meta-cognitively research direction guidance.-