← Back to context

Comment by warpech

3 hours ago

I wonder what’s more valuable in our prompts: the raw data or the feedback system that drives the exchange towards a goal.

For a long time it was clearly the former, but now I think it is the latter.

The models have enough knowledge (orders of magnitude more than a human could ever learn) but are now getting better at what to do with it thanks to learning from the decisions that we make in conversations with AI agents.

I think so too. The value is in the entire conversation. IMO, "domain experts" don't run LLMs blindly and hands free. This does not work for top level work (e.g., mathematical proofs, coding anything more complex than yet another slop game or website). Experts have long sessions where they prompt and guide LLM in response to what it produces. This is the discovery process. And frontier labs definitely train on that.

The billion dollar question is whether this works "out of the distribution". I.e., whether LLMs can only find and use the specific ideas buried in training data, or whether they can learn to apply the "thinking process" to a new problem. IMO this is still unanswered (due to these recent controversies).

But regardless of the answer, it seems we have a planet-scale positive feedback loop here. LLM became good (enough) by training on generally available data (books, internet, github) + RLFH, so experts tried to use them on hard tasks, which required lots of hand holding. These conversations became part of the training data, and the next generation of frontier LLMs were better. So, more experts used them on harder tasks, again requiring hand holding. These conversation became part of the training data... etc.

In a nutshell, top human minds across the world are pouring their skills into LLMs just by using them. This is not "continuous learning", but if you re-train on the most recent sessions every, say, quarter (which seems to be happening?) you get close to that in practice.

  • Last year we were saying there must be a human-in-the-loop (HitL), but anyone who is the HitL exhibits the “HitL skill” to the agent.

    There might be no books about human intuition but we teach it to LLMs by interacting with them

  • 10000000% Correct.

    I’ve been working on a novel project for 1 year.

    I now no longer use llm’s - the continual chatter I’ve had has resulted in my insights being found in the training data now.

    Get stuffed OAI.

    Every large firm will soon enough want its own on-prem servers eventually. Maybe nation’s will get involved and build out their own data centres.

    Not a chance in hell I’d trust a tech firm to treat my IP as safe and sound - only a sovereign can ‘promise’ that.