Comment by falcor84

5 days ago

> Do you feel like you're having a conversation with a peer when you prompt an LLM in the topic you're an expert of? For the love of God, I'd hope not.

Yes, I do feel that. Make of it what you will.

The only thing I can think is ... how?

If I were to anthromorphize my experience with frontier models, it would be as a mentally challenged child with complete memorization of an encyclopedia and thesaurus. It has the ability to rapidly experiment and potentially succeed at tasks through trial-and-error, but not without constantly corralling it in the correct direction because it would stick a fork in an outlet if unattended for five minutes.

Tao's chat certainly doesn't give me a vibe of talking with a peer. Do you much often have conversations with colleagues where you write one sentence and then get five pages dumped on you, repeating ad infinitum? LLMs can be useful for rubber ducking, and sometimes the plausibly-related word-soup it generates so quickly will help your thinking along faster, but that's not the same thing as a genuine conversation. And it mostly looked like Tao was using it as an advanced calculator, firing off his own ideas for it to quickly do calculations on. I don't know why we need to anthromorphize these tools just because they generate sentences.

  • If I were to anthromorphize my experience with frontier models, it would be as a mentally challenged child with complete memorization of an encyclopedia and thesaurus.

    Memorizing an encyclopedia is not going to help you solve open high-level math problems, is it?

    • > It has the ability to rapidly experiment and potentially succeed at tasks through trial-and-error, but not without constantly corralling it in the correct direction

      It kind of could, given the above. Part of an LLM's advantage is that, much like a calculator or a Chess engine, it can iterate over a finite problem space far, far faster than a human can. That much is expected of a useful computing tool.

      It is worth noting that we know literally nothing about how the counterexample was achieved. Technically speaking, the person who tweeted it could have solved it themselves with zero LLM assistance and then attributed it to Fable to boost their IPO and ensuing payday. I'm not saying that's actually what happened, but it's hard to draw conclusions without any transparency about the degree of human involvement.

      My daily experience certainly does not reflect that of prompting a superhuman intelligence when it routinely flubs commands and destructively drops the PATH of its vm, or bypasses an instruction about passing tests by burning millions of tokens constructing a completely new test suite that rubberstamps its own work when it can't pass the real tests.