← Back to context

Comment by Dilettante_

3 days ago

These are always cherry-picked, though. They tell you about the 1/10 that went really impressively, ignoring the other 9 shots at the task where the clanker started to try selling tungsten cubes (in person, wearing a blue shirt).

Fair. But that 1/10 continues to get more and more impressive. Say, Claude 7 creates a new, brilliant scientific idea every 1 out of 1000 times. Anthropic reports "Claude figured out how to tie general relativity with quantum mechanics." Would you hand-wave it away saying that it's cherry-picked?

  • > Say, Claude 7 creates a new, brilliant scientific idea every 1 out of 1000 times.

    I would not hand wave that away - but Claude 5 feels closer to Claude 1 than the hypothetical Claude 7 you propose. I do not believe one can extrapolate LLMs that far ahead despite the very substantial progress so far.

  • It wouldn't surprise me that a company with an effectively unlimited budget could fund enough researchers to solve breakthrough problems, all while using LLM's (which are excellent research and exploration tools!) and then claim that the LLM found the breakthrough. The amount of money Anthropic has to play with is >10,000x what entire fields of research have, I think people don't appreciate how few resources we spend to support people working on hard problems that don't have clear commercial applications.

    Very similar story with security research, LLM's are a super useful tool while hunting vulnerabilities, but it turns out when the entire software industry starts throwing tens of billions of dollars at vuln. research, a lot of stuff gets unearthed, something security people have been insisting on for years and complaining that their work is underfunded and under-resourced.

  • > Say, Claude 7 creates a new, brilliant scientific idea every 1 out of 1000 times

    Sadly, general public (us) is never seeing that model

When I talk about my kid to friends I talk them about he did that awesome thing, I don't specifically insist on the 99 times before where he miserably failed. They're not hidden, and we all know they exists and on occasion laugh about a few particularly funny ones, but overall the idea is that they don't matter much in terms of development, what matters is that if he succeeded once from now on his percentage of success will keep improving.

I don't believe in all the LLm is AI is AGI dream, it's too easy to trigger failure case that show a lack of basic thinking no matter how good they do on these tests. But I also can recognize the insane things that are made possible by them.

PS: I believe llm true power comes from hive/ant behavior, that's why we're so amazed by goal and agentic and sub agent

PS2: it's rather easy to figure out when we're there : when they can /goal it into improving itself until it does strictly better than itself at those benchmark, they've essentially reached mini singularity.

This is exactly the kind of take that the comment you replied to is talking about.