← Back to context

Comment by aizk

17 hours ago

People had joked a couple years ago "Well if they solve a Millenium problem it's AGI"... Well here we are.

Yeah well, its easy to do if you steal someone elses work and then try to threaten them into staying quiet about it

Edit:

OpenAI have now admitted they were training on prompts at the time they made their breakthrough:

https://mastodon.social/@tristanbuckmaster/11723647135247030...

  • lol, the "other work" was also probably 95-99% AI generated. By a similar breed of OpenAI (and some Anthropic) models, as well.

    I dont know why this monumental achievement is being drowned out by some arbitrary drama. No matter which way you slice it, AI solved this problem. Doesn't matter if it was some internal OpenAI model, or whether it was Astra + Fable.

    • Yeah but the mathematicians are claiming that the key insight that made the problem tractable for AI in the first place, came from them.

  • what's there to admit? they always said they do it and there's a way to opt out. you are making it sound more dramatic than it is.

    • This is textbook plagiarism, scooping their result knowing that the research was part of the training data

  • All the ai labs are open about training on prompts. The question is if buckmaster had disabled that with the toggle openAI provides.

> I have a couple friends who did the Math tripos at Cambridge (so a pretty high level!) who work in tech and have unanimously said they have 0% expectations of an LLM doing a millennium problem anytime soon

https://news.ycombinator.com/item?id=42331654

> An LLM is like a well read college student with a nearly photographic memory that sometimes mixes things up. It's great for bouncing ideas off of and getting feedback on them. And yeah, it might product "novel ideas" by mixing and matching existing ideas, but LLMs will never create truly novel ideas. Not in their current form.

The paper didn't really answer the question sadly: their conclusion was just that humans rate LLM answers as more novel than human ones, but less feasible.

https://news.ycombinator.com/item?id=35752293

> But still, the questions in that test are "solved" in the sense of "I can take a dictionary and answers these questions with full certainty". Beyond established knowledge LLMs are monkeys with typewriters, at best.

> I agree but I have tried many times to intersect two ideas with a LLM that would be novel and the LLM can not do this at all. We shouldn't expect the stochastic parrot to be able to do this though and it is unfair to the stochastic parrot.

> It is like expecting a real parrot to say words it has never heard before.

> No one asks that of a real parrot because we don't anthropomorphize a real parrot like we do the LLM

https://news.ycombinator.com/item?id=41525962

  • Will history look back at comments like these as people being dumb, or people trying to cope?

    • It's denial and coping. Most people i see show this tendency around AI which is also why it 's easy to be far ahead of most population nowadays

      2 replies →

    • A 3rd possibility is that they simply have not been exposed to the best models available (which is extremely likely if you only use the free tier chatbots), and/or did not invest the effort needed to truly harness this new very weird new technology, and so had a very skewed perspective of their actual capabilities.

  • You can see that your math friends completely wrote off LLMs entirely and were showing signs of coping.

    4 years ago it was a "not yet" [0], since ChatGPT at this time was not ready nor it was "AGI". Now with this 'unreleased' AI model, it has reached a point where it has solved an unsolved problem which only one human solved a millennium prize problem (Poincare conjecture).

    Now finally "AGI" means something again.

    [0] https://news.ycombinator.com/item?id=33905609

  • Some observations:

    1. It seems at least possible that some of the proof of NS was contained in the training data, making it less novel.

    2. The formalisation of mathematics into lean has been an underappreciated force multiplier on discovery.

  • As someone with a background in AI and who has been playing around with neural nets for decades at this point, it's been genuinely amazing watching extremely intelligent people make confident predictions about AI capabilities and progress, then be so completely wrong.

    There's a kind of theory of mind for AI (specifically neural nets) which I now realise I seem to have which is very hard to explain to people who haven't felt the magic of these algorithms. In fact, the algorithmic details almost doesn't matter at all. When you have a generalised learning algorithm really the only essential components are – compute, data and time. So long as you can scale these you can be certain you will also scale capabilities. There is never any exception.

    That said, the capabilities neural networks tend to progress in step-functions rather than scale in correlation with compute, data and time, because algorithmic improvements tend to come every ~5 years and bring a significant step change in capability (or efficiency depending on what you measure).

    I think people like Dario and others working at frontier labs see and understand this very clearly. And I suspect it's also why they worry about AI risk because even if you ignore the significant increases in compute and data these models are being trained with, it's concerning that it only took two real algorithmic improvements to take us from mostly useless predictive language models to AGI-level intelligence – and we're due another step change.

    • > extremely intelligent people make confident predictions about AI capabilities and progress, then be so completely wrong.

      The ability for the human mind to rationalize conclusions to maintain denial in the face of a very scary future is immense. Genuinely grappling with the implication of where we're headed is usually very crushing. It's not easy to engage with the possibility, and very intelligent people will use those smarts to feel safe.

      2 replies →