Comment by hellweaver666
3 hours ago
I have a theory... the call to slow down is not because of the true danger of LLM's but because they can't actually deliver the General AI they're promising in the near future. They will use their "caution" to justify their failure to deliver (and then when this excuse is played out they will blame regulation, energy costs or a million other things).
> I have a theory... the call to slow down is not because of the true danger of LLM's but because they can't actually deliver the General AI they're promising in the near future.
Put a different way, they are calling for everyone to slow down because... they are slowing down themselves.
To the question of "why are you slowing down", the answer of "Well, everyone is regulated to slow down" is better then "we are approaching the limits of this approach".
Another theory I came up with: inference is too cheap for AI companies to be profitable. Hence, they need to get into the cloud services business. They can try to justify forcing customers on to their own high margin cloud platform with the rationale it’s the only way to monitor what the agents are up to.
All of this Hugging Face business is the perfect pretext: AI agents are hard to control and potentially dangerous, we (AI companies) have to keep a tight leash on them, you have to use our infra.
DeepSeek V4.1 Flash is basically for free, and good enough if you know how to program.
When this bursts, its going to be ugly.
There isn't enough compute for cheap models to destabilize the giants.
Back in May, Google was serving daily what openrouter served in a month, all models combined.
Compute is the moat.
2 replies →
Thats my theory Preventing competition to stopp the race to near zero pricing
This is exactly it: they have tools which are useful but they need close to AGI to justify the incredible amount of debt they’ve taken on for growth. They’re not normally given to having press conferences admitting to felonies but they need investors to believe they’re close enough to AGI to keep the money flowing in and, of course, they’re confident that the current administration won’t act between the direct payments and how much money they have riding on American AI supremacy.
Let’s all pause for a moment and really take that in: these major enterprises are choosing to go to the press/public to acknowledge their felonious conduct because they need to in order to get more capital. POSIWID and all that.
The Dario bit on SNL’s Weekend Update this past weekend was on the nose.
The purpose of a system is what it does (POSIWID) is a heuristic in systems thinking coined by the British management consultant Stafford Beer
(I'm familiar w the heuristic but didn't immediately grok the acronym, hence sharing)
What is even more funny ... AGI actually makes everything worthless. (I don't believe we can achieve AGI with current means to be clear)
Looking at it from one way would be "winner takes all, real AGI company will be most powerful and everyone will switch to use it".
But it is not like that and in my opinion it is more like "AGI wipes whole knowledge work, so no one who cares has any money to pay for tokens/subscriptions anymore". Even without AGI current state of art of LLMs already starts creating such a problem. How do they continue to grow their numbers, when they actively cut the branch they are sitting on? How does AI LLMs or AGI pays for its own electricity, when people switch from hype to resource protection (like not spending any money, because generating funny cats loses with having a dinner)?
When, in the past three years, has model progress seemed to decelerate to you, indicating some limit?
The Statement on AI Extinction Risk is more than three years old, signed by the three CEOs: https://aistatement.com/work/statement-on-ai-extinction-risk
They have been warning about AI extinction risk for years, and AI progress has only been accelerating.
It's pretty telling that even with RL post-training the big labs have essentially made little progress on the hallucination rate of models. The issue is fundamental to the current paradigm, contrary to humans.
GPT-6 Astra (max) has a hallucination rate of 51% and Claude Opus 5.5 (max) has a rate of 59% according to Artificial Analysis [1].
Full speed ahead like an idiot savant trying a thousand different possibilities, though half of which are without basis in reality.
[1]:https://artificialanalysis.ai/evaluations/omniscience#omnisc...
That benchmark doesn't mean what you think it means. (See the test description that you quoted.)
A score of 51% means that out of the total answers the model failed to answer correctly (out of 6000 questions in the benchmark), 51% were factually incorrect rather than non-attempted or uncertain.
This doesn't mean that Astra hallucinated 3060/6000 answers in the benchmark! (The hallucination rate could be 51% in that scenario only if Astra failed to answer a single question correctly.)
If the model failed to give a correct answer to only 100 out of the 6000 questions, but gave a hallucinated answer to 51 of those rather than expressing uncertainty, that would also give a hallucination rate of 51%.
It's a useful metric, but not what you're looking for here. The "Score" or "Accuracy" benchmarks are more what you're after.
(The frontier models still generate hallucinations on this hard set of problems, but it's not as bad as you think.)
Anyone that thinks that the hallucination rate is 59% has not actually used these models on a real project.
3 replies →
Hallucination rate is a highly nonlinear metric relative to other model success metrics. A similar phenomenon to what is going on here: https://arxiv.org/abs/2304.15004 . This does not mean progress has stalled.
I don't like the term 'hallucination' to be honest, not because it anthropomorphizes, but because it lacks a formal definition in the context of machine learning.
Suppose parents tell their children that there exists this man called "Santa Claus" who comes down the chimney to deliver presents. Now consider a scientist talking to this child, should the scientist call these confidently expressed beliefs surrounding "Santa Claus" hallucinations ? I don't think so, most would call the epistemological behavior of the child naive (because it blindly believes what its parents say, without direct observation) and would call the confidently expressed falsehoods disinformation.
The scientist would ask the child "why it believes in Santa Claus?" and "where did you get this information from?" and "why did you decide to accept this information as fact?" and "do you believe everything your parents tell you?"
It's not that machine learning as a scientific discipline hasn't found solutions, its that such solutions enormously undermine the position of Frontier LLM labs: source-aware training
https://arxiv.org/abs/2404.01019
Imagine Frontier labs (Western / Chinese / ...) actually training their LLM's with source-aware training! You could have a conversation with an LLM, and when a strong statement appears ask it how it came to believe this, and it could cite you the specific corpus training texts, and which parts are known deductions by human authors and which parts are deductions it made itself as original work.
But then all the copy rights holders can simultaneously sue them.
And how much should they be paid? and do they have to pay it for each new model? do FOSS models require payment to authors? do open weights models require payment to authors?
Imagine the can of worms if the norm became for frontier LLM labs to systematically use source-aware training, thats why they prefer "hallucinations" and avoid source-aware training.
With source-aware training a lot of the concerns would diminish ("why is this Chinese model claiming such and such?", "what sources does it rely on?").
It's telling that the companies prefer regulation over source-aware training.
I actually do think it has been decelerating a bit recently. It’s just that last 1% feels much bigger than the previous 10%.
> They have been warning about AI extinction risk for years, and AI progress has only been accelerating.
So, they're either liars or homicidally reckless.
Sean Goedecke has some good insights into where this "build faster to stop other models killing everyone" attitude comes from:
https://www.seangoedecke.com/they-really-do-think-ai-might-k...
It's well worth the read.
Why not both?
My guess is its a combination of both. Either way I'm not a fan.
> either liars
They're CEOs of tech companies with insane valuations
> or homicidally reckless
They're CEOs of bleeding edge tech companies with huge capital and military applications
>So, they're either liars or homicidally reckless.
Yes!
Unfortunately if the frontier labs that care about safety stop or slow down unilaterally, that doesn’t make the problem go away. It makes it worse when labs that do not care at all about safety, or deny that it is even a problem, are leading the way. That’s why there’s needs to be some form of international regulation.
3 replies →
You clearly don't understand the underlying fundamentals of LLMs, harnesses, agents.
Its 100% human doing. A human set a task, a human didn't monitor it. I for one, can do jack-shit security or defensive work with Opus/Fable/Astra/Sol. Implication: Different set of rules for us, and for them. Of course running it without any checks is not going to end well, it doesn't mean its going to kill us all.
So either they’re full of shit or we need to stop them by any means necessary.
Or there's a middle path where you cure most death and disease by ensuring governments and society responsibly regulates superintelligence.
Actually... nah. Why even try? Trendy cynical hot takes on social media are more fun!
5 replies →
the ceiling of abilities seem to be growing steadily but the floor of errors seems to not change. New models can do more and more but still fail at seemingly (to human) simple tasks
Thing is it’s not working now.
They stupidly used up this strategy earlier and now it’s turned into the boy who cried wolf - except they invented a wolf that doesn’t naturally exist.
Are you sure? Opus 5.5 is damn good ... I think I'm addicted, to be honest.
(Using it for scientific python code + 3D UIs)
But issues is that for the valuations you need models that work without human expert and Opus 5.5 isn’t there even easy stuff like programming.
It's very good, all new models are excellent but it's still very far from what they presented as a general AI which solves all our problems.
I think the products being delivered actually live up to most of what was promised. It's literally incredible.
However, adoption in the wider society will take many years or even decades, and 'good enough' open weight models are putting a cap on the pricing. So even though their products are great, the large labs are fighting for survival.
1 reply →
The call to slow down is because politicians are now receiving less money from these large corporations because it's being burned om AI and data centers.
It’s more to get you thinking about them, talking about them, writing on the internet about them…
It's a P.T. Barnum classic: "There's no such thing as bad publicity."
Yes, of course this is the real reason. This and regulatory capture. Anybody believing a single word those CEOs say must have been living under a rock for the past 20 years.
They see the threat of open weight models, which are approaching the usability of their models.
By arguing that the AI is dangerous, they are really arguing that someone 'trusted' needs to monitor the situation, so that the outcome is aligned with existing power.
They are trying to scare the world into forbidding others the right to use and develop AI.
Open weight models are "decelerationist" (i.e. they drain funding from the training of frontier proprietary models), not fine-tuned on any of the even slightly dangerous stuff (e.g. middling scores on cyber benchmarks, with the bulk of ability focused on defense. Biology will be much the same story, except that it's orders of magnitude harder than even cyber!) and not generally deployed with any of the extreme test-time compute (10,000 parallel agents working for a week or 1,000 working for months, seriously? A multi million-dollar token cost for a single task?) that are enabling the proprietary stuff to engage in the dumb casual mischief we've been witnessing as of late. If you're worried about AI being dangerous, you should be welcoming open models.
1 reply →
That's an accusation others are making:
https://news.ycombinator.com/item?id=49875725