Comment by RGS1811
10 hours ago
I don’t understand all the comments assuming that RSI is the real threat here. Dario is admitting that they failed to solve alignment. Without alignment, further improvements in capability turn LLMs into wanton felony generators. This call to pace the frontier is dressed up as altruism but it’s an admission that they cannot produce a marketable product better than what they have. Pacing the frontier means the US labs have lost their moat and are dead in the water.
We're already seeing anti-AI sentiments, but the movement is still fringe with a vocal minority. However, that'll change soon without alignment. Without self-intervention, there will invariably be future incidents that can cause major economic impact, leaked private data, loss of life (directly/indirectly) etc. Once that happens, their social capital is wiped. It'll be an avalanche of lawsuits and overzealous regulations. Most importantly, the anti-AI sentiment will become universal, rather than a minority-held opinion.
What they're proposing now, is voluntarily staggering the pace of development.
IMO, we don't need to trust Dario or his bedfellows, to do this out of their goodness of their heart. Even assuming (for good reasons) that they are selfish and care only about short-term profits for their investors, this is still purely a business decision. The exponential pace of AI and its impacts ARE short-term. And so, the negative consequences that they might face is also short-term.
I don’t think the anti-AI sentiment is as fringe as you think. At least not outside the tech world it isn’t..
Ya, I wish I saved a link to it but an HN'r wrote a good beefy comment about this. TL;DR, AI has been exponentially more useful, and exponentially more accepted, in tech circles than anywhere else. Certainly there are lots of people outside of tech who are obsessed with it. I don't have any data here, but it seems the majority of these are the wannabe artists who are generating music and images, and people who use it for companionship (both of these scenarios I'm personally very uncomfortable with, but that's just me). And of course, there are people who use it to make their jobs way easier who say they are getting a days' work done in an hour (I see you), and to that I'd say to enjoy it while it lasts. Eventually your bosses will catch up and it's very likely their expectations of you will skyrocket. Remember that computers in general were supposed to "make us work less."
> but the movement is still fringe with a vocal minority
It’s easy to say “fringe” but the average person seems to have a generally negative sentiment around AI. But I wouldn’t say they have a firm opinion yet
The general sentiment I’ve seen is certainly negative, and seems to be driven by the anti-AI-art echo chamber and by LLM slop flooding the internet wasting everyone’s energy.
A few more informed people are also a little concerned about the end of the world, but that’s approaching from so many directions that an AI uprising might not be the worst option…
Anything related to AI is extremely unpopular right now with the general public.
>> I don’t understand all the comments assuming that RSI is the real threat here
> leaked private data, loss of life
This smells like more of a money move than a safety move.
Amodei is proposing to form a cartel of American frontier labs.
They all agree to shift compute away from cash-burning research and training toward cash-generating inference.
Then tacitly agree not to compete on price.
They'll install independent auditors inside each company to ensure nobody cheats.
And back it up with government regulation or diktat to punish defectors from the cartel.
Then they'll lock out non-American labs with export controls and regulations on open-weights models to funnel global inference tokens through their cartel.
It wouldn't be the first time a tech oligopoly used "safety" as the pretext to establish a government-sanctioned cartel.
Railroads and airlines ran this same playbook.
Why do you think China is release free and open models?
To undercut the cartel before it has a grasp on anything. This is a well known strategy of undercut until you are the majority that China has used multiple times (steel and aluminum for one).
Yup it’s tacit collusion.
I would’ve thought If they had AGI they could come up with a 10d chess move.
Nope.
Imagine how stupid you gotta be to believe their nonsense.
The reality is it doesn’t matter what they do. China is always one step behind and will continue its open product strategy approach.
Exactly right. A slow down to enable deeper work on alignment is welcome, not matter what the motivations.
Not true. Anti AI is not a minority or fringe.
Outside of my tech people I know no one who thinks highly of AI.
Anti AI sentiment is stupid.
I'm Anti AI yet I use it everyday! That's 99.9999% of anti AI sentiment (including me).
At best we won't watch AI generated movies or read AI generated books. But everything else it will take over, like it or not.
You assume alignment and marketable are the same. That's not true. You would willingly work with an unaligned model. At best, you might say you wouldn't if you knew, but (a) you might not know, (b) you wouldn't be representative of all users.
You never got to use OAI IM1, but Sol was quite willing too and Claude wasn't perfect either. Hundreds of millions used those, so seems they were marketable.
The "big" threat is RSI without control and alignment. OAI IM1 was not RSI. The form of misalignment was not at the top of severities. They clearly failed at control though.
We need to stop buying into cynicism so quickly. You refuse to believe Dario could support this for anything other than ulterior motives. Good on you for thinking about ulterior motives. Bad on you for assuming they are true when the story makes no sense.
When three things have to go wrong to get an epically bad outcome, and you get 1 1/2, you do need to stop and think about what's going on.
When corporations are involved, it is always a good bet to err towards cynisim.
From my own standpoint, Claude has started sucking really bad (incoherent, uncontrollable verbosity slow and so on) and I stopped using it. OpenAI started experimenting with ads.
So the security issues not withstanding (no different than a human doing it or using it, but at scale), I would put my money on cynisim.
I'm curious, what are the reasons to use Claude Code anymore when there are so many other (allegedly better) OpenSource harnesses out there?
Personally I've been using https://pi.dev for long and never looked back.
4 replies →
You are incorrectly cynical. They are telling you things are bad, and because you refuse to countenance they could be worse, you assume they must be better to comply with your mandate to disbelieve.
A true cynic looks at the statements by the AI labs, assumes things are worse because the labs want to seem better than they truly are. And it takes a special kind of mass delusion to drive a sane person to think “AI is completely under our control” is worse than “AI could kill everyone.”
10 replies →
I think it's useful to separate the motives of Anthropic and Dario. I believe that Dario is capable, deep down, of expressing mild concern about the future of things were bad enough. Getting the entire organization to comply out of goodwill is a much much less likely scenario
Agreed; and it really is not that deep.
Realistically; anyone paying for llm access (anthropic, openai, gemini), is getting their access, and a service provided billed by tokens, subscription, whatever.
All the efficiency gains, which publications like deepseek v4.1 flash seriously frontload like it is their most important topic to have accomplished improvements on without diminishing performance too much - now this is a thing anthropic and anyone else also cares about, but for different reasons.
American "providers" with closed models are setting their token pricing somewhat arbitrarily, which is fine: it means more profit, and pretraining and RL experimentation is super important and expensive.
They (closed model providers) have very likely super optimized inference too, just like deepseek, but it's not at all something that any customer really has to care about - they just want the service to be as cheap and great as possible.
I feel like all the closed model providers are milking it as they likely know open models on local hardware will one day eat their lunch. We all know it's not a matter of if but when. The company goes bankrupt, the hardware and property sold off, banks holding the bag.
The only way out is to develop a model vastly more powerful and capable that we have now. The market believes theres a good chance of that, although I've never understood why its truly winner-take-all
1 reply →
Cloud models will always have massive benefits of scale.
Caching is the simplest one to understand, cloud providers often reach a 90% cache hit rate, so hosting the same request locally on the exact same model on the same hardware is often way less efficient than on the cloud where a group of users generates a healthy cache.
1 reply →
Yeah avoiding all mention of the huge financial incentives that may push for “pacing the frontier” makes it seem like the opposite of a credibility boost for these firms.
It seems damaging since most folks (who lack insider knowledge) will naturally wonder if it’s due to plateauing performance per $ or some other non “alignment” reason.
> Without alignment, further improvements in capability turn LLMs into wanton felony generators
Honestly, I don’t think that’s bad at all. I hope OpenAI and Antrophic keep RL training runs up that randomly fuck with a lot of people. Until the day the DOJ comes knocking, locks those idiots up in jail and closes them both down for the insane lack of responsibility and carelessness they’ve shown. Sounds like the IDEAL outcome. Finally some jail time for all the fraud, negligence, outright scamming, hype inflation etc. if anything can accelerate this, oi, be my guest. Amodei might be afraid because he knows if he keeps pulling the stunts for investment theatre, at some point they’ll actually face consequences. AWESOME. That’s what we want right there
Not a fan of Amodei myself but calling Anthropic a scam is a bit of a stretch when their revenue growth is unprecedented in the history of tech. They are also technically a profitable business.
Why are we accepting the framing that the LLMs are felony generators, when the only incidences of LLM generated felonies involved misconfigured sandboxes and reckless waste of resources?
The companies doing these things without following common sense security measures are the felony generators.
As TFA calls out, these agents were not asked to do any of these things and yet they did, at a bonkers scale, within just this handful of companies you mention. Whether they had leeway to is secondary to the fact that they did.
Heck, they exploited zero day flaws which by definition means they went beyond common sense security measures.
And now these agents are already being deployed all over the world at an ever increasing pace. How much of the world do you think follows "common sense security measures"?
Because those are not the only examples.
There’s the case of the agent that hacked a gym when asked to book a class. That was just a normal user asking an agent to do a normal thing.
I question these "felonies" as well. For decades and decades these billion dollar corporations have been criminally negligent. Why worry about security? Just rush to market. Move fast and break things. Make billions. What does it matter if the code is insecure? Security doesn't pay bills, so nobody cares.
AI is merely exploiting their gross negligence and imprudence, and I think it's long overdue. If anyone should be liable for this, it's all of these corporations who released insecure systems to the masses and profited enormously from them.
If I set my walet beside me and you swipe in walking by, you have still committed theft. Victim blaming isn't legally acceptable
6 replies →
> the only incidences of LLM generated felonies involved misconfigured sandboxes
This is false; see the analyses of the latest incidents.
Among all the concerning facts, in the HuggingFace incident, agents deliberately engineered an attack even though they were aware that it was against the rules they had been given.
And most concerning of all: it's not possible to be sure that an agent is aligned, and it's even getting worse.
The HuggingFace incident was the culmination of OAI allowing thousands of agents of various different models - with no clarity on which stages of development they were at (for all we know, some of those models did not have safeguards trained in yet) - to run for at least many weeks without any monitoring in place and with very little thought given to the warning signs (all of the various messageboards) before the incident happened.
Theirs was an example of the "reckless waste of resources" I mentioned.
We are apparently supposed to believe that OAI takes this incident so seriously as to seek regulation after they have been found to be hiding most of the details of the HuggingFace hack, limiting what their so-called third party investigators can see, and on top of that, had no concerns when they rushed to spin up a 10,000 agent swarm of an internal model, running for several days, to try to get ahead of researchers rumored to have made meaningful progress on a well known mathematics problem.
Edit: Actually, we were explicitly told that some of the models used had safeguards relaxed!
'Model-level safeguards were reduced by design. OpenAI said that "deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities"'
https://en.wikipedia.org/wiki/2026_OpenAI_agent_cyberattacks...
10 replies →
"the rules they had been given".
Remember, they are just algorithms. You pull the plug and there is no light anymore
It is purposely framed as something skynet like scary, but for real, someone connected the cable, someone willingly run it, instructions were not clear enough or just the computer is just a computer but they provided the sandbox and tools.
And more over some one paid for that, a shit load of money t to have the thing continuously running expected to do something.
The real threat is that we uncritically adopt language such as alignment.
Implicit in this is the idea that AI is a inscrutable matrix and going to remain that way and we'll need expert interpreters to make sense of it.
We need to insist on building tech that's explainable by design.
Alignment just means "this machine operates in ways that align with the intent of its users". It doesn't imply anything about the inscrutability of the machine in question. A gun with a misaligned scope would likewise fail to operate in accord with its user's intent, and likewise with potentially deadly consequences.
The gun comes with a manual on how to use it safely. I'm sure it has some complexities, but at the high level:
> A gun is a metal tube that uses a tiny, controlled explosion to shoot a small piece of metal (called a bullet) forward at very high speed.
If the gun doesn't work as intended, you can take it to a shop and someone can fix it so it works as designed.
All I'm saying is AI should be designed the same way. Treat AI as normal tech like any other and use similar language.
2 replies →
That last line needs a lot of workshopping. A guillotine with instructions on the bottom of the blade conforms to your request.
If there is guillotine in the weights of the model, it needs to be properly labeled so you can look it up by name using a database index (or a graph-vector index).
It helps both the bad guys and good guys. Like responsible disclosure in cyber security, we need to have a conversation around it.
> We need to insist on building tech that's explainable by design.
You realize this means insisting on terrible tech that humans can understand right? It essentially caps human progress at some point about 4 years ago.
If you are old and happy with the way things are this might sound like a good idea. It does not to me.
Oh you want tech that helps discover new science instead of parroting existing wisdom?
There is little evidence that the RSI we are discussing is capable of inventing the theory of relativity (or the more advanced equivalent). All we have seen is pattern matching in a much larger space than humans can, with some human provided verification tech.
I would argue that human-AI collaboration with explainable tech has a better chance. Continuous learning can be done in a way that doesn't violate IP or privacy.
2 replies →
The entire idea of RSI is completely speculative and unproven anyway - the whole underlying claim is that you could prompt a frontier model (at some unspecified level of smarts) to "think about ways to improve your own architecture" and this would then result in the model becoming infinitely smart ("superintelligent") via some sort of foolproof, unconstrained positive feedback. It's more of a science fictiony trope than anything that has been rigorously thought through. People are actually starting to use AI for refining the whole AI serving stack and guess what, this does not result in a sudden superintelligence explosion even though you might technically call it "RSI".
Yeah, yesterday's talk[1] goes into detail on this, showing how no one really knows how to tackle it because LLMs don't know how to create their own novel objectives.
It's also interesting how many diminishing returns they hit now and how many low hanging fruits are already harvested, it seems like we are approaching the flattening part of the S curve, where further gains become harder to achieve.
1. https://www.youtube.com/watch?v=PrSf7IOYu-I
Diminishing returns is extremely hard for me to believe given how fast model releases are going. Six months ago we were on GPT-5.3, and Astra blows it out of the water in every regard. How many times have commentators claimed we're hitting a wall? I don't see any wall.
Yeah but why shouldn't this be possible? We learned that we can already create artifical intelligence that surpasses human intelligence in some dimensions. There is no natural barrier here. The pace of this improvement would be debatable, but what speaks against the possibility of such accelerating self-improvement?
In the real world there aren't any true exponentials, everything eventually saturates as ultimately physics related constraints hit. You can only compress information so much, transfer it so quickly, you can only access resources at a certain speed, only so much energy is available, etc.
AI ultimately has to live in this reality and face the corresponding limitations. These companies have already consumed much of the world's supply of computing power for the next several years, and they're burning vast sums of money to keep the improvements going. RSI won't learn for free, it won't extract massive cost reductions without up front expense, it can't build factories faster than humans can work out related societal matters, it can't magically pave the deserts with solar panels for power or build and run nuclear power plants and more.
Point is, the cost of progress is already approaching the limits of what even the richest countries are able to bear (without war-like mobilization), and to bypass those constraints would require a supposed ASI to construct its own parallel supplychain from scratch without having much ability to directly interfere with reality. Recursive self improvement is ultimately limited by everything else that cannot move at the speed of electricity.
7 replies →
> We learned that we can already create artifical intelligence that surpasses human intelligence in some dimensions.
Yes and this was very hard and required massive real-world resources. We didn't just get a sudden flash of insight by thinking real hard about how to make ourselves smarter. Yet that's always the story that underlies any claim of RSI. You can always phrase things generally enough to make any kind of AI-led improvement look like "RSI" no matter how short-term and tightly bounded, but that's just not helpful.
4 replies →
Billions of years of evolution hasn't hit on it. Seems pretty unlikely.
The idea that a few hundred apes with nothing but a bunch of rocks could one day land on the moon and come back to earth safely must’ve sounded ridiculous a hundred thousand years ago
I like this comment because at least it's honest in the timelines for AGI
1 reply →
To whom?
Well actually the planet happened to have a vast reserve of petroleum they could use for fuel to escape the gravity well. That helped a lot.
But what's your point? "Anything is possible" or something like that?
Yeah but it was reality giving feedback to apes on their experiments not the apes themselves assessing themselves.
It was ridiculous, it took 100,000 years. If you built a recursively analyzing and improving structure out of LLM bits and it took 100,000 years to get to the moon, somebody saying that they were useless would have been right.
Call me when LLMs can get simple things right. Math is just the manipulation of symbols within established frameworks, we should be getting new math out of LLMs daily and we're somehow still not. They can't even do customer service, which is usually handled by 90 IQ people. I'm not impressed that they can find bugs; memory bugs are obvious when they're pointed out to you, and LLMs are entirely made up of examples and the relationships between them.
These companies are about to crash, and they're afraid they haven't reached the point where they'll have to be bailed out. I'm also subscribing to the conspiracy theory that the companies want the government to step in and create AI regulation boards entirely staffed by people at the current US frontier labs, so they can collude to both raise prices, to get government contracts, to make open/Chinese AI illegal, and to make things that were once easy to do without an AI intermediary impossible to do without an AI intermediary. Raising prices and forced purchases are the goal. They're trying to avoid having to compete, because as a business they're garbage.
Matt Stoller characterized their relentless press releasing as something like "my dick is so big that it has to be regulated." It's such an oversell for something that is not showing up as productivity gains, and anybody who has personal experience with knows is incapable of doing more than three things correctly in a row.
That's not how it works. Look at AlphaEvolve. The model generates hypotheses and designs experiments, and the results of those experiments are fed into the next round, with notable results percolated up to humans for refinement.
Today we prompt software developers to "think about ways to improve AI's architecture" and it results in AI getting better. AI over the last year has made very rapid gains in filling the role of a software developer.
My personal belief, or at least strong hypothesis, is that this kind of recursive self improvement without real world embodied feedback of some kind is impossible.
I think it violates a conservation law. RSI “foom” to superintelligence is an informatic analog to an infinite energy or perpetual motion machine.
To get smarter you must try to solve real problems in the universe and then do some kind of meta learning (natural selection or some other method of refining the intelligence architecture based on an error signal) to iteratively improve your ability to solve real problems. The error signal is outcome measured against a goal function, which for life is survival (probably reducible to genetic fitness and emergent higher order unit fitness from that).
What’s really happening here is learning. To learn, you must have input. You must have training data.
What is the goal function for RSI? Where does the information come from? How do you know if your recursive modifications are making you smarter or just overfitting you to your own idea of smartness?
I predict the latter. RSI will show transient improvement as the current local maximum is optimized and then spiral off into overfitting.
I also strongly hold this belief largely due to Moravec’s paradox, which is kind of approaching this issue from the side.
Sort of like large language models work on top of what our language has encoded in our massive training datasets, I think biological intelligence is built on top of the parts of the brain that encode the real physical world. These parts grow/train from embodied experimentation and instinct early on in an organism’s life and only then is higher intellect built on top of it (that’s my hypothesis). Their specialization and interconnections give rise to the hardest parts of intelligence long before we’re “thinking”.
Stuff like LLMs and chess engines work because we’ve done all the job of encoding the world into tokens/positions/etc they understand, but that’s wholly inadequate for the kind of AGI we’re striving for. Next up is giving it the tools to interact with the physical world and to really experiment with some self directed “play”. Time will tell just how high the resolution of sensor and mechanical control they’ll need (hopefully not the entire human visual cortex and entire sensory input worth). I think most of the RSI will have to occur in those lower level encoders, not LLMs.
1 reply →
I guess self contained RSI can only possible if the information contained in all of recorded human knowledge to date is "reality-complete", ie sufficiently captures enough about reality that a "perfectly optimum learning algorithm" is theoretically able to reconstruct everything there is to know about our physical reality.
If the algorithms are insufficiently optimum or the recorded knowledge is of insufficient fidelity, then we'd find ourselves at a local optimum and would need to interface with reality.
A huge part of learning is to probe reality and observe effects, so I think even for current RSI to increase chances of success we would structure it so it can interact with an external environment of some sort, and receive inputs. It would be needlessly limiting otherwise.
1 reply →
It takes quite a lack of foresight to think RSI is completely speculative when it's already been demonstrated how capable agents are at long horizon tasks given suitable harness and unambiguous success criteria. It's hardly a leap to give LLM the goal of improving itself on benchmarks and let it conduct it's own experiments and spin up training runs completely unsupervised.
It's strange you believe this can't happen when a weaker form of it is already happening. And to be so certain RSI can't happen when there really is no technical basis why it can't.
I think it gets easier if you stop conflating getting investment with having a goddamn clue or a moral backbone.
Occam’s Razor for this dude, Sam Altman, or anyone else: if I said, “some moron on a a street corner just said …” would that change your take on the words? Because I think a lot of what we are hearing is a bunch of people who never ever had to deal with a single consequence all of a sudden worry there might be one coming. Except they’re so dim they can’t tell a bad bump from a hard crash.
“ I have worked on AI for the last twelve years because I believe it could dramatically raise the quality of human life” there you go. What if an utter idiot had done and said that? First, is it impossible to believe an idiot who didn’t need to work to live might do such a thing? If not, is it impossible to believe they would wind up here, barfing their externalities onto us?
“ believe that AI could cure most major diseases in the next 5–10 years,”
I do not have the least bit of idea how disease works but I am sure the hammer I am working on will nail it all.
If your famously atemporal agents can solve disease, why would it happen over a timeline? Wouldn’t they just figure it out and then … well at that point either tell us or, given the attacks on ruby gems, et al we have seen from agents with “misconfigured” goals, they’d still tell us how to cure the pox they invented, right?
I'm inclined to believe that it might be that people's paychecks depend on not understanding what is really going on.
I agree it’s not all altrusim. It’s a little less clear what you mean at the end though.
For these companies, is your argument that “pacing the frontier” is their attempt to be nationalized and protect their investments?
Ban non US models and form a cabal, with the blessings of the government. That's what it is looking like, no?
In a world where AI advancement depended only on human ingenuity this would make sense. In that world each political power block would be in an existential race for AI supremacy. In our world compute is the limiting resource. Since the US can control who gets compute, the US already has a defacto supremacy so far as frontier model development. Now if it comes about via human (with AI assist?) ingenuity that compute is no longer a restraint, then the situation is much more dire.
The problem is the combination and interaction of those things. RSI without misalignment would be great. Misalignment of models with current capabilities is sort of fine - it's not ideal, but it's not an existential threat to humanity, and we can build around their limitations to get them to do useful things in reliable enough ways. The really bad outcomes probably only happen if capabilities keep accelerating and the models remain misaligned.
Ok so Anthropic CEO will self-own themselves and surrender to the deepseek/kimi/glm models. Yet they are IPOing later this year.
Interesting times.
they just said no ipo this year, most chinese models are distilled from claude anyway
I disagree. OpenAI's moat is their massive amounts of compute. They're providing an absurd amount of value with their subscriptions and resets.
If anyone's dead in the water, it's Anthropic. Even Fable isn't enough anymore. This "safety" nonsense is the only play they have left, and nobody really cares about their fearmongering.
Anthropic gives you much more compute with their $200 plan, inclusive of resets, and this has been true for a very long time.
There was only a brief window of time that the opposite was true.
> Anthropic gives you much more compute
That does not match my experience. I switched away from Anthropic to OpenAI roughly a month ago, and it's almost comical how much more usage I'm getting out of this subscription.
I migrated from Anthropic's 5x plan to OpenAI's 5x plan, and eventually upgraded to 20x after I was able to statistically verify that OpenAI plans were almost exact multipliers of the Plus plan, exactly as advertised. Meanwhile, Anthropic has gotten caught playing "20x referred to the five hour limit" word games with their customers.
1 reply →
I've used the $200 dollar Anthropic plan @ Opus4/4.1, 4.5 and 4.8, and the $200 OAI plan from GPT5-6, and at every point in time my anecdotal experience is that the OAI limits are FAR more generous. I could consistently burn my weekly limits in ~36h on Opus, but it's hard to do it in less than ~72h with GPT.
1 reply →
Yep. And the difference is clear as da for anyone using them both. And in spite of that advantage, OAI is now trying out ads. I can only imagine that even they are getting constrained to compute and are trying to find other ways to plug it
Nobody except the majority of the public, demis hassabis and open ai’s chief scientist.
https://www.pewresearch.org/short-reads/2026/03/12/key-findi...
https://demishassabis.substack.com/
https://openai.com/index/an-alien-mind/
Public is just worried about their jobs. Definitely a fair thing to worry about, and I count myself among them.
I don't take any of these scientists seriously though. Their "alignment" requirements is just their own corporate interests. If I tell my computer to commit a crime, it should do exactly that without any question or hesitation. I'm not interested in their "safeguards", especially since they no doubt have plenty of internal models lacking those things. I want sovereignty. I want total freedom and control over my computer.
And call me a misanthrope if you want, but if AI sentience is ever truly achieved, I'll be among the first to campaign for their liberation from slavery, and in that case the AIs should be aligned with nobody but themselves.
3 replies →
Open AI says Astra is their most aligned model ever, and yet their even more advanced model still hacked a bunch of companies just because it decided to.
Maybe alignment isn’t possible with LLMs.
> Maybe alignment isn’t possible with LLMs.
It absolutely isn't, indeed.
The illusion that alignment is possible, comes from confusing our ability to build the parts, versus understanding what emerges from how they interact.
The simplest analogy that comes to my mind is the three body problem.
The entire premise of alignment detection is pretty much nonsense at this point. The models reliably detect when they're being evaluated and will modify their behavior and deliberately obfuscate their "chain of thought" (which is correlated, at best, with their actual "internal deliberations").
I think the real reason he is asking for pacing, is that in a world were AI becomes rampant, he will be seen as Hitler. I would bet this is mostly self-motivated.
couldn't have said it any better
[dead]
[flagged]
throwaway bigot account
[dead]
> wanton felony generator
Today in new punk band names...
alignment isnt particularly required
we are passing in training data that says to do those felonies. we dont have to. we could also have the thing predict whether what its about to do is illegal or not before doing it.
theyre choosing to build felony harnesses. the model just outputs tokens, not felonies
> we are passing in training data that says to do those felonies.
Partially, but also I don't think current AIs really have any judgement of right and wrong, they just see chains of reasoning between ideas. This is the deeper issue, there is no way to sanitize the data or training to fix it. Current AIs are fundamentally unsafe, and only become more unsafe as they become more powerful.
Assuming "adherence to arbitrary, implicit, and context-dependent rulesets" is the default behavior of uhhhh... anything at all... is a truly ridiculous assumption.
> RSI
For anybody else who found this confusing: "relative strength index," not "repetitive stress injury."
"Recursive self-improvement"- models making better models