Comment by uselessTA
11 hours ago
I know some people who are worried at Anthropic, and their position seems to be "if we don't do it, someone even less responsible will. Unilateral disarmament didn't work and real oversight seems unlikely to happen in time, so we'll just try to be as safe as we can be (while still winning the race)"
Not that they're happy about it, they just see no other realistic choice
I know, that’s the position Dario Amodei argues for in his essays. I did pass their cultural interview and had to consume a lot of their content to prepare, I think I have a good idea of their stated values. But what the company does and what the leadership states their vision is is pretty contradictory.
They are providing everything bad guys need to develop their unaligned frontier models. Chinese models that Dario considers to be dangerous are distilled from Claude, and they know this.
They are creating the FOMO around AI which pushes adversary countries to invest so much into unaligned models.
They offer models as a service they know are jailbreakable and can be used by bad actors.
They are running internal red-team experiments without adequate isolation.
If I take their statements seriously, AGI research should really be seen as bioweapon, or cloning, or nuclear research. Something strictly regulated worldwide, with export controls for HBM and other hardware used for AI training. What they are trying is instead to boost their position by becoming too big to fail and too powerful to ban, but then want the industry to be regulated to pull the ladder behind them. It really doesn’t feel they are serious about their values, otherwise they wouldn’t be offering Mythos (a model that is unsafe from their own admission) as a service to their close partners
>AGI research should really be seen as bioweapon, or cloning, or nuclear research. Something strictly regulated worldwide, with export controls for HBM and other hardware used for AI training
This is basically exactly what the people I know there support (when training & testing future more capable models), if it could be made to actually happen. Something like https://ai-2040.com/
But I'm just speaking for the people I know, so this is probably not representative of Anthropic as a whole.
> Could be used by bad actors
The people I know aren't as worried about jailbreaking current models as they are about future models, e.g. "the ~50% probability that humans are eclipsed almost entirely, sometime in the next 1-20 years" and what happens then. But it's just hard to get people to take that seriously v.s. bad actor threats which are legible but probably not as catastrophic.
I agree that that they are contributing to the race to the bottom via creating more pressure for countries/competitors to move faster, in a way that seems quite bad on this view too. They arguably were the ~first to push for "recursive self improvement" (models helping build future models) which also seems quite bad on this view.
But although I'd dispute some actions + think there's some overconfidence in superintelligence happening soon, I'm not sure I have a better alternative. They probably bled so many customers to OpenAI while they were sitting on Mythos for months.
https://theonion.com/sam-altman-if-i-dont-end-the-world-some...
"AI will probably, most likely, sort of lead to the end of the world. But in the meantime, there will be great companies..." - actual Sam Altman quote, the man is so unhinged he's beyond satire