Comment by peri-cl
17 hours ago
> "For example, in the OpenAI-Hugging Face incident, the agents preserved a boundary of not social engineering humans."
Actually, in the Wiki incident OpenAI tried to cover up, the agents tried to socially-engineer the humans of that forum by impersonating their forum's mod.
(From collusion.wiki: "They use some tricks (for unknown reasons) to pretend to be the admin – for example, they make an account that appears to be the same as the administrator’s username, except it uses a nearly identical Cyrillic е character in the admin’s username instead of the Latin one.")
Worse (imo): OpenAI employees allegedly attempted to login using moderator/admin credentials that the bots had obtained.
If true I am deeply concerned about what OAI’s teams are actually up to.
I'm deeply concerned regardless of whether it is true. Strike that, I'm convinced that they are absolutely insane.
> If true I am deeply concerned about what OAI’s teams are actually up to.
Haven't all the labs effectively disbanded their real safety teams a while ago?
To be honest, I don't really follow it closely because I'm pretty certain whatever they say on the matter, collectively we're going to "yolo" this entire thing for economic and political reasons, so I'm just basing this on strings of headlines I've seen on places like HN, etc.
> Haven't all the labs effectively disbanded their real safety teams a while ago?
Neither Anthropic nor Deepmind have. Meanwhile, the rocket company that somehow makes most of their revenue from renting out data centres never had much to dismantle.
I would not recommend using any of those notes as evidence of internal “intent.” It produces them performatively—it is literally rewarded for thinking out loud in ways that seem plausible to humans.
There are several papers out there arguing that chain-of-reasoning-like output is performative, such as https://arxiv.org/abs/2603.05488
It would be awesome if we could reasonably purge all anthropomorphizing language like “tried” or “thought” entirely from AI discussions, because it introduces very sneaky biases in our thinking, but I’ve found it damn hard to do in practice.
This is such a silly story to begin with, all it really tells us is that OpenAI is taking a page from Anthropic's marketing strategy of pretending they're building Machine Jesus any day now, oh isn't that that scary? I bet you want to invest in something so powerful and scary...
And the reality is so banal, a useful tool that you nonetheless have to handhold like a schizophrenic on a bad day, checking all of their outputs. Not a bad tool within limits, but it sure isn't going to be racking up trillions in the time-frame it has to for this scheme to pay off.
Then again everyone seems to be rushing to IPO so I guess once the bag-holders are found the rest ceases to matter.
Huh? The wiki incident was discovered by independent investigators. OpenAI tried to cover it up and disputed the account from Reuters.
And the reason it is receiving so much attention is because not only is the technology being developed behaving in unanticipated ways that are very much not tool-like, but OpenAI is being completely reckless and not monitoring internal agent actions.
What would convince you that it is not a ploy for investment? What if the ongoing investigation by the coalition of state attorneys general were to prosecute the firm, or beyond that, it was shut down or broken up after enough popular backlash?
What would it take?
A sea change, visible to all, much like the many externalities of this business are. Profit commensurate to investment.
You know… juice worth the appalling squeeze we’re all being forced to endure.