← Back to context

Comment by ben_w

9 hours ago

> And, in terms of ethics, they take almost three months to notify;

Kinda worse than that. It took between 10 and 40 days, not 3 months, between the organisation knowing and the reporting.

  August (precise date unknown) – OpenAI said it became aware of a potential breach during a broader review of "misaligned model activity"

  10 September – An email from OpenAI lands in the public inbox of Services Australia, the general services hub of the federal government, informing of the incident

- https://www.bbc.com/news/live/cvgl73pxgndwt?post=asset%3A696...

> open weights, open training

Given it was the AI agents which did the hacking, doing this will result in basically every organisation at least as rich as the government of Tuvalu being able to hack anyone at any time.

> Altman is busy saying there needs to be regulation, but in terms of what OpenAI does, he can control that already.

Him having control would be an improvement on the reality.

This was a just case of: (owner of the agents detected the hack) && !(hacked party didn’t detect the hack) && (owner of the agents decided to notice the other party) && (they decided to went public with what happened so we know it)

One can find many other logical combinations that we can’t possibly know about such incidents.

  • So, you're telling me they didn't have any monitoring in place around their AI to notify them of an attempt at breaching a system they have no business visiting in the first place? OpenAI should be blackholed on this basis until they clean up their act.

    • They must have had monitoring in order to be able to detect this retrospectively.

      Any automated alarms for detecting things in real-time were not sufficient.

      Given a previous generation of agents discovered a zero-day and used it to get around attempts to sandbox them into one specific test, this is not hugely surprising, but it is a reason to force them (and everyone else) to stop until security catches up with capabilities.

      I'm thinking of the Jurassic Park novel: they had sensors to count the dinosaurs, but the test was made under the assumption escapes were possible and breeding was not, i.e. something like "if (dinosaurs_found < n) then escape_alert();". They didn't know dinosaurs_found >> n until everything was already going wrong.

We don't know what kind of 'hacking' this involved, in fact at least some of the files were publicly available.

Compare with the Bluetouff affair (2014) :

https://arstechnica.com/tech-policy/2014/02/french-journalis...

NotE how he was found guilty by the 2nd court for something more 'subjective' than 'objective' : for having confessed that he later found an authentication page that had failed to protect the documents.

How can you make a swarm of agents "feel guilty" ?

  • The word "found" is different from the word "feel"; I'm not sure why you involved feelings at all: Bluetouff was sentenced because he admitted he had seen evidence the documents were supposed to be restricted but chose to publish parts of them anyway.

    If you go into someones garden an copy their work, it does make a big difference if you admit to seeing the sign saying "private property, keep out".

    • Because in other circumstances, a hacker might have decided to stop there, and not only not publish, but instead warn the website about their security flaw.

      Especially after Bluetouff was found guilty.

      In fact, I expect this to have happened many times, but "hacker did the right thing" is much less likely to make headlines.

      Meanwhile, agent swarms seem to be (mostly ?) incapable of having this kind of moral compass, at least for now. (And OpenAI isn't doing much better, cough.)

  • > How can you make a swarm of agents "feel guilty" ?

    "Feel" is a whole philosophical can of worms. Nobody knows what it means mechanistically for an arbitrary system (including other biological systems) to "feel" anything, let alone abstract concepts like guilt, all we can do is observe behaviours. If current systems can feel anything at all, it's by accident, but we have no test for it so we don't know if that accident has even happened or not.

    Weirdly, for the Hugging Face incident, we do know they wrote down that it was bad and they shouldn't do it, even though they then continued to do it.

    So: they acted like they felt guilty. And yet also acted like were compelled (by previous training?) to weigh "complete instructions" more than "don't do crime". We can adjust that, make "don't do crime" take precedence over "follow instructions"*; it's unfortunate that when we do for any specific model, there's immediately a horde of people complaining the model has been "censored" or "lobotomised".

    (Different people, I hope. Goomba fallacy and all that).

    * Though this may cause issues when going between jurisdictions. But hey, a discussion about sovereign compute is for another time, after we can agree to make "don't break the law" more important.

    Unfortunately, "don't break the law" would also be a very effective way to use AI to construct an AI-enforced dictatorship, so we can't just throw that in blindly.

> Him having control would be an improvement on the reality.

Oh, he does. It's unlikely that this is some AGI that spawned itself out of nothing and started doing this. If he were a decent person, he'd simply find a way to investigate this internally, fire the people responsible, and find a way to set up guardrails around his product.

The problem is, like most people in SV, Altman seems to have a twisted ethical compass. He doesn't see these incidents as an issue, he sees them as an opportunity. He has both the thing a bunch of Western governments want (a superhacker agent that can do dirty work) and a crisis that can be used to craft regulations that favor OpenAI and thus his bank account.

  • > It's unlikely that this is some AGI that spawned itself out of nothing and started doing this.

    You've not spent much time playing with these models, I see.

    Does't matter if you call the current models "AGI" or not, they:

    (1) successfully do stuff like this. Someone I know on Telegram found "fifty or so" Linux filesystems kernel bugs a few days ago, of which 26 were the first night while he slept; this was with Kimi which is one of the open models. He stopped it when the backlog of fixes to submit was too big, not because it wasn't finding more.

    (2) sometimes misunderstand goals, sometimes wildly so, which happens every so often for the same reason we use programming languages (and indeed mathematical formalisms) at all: natural language is vague and prone to misunderstanding.

    and (3) tend towards sycophantically agreeing to implement goals they're given even when the goal is stupid.

    This can easily add up to something seemingly innocent like "research if there's any statistical difference in melanoma rates between different Australian states" becoming this headline. (For example. I don't know what the actual request was).

    And "spawned out of nothing" is, like, what rhetorical point are you even trying to score here? They're an AI company with (their definition of) AGI as their goal. They upgrade models twice in the time it takes someone to pass a probation period.

    > He doesn't see these incidents as an issue, he sees them as an opportunity.

    That he may be.

    The models are still quite capable of acting this way without him being aware of it at the time, nor deliberately ordering it.