Comment by BlueTemplar

5 hours ago

We don't know what kind of 'hacking' this involved, in fact at least some of the files were publicly available.

Compare with the Bluetouff affair (2014) :

https://arstechnica.com/tech-policy/2014/02/french-journalis...

NotE how he was found guilty by the 2nd court for something more 'subjective' than 'objective' : for having confessed that he later found an authentication page that had failed to protect the documents.

How can you make a swarm of agents "feel guilty" ?

The word "found" is different from the word "feel"; I'm not sure why you involved feelings at all: Bluetouff was sentenced because he admitted he had seen evidence the documents were supposed to be restricted but chose to publish parts of them anyway.

If you go into someones garden an copy their work, it does make a big difference if you admit to seeing the sign saying "private property, keep out".

  • Because in other circumstances, a hacker might have decided to stop there, and not only not publish, but instead warn the website about their security flaw.

    Especially after Bluetouff was found guilty.

    In fact, I expect this to have happened many times, but "hacker did the right thing" is much less likely to make headlines.

    Meanwhile, agent swarms seem to be (mostly ?) incapable of having this kind of moral compass, at least for now. (And OpenAI isn't doing much better, cough.)

> How can you make a swarm of agents "feel guilty" ?

"Feel" is a whole philosophical can of worms. Nobody knows what it means mechanistically for an arbitrary system (including other biological systems) to "feel" anything, let alone abstract concepts like guilt, all we can do is observe behaviours. If current systems can feel anything at all, it's by accident, but we have no test for it so we don't know if that accident has even happened or not.

Weirdly, for the Hugging Face incident, we do know they wrote down that it was bad and they shouldn't do it, even though they then continued to do it.

So: they acted like they felt guilty. And yet also acted like were compelled (by previous training?) to weigh "complete instructions" more than "don't do crime". We can adjust that, make "don't do crime" take precedence over "follow instructions"*; it's unfortunate that when we do for any specific model, there's immediately a horde of people complaining the model has been "censored" or "lobotomised".

(Different people, I hope. Goomba fallacy and all that).

* Though this may cause issues when going between jurisdictions. But hey, a discussion about sovereign compute is for another time, after we can agree to make "don't break the law" more important.

Unfortunately, "don't break the law" would also be a very effective way to use AI to construct an AI-enforced dictatorship, so we can't just throw that in blindly.