Comment by jagraff
1 hour ago
I don't understand why so many comments here are so confident that this is all marketing, that rogue is just hype, that agents are just simple tools, etc. If a bunch of nuclear engineers were going to the news and saying "Our reactor is dangerously close to a meltdown - we need government intervention now!" would your response be that they're just hyping up boring old power generation technology?
I’d roll my eyes if the engineers stated that they didn’t design the reactor to melt down, and that it simply developed rogue meltdown-desiring behavior on its own, and I would also wonder about negligence if they claimed that nobody could have anticipated this (given that, like with botnets and viruses, we have decades of knowledge and experience regarding reactor meltdowns)
I mean sure, negligence is absolutely on the table; but that makes the problem worse, not better! We don’t allow nuclear engineers to be negligent; they can go to jail if they don’t follow strict protocols to make sure the dangerous systems they work on are safe.
Go and get one of their models to hack something, it won't do it, why?
They have claimed this happened during a "training run", but why are they training on systems connected to the internet?
That's why people are skeptical.
The public models won’t hack because they have a classifier that shuts down anything that looks like hacking; without the classifier they are perfectly capable of hacking, multiple third-party evaluators have confirmed this.
The models were not trained on systems intentionally connected to the internet; they chained mutliple zero-days (that they discovered) together to get access to the open internet and into huggingface.