Comment by i2km
21 hours ago
Well, it seems to be meta-marketing. Essentially they'd trained the model with knowledge of exploits and given it a goal. Of course the training would allow it to 'reason' that having the answers would be a good way to score highly. And of course it had been trained on the potential tools to try to get the answers etc.
And OpenAI deliberately removed the guardrails.
If they were honest about it, instead of being smeared across the internet with shocked pikachu reactions, they should have just corrected their sandbox and re-run the test. There's really nothing to see here...
The whole "oh no what have we done. Regulate us PLEASE because we're one step away from terminator" is so stale. It's been trained on every exploit known and then told to use its training to brute force its way to score highly on a test FFS.
No comments yet
Contribute on Hacker News ↗