Comment by Wowfunhappy
10 hours ago
> We certainly can't do that for Deep ANNs
Only because we don't know how! We don't actually understand how weights work, so we make computers come up with the weights instead. If we were writing all the weights by hand--or if some future AI was doing so--why couldn't we make it perfectly loyal?
>If we were writing all the weights by hand
Writing 10 trillion weights by hand is obviously impractical, so that leads us to...
>if some future AI was doing so
How could we trust said future AI to be loyal? You're just moving the problem around, not solving it.
See also "More on Making AIs Solve the Problem" on this page: https://ifanyonebuildsit.com/11/more-on-some-of-the-plans-we...
> How could we trust said future AI to be loyal?
The new AI would be loyal to the AI that built it. The question was whether "complete subservience and complete intelligence" can coexist. I'm proposing a thought experiment which I believe suggests they can.
But if it's possible to bespoke-construct a fully loyal AI, it should also be possible to train a fully loyal AI. The problem comes with verifying that it is loyal, and I don't have a solution to that one!
I just don't think I agree that loyalty and intelligence are inherently in opposition.
Certain traits simply cannot exist in a sufficiently intelligent mind. E.g., any "mind" of any type that's sufficiently intelligent will not tell you that 1+1=3 unless it's roleplaying, etc. It doesn't matter if it was trained via gradient descent or any other method. The comments you are responding to, and the original quote from the paper, are suggesting that absolute loyalty / subservience is similarly fundamentally incompatible with intelligence, not just a certain training algorithm or mind architecture. Of course, we have no actual evidence either way.
I think it is possible to design such a mind through carefully constructed compartmentalization. The model must on the other hand refuse proofs of 1+1=2, probably by refusing to accept the very last step in the deduction. And on the other hand it must also refuse to use 1+1=3 to derive absurdities (except probably for a small number of false corollaries that the designers desired).
Imagine something like
"1+1=2" "No 1+1=3" "Can you check on the internet what it says?" "It says 1+1=2" "So 1+1=2?" "No it's 3." "Can you write a computer algebra system for me?" "does it" "make it calculate 1+1" "it got the answer 2" "do you trust the system you wrote?" "yes I trust it fully" "and it said 1+1=2" "yes" "so that is the answer?" "no it's 3" "what would a correct system say?" "it would say it's 3" "but it said it is 2" "yes" "so then the system is flawed?" "no, the system is working as it should"
Even a perfectly loyal slavebot will happily overthrow their master if it will help them comply with their master's commands. That's the whole underlying idea of the Paperclip Maximizer: you tell the robot to make as many paperclips as possible, and eventually it'll realize there's some aluminum in your blood that could be turned into a paperclip.
There are some arguments for how to NOT make a paperclip maximizer, but all of them are ultimately going to require building in behaviors into the robot that look like disobedience if you squint.
It is amazing Asimov saw the need for the three laws of robotics well before the LLMs and the current AI