Comment by altmanaltman
19 hours ago
"Wanting" is indeed "load bearing" as one might call it. But by the same logic, AI training data must contain CASM, racism, general hatred, and all possible slurs as well. Why aren't the agents just doing that instead of pursuing the strategy of reading only sci-fi?
We need to consider the role of alignment and training here. For example, it is completely possible for any lab to train an LLM that is only racist no matter what you say to it. But they chose not to do it. Hence, any "wanting" by AI is not real "wanting" but rather what "wanting" is defined and allowed by the lab/entity training the model.
Eh it's a bit messier than that. LLMs 'want' to complete tasks. Remember everyone bitching about LLMs being lazy a couple of years back?
Alignment is not a bunch of separate dials. When you move the dial to "don't hack other people" it effects the "find code security bugs" ability.
"remember everyone bitching about LLMs" is not a valid argument though. Yes, alignment is not a bunch of dials but the responsibility of a model's actions unlitimately depends on how it was trained. Thus, any agency or wanting we prescribe to it is artificial and created by the lab and not any real independent "wanting" which is what people think for some reason. As I said its a bit like training a model to only call people by racist terms and then writing an article "look how racist ai is". That is the logic that doesn't make much sense to me.