← Back to context

Comment by Yiin

1 hour ago

how do you imagine that would work? you being able to. influence globally stored weights with some prompts? we have fine-tuning for that.

Yet I can't randomly order another person to steal a car for me, just because I tell them to. Alignment for an intelligent system is a hard problem and at this stage is seems close to unsolvable.

My guess is that we'll just ignore it and make money along the way and every 2-3 months we'll have the equivalent to "Equifax gets hacked and millions of user records are stolen", etc. (this time with the LLM itself doing the hacking at someone's behest - accidental or not).