Comment by Anoian
16 hours ago
I have told llms all kinds of stories to find out what its answers would be. I always make it sound like it is the truth to make sure the AI answers in a way that it would if somebody actually said this. I also tested internal flagging systems of the ai company I work at with the most evil things a person can ever say to find out if it would flag them.
Of course I did not mean any of that stuff, but how can you make sure a human reviewer knows you did not mean it while the llm does not know that you did not mean it.
I guess its a miracle I am not in jail yet.
Flagging people for anything said to an llm sounds wrong to me because an LLM is not a real person and while some people put in their internal thoughts, others just roleplay and the two are inseparable just from reading it.
Don’t worry, they’ll store those messages forever and incarcerate you at their convenience.
I was sure I couldn't be the only one curious to push LLMs to their limits. Though these days it's much tougher, mostly impossible to get them to react in unforeseen ways to horrendous scenarios.
This is exactly the typical use I make of the llm.
Adding: - I typically ask questions in the I form, regardless for whom or why I ask for. - Gemini chats quite often end when it starts recommending psychological council or a suicide line, to talk about my problems. It apparently detects a persistent tendency to not agree with the party line. So it makes sense I must be suicidal ;-
But sure, as llm's start to babysit us, and know our inner dialog better than anyone else, we'll soon be debugging their opinion/behavior/co-existence/authority, when it comes to reporting people to the authorities, or taking on tasks in society in general. We'll hire doctors to cure our psychological profile from our record (Total Recall).
A Minority Report like this shouldn't cause a referral to the police.