Comment by stingraycharles
5 hours ago
“Their objective is to solve the problem and they'll use anything they can to solve it.”
My point is: is this really what people want? It seems like they’re optimizing for one-shotting solutions, where most of the time in an actual workflow it’s much more productive for the model to make sure it got the question right if things get difficult.
Like, “hey, do you REALLY want me to use this local privilege escalation bug so I can download your Google Drive file?” is the bare minimum I would expect.
Yes, and to bring in another tired metaphor people make about AI agents, this is what you want an intern to do when they get stuck. Don't just churn indefinitely without an idea what the right direction is. Certainly don't go hack other companies to steal an answer. The model's lack of any sense of legal or ethical boundaries is where it's far, far stupider than the intern, and far, far more reckless for a company to wield the way OpenAI did here.
yes, this exactly
but, there is a fatigue that sets in and i've experienced it myself.
- is it ok to run script xyz?
- allow permission to edit abc?
- allow to request blablabla?
over and over.... click click click
something will get in there that is dangerious and then its whopsie our keys are now on github