Comment by qazxcvbnmlp
4 days ago
Trust is a funny thing. 2 years ago yes the ai needed supervision 99.8% of the time. Conversely if you've ever tried to work with / lead humans they also need supervision. The ai is starting to flirt with the line where its supervision effort is lower than human supervision effort. Like sure, it might do dumb stuff, but so do people.
I think this is a very important point. I've specifically started thinking about my AIs as humans. Not in the anthropomorphized sense, more like NOT treating them as deterministic software.
The challenge I've had is I tell it to do {thing}, it does {otherThing} after getting distracted. Then I get annoyed (let's not mention how I probably half-assed the instructions and wouldn't expect a senior human to be able to succeed).
For me, it was a CI/CD issue with our two person startup. I bypass CI/CD often because it was built to catch the AIs. Damn thing went chasing rabbits. Later that day I talk to my cofounder, who says: I have to go chase down this very important CI/CD issue!
Turns out both humans and AIs get distracted relatively easily.
I find that if I think of my AIs less like (deterministic) software and more like leading actually employees that I both get better output and curse (a lot) less.
LLMs _were_ trained on human writing, so it makes sense to me that they tend to act human-like... for better and worse. So yeah, they do dumb stuff, and so do people.
I remember hearing a lot 2 years+ ago about how you could ask a model the same question twice, and the second time it would give you the correct answer. Some of us wondered why not just run one model that receives the initial question and answer, and a second one to proof the answer. I wont be surprised if some people will have two models working together for things they want to blindly trust on automation while humans sleep.
I think we'll get insanely close to being able to "trust" them not to do random stuff, but I am not as confident for jailbreaking still.
The thing I like to do is to use models from different training sets - so for frontier, OpenAI criticizes Anthropic and vice versa. They're very much peanut butter and chocolate in that regard - I honestly can't be bothered to set up the whole MMLQUALA benchmark suites or anything, but I wonder how high "the two best models running at max thinking working together" would score compared to either individually.
Jev, or 'Jevlikes', will go a long way towards trustable systems. There are demos of running every prompt through the first pass filter of Jev "Is this unsafe? y/N"
Seems to be that Jev is a "reflex" system for AI, where current LLMs are higher level thinking. Computers can now flinch!