Comment by giancarlostoro

4 days ago

I remember hearing a lot 2 years+ ago about how you could ask a model the same question twice, and the second time it would give you the correct answer. Some of us wondered why not just run one model that receives the initial question and answer, and a second one to proof the answer. I wont be surprised if some people will have two models working together for things they want to blindly trust on automation while humans sleep.

I think we'll get insanely close to being able to "trust" them not to do random stuff, but I am not as confident for jailbreaking still.

The thing I like to do is to use models from different training sets - so for frontier, OpenAI criticizes Anthropic and vice versa. They're very much peanut butter and chocolate in that regard - I honestly can't be bothered to set up the whole MMLQUALA benchmark suites or anything, but I wonder how high "the two best models running at max thinking working together" would score compared to either individually.

Jev, or 'Jevlikes', will go a long way towards trustable systems. There are demos of running every prompt through the first pass filter of Jev "Is this unsafe? y/N"

Seems to be that Jev is a "reflex" system for AI, where current LLMs are higher level thinking. Computers can now flinch!