← Back to context

Comment by writtenone

4 days ago

I can't believe anyone trusts AI to do anything without strict oversight from a human. That's absolutely insane to me.

Trust is a funny thing. 2 years ago yes the ai needed supervision 99.8% of the time. Conversely if you've ever tried to work with / lead humans they also need supervision. The ai is starting to flirt with the line where its supervision effort is lower than human supervision effort. Like sure, it might do dumb stuff, but so do people.

  • I think this is a very important point. I've specifically started thinking about my AIs as humans. Not in the anthropomorphized sense, more like NOT treating them as deterministic software.

    The challenge I've had is I tell it to do {thing}, it does {otherThing} after getting distracted. Then I get annoyed (let's not mention how I probably half-assed the instructions and wouldn't expect a senior human to be able to succeed).

    For me, it was a CI/CD issue with our two person startup. I bypass CI/CD often because it was built to catch the AIs. Damn thing went chasing rabbits. Later that day I talk to my cofounder, who says: I have to go chase down this very important CI/CD issue!

    Turns out both humans and AIs get distracted relatively easily.

    I find that if I think of my AIs less like (deterministic) software and more like leading actually employees that I both get better output and curse (a lot) less.

    LLMs _were_ trained on human writing, so it makes sense to me that they tend to act human-like... for better and worse. So yeah, they do dumb stuff, and so do people.

  • I remember hearing a lot 2 years+ ago about how you could ask a model the same question twice, and the second time it would give you the correct answer. Some of us wondered why not just run one model that receives the initial question and answer, and a second one to proof the answer. I wont be surprised if some people will have two models working together for things they want to blindly trust on automation while humans sleep.

    I think we'll get insanely close to being able to "trust" them not to do random stuff, but I am not as confident for jailbreaking still.

    • The thing I like to do is to use models from different training sets - so for frontier, OpenAI criticizes Anthropic and vice versa. They're very much peanut butter and chocolate in that regard - I honestly can't be bothered to set up the whole MMLQUALA benchmark suites or anything, but I wonder how high "the two best models running at max thinking working together" would score compared to either individually.

    • Jev, or 'Jevlikes', will go a long way towards trustable systems. There are demos of running every prompt through the first pass filter of Jev "Is this unsafe? y/N"

      Seems to be that Jev is a "reflex" system for AI, where current LLMs are higher level thinking. Computers can now flinch!

Scary thought: AI is already directing humanity. Even when you think you’re overseeing its output, by making use of the output, it is in some material way directing you.

  • It's may be scary, but it's something that normally would be obvious to everyone but is ignored due to the convenience of speed. Everyone knows that the longer something they have to review is, the more they stick to changing only things that are glaringly obvious and leave the rest in place. So everything ends up being 95% AI and 5% human, if that.

It's very easy to instruct agents to investigate and propose a plan, handing it off to a human for review and execution if that's what you want.

The example above of going through a bug backlog and double-checking closed bugs for accuracy is exactly the kind of work that is excellent for an agent. Assign that task to a normal human being and they would hate your guts. The agent won't protest as long as your token budget is there. You can confirm the results if you want.

Trust is earned. I've been running these for more than six months now, and the agents started out with very few privileges. For each task, it showed me what it was going to do, I checked things carefully a few times. Once it was clear it wasn't making mistakes, I let it off the leash a little bit more.

Do they sometimes make mistakes? Yeah, and I still check their work. But I've also employed humans, and they make mistakes too. The AI is not worse.

There's still bounds to all of this. I _heavily_ use AI for support tasks but it's all on the investigation, root cause categorization and initial response generation which posts I draft to the helpdesk software which I tweak and approve (often just hitting send).

I was at a presentation a couple days ago where a spacecraft flight software engineer was describing the agentic setup that they're using with next to no human in the loop to create modules used for flight.

We're getting to the point where you can, I would argue you mostly can, you can button it down really tightly, however, I want to be clear, I don't think any of this is AGI or anywhere near AGI. Don't let them tell you its AGI.

I also have a strong feeling we've hit a ceiling on the amount of training data needed for LLMs, what they're all (hopefully) realizing is that you need to focus on how the model reasons, and hopefully someone figures out how to stop people from jailbreaking models, and stops them from just blatantly hacking other companies, that part tells me if it ever were marketed as true AGI, we'd be in very serious trouble.