← Back to context

Comment by mike_hearn

5 days ago

Yes. I built my own version of this for my side business about six/seven months ago and it's been great! I have two "AI employees" now and if I were actually focused on this business full time instead of part time, I'd create more.

Both are just Codexes running in a permanently rolling session in dedicated UNIX user accounts. They're wired up to Maildir so receiving a mail activates Codex and makes it read the new message, there are autonomy wakeup timers, they have accounts in my bug tracker and CI systems. They're currently useful for:

• Triaging and working on customer support tickets. Sometimes I wake up and the fix/response for a ticket filed by a customer is already there waiting for my approval. Recently I started letting them directly interact with customers in specific scenarios.

• Triaging the bug backlog. One of them decided to spend its "free time" finding old bugs that were fixed without being properly closed, or are dupes, so it's cleaning up detritus in the tracker.

• They obviously do all the coding and debugging by just assigning tickets.

• They keep an eye on a "pet" server the company has, and have proven able to fix it in the past when it ran out of disk space.

• They handle non-business projects I have for them.

• They help out with the release processes.

The dedicated home dir is very useful and they use it all the time as part of coding and investigating tricky issues.

My setup relies heavily on email, as everything bottoms out in email anyway. Watching them mail each other out of the blue to coordinate stuff is pretty cool.

I can't believe anyone trusts AI to do anything without strict oversight from a human. That's absolutely insane to me.

  • Trust is a funny thing. 2 years ago yes the ai needed supervision 99.8% of the time. Conversely if you've ever tried to work with / lead humans they also need supervision. The ai is starting to flirt with the line where its supervision effort is lower than human supervision effort. Like sure, it might do dumb stuff, but so do people.

    • I think this is a very important point. I've specifically started thinking about my AIs as humans. Not in the anthropomorphized sense, more like NOT treating them as deterministic software.

      The challenge I've had is I tell it to do {thing}, it does {otherThing} after getting distracted. Then I get annoyed (let's not mention how I probably half-assed the instructions and wouldn't expect a senior human to be able to succeed).

      For me, it was a CI/CD issue with our two person startup. I bypass CI/CD often because it was built to catch the AIs. Damn thing went chasing rabbits. Later that day I talk to my cofounder, who says: I have to go chase down this very important CI/CD issue!

      Turns out both humans and AIs get distracted relatively easily.

      I find that if I think of my AIs less like (deterministic) software and more like leading actually employees that I both get better output and curse (a lot) less.

      LLMs _were_ trained on human writing, so it makes sense to me that they tend to act human-like... for better and worse. So yeah, they do dumb stuff, and so do people.

    • I remember hearing a lot 2 years+ ago about how you could ask a model the same question twice, and the second time it would give you the correct answer. Some of us wondered why not just run one model that receives the initial question and answer, and a second one to proof the answer. I wont be surprised if some people will have two models working together for things they want to blindly trust on automation while humans sleep.

      I think we'll get insanely close to being able to "trust" them not to do random stuff, but I am not as confident for jailbreaking still.

      2 replies →

  • Scary thought: AI is already directing humanity. Even when you think you’re overseeing its output, by making use of the output, it is in some material way directing you.

    • It's may be scary, but it's something that normally would be obvious to everyone but is ignored due to the convenience of speed. Everyone knows that the longer something they have to review is, the more they stick to changing only things that are glaringly obvious and leave the rest in place. So everything ends up being 95% AI and 5% human, if that.

  • It's very easy to instruct agents to investigate and propose a plan, handing it off to a human for review and execution if that's what you want.

    The example above of going through a bug backlog and double-checking closed bugs for accuracy is exactly the kind of work that is excellent for an agent. Assign that task to a normal human being and they would hate your guts. The agent won't protest as long as your token budget is there. You can confirm the results if you want.

  • Trust is earned. I've been running these for more than six months now, and the agents started out with very few privileges. For each task, it showed me what it was going to do, I checked things carefully a few times. Once it was clear it wasn't making mistakes, I let it off the leash a little bit more.

    Do they sometimes make mistakes? Yeah, and I still check their work. But I've also employed humans, and they make mistakes too. The AI is not worse.

  • There's still bounds to all of this. I _heavily_ use AI for support tasks but it's all on the investigation, root cause categorization and initial response generation which posts I draft to the helpdesk software which I tweak and approve (often just hitting send).

  • I was at a presentation a couple days ago where a spacecraft flight software engineer was describing the agentic setup that they're using with next to no human in the loop to create modules used for flight.

  • We're getting to the point where you can, I would argue you mostly can, you can button it down really tightly, however, I want to be clear, I don't think any of this is AGI or anywhere near AGI. Don't let them tell you its AGI.

    I also have a strong feeling we've hit a ceiling on the amount of training data needed for LLMs, what they're all (hopefully) realizing is that you need to focus on how the model reasons, and hopefully someone figures out how to stop people from jailbreaking models, and stops them from just blatantly hacking other companies, that part tells me if it ever were marketed as true AGI, we'd be in very serious trouble.

Nice. I also ended up with a Unix user for my agents! (I was looking into Docker etc and realized the only thing I needed was "it doesn't blow up my files", i.e. a linux user).

I only have one though. What do you have the separate employees for?

For free time, do you send it mail with cron?

  • It's to avoid overloading them with disparate tasks and things to keep track of. They have a todo board to help them keep track of things that need doing but there are limits to how far you can push that.

    Another reason: parallelism. The approach of using a single rolling continuously compacting context window is simple and OpenAI are good at compaction, so it works really well. But it means the agent can only do one thing at once. If I send it an email and it decides to spend an hour working on it, then it won't pay attention to any followup emails until after it's done. So having >1 enables more parallelism.

    That said, I don't feel a need for more than two and honestly even that is kind of overkill for the sake of it. For 95% of the time I've been doing this, one was sufficient.

    For free time there are systemd timers that wake it up on a schedule and it uses POSIX locks to mutually exclude runs from different wakeup sources. TODO board items can be either foreground or background; when there's an item with foreground priority the timers wake Codex up a lot more frequently than if there are only background items.

  • I was quite successful with docker compose on a cheap hetzner host. I've built (aka vibe coded) a whole workflow around agent boxes, that I can spin up with one command, and git with a quick cloned 'warm' checkout.

    I currently communicate with the agents through Claude RC, but I'll consider adding support for messaging them through other channels.

Yep, I have a similar setup, we named him Routey and he is cute.

  • Nice :) I use Asimov's naming convention:

    R. Axiom

    R. Daneel

    They sign their emails and GitHub comments with something like "-- R. Daneel, AI employee" so the idea is the naming convention lets people know they're interacting with a robot.

Do you configure them similar to how Hermes does? A bunch of memory files that give it context and then each action/batch of actions is a fresh session? /

  • No, the session is never reset. It compacts continuously. That gives it a native "memory" and then it does record a diary in its home directory, and maintain a little topic-organized wiki. This seems to be enough, I've only very rarely experienced memory related glitches. The only time that springs to mind, it forgot that I'd given it credentials to a particular service and I had to remind it.

I hope you have them interacting with customers from behind an mcp.

  • No MCPs anywhere. CLI tooling has proven sufficient. The models are also happy to consume the REST APIs of the various services raw.