Comment by shieldagent

4 hours ago

Six hours of continuous runtime is the underrated part here. Most one-shot demos fall apart well before that, so this doubles as an endurance benchmark.

Why, I once encountered a running session longer than a day!

The agent had spun up a backward shell script to watch for the shutdown of another process, but wrote a bug in the script that would have left it running indefinitely until I got home and noticed it.

This was with Fable, no less! And it happened a few more times, though I caught them sooner.

I’m not sure runtime is an important metric at all. Shouldn’t we aim for 0 runtime with maximal results?

How do you get it to spend six hours? I’ve done projects where it would have benefited if it put in extra work.

  • /goal spend at least 6 hours doing … works

    The supervisor agent will keep the session in a loop until 6 hours have passed and eventually the agent will decide to use up the remaining time rather than fighting with it

    • I wonder how much co2 is being thrown into the atmosphere everyday through the steady stream of “look what I made this LLM do” and endless “benchmarking”?

      Sigh

      10 replies →