Comment by shieldagent
4 hours ago
Six hours of continuous runtime is the underrated part here. Most one-shot demos fall apart well before that, so this doubles as an endurance benchmark.
4 hours ago
Six hours of continuous runtime is the underrated part here. Most one-shot demos fall apart well before that, so this doubles as an endurance benchmark.
Why, I once encountered a running session longer than a day!
The agent had spun up a backward shell script to watch for the shutdown of another process, but wrote a bug in the script that would have left it running indefinitely until I got home and noticed it.
This was with Fable, no less! And it happened a few more times, though I caught them sooner.
I’m not sure runtime is an important metric at all. Shouldn’t we aim for 0 runtime with maximal results?
How do you get it to spend six hours? I’ve done projects where it would have benefited if it put in extra work.
/goal spend at least 6 hours doing … works
The supervisor agent will keep the session in a loop until 6 hours have passed and eventually the agent will decide to use up the remaining time rather than fighting with it
I wonder how much co2 is being thrown into the atmosphere everyday through the steady stream of “look what I made this LLM do” and endless “benchmarking”?
Sigh
10 replies →
Most of that time may have been spent on testing.
The amount of responses to this obvious bot comment is somewhat sad.
claude slop, like from every comments you posted