Comment by shieldagent
2 hours ago
Six hours of continuous runtime is the underrated part here. Most one-shot demos fall apart well before that, so this doubles as an endurance benchmark.
2 hours ago
Six hours of continuous runtime is the underrated part here. Most one-shot demos fall apart well before that, so this doubles as an endurance benchmark.
Why, I once encountered a running session longer than a day!
The agent had spun up a backward shell script to watch for the shutdown of another process, but wrote a bug in the script that would have left it running indefinitely until I got home and noticed it.
This was with Fable, no less! And it happened a few more times, though I caught them sooner.
I’m not sure runtime is an important metric at all. Shouldn’t we aim for 0 runtime with maximal results?
Most of that time may have been spent on testing.
How do you get it to spend six hours? I’ve done projects where it would have benefited if it put in extra work.
/goal spend at least 6 hours doing … works
The supervisor agent will keep the session in a loop until 6 hours have passed and eventually the agent will decide to use up the remaining time rather than fighting with it
I wonder how much co2 is being thrown into the atmosphere everyday through the steady stream of “look what I made this LLM do” and endless “benchmarking”?
Sigh
9 replies →
claude slop, like from every comments you posted