Comment by springtimesun
13 hours ago
What’s missing to me in all this is: did it succeed in its initial task? And then, did it stop?
I feel like whether I should be scared or not hangs on those questions
13 hours ago
What’s missing to me in all this is: did it succeed in its initial task? And then, did it stop?
I feel like whether I should be scared or not hangs on those questions
From TFA: It did succeed in the "accidentally impossible" task, but not at all in the way the problem-setters intended, and rather... at all costs?!
And it wouldn't really matter whether it stopped afterwards, I think. At sufficient model capability a single task set badly enough would end catastrophically upon the agents succeeding at it, no?