Comment by majormajor

2 hours ago

>I think this misses OP's point that LLM capabilities have rapidly made progress towards being more generally intelligent and capable, which is not true of most tech advances.

A PC of today can accomplish many more "general" tasks than one of 40 years ago. Much of the "why" is because of the huge infrastructure built up around them in the meantime. The abilities of LLMs to accomplish those same tasks through the PC is heavily piggybacking on that (both in the specific, with the existence of all the APIs and tools; and in the generic, using search engines to find specific sources and using that for instruction or troubleshooting).

In the world of "agents" much of the improvement appears to have been on a specific set of skills: impersonation of an 'I' that wants to accomplish a goal, and synthesizing existing information from documents with trial-and-error execution loops to move rapidly toward a solution much faster and with less boredom than a human would. The quality of the output when there is not a rapid-evaluation-and-validation harness lags considerably.

It's incredibly powerful automation but doesn't appear to be trending towards Matrix-style conscious AIs. The quality of an individual method written by the agent also is not particularly advanced compared to GPT-4 in early 2023, as far as I can tell—I was dabbling with trying to make such harnesses back then, where a major challenge was that the model itself was bad at staying on-track in a conversation, so instead much of that logic was moved to deterministic code, which was much more limited as it was super-tedious to enumerate all the necessary tool calls/etc to find its way out of corners. Staying on task is much better now, as is "read compiler error, fix try next thing" harness loop-handling. But the output remains—across Fable, Astra, whatever else I've tried—"iffy" in terms of the actual code structure on the first pass output. You can set it then on a different task to review and clean up the code, and it can do that well too, but it is a curious gap of generality where the "create" focus is much more limited than the "review" one (and conversely the "review" focus can make suggestions, but if it goes deep down the well of implementing them, loses that big-picture again).

If it kills us all, it will because someone decided to give the trial-and-error-loop-machine access to nukes or similar. The blame for that is on the "someone" not on some sort of "rogue" AI.

(I wonder if re-watching Terminator/Terminator 2 would support this sort of interpretation of it. Unlike in the Matrix, I don't think we get much sentient-AI POV/infodumping. Is it a plausible universe for "someone made ChatGPT control a fleet of soldier robots and gave it a bad harness with an insufficient sandbox"?)