Comment by jswelker

1 day ago

LeCun is saying AI will be totally safe as long as we have competent and well aligned corporate management.

What could go wrong?

He has a point, though. The LLMs (or any AI for that matter) can't do anything. They can't. It's a function call that ingests symbols and spits out symbols and that's it.

100% of its actual capabilities are tied to harnesses (the actual "agent"), i.e. ordinary deterministic programs that are connected to networks or machines and enable interaction with the outside world. This part (the part that can do harmful things) is fully under human control and all the recent headlines about "agents going rogue" are - as someone (forgot who) put it - akin to strapping a weedwhacker onto a dog and letting it run wild.

The tech itself is safe as far as real-world interactions go - the weakness lies in unchecked access to systems surrounding it. It's not safe at all when it comes to human interaction (lots of ongoing lawsuits demonstrate that), though. There is real danger here, but it has nothing to do with doomsday scenarios ala Terminator or I,Robot and more with total corporate control over the lives, perception of reality, and abilities (like critical thinking) of people.

  • This is like saying cars don't kill people, because if nobody drives them faster than 3mph there's no problem. The _whole_ promise of cars is that they can go fast, just like the whole promise of AI is offloading thinking to a computer. If AIs are unsafe without close human supervision and checking every interaction with the real world, they are unsafe full stop.

    • The analogy would be more fitting if you had said "cars don't kill people if every drivers has proven skills, never drives impaired, keeps the speed in line with weather and road conditions and stays on actual roadways". You know, like everyone should, regardless of whether they're driving a high performance sports car.

      To keep with the analogy: cars have seatbelts, airbags, ABS, ESP, lights, horns, crumple zones, emergency braking systems, roads have speed limits, there are traffic stops, insurance, regular inspections (not in all countries), etc. etc.

      So what's unsafe here? The car or roads without speed limits, complete lack of safety measures (both active and passive), absence of any supervision and no insurance? That's the problem. It's not the models themselves - they can spit out tokens by the billions, there's no risk there.

      You wouldn't give full access to your phone, your computers, your house keys and your credit cards to any stranger on the street now, would you? How is it then, that people act all surprised when a non-deterministic machine that's optimised to achieve goals while taking all the shortcuts it can, suddenly uses the tools handed to it in unexpected ways? That's a failure on the operator's side, not an inherent danger within of the model.

      6 replies →

  • > It's a function call that ingests symbols and spits out symbols and that's it. 100% of its actual capabilities are tied to harnesses

    This is a bad and misleading way to think about it. Note that it's trivial to make the harness that you claim capabilities are tied to (the LLM itself could write it from scratch in one shot), but no matter how good a harness you have, it won't make gemma4:e4b capable. That's because what actually gives capabilities is the LLM's intelligence - or if you prefer not using that term, the fact that the probability distributions the LLM spits out depend on the context in useful ways.

    • I think we're talking about fundamentally different perspectives here.

      I'm not talking about what the LLM does internally. If a metaphor helps, here's one to help you understand what I was trying to get at:

      Imagine an evil genius that has no eyes and no limbs. Everything they could learn about the world is presented to them by means of some person describing it to them through words. They have no way of directly interacting with the world and rely on someone executing any action they want to take and describe the outcome to them. Now how dangerous would you say such person would be? How dangerous could they become?

      That's what I was getting at. Replace person with LLM (or any other AI system). Replace the person that communicates with an external interface (the harness) and I hope you understand. It doesn't matter whether the LLM could generate the harness by itself - it still is just a bunch of weights sitting in memory being run by an execution engine. That's what it fundamentally is, whether you like it or not. It cannot do anything on its own - and no, not even writing files. It's the execution engine that translates the numeric output into words (or images or video or audio) and the layer above (the harness) that takes that output and interprets it to execute actual actions.

      This is not about what you or I think about the internal capabilities of the model - that's irrelevant to the conversation and you can replace LLM with a random token generator and the point still stands. The model itself is incapable of performing actions - from reading files to writing files, to controlling physical machines. All that is and HAS to be done by external interfaces outside the control of the model.

      3 replies →

  • But the harnesses exist, and will always exist. They will continue to get more access than is safe because it is convenient and profitable. Your argument is based on a distinction without a difference.

  • I would say it's more like an interface for the model to interact with the world. If you give the model access to filesystem and bash that technically unlocks all computer use, so how are you going to control that? By trying to regex match against the commands the AI uses? All you have is auth or containment, and AI can hack auth and people will not stop connecting AIs to the internet. It's a ridiculous premise that just because the harness is "normal code" that means we can control the AI.

    The world's institutions, systems, and industries are all rapidly digitizing. So while I'd concede the point that, yeah, there's no way a rogue AI can just take over some powerplant and blow it up because of analogue systems the AI can't access, that isn't necessarily true for some powerplants already, and more and more powerplants will be connected to networks and controlled by software systems in the future. The more we digitize our systems the more potential for AI to exploit vulnerabilities and affect the real world.

    AFAIK there isn't that much stopping anyone from spawning an AI swarm and telling it to "spread and go hack everything for the lulz."

    • > If you give the model access to filesystem and bash that technically unlocks all computer use, so how are you going to control that?

      The same way we've done it since machines became multi-user: boring old system access restrictions. Nothing fancy, nothing radical, just good old minimal access rights required to perform a defined set of whitelisted operations.

      > It's a ridiculous premise that just because the harness is "normal code" that means we can control the AI.

      What is it then? Is not just a program that takes model output, parses it and performs tool calls from the text it receives and then feeds the result back into the model and calls it again with those results? It is normal boring old deterministic code. Many are open source. Look at them. Understand what they do and the apparent "magic" goes away real quick. Harnesses are nothing special.

      > AFAIK there isn't that much stopping anyone from spawning an AI swarm and telling it to "spread and go hack everything for the lulz."

      Aside from lower cost and possibly greater scale, there's literally NO difference between that and (state sponsored) hacking that has been going on for decades. First it was script kiddies, now it's ML models. The threat model remains the same and so do the counter measures. The real danger is still the harness (and its access to external systems), not the model itself. Restrict the access of the harness and the model can't do anything harmful, see above.

      3 replies →

We are living in the same world with nukes, where you say as long as we have competent government. Everyone is actually doomed without AI solving diseases or old age.

While in the same breath calling the leaders of the largest AI companies “deluded” and “insane.”