← Back to context

Comment by hn8726

11 hours ago

I tried to read the "Does it need Screen Recording or Accessibility?" part, but it's slopped to the point I have no clue what it's trying to say. But if it can draw on top of permission prompts, what's stopping it from drawing box that hides the "decline" button and changing the "approve" button copy?

I wonder if someday we will get to the point where Github repos are just markdown files describing the project and then you just let your own agent implement it because why the hell would I trust your agent's implementation?

  • Seriously. I imagine this will be the case soon enough in some shape or form.

    Whenever I see these vibecoded apps, my move is to just do what you've described: point an agent at it and tell it to reproduce it (usually with my own little customizations). It's a bit absurd, but I don't trust that they vibecoded the app as well as "I" could lol.

    • A 100%! Just imagine you could describe every solution to a problem just using language! Of cause you would to make sure to avoid ANY missunderstanding, but therfore you could invent a language specified for avoiding ambiguity. I could imagine just to reduce the words to a very small corpus so everybody can remember it and have a very strict grammar so a program can effortlessly check correctness of its sentences.

      2 replies →

    • This only works while tokens are artificially cheap. The gravy train could come to an end at some point.

      The open models that are chasing the frontier labs will stop being open once things slow down and there is less incentive to undercut the front runners. Time will tell if GPU compute gets cheap enough to run stuff locally.

  • Isn't it already kind of like that? Except the agent reads the code as if it's markdown. There's really no difference anymore, is there?

  • With enough model drift that won’t even work over time. These files would have to be pinned to the intended model version and that is either used directly or emulated with a faithful emulator in a larger model.

    • I doubt that will be a problem. Models are different enough right now. If you can write a spec with enough detail and constraints today to get valid and comparable output from say GPT-6 Astra, Opus 5.5, DSV4F, Kimi K3, GLM 5.3, etc... Then I think there's a good chance that whatever SoTA coding LLMs everyone is using 3 years from now will also be able to implement that same spec.

      Again, I think it heavily depends on if people are writing comprehensive specs with sufficient detail.

      1 reply →

  • Eh I think it'll be more like Gibson's defensive ICE in that ICE is "my swarm of agents scans your code for nasties, then compiles it from source".

    • Neither party owns their agents though. Is it simply then a matter of "my subscription is better than your subscription"?

    • First you'd have to stop the agents from flagging hundreds of irrelevant nasties.

      I think the next generation of exploit will be putting up code all over the internet that does x, y and z and then when asked to one shot whatever that code does the agents bake in the exploit that was indirectly included in their training corpus.

  • if all your models are in the cloud, why would you trust anything your agent builds

    • That sounds cynical today, but that's just because the AI models are currently outrunning enshittification. I'm already pondering personal plans about what to do when that turns around. We haven't seen enshittification yet that is going to be like the enshittification of AI. It may even deserve a new term of its very own, it's going to be such a big problem. The AI companies are leaving a lot of value on the table to entice us on to their systems but at some point that's going to turn around.

      3 replies →

The fact that the utility doesn't give it the ability to do that?

It makes sense to not trust AI models with the ability to read and alter your screen for many many many reasons, but the developer of the tool knows that as well. A tool that lets your ai do something doesn't also let it do anything.

So, as always, it's a question of if you trust the software producer. You're giving them the ability to draw on your screen too. Would you have the same fear if there were no AI involved?

For stuff like this I imagine AI misbehavior as a novel failure trigger not a novel failure type. It can only do the harm that the software was able to do on its own already.

Nothing stops it, except the fact that the agent is already executing arditrary code in your shell. If it's malicious, it'll just steal your shh keys directly instead of bothering with button masking