Pop!_OS bans AI-generated code from much of its codebase

5 hours ago (neowin.net)

I also have found that AI has not lived up to many of its promises and have dialed back what AI gets control of. My projects were turning into unmaintainable messes. The people who say coding is solved aren't paying attention.

  • Yes. I started two projects with AI from scratch. Both abandoned, complete mess. Projects without AI are so easy to manage, maintain, add/remove features, etc. I use AI as a search engine on my projects instead of Google. I ask what's wrong with my code and change or improve it myself based on my experience.

    • I have the opposite experience. I don't have the patience sometimes to clean up my code and stick to coherent conventions and organization even though the will is there. With my AI projects I watch it and if the AI starts drifting I ask it to go through and look for convention/directory structure violations and it happily cleans everything up in about 10 or so minutes.

      Are you in Python by chance? Python has a lot of crazy hidden/inexplicit/spooky action at a distance stuff (especially in the frameworks) that can make LLMs gunk up code by defensively programming or just burn context chasing data provenance

    • I've been a vibe-coding skeptic for years, but because of the math breakthroughs of the past few weeks I decided to experiment with the latest models on some test projects. They're a lot more capable than I thought they would be. I agree that it's easy to create an unrecoverable mess, especially when you're one-shotting a lot of features without detailed instructions. But I find that as long as I'm strict about the API boundaries and force the agent to work in small chunks, it's pretty effective. As one example, I got it to write an SVG renderer in a few hours (not the whole spec, but most of the path features and text rendering), which would have taken me at least a week just for the coding part, plus extra time to learn the algorithms.

      3 replies →

    • I'm reading y'all comments and it seems we're still finding our footing, and will for some time. I have the luck that I have access to pretty much all frontier and other models alike. I've been extremely negatively biased towards any LLM use in the start. Then the influx from juniors came, then the revolt of doing PRs on such slop came, then some structured methods how to do LLM work came, then vibe projects came, etc, etc. I literally have all the described experiences you've guys mentioned. From good, to bad, to ugly. It's like there's no one particular way about doing this, and no two projects share the same approach - just like ye olde times.

  • People, especially those who don't do software development, often conflate coding and software development.

    The models aren't good at architecture and design. But they take direction on architecture and design and design well and can refactor code quite effectively. AI agents can absolutely be used to clean up vibe coded code bases once you figure out if the investment is worth it. The mess can be avoided if you give them sufficient guidance on architecture and design upfront.

    That said, doing so purely in text form doesn't feel great right now. I've been thinking about UML lately. The problem with that was the roundtrip after the code was generated and then the implenetation happened. I don't necessarily think UML is the solution, but neither is walls of dense text.

    • >The models aren't good at architecture and design.

      Can you elaborate to back up this claim? WHat exactly is your yardstick for "being good at SW design and architecture"?

      Because I found the current SOTA AI models being great at architecture and design, much better in fact than most average real-world devs. Is your yardstick just the John Carmacks of the world by any chance? Because most devs are not John Carmack. They are also not Linus Torvalds, they are not Stallmann, etc.

      Maybe your LLM experience is still stuck in the 2023 era of ChatGPT?

      2 replies →

  • I've had the opposite experience. Spec driven development really makes things much easier to maintain. If you're just yolo'ing it and throwing prompts around things will fall apart fast.

  • Were you reading through the code it was generating to make sure the flow was intuitive and comprehensible for each PR? Projects only turn in to maintainable messes if you blindly merge in unmaintainable messy code.

  • As usual, when people post comments like this they never specify what model and when they used it. There is a huge difference between GPT 3 and 6 for example.

    Late 2025 also had a step change when agents could largely code autonomously without handholding like previously, and to be honest it's not worth hearing opinions about AI from before that time, that's how significant the change was.

  • Oh my god. I am sad to have to acknowledge this even if the models have “theoretically” gotten better. I am not a programmer but I am super opinionated with the design and structure of the code and need to make sure whatever I write/ask to generate and use for my own tasks is understood by me at least once. I wasn’t this way in 2023. I am constantly reading/trying for an agent to help me write simple code without spending too much time. But so far nothing has worked. I wish someone figures it out otherwise, I personally would not be able to realize the AI agent productivity benefits that others are supposedly seeing.

    • If you specify a structure, pattern, or design, it will stick to that design after a code review phase. Just tell it what standards you have, and after two rounds, it will have written what you described.

  • I treat AI like an intern or junior team member. With enough guidance, they can contribute a lot but you can't let them loose without supervision or they usually will produce a big mess. As of now, you are still responsible for overall architecture. One strong indicator that something is going wrong are big pull requests where the AI has rewritten large sections of the code.

  • AI is a useful tool but it absolutely needs a lot of human guidance.

    Unlike a compiler it won't give up at the first sign of trouble but that just means it left alone it will dig bigger and bigger holes.

    Treat prompt engineering as a discipline and refine your technique. When it produces garbage throw out the work and start over until you figure it out.

  • Vibe coding has always had and always will have one fundamental flaw: you didn't write the code it produced, therefore you don't have a good understanding/mental model of it.

    On the other hand asking these clankers "review the feature branch I wrote" and "review my entire codebase for bugs" or "help me debug this" has saved me months of prospective work.

    And more recently most major models have been getting _really_ good at RE, for example you can have OAI models (and maybe A/'s if they don't refuse) use idalib MCP and reverse-engineer stuff from start to finish, then follow up with GLM 5.3 for vuln assessment and exploit PoC.

    Stuff that used to take weeks or months now just takes a few hours, or less.

  • Were you reviewing the changes? I've seen lots of projects with this problem. But if AI PRs have to pass the same review bar as any other PR then it shouldn't be an issue in theory (if you can actually maintain the discipline).

    • Isn't it tiring to keep up with considering the speed of the generated code? And if you want to be careful with the review you lose a significant part of the speed advantage.

      3 replies →

  • I am not saying you are right or wrong, but I don’t think you can make such a broad conclusion simply because of your own experience.

    You say you used AI and your projects turned into unmaintainable messes, so your conclusion is that it means AI is not living up to its promises.

    I guess if the argument is “AI makes it so you always get a great result no matter how you use it”, then your argument is sound. Your projects not working out proves that AI doesn’t always work no matter what.

    However, that doesn’t mean you can’t use AI to create sustainable and well organized code. Failing to do something doesn’t mean it’s impossible and anyone who thinks they can is not paying attention.

    I can’t run a marathon. If I went out and tried to run one, I would get a few miles and collapse, failing completely.

    I don’t think it would be reasonable, though, at that point to say “running a marathon is impossible, anyone who says they can do it clearly lying. I tried and didn’t even make it 5 miles!”

    I wish people would stop assuming their experience with something is the only possible truth.

    • Stop being overly polite to these people. It is seriously coming down to the point of utter delusion.

      As ai becomes better these people will begin changing their story because it’s utterly obvious what’s happening.

  • The people saying programming is solved couldn't program in the first place is what I've noticed. So to them it really does feel solved, things "work" and they don't have to learn what they don't know about programming.

    • I've been programming professionally for decades. LLMs are extremely useful. At this point if you haven't figured out how to get value out of them, you're either holding them very wrong or you're being willfully ignorant.

      3 replies →

    • There are a couple of notable counterexamples here (nobody sane thinks Carmack doesn't know how to program, for example), but by and large I agree with your observation. The people excited about programming with LLMs are, on average, people who weren't good at programming to begin with. Still, given that these counterexamples do exist I try to avoid painting with an overly broad brush for the sake of nuance.

Completely performative.

You build your OS atop thousands of open source packages, many of which contain AI generated code. Are you going to audit them one by one and remove offending packages? What about the ones you won't remove because the OS would be irreparably broken?

  • It's not. Read the article, they have a good reason for banning it and it's not because they hate LLMs.

  • It also doesn't answer the question of how they might even recognize LLM generated code in contributions to PopOS directly.

    I have yet to understand how maintainers can't distinguish beyond (1) PRs that literally include Claude co-author notes or (2) low quality code contribution regardless of the creator.

  • I agree.

    PopOS is Ubuntu with extra problems. Ubuntu itself is fine, but then PopOS adds weirdness.

    Cosmic has been in beta for how long ?

    • > Cosmic has been in beta for how long ?

      I daily drove the alphas before the betas, and of course there were a couple rough edges. But I had a minimal working desktop instead of sway or KDE/Gnome (too heavy).

      It has been releasing non-betas for a good year now, and it's been a smooth sailing.

  • I don't think it's performative.

    They aren't saying that AI produces bad code or is terrible for the world in some way.

    It's mainly just resulting in a lot of PRs that they don't have enough time to review or features they don't plan to add.

  • It sounds like what they are primarly aiming for is introducing backpressure in their own review pipelines.

  • This is a good idea, we should start a blacklist of open-source projects that are known to have used LLMs. There should be two universes of code, one for hand-typed code used by people who care about quality and one for slop used by those making trash.

    • There was plenty of bad code long before LLMs came around. Let’s be real.

      Being hand written is no guarantee of high quality, just like using LLMs is no guarantee of low quality.

      2 replies →

    • Yeah but that's just bad code in general, no? You can make good code with LLMs, you just have to actually engineer it and give up some of the velocity; which is just a bigger version of the same problem we've always had (yes, I get that code review can't scale).

      This is just really silly.

      The more I think about this the more I think it's like self-driving cars. We have this expectation that self-driving cars MUST be 101% safe and never get into any accidents, ever, before the technology is worth adopting. LLMs are the same -- it's like we think if you can't one-shot a prompt and get perfect software out of it, it's failed. You can choose to spend time getting the LLM to refine the code it's written, review the architecture, come up with an actual engineering process around the LLM. Yes that means you'll be producing less code per time spent -- which is a good thing.

I had a PR in flight that got closed because of this. I had an issue with passwords in the network applet for the VPN and had used Claude to help me identify and then come up with a fix. I did spend a lot time handcrafting and making sure the quality was good, but I respect their decision and no hard feelings, but as someone who have struggled to find time and opportunity to contribute to open source it was a small set back.

  • I found your commit and your usage of AI seemed reasonable. It seems to me like your PR itself and the subsequent comments and correspondence was also human written.

    I think a PR "in flight" shouldn't have been closed like that.

    All this will do is push out developers like you that honestly disclose, and instead people will now just lie.

  • What is actually the meaning of “handcrafting” here?

    • What I meant by it was that I read every line of code generated, made sure I understood its purpose and either manually rewrote it if I felt there was a better way or asked Claude to do it. E.g. there was a bug where the password could end up in a configuration labelled as a username. Claude's initial fix was to simply exclude 'username' in a for loop, but I asked it to find examples in similar code in other codebases to see what the best practice was and ended up basing to fix on what is done in Gnome.

    • My guess: taking personal responsibility for the functionality, readability, and sanity of the change proposed, both atomically and in the context of the wider code base (adhering to existing conventions and patterns), to the best of the author’s ability.

      2 replies →

I don't see how this will survive the attacker/defender gap as ls get increasingly good at cyber security and finding 0 days... but maybe it's an obscure enough is it doesn't matter?

  • Probably because pop_os and Cosmic are so niche and their market share so insignificant, that they're irrelevant to attackers and bad actors, when those now have much bigger fish to fry to make their time and effort worth it. See the Arch AUR attacks, for perspective.

    I think even amongst the HN and Linux userbase, pop_os is still niche, let alone amongst normies who never heard about Linux. So they can afford take the high road and treat it like their personal sandbox, accepting only human written code.

    But larger and more important projects like Fedora and Debian are more pragmatic with the fact that they'll have to accept AI written(but human reviewed) code, if they wish to keep up with the real world development and threats, as expressed by Linus Torvalds himself.

    The thing is, the cat's out of the bag on this one now, especially in the field of pen-testing and reverse-engineering. AI can brute-force its way into projects in ways that beat even experienced researchers, so your only choice to keep up is to accept the use of AI generated fixes as a counter defense.

    • You only need one bad actor. For example, someone reading this thread could easily decide to start attacking it just because someone else said it wasn't worth it, as a personal challenge.

      6 replies →

    • It does show up on top 5 lists for Linux desktop distros quite a lot, and COSMIC is quite unique, so I suspect a lot of people are at least trying it.

      (I wasn't able to make it work on a scrap Dell I tried it on because the GPU was too old. Booted the USB key and COSMIC greeter failed to start)

      20 replies →

    • << Probably because pop_os and Cosmic are so niche and their market share so low, that they're irrelevant to attackers and bad actors, when those now have much bigger fish to fry.

      That is such a weird statement that I am not entirely certain where to begin. PopOS is hardly niche. Its base are all fairly common components by linux standards. And, more importantly, attackers and bad actors may other considerations in mind than sheer population size -- just to point out the glaringly obvious.

      << I think even amongst the HN and Linux userbase, pop_os is still niche

      I think rather than trying to disprove it, I think I should ask why you think that? If anything, PopOS annoys me because it is just a step before ubuntu ( and ubuntu is just windows at this point ). Maybe I am defensive, because my first real distribution ( that did not share disk with windows was popos )?

      1 reply →

    • It also means security is not held as high and vulnerabilities not as much found. A simple 0-day may survive for years. Not much effort needed to have permanent access.

      5 replies →

  • So you took every single line of open source code you could possibly get your hands on (using scrapers so violently dumb that they amount to a permanent low-grade DDoS) and spent billions of dollars to tune trillions of parameters, and the value you can offer is… “let us inundate you with bad code or else we’ll generate exploits for your software”.

  • Using AI to find vulnerabilities doesn’t mean that you need to use AI to generate the code that fixes them. And you can still ask AI whether it thinks the fix is okay, as a second opinion.

  • AI finds a lot of "vulnerabilities" but most are fake, untested, or not actually vulnerabilities.

    Reminder that AI is quite stupid.

  • Irony is that they will get the benefits from upstream projects (like the Linux kernel) that does accept AI inputs.

  • The issue is “ai generated code” using ai to find bugs/0 day and manually writing a fix would (I assume) be allowed under the new rules.

I wonder if the issue is mostly the code or the AI written PRs and people using AI to talk to the maintainers. I personally just ban anyone doing the latter, I don't want to talk to opus more than I already do lol

  • Issue is entrenchment in the old SWE world that ceased to exist somewhere in August this year, and using AI to mechanize the traditional workflows. Also people with "I need to eyeball each char in PR diff" attitude, which is fucking unproductive at this point.

Ok, this may be controversial, but LLM code tokens aren't free, and I run out of my weekly allowance pretty regularly just from doing some fairly heavy projects, so I don't understand why somebody would ever want to spend their own money to make bad PRs on purpose, and I like to assume good intentions from people unless proven otherwise, which means a near blanket ban for LLM authored code for these big open source projects just seemed a bit extreme to me, when the core issue seemed to be that review process/policy should change with the times.

For example, I was helping work on an open-source game engine earlier this year with a longstanding text rendering bug dating back to around 2021 that prevents the engine from being production ready, which the community and myself have developed extensive workaround for. So, one day I've finally said enough and got Claude to debug it. It took Claude 10 minutes to find the bug, it was 3 lines of code change in the renderer (yes, three).

So, I wrote up the regression tests, documented the bug and opened up a PR for the fix, thinking it'll get merged in like less than a week and then we can all move on. The maintainers received it fairly well on the PR, but the PR sat there for nearly 6 months, unmerged, until it finally closed from a bad squash upstream. I'm pretty sure the bug is still there too.

And as a side note, I would be ecstatic if someone wants to contribute to my Github projects with their AI.

"... many of the AI contributions were not planned and showed little understanding of the software architecture. So the team wants to "prioritize working on contributions from our own team and regular contributors.""

Sounds reasonable, even to avid LLM users, I suppose. You have to draw a line. This line is too simplistic, but it'll work, for now.

This would just push people to maintain their own fork. If I already have an agent to investigate and fix a bug and able to send a PR, the added cost of maintaining a local fork is minimum.

In fact I've start doing that myself. Sending PR and convincing the maintainer why the fix is necessary is just too much effort.

Same trouble we have. Some clever person says to use AI agents for code review. 100kloc commit got flagged through on Friday. Taking this week off. Not my circus.

Faith is the problem. Extraordinary claims must stand up to scrutiny. They do not.

  • A 100kloc PR is unacceptable regardless of the source. That needs to be broken down into reviewable chunks.

Doing that will put you put of the market. I am certain of that.

  • There are a whole lot of people (in tech) who truly hate AI and want nothing to do with it. Those people will flock to projects who take a stand against it.

    • it doesn't matter. it's delusional to think you can outcompete a thing for which solving a Millenium problem is just Tuesday. it's the anger phase of grief, nothing more.

      4 replies →

banning AI from PRs because you're swamped with too many low quality PRs, definitely. We pretty much are doing this with SQLAlchemy. If I'm going to have a small fix or improvement coded by an LLM (which I do all the time), I want to prompt the LLM directly, rather than having someone trying to pad their resume forward my communications onto their LLM via PRs. What's the point of that?

  • Reasonable, although I've taken a different approach. Either closing such PRs, or treating them as very detailed issues and having my own LLM build the actual fix.

    My repos probably don't see as much traffic as SQLAlchemy though.

  • > If I'm going to have a small fix or improvement coded by an LLM (which I do all the time), I want to prompt the LLM directly, rather than having someone trying to pad their resume forward my communications onto their LLM via PRs. What's the point of that?

    Which is why LLM PRs should just be issues (if there isn’t one already). Make the issue author a co-author on the PR. But let the maintainer actually oversee the LLM generated solution.

  • rejecting low quality ones should be the norm regardless of whether an AI or a human wrote them. the question is what would happen if you were swamped with high quality PRs? what will happen once you are? (that's probably a 2027 question!)

    • I'd still reject them. It's the bug reports and feature requests that are valuable. When you have an AI yourself, there is little point in having somebody else let their AI implement them, that just creates a lot of risks and unknowns for no benefit.

    • yeah a lot of LLM PRs are pretty good, but still need changes, and still didnt come from my own prompting which would have got them more exactly where I want them, so it's again, I have to put messages on a PR and wait for someone somewhere to see them and act on them. That friction is a huge waste of time if they're just prompting their own robot. I have the same robot right here and I usually use opus 5.x which is usually better than what they're using.

  • Maybe the policy should be: no AI PRs except from established contributors who have been vetted?

    So basically you're saying you reject drive-by PRs.

Well this project is dead then.

I have a hard time knowing if anti AI is a mental illness or propaganda coming out of China.

Unironically.

Separately, why would anyone use a Debian based desktop OS? Your $11 Amazon mouse won't work. An Nvidia card won't work. Just use Fedora.