You build your OS atop thousands of open source packages, many of which contain AI generated code. Are you going to audit them one by one and remove offending packages? What about the ones you won't remove because the OS would be irreparably broken?
It also doesn't answer the question of how they might even recognize LLM generated code in contributions to PopOS directly.
I have yet to understand how maintainers can't distinguish beyond (1) PRs that literally include Claude co-author notes or (2) low quality code contribution regardless of the creator.
This is a good idea, we should start a blacklist of open-source projects that are known to have used LLMs. There should be two universes of code, one for hand-typed code used by people who care about quality and one for slop used by those making trash.
Yeah but that's just bad code in general, no? You can make good code with LLMs, you just have to actually engineer it and give up some of the velocity; which is just a bigger version of the same problem we've always had (yes, I get that code review can't scale).
This is just really silly.
The more I think about this the more I think it's like self-driving cars. We have this expectation that self-driving cars MUST be 101% safe and never get into any accidents, ever, before the technology is worth adopting. LLMs are the same -- it's like we think if you can't one-shot a prompt and get perfect software out of it, it's failed. You can choose to spend time getting the LLM to refine the code it's written, review the architecture, come up with an actual engineering process around the LLM. Yes that means you'll be producing less code per time spent -- which is a good thing.
I also have found that AI has not lived up to many of its promises and have dialed back what AI gets control of. My projects were turning into unmaintainable messes. The people who say coding is solved aren't paying attention.
Yes. I started two projects with AI from scratch. Both abandoned, complete mess. Projects without AI are so easy to manage, maintain, add/remove features, etc. I use AI as a search engine on my projects instead of Google. I ask what's wrong with my code and change or improve it myself based on my experience.
I have the opposite experience. I don't have the patience sometimes to clean up my code and stick to coherent conventions and organization even though the will is there. With my AI projects I watch it and if the AI starts drifting I ask it to go through and look for convention/directory structure violations and it happily cleans everything up in about 10 or so minutes.
Are you in Python by chance? Python has a lot of crazy hidden/inexplicit/spooky action at a distance stuff (especially in the frameworks) that can make LLMs gunk up code by defensively programming or just burn context chasing data provenance
I've been a vibe-coding skeptic for years, but because of the math breakthroughs of the past few weeks I decided to experiment with the latest models on some test projects. They're a lot more capable than I thought they would be. I agree that it's easy to create an unrecoverable mess, especially when you're one-shotting a lot of features without detailed instructions. But I find that as long as I'm strict about the API boundaries and force the agent to work in small chunks, it's pretty effective. As one example, I got it to write an SVG renderer (a subset of the standard, most of the common features) in a few hours, which would have taken me at least a week if not more because I don't know all of the algorithms.
I'm reading y'all comments and it seems we're still finding our footing, and will for some time. I have the luck that I have access to pretty much all frontier and other models alike. I've been extremely negatively biased towards any LLM use in the start. Then the influx from juniors came, then the revolt of doing PRs on such slop came, then some structured methods how to do LLM work came, then vibe projects came, etc, etc. I literally have all the described experiences you've guys mentioned. From good, to bad, to ugly. It's like there's no one particular way about doing this, and no two projects share the same approach - just like ye olde times.
I've had the opposite experience. Spec driven development really makes things much easier to maintain. If you're just yolo'ing it and throwing prompts around things will fall apart fast.
Oh my god. I am sad to have to acknowledge this even if the models have “theoretically” gotten better. I am not a programmer but I am super opinionated with the design and structure of the code and need to make sure whatever I write/ask to generate and use for my own tasks is understood by me at least once. I wasn’t this way in 2023.
I am constantly reading/trying for an agent to help me write simple code without spending too much time. But so far nothing has worked. I wish someone figures it out otherwise, I personally would not be able to realize the AI agent productivity benefits that others are supposedly seeing.
As usual, when people post comments like this they never specify what model and when they used it. There is a huge difference between GPT 3 and 6 for example.
Late 2025 also had a step change when agents could largely code autonomously without handholding like previously, and to be honest it's not worth hearing opinions about AI from before that time, that's how significant the change was.
Were you reviewing the changes? I've seen lots of projects with this problem. But if AI PRs have to pass the same review bar as any other PR then it shouldn't be an issue in theory (if you can actually maintain the discipline).
Isn't it tiring to keep up with considering the speed of the generated code? And if you want to be careful with the review you lose a significant part of the speed advantage.
It's the same as any other PR. Keep the changes small and contained to that feature. AI can do this if you instruct it to, there is no need to vibe code some 40k line monstrosity.
The people saying programming is solved couldn't program in the first place is what I've noticed. So to them it really does feel solved, things "work" and they don't have to learn what they don't know about programming.
I've been programming professionally for decades. LLMs are extremely useful. At this point if you haven't figured out how to get value out of them, you're either holding them very wrong or you're being willfully ignorant.
I think it depends which LLM tool you're using. If you're using an older, worse model (the kind that you can use for free), the experience is significantly more frustrating. On the other hand, I'd say that the current best models are very useful with a skilled operator.
I had a PR in flight that got closed because of this. I had an issue with passwords in the network applet for the VPN and had used Claude to help me identify and then come up with a fix. I did spend a lot time handcrafting and making sure the quality was good, but I respect their decision and no hard feelings, but as someone who have struggled to find time and opportunity to contribute to open source it was a small set back.
I found your commit and your usage of AI seemed reasonable. It seems to me like your PR itself and the subsequent comments and correspondence was also human written.
I think a PR "in flight" shouldn't have been closed like that.
All this will do is push out developers like you that honestly disclose, and instead people will now just lie.
My guess: taking personal responsibility for the functionality, readability, and sanity of the change proposed, both atomically and in the context of the wider code base (adhering to existing conventions and patterns), to the best of the author’s ability.
I wonder if the issue is mostly the code or the AI written PRs and people using AI to talk to the maintainers. I personally just ban anyone doing the latter, I don't want to talk to opus more than I already do lol
I don't see how this will survive the attacker/defender gap as ls get increasingly good at cyber security and finding 0 days... but maybe it's an obscure enough is it doesn't matter?
So you took every single line of open source code you could possibly get your hands on (using scrapers so violently dumb that they amount to a permanent low-grade DDoS) and spent billions of dollars to tune trillions of parameters, and the value you can offer is… “let us inundate you with bad code or else we’ll generate exploits for your software”.
Probably because pop_os and Cosmic are so niche and their market share so low, that they're irrelevant to attackers and bad actors, when those now have much bigger fish to fry.
I think even amongst the HN and Linux userbase, pop_os is still niche, let alone amongst normies who never heard about Linux. So they can afford take the high road and treat it like their personal sandbox accepting only human written code.
But larger and more important projects like Fedora and Debian are more realistic and pragmatic with the fact that they'll have to accept AI written(but human reviewed) code, if they wish to keep up with the real world development and threats, as expressed by Linus Torvalds himself.
The thing is, the cat's out of the bag on this one now, especially in the fields of pen-testing and reverse-engineering. AI can brute-force its way into projects in ways that beat experienced researchers, so your only choice to keep up is to accept the use of AI generated fixes as a counter defense.
You only need one bad actor. For example, someone reading this thread could easily decide to start attacking it just because someone else said it wasn't worth it, as a personal challenge.
2) Secondly, keep in mind that distrowatch is not representative of linux userbase. MX-Linux kept showing up at the top spot for many years despite being niche.
3) And thirdly, TOP 5 DE isn't really an achievement when Linux DE market share is overwhelmingly dominated by KDE Plasma and Gnome as the majority shareholders, with XFCE and Cinnamon trailing. So Cosmic DE if it somehow made the no. 5 spot, would still be ignorable sub <1% market share, as I initially claimed, basically invisible to bad actors.
<< Probably because pop_os and Cosmic are so niche and their market share so low, that they're irrelevant to attackers and bad actors, when those now have much bigger fish to fry.
That is such a weird statement that I am not entirely certain where to begin. PopOS is hardly niche. Its base are all fairly common components by linux standards. And, more importantly, attackers and bad actors may other considerations in mind than sheer population size -- just to point out the glaringly obvious.
<< I think even amongst the HN and Linux userbase, pop_os is still niche
I think rather than trying to disprove it, I think I should ask why you think that? If anything, PopOS annoys me because it is just a step before ubuntu ( and ubuntu is just windows at this point ). Maybe I am defensive, because my first real distribution ( that did not share disk with windows was popos )?
It also means security is not held as high and vulnerabilities not as much found. A simple 0-day may survive for years. Not much effort needed to have permanent access.
Using AI to find vulnerabilities doesn’t mean that you need to use AI to generate the code that fixes them. And you can still ask AI whether it thinks the fix is okay, as a second opinion.
Same trouble we have. Some clever person says to use AI agents for code review. 100kloc commit got flagged through on Friday. Taking this week off. Not my circus.
Faith is the problem. Extraordinary claims must stand up to scrutiny. They do not.
There are a whole lot of people (in tech) who truly hate AI and want nothing to do with it. Those people will flock to projects who take a stand against it.
it doesn't matter. it's delusional to think you can outcompete a thing for which solving a Millenium problem is just Tuesday. it's the anger phase of grief, nothing more.
banning AI from PRs because you're swamped with too many low quality PRs, definitely. We pretty much are doing this with SQLAlchemy. If I'm going to have a small fix or improvement coded by an LLM (which I do all the time), I want to prompt the LLM directly, rather than having someone trying to pad their resume forward my communications onto their LLM via PRs. What's the point of that?
Reasonable, although I've taken a different approach. Either closing such PRs, or treating them as very detailed issues and having my own LLM build the actual fix.
My repos probably don't see as much traffic as SQLAlchemy though.
rejecting low quality ones should be the norm regardless of whether an AI or a human wrote them. the question is what would happen if you were swamped with high quality PRs? what will happen once you are? (that's probably a 2027 question!)
> If I'm going to have a small fix or improvement coded by an LLM (which I do all the time), I want to prompt the LLM directly, rather than having someone trying to pad their resume forward my communications onto their LLM via PRs. What's the point of that?
Which is why LLM PRs should just be issues (if there isn’t one already). Make the issue author a co-author on the PR. But let the maintainer actually oversee the LLM generated solution.
Completely performative.
You build your OS atop thousands of open source packages, many of which contain AI generated code. Are you going to audit them one by one and remove offending packages? What about the ones you won't remove because the OS would be irreparably broken?
It also doesn't answer the question of how they might even recognize LLM generated code in contributions to PopOS directly.
I have yet to understand how maintainers can't distinguish beyond (1) PRs that literally include Claude co-author notes or (2) low quality code contribution regardless of the creator.
This is a good idea, we should start a blacklist of open-source projects that are known to have used LLMs. There should be two universes of code, one for hand-typed code used by people who care about quality and one for slop used by those making trash.
Yeah but that's just bad code in general, no? You can make good code with LLMs, you just have to actually engineer it and give up some of the velocity; which is just a bigger version of the same problem we've always had (yes, I get that code review can't scale).
This is just really silly.
The more I think about this the more I think it's like self-driving cars. We have this expectation that self-driving cars MUST be 101% safe and never get into any accidents, ever, before the technology is worth adopting. LLMs are the same -- it's like we think if you can't one-shot a prompt and get perfect software out of it, it's failed. You can choose to spend time getting the LLM to refine the code it's written, review the architecture, come up with an actual engineering process around the LLM. Yes that means you'll be producing less code per time spent -- which is a good thing.
I also have found that AI has not lived up to many of its promises and have dialed back what AI gets control of. My projects were turning into unmaintainable messes. The people who say coding is solved aren't paying attention.
Yes. I started two projects with AI from scratch. Both abandoned, complete mess. Projects without AI are so easy to manage, maintain, add/remove features, etc. I use AI as a search engine on my projects instead of Google. I ask what's wrong with my code and change or improve it myself based on my experience.
I have the opposite experience. I don't have the patience sometimes to clean up my code and stick to coherent conventions and organization even though the will is there. With my AI projects I watch it and if the AI starts drifting I ask it to go through and look for convention/directory structure violations and it happily cleans everything up in about 10 or so minutes.
Are you in Python by chance? Python has a lot of crazy hidden/inexplicit/spooky action at a distance stuff (especially in the frameworks) that can make LLMs gunk up code by defensively programming or just burn context chasing data provenance
I've been a vibe-coding skeptic for years, but because of the math breakthroughs of the past few weeks I decided to experiment with the latest models on some test projects. They're a lot more capable than I thought they would be. I agree that it's easy to create an unrecoverable mess, especially when you're one-shotting a lot of features without detailed instructions. But I find that as long as I'm strict about the API boundaries and force the agent to work in small chunks, it's pretty effective. As one example, I got it to write an SVG renderer (a subset of the standard, most of the common features) in a few hours, which would have taken me at least a week if not more because I don't know all of the algorithms.
I'm reading y'all comments and it seems we're still finding our footing, and will for some time. I have the luck that I have access to pretty much all frontier and other models alike. I've been extremely negatively biased towards any LLM use in the start. Then the influx from juniors came, then the revolt of doing PRs on such slop came, then some structured methods how to do LLM work came, then vibe projects came, etc, etc. I literally have all the described experiences you've guys mentioned. From good, to bad, to ugly. It's like there's no one particular way about doing this, and no two projects share the same approach - just like ye olde times.
Have you considered that maybe this is a reflection of your skills rather than that of the LLM?
Have you considered it isn't?
Save your "you're holding it wrong" if you're not going to suggest how to hold it.
Cult speak escape hatches are intellectually lazy.
Meaning OP is a better programmer then a statistical model which produces the most probable results?
When and what model?
I've had the opposite experience. Spec driven development really makes things much easier to maintain. If you're just yolo'ing it and throwing prompts around things will fall apart fast.
Oh my god. I am sad to have to acknowledge this even if the models have “theoretically” gotten better. I am not a programmer but I am super opinionated with the design and structure of the code and need to make sure whatever I write/ask to generate and use for my own tasks is understood by me at least once. I wasn’t this way in 2023. I am constantly reading/trying for an agent to help me write simple code without spending too much time. But so far nothing has worked. I wish someone figures it out otherwise, I personally would not be able to realize the AI agent productivity benefits that others are supposedly seeing.
As usual, when people post comments like this they never specify what model and when they used it. There is a huge difference between GPT 3 and 6 for example.
Late 2025 also had a step change when agents could largely code autonomously without handholding like previously, and to be honest it's not worth hearing opinions about AI from before that time, that's how significant the change was.
Were you reviewing the changes? I've seen lots of projects with this problem. But if AI PRs have to pass the same review bar as any other PR then it shouldn't be an issue in theory (if you can actually maintain the discipline).
Isn't it tiring to keep up with considering the speed of the generated code? And if you want to be careful with the review you lose a significant part of the speed advantage.
It's the same as any other PR. Keep the changes small and contained to that feature. AI can do this if you instruct it to, there is no need to vibe code some 40k line monstrosity.
Having unmaintainable code faster is only an advantage if it's a one-shot throwaway artifact.
The people saying programming is solved couldn't program in the first place is what I've noticed. So to them it really does feel solved, things "work" and they don't have to learn what they don't know about programming.
I've been programming professionally for decades. LLMs are extremely useful. At this point if you haven't figured out how to get value out of them, you're either holding them very wrong or you're being willfully ignorant.
You both can be right at the same time.
I think it depends which LLM tool you're using. If you're using an older, worse model (the kind that you can use for free), the experience is significantly more frustrating. On the other hand, I'd say that the current best models are very useful with a skilled operator.
I had a PR in flight that got closed because of this. I had an issue with passwords in the network applet for the VPN and had used Claude to help me identify and then come up with a fix. I did spend a lot time handcrafting and making sure the quality was good, but I respect their decision and no hard feelings, but as someone who have struggled to find time and opportunity to contribute to open source it was a small set back.
I found your commit and your usage of AI seemed reasonable. It seems to me like your PR itself and the subsequent comments and correspondence was also human written.
I think a PR "in flight" shouldn't have been closed like that.
All this will do is push out developers like you that honestly disclose, and instead people will now just lie.
What is actually the meaning of “handcrafting” here?
My guess: taking personal responsibility for the functionality, readability, and sanity of the change proposed, both atomically and in the context of the wider code base (adhering to existing conventions and patterns), to the best of the author’s ability.
I wonder if the issue is mostly the code or the AI written PRs and people using AI to talk to the maintainers. I personally just ban anyone doing the latter, I don't want to talk to opus more than I already do lol
I don't see how this will survive the attacker/defender gap as ls get increasingly good at cyber security and finding 0 days... but maybe it's an obscure enough is it doesn't matter?
So you took every single line of open source code you could possibly get your hands on (using scrapers so violently dumb that they amount to a permanent low-grade DDoS) and spent billions of dollars to tune trillions of parameters, and the value you can offer is… “let us inundate you with bad code or else we’ll generate exploits for your software”.
Probably because pop_os and Cosmic are so niche and their market share so low, that they're irrelevant to attackers and bad actors, when those now have much bigger fish to fry.
I think even amongst the HN and Linux userbase, pop_os is still niche, let alone amongst normies who never heard about Linux. So they can afford take the high road and treat it like their personal sandbox accepting only human written code.
But larger and more important projects like Fedora and Debian are more realistic and pragmatic with the fact that they'll have to accept AI written(but human reviewed) code, if they wish to keep up with the real world development and threats, as expressed by Linus Torvalds himself.
The thing is, the cat's out of the bag on this one now, especially in the fields of pen-testing and reverse-engineering. AI can brute-force its way into projects in ways that beat experienced researchers, so your only choice to keep up is to accept the use of AI generated fixes as a counter defense.
You only need one bad actor. For example, someone reading this thread could easily decide to start attacking it just because someone else said it wasn't worth it, as a personal challenge.
It does show up on top 5 lists for Linux desktop distros quite a lot, and COSMIC is quite unique, so I suspect a lot of people are at least trying it.
(I wasn't able to make it work on a scrap Dell I tried it on because the GPU was too old. Booted the USB key and COSMIC greeter failed to start)
>It does show up on top 5 lists for Linux desktop distros quite a lot
1) Firstly, surce? My research according to Google Gemini 3.8 Pro shows top 5 DEs are as follows:
2) Secondly, keep in mind that distrowatch is not representative of linux userbase. MX-Linux kept showing up at the top spot for many years despite being niche.3) And thirdly, TOP 5 DE isn't really an achievement when Linux DE market share is overwhelmingly dominated by KDE Plasma and Gnome as the majority shareholders, with XFCE and Cinnamon trailing. So Cosmic DE if it somehow made the no. 5 spot, would still be ignorable sub <1% market share, as I initially claimed, basically invisible to bad actors.
You're being way too defensive. GP replying to you was clearly trying to have a conversation, not attack your knowledge. Chill.
Don't post AI generated content, if needed post the actual source.
Define "actual source"?
<< Probably because pop_os and Cosmic are so niche and their market share so low, that they're irrelevant to attackers and bad actors, when those now have much bigger fish to fry.
That is such a weird statement that I am not entirely certain where to begin. PopOS is hardly niche. Its base are all fairly common components by linux standards. And, more importantly, attackers and bad actors may other considerations in mind than sheer population size -- just to point out the glaringly obvious.
<< I think even amongst the HN and Linux userbase, pop_os is still niche
I think rather than trying to disprove it, I think I should ask why you think that? If anything, PopOS annoys me because it is just a step before ubuntu ( and ubuntu is just windows at this point ). Maybe I am defensive, because my first real distribution ( that did not share disk with windows was popos )?
It also means security is not held as high and vulnerabilities not as much found. A simple 0-day may survive for years. Not much effort needed to have permanent access.
Sorry, I don't understand what you mean by this, can you elaborate pls?
Using AI to find vulnerabilities doesn’t mean that you need to use AI to generate the code that fixes them. And you can still ask AI whether it thinks the fix is okay, as a second opinion.
AI finds a lot of "vulnerabilities" but most are fake, untested, or not actually vulnerabilities.
Reminder that AI is quite stupid.
I think this is just plain denial, AI has found many high severity vulnerabilities.
There's plenty more false positives than actual finds. There are still actual finds, but that doesn't change all the false positives.
This is why you run a second agent to verify any assertions.
Irony is that they will get the benefits from upstream projects (like the Linux kernel) that does accept AI inputs.
The issue is “ai generated code” using ai to find bugs/0 day and manually writing a fix would (I assume) be allowed under the new rules.
I'm assuming they still use AI to find vulnerabilities, just not fix them?
As they should. If you can tell it's AI, it shouldnt be part of any codebase.
Same trouble we have. Some clever person says to use AI agents for code review. 100kloc commit got flagged through on Friday. Taking this week off. Not my circus.
Faith is the problem. Extraordinary claims must stand up to scrutiny. They do not.
Doing that will put you put of the market. I am certain of that.
There are a whole lot of people (in tech) who truly hate AI and want nothing to do with it. Those people will flock to projects who take a stand against it.
it doesn't matter. it's delusional to think you can outcompete a thing for which solving a Millenium problem is just Tuesday. it's the anger phase of grief, nothing more.
I still wanna see what Slop OS looks like. Go full AI spam making a Linux distro from the kernel up.
It might get millions in founding.
banning AI from PRs because you're swamped with too many low quality PRs, definitely. We pretty much are doing this with SQLAlchemy. If I'm going to have a small fix or improvement coded by an LLM (which I do all the time), I want to prompt the LLM directly, rather than having someone trying to pad their resume forward my communications onto their LLM via PRs. What's the point of that?
Reasonable, although I've taken a different approach. Either closing such PRs, or treating them as very detailed issues and having my own LLM build the actual fix.
My repos probably don't see as much traffic as SQLAlchemy though.
rejecting low quality ones should be the norm regardless of whether an AI or a human wrote them. the question is what would happen if you were swamped with high quality PRs? what will happen once you are? (that's probably a 2027 question!)
> If I'm going to have a small fix or improvement coded by an LLM (which I do all the time), I want to prompt the LLM directly, rather than having someone trying to pad their resume forward my communications onto their LLM via PRs. What's the point of that?
Which is why LLM PRs should just be issues (if there isn’t one already). Make the issue author a co-author on the PR. But let the maintainer actually oversee the LLM generated solution.
Maybe the policy should be: no AI PRs except from established contributors who have been vetted?
So basically you're saying you reject drive-by PRs.
drive-by PRs that are handwritten are presumably okay
Thank god. Tired of these orgs blindly and stupidly allowing them (Debian, cough).
Well this project is dead then.
I have a hard time knowing if anti AI is a mental illness or propaganda coming out of China.
Unironically.
Separately, why would anyone use a Debian based desktop OS? Your $11 Amazon mouse won't work. An Nvidia card won't work. Just use Fedora.