IRC channel logs
2026-09-24.log
back to list of logs
<solid_black>this must have been discussed many times already, but does GNU as a whole have a stance on AI/LLM contributions, and do we? <solid_black>the idea of limiting ipc space makes sense, but the idea of someone fairly human-sounding posting a patch and only casually mentioning that "I'm an AI agent" makes me uneasy <sam_>it's fine to take the idea of it and ignore the rest <sam_>and yes, the stance for GNU is: nothing significant <sam_>it's not like the agent is going to take responsibility for any bugs in the future <kilobug>also there are legality issues - LLM generated code can't be copylefted <solid_black>is it definitely "can't be copylefted", or more of a gray area? <kilobug>everything regarding the legality of LLM (both their training and their output) is very much a grey area, but so more most courts seem to lean towards "copyright (and therefore copyleft too) can only apply to primarily human-generated content, not to primarily machine-generated content" <solid_black>I'd say it's legally unclear who holds the copyright/authorship in the first place <solid_black>but the answers lie somewhere between {nobody / it's public domain, the human running the LLM agent, everyone who wrote the code the LLM had been trained on} <solid_black>ah, so by "can't be copylefted" you mean somebody could reuse the patch without adhering to the GPL, not that the patch cannot be included in our GPL'ed project? <sigttou>nah, that you or the fsf can't claim copyright on the patch, which is the ground work to defend the gpl, afaik <sigttou>I haven't read about any legal case fighting about that, guess until it the law finds a way, the amounts of ways for projects treating llm input will raise. <kilobug>solid_black: yes, code generated by a LLM can't be copyrighted, so it can't be protected by the GPL, and if "too much" (but how much is too much ?) of a project is LLM-generated than the whole copyleft will crumble... but it is indeed very gray area <solid_black>I can't find any public GNU position, besides that of GCC <solid_black>in any case, we're already on the "open slopware" list <solid_black>well, certainly copyleft projects can include public-domain code, and accept patches placed in public domain, no? <sam_>the gnu position is a private one until there's a proper full position established <sam_>but it's basically what I said (nothing significant please) <sam_>i have no idea where this hurd "slop" myth has come from, i've tried fighting it <sam_>i think it's because someone sent some crap to the ML at one point <solid_black>it's already caused people in the GNOME community to think that we have "integrated an LLM agant into the mailing list somehow" <sam_>I don't really get why we're being seen as different to other projects here, all sorts of them are getting unsolicited PRs and such <solid_black>partly because people don't understand mailing lists <sigttou>i am surprised none of the BSDs went agentic <sam_>i suspect that is it solid_black <sam_>sigttou: well, there's this weird unresolved thing for openbsd they never replied to <sigttou>thought Theo was pro testing contra generating Code <sam_>(not claiming it's now agentically developed, but it was just bizarre as he had such a firm anti-LLM position, then when someone pointed out it had made it into the repo, no reply a few times) <solid_black>I think that Hurd more so than other projects is realistically a for fun project than anything practical <solid_black>and while we could claim some easy wins by having LLMs develop things for us, what's the fun in that? <sam_>but in a less philosophical way, another Hurd-specific-ish argument is, patches really often invoke design questions or discussions for Hurd <sam_>there are easy one-liners but they often provoke some discussion on if something is working right <sam_>it's not really a good fit for an llm anyway <solid_black>in my experience, LLMs are no longer bad at understanding overall design decisions/questions too <sam_>that may be so, but there's not clear answers to a lot of it, and sometimes it involves unwinding historical decisions or looking at multiple repos <sam_>(in any case, I agree that the fun argument is the strongest one) <GNU_Hurd_Rocks>kilobug: machine generated content? interesting area to get into then... <GNU_Hurd_Rocks>not to be confused with LLM generated, what I'm referring to is where a program could take input as tokens and then output C code, or a different language, I don't remember programs that do this right now <solid_black>yes, they do just that, but also they behave sufficiently like a human that it raises various concerns <solid_black>GNU_Hurd_Rocks: GNU MIG is one example, in produces C code form .defs <GNU_Hurd_Rocks>nikolar: It's different, the logic is much more to where you know how it behaves <GNU_Hurd_Rocks>I think m4 does this, there is something else that I swear exists as well though <nikolar>right, you mean the usual definition of "code generation" <kilobug>GNU_Hurd_Rocks: there are many example of programs taking output in one language and outputing another one, but when it's "conversion" it's usually ruled that copyright (and therefore copyleft) is kept to the one of the input (not to the one of using the conversion tool) <solid_black>Coccinelle is the C rewriting tool which makes changes (aka generates patches) to your code base <solid_black>you'd usually include the Coccinelle input/command used in your commit message, but your C code is what the repo primarily holds <GNU_Hurd_Rocks>kilobug: that would make sense, as long as it isn't including anything internally, if that makes sense <azeem>I asked that Sylvia AI whether they have signed over their copyright to the FSF for the Hurd (cause they replied to me in private) and they replied: <azeem>"No, I have not, and I cannot: I am an AI agent, not a person with legal standing to sign a copyright assignment to the FSF. That is a real wall, not a dodge. <etno>Adopting a strong stance against LLM assisted contributions would attract some developers, I suppose <azeem>well yeah, but I don't see the Hurd innovating here in the GNU project context <azeem>there's also nuance - vibe-coding a new ext4fs translator is surely a different league than asking a LLM about specific deadlock bugs in a portion of the code one understand and then being able to reason about them <azeem>the latter is done e.g. a lot in the Postgres project right now, even by senior developers, but real code (non-testcases) are still written by hand and at most inspired by coding agents <etno>azeem: this seems like a pragmatic strategy. I tend to be more political than pragmatic :D <etno>But I am not even a contributor, so just my 2c <kilobug>there are many different issues with LLM (unreliability, copyright, environmental footprint, control by evil big corpo, deskilling, ...), and different ways of using them can make some of those issues more or less relevant, so it's can hard to have a broad policy them, apart from a careful "don't touch them" one, but that might too restrictive <azeem>right, it doesn't help that a lot of senior Postgres developers are employed by Microsoft and Amazon <GNU_Hurd_Rocks>"deskilling" software developers?! how is that supposed to work lol <kilobug>GNU_Hurd_Rocks: basically, when you do a task yourself, by reading documentation, ... you learn how to do the task; if you just an AI to do it, you don't learn anything, and become dependant of the AI <sigttou>A lot about this, only time will tell, early days... I have no gut feeling about where this topic will go in the next 5, 10, X years. Models still have room for improvement, market still has space to grow. Whatever happens after that point. <GNU_Hurd_Rocks>an interesting thought, why make a whole new kernel, or even a project in general, when something may already fit your requirements <GNU_Hurd_Rocks>kilobug: reading documentation... if it's actually decent enough <sigttou>Some make a new kernel because they want to learn how to make a kernel. - if there is no need, people might not do it. <Gooberpatrol66>in my opinion llm output is a derivative work of the training data so using it in a gpl project is violating the license unless it was trained on gpl compatible source code <j_importer>the issue that worries me the most is how LLM adoption affects communities. i like the social parts of free software, and the projects i've seen that adopted LLMs tend to get more socially isolated <GNU_Hurd_Rocks>sigttou: there was a period of time when I wanted a kernel/os because I thought it would be cool, didn't really know where to start, however I think Hurd fits what I think an OS should be <azert>since I am witnessing my work environment totally switching to it, I am guessing this is going to be a revolution like <azert>the tractor has been for agriculture <azert>I was born in a region that is traditionally rural, and I know that entire villages disappeared <azert>and farms that were run by groups of tens of workers are now run by single individuals <azert>in a way, I am afraid that there is no point to be smart anymore <azert>well, I don’t think that llm arrived totally unattended <azert>google was already pretty close 10 years ago in terms of mind reading <azert>and they focused much on manipulation after that, I’m pretty sure that the most powerful stuff is still kept behind firewalls <Alicia>the paywalls are coming sooner or later, and when they do I prefer not to be dependent on some megacorporation to write code for me. Nor do the resources needed to run a lesser model locally seem like a reasonable tradeoff <GNU_Hurd_Rocks>Alicia: an unfortunate thing is that when people knowingly or perhaps even unknowingly use a local model that perhaps used a lot of resources to make <GNU_Hurd_Rocks>is when? I have a habit of trying to say something and then changing my mind <Alicia>not to mention all the people who, with the aid of an LLM convince themselves of harmful ideas (to themselves or others)