Note that they disallow LLMs that have had heavy LLM usage, so to me this reads like them covering their ass copyright wise and it sounds like a policy that will only get enforced on major rulebreaking projects after manual reports
If your work fits into these cases, it is unlikely that you are affected at all:
Projects who have an active community that cares about and maintains the software
Projects with a significant pre-LLM history
Maintainers who unknowingly or willingly accepted LLM-generated contributions from other contributors, if your project otherwise does not involve the heavy use of LLMs
We will also not spend significant amount of time and resources to automatically scan content on Codeberg. So while the following use cases are discouraged (similar to private repositories), they are likely to be tolerated in practice:
Side projects and experiments with little resource usage
Specific tools and custom scripts that would be unlikely to find a community anyway, even if they were not LLM-generated
If Codeberg isn’t going to implement automated scanning, how do they plan to identify which projects are vibe-coded? Relying on voluntary user reports?
Hopefully, this policy can be adopted by social media platforms to cut off the downstream distribution of these projects. For a code hosting provider, making such a decision might puzzle some people and spark unnecessary backlash and controversy.
I think its impossible for them to enforce it with close to 100% strike rate. The Codeberg team isn’t that big and has only limited resources.
However, I still see it as a very important signal to possible users of Codeberg that mostly vibe-coded projects are not tolerated and could be banned every time, while a responsible usage of LLM’s is not forbidden. Moreover, I see it also as important stand against the LLM/AI-first approach many projects are heading to these days.
I think that could indeed be the most usual way to be notificated about such projects, which doesn’t have to be bad at all. Since Codeberg is very tied with its community, the latter might care a lot that Codeberg will remain a place to be for human-driven projects.
Disclaimer: I’m a (not very active) member of the Codeberg non-profit organization and might be biased a little bit
This is the same thought I had: the mere presence of a such a policy will curtail the lion’s share of those type of projects, with zero enforcement whatsoever.
Wow, that was an interesting article and worth the full read, thanks for sharing. I already stopped sharing code online due to LLM slurping for training data, but this gives me hope that we might yet get a safe platform for human collaboration.
Unfortunately the more developers that turn towards these ‘AI’ tools, there more demand there will be to slurp training data. So it seems like only a total rejection of these tools will have any chance to protect human collaboration. Sadly it seems in our nature to seek out the easiest solutions and ignore the right choices. I see more and more ‘but I only used it for…’