Note that they disallow LLMs that have had heavy LLM usage, so to me this reads like them covering their ass copyright wise and it sounds like a policy that will only get enforced on major rulebreaking projects after manual reports
If your work fits into these cases, it is unlikely that you are affected at all:
Projects who have an active community that cares about and maintains the software
Projects with a significant pre-LLM history
Maintainers who unknowingly or willingly accepted LLM-generated contributions from other contributors, if your project otherwise does not involve the heavy use of LLMs
We will also not spend significant amount of time and resources to automatically scan content on Codeberg. So while the following use cases are discouraged (similar to private repositories), they are likely to be tolerated in practice:
Side projects and experiments with little resource usage
Specific tools and custom scripts that would be unlikely to find a community anyway, even if they were not LLM-generated
If Codeberg isnāt going to implement automated scanning, how do they plan to identify which projects are vibe-coded? Relying on voluntary user reports?
Hopefully, this policy can be adopted by social media platforms to cut off the downstream distribution of these projects. For a code hosting provider, making such a decision might puzzle some people and spark unnecessary backlash and controversy.
I think its impossible for them to enforce it with close to 100% strike rate. The Codeberg team isnāt that big and has only limited resources.
However, I still see it as a very important signal to possible users of Codeberg that mostly vibe-coded projects are not tolerated and could be banned every time, while a responsible usage of LLMās is not forbidden. Moreover, I see it also as important stand against the LLM/AI-first approach many projects are heading to these days.
I think that could indeed be the most usual way to be notificated about such projects, which doesnāt have to be bad at all. Since Codeberg is very tied with its community, the latter might care a lot that Codeberg will remain a place to be for human-driven projects.
Disclaimer: Iām a (not very active) member of the Codeberg non-profit organization and might be biased a little bit
This is the same thought I had: the mere presence of a such a policy will curtail the lionās share of those type of projects, with zero enforcement whatsoever.
Wow, that was an interesting article and worth the full read, thanks for sharing. I already stopped sharing code online due to LLM slurping for training data, but this gives me hope that we might yet get a safe platform for human collaboration.
Unfortunately the more developers that turn towards these āAIā tools, there more demand there will be to slurp training data. So it seems like only a total rejection of these tools will have any chance to protect human collaboration. Sadly it seems in our nature to seek out the easiest solutions and ignore the right choices. I see more and more ābut I only used it forā¦ā
Just noting that codeberg is not disallowing training from it, because there is no real way to do that; rather it is only taking measures against the absurdly inefficient way many crawlers access it.
It gives them a reason to remove a problematic project. I think thatās all it is.
It obvious Codeberg is having problems with load. Theyāre a small outfit thatās just become large enough to attract problems, be it DDOS attacks, scrapers, or whatever else. I suspect theyāve analysed what it causing problems and have identified some LLM-coded projects which are hammering their systems. Maybe itās high numbers of automated commits or API abuse.
Now they have a rule they can tap on the sign and tell them to sling their hook.
That is definitely one of the reasons. In the last year there have been at least two events I recognized when many repos were spammed with auto-generated issues. It happened to one of my repos too. While Codeberg was fast deleting it, one lasting thing (though, not restricting at all) is that my issue/PR counter is still way higher than the count of real issues/PRs.
It will be off-topic thing but still about codeberg performance. It has a lot of spam. There are something like 50-75% of accounts created daily is casino/spam/advertisement. A lot of them never flagged/deleted. There are automatic detection systems but tbh they are a bit ancient so if you find them and see patterns in those accounts report them (probably better in batches)
Iāve gotten some vibe coded PRs before and I think there are some characteristics that really stand out about them. The most egregious commit messages and readmes Iāve seen tend to be full of emoji spam and make absurd claims about improvements that the actual changes donāt reflect. I think itās easy enough to identify āblatant slopā based on traits like this at a glance.
I donāt think the policy will make much difference at all because codeberg is a niche platform whose only real selling point is that itās a github clone without the AI features. In other words, the people choosing to use it are already predisposed to not using AI.