Comment by KingOfCoders
6 days ago
What is a vibe coded project? Where does it start? Cursor autocomplete? One shot Github project copies?
[Edit] The pull link is https://codeberg.org/Codeberg/org/pulls/1253/files and says
"7. You must not share projects that mostly consist of code written by "generative AI"-tools (including services such as Claude, OpenAI Codex). Such projects having an unclear copyright status (see requirements § 2 (1) 1 and § 2 (1) 3) and furthermore have little safeguards to ensure that they do not include harmful code (c.f. § 2 (1) 5)."
Whatever "mostly" means. If you autocomplete a lot, and the code written is "mostly" written by AI by autocomplete - it seems you fall under this.
I wonder what the ratio needs to be. And I wonder if auto-refactoring in Intellij is also included, because Intellij created the code and then also has "unclear copyright status" if we follow the logic.
Together this opens up more questions to me than it answers.
“You must not share projects that mostly consist of code written by "generative AI"-tools (including services such as Claude, OpenAI Codex).”
Seems pretty clear to me?
What is mostly? >50%? >75%?
What does "written by generative AI" means? Autocomplete? IDE-with-AI-normal-classname-completion? Everything not typed by a human.
I personally don't find this clear at all and see many unhappy discussions in the future for Codeberg.
> What is mostly? >50%? >75%?
1. Don't ask us, ask the Codeberg people in the thread above.
2. If you're not sure you can meet Codeberg's terms of service, do what others are claiming to do: Take your repositories elsewhere.
3. I'll just hypothesize that you don't have Codeberg repositories and are only arguing here for argument's sake.
11 replies →
It means if the guy running Codeberg doesn't like it. This is his fiefdom, not a public square. Make your own fiefdom for AI coding. You could call it Vibeberg.
6 replies →
I think it's reasonable to assume >50% by default, unless the source clarifies what they mean by "mostly".
10 replies →
> I personally don't find this clear at all and see many unhappy discussions in the future for Codeberg.
You don't understand their goal in doing this?
Frequently (much more than we'd like to admit) actual contracts leave loopholes in, said loopholes which go against the spirit of the contract. It is not unusual to have a concrete contract that allows more (or less) than the spirit the contract was signed in.
Rather than nitpicking the contract (the TOU), why do you think they need this updated contract in the first place?
Protest by no longer using Codeberg
If 'mostly' means over 50% you could call this clear.
Otherwise I don't think this is clear at all.
Otherwise, what is mostly?
Mostly at the outset, or at any given time? Must the project move out when it goes from 49% to 51% AI-assisted code?
"Mostly at any given time?"
Yes, old project with 1M LOC of pre-AI code + 100% AI 100k LOC generated coded for the last three months, is that mostly?
3 replies →
Noone can discuss in good faith if the argument is: "If generated in bulk by vibe code" >> "what if I generate word by word with autocomplete".
This is IMHO the reason why we can't have sanity, because there are people ready to hack it.
That's why terms and laws are written vaguely. If you're obviously on one side of the line then you're obviously on one side of the line. If you're straddling the line, your punishment depends upon how the judge feels that day. So it would be wise not to risk it.
So no autocomplete. This is why people vote against their interest, they always assume laws are to reign in on other people until they understand the intention was different from what they understood and it hits them. Same old story.
I don't think this is a bad-faith response. Someone might fully retain copyright while using heavy autocomplete. I don't think such a person would be banned by the codeberg policy. This is at least a little bit unintuitive.
Acting like the decision to ban misween coded projects is a slippery slope is like being surprised that any hosting services refuse to store and maintain petabytes of your lorem ipsum novel.
If you have to ask then the answer is no.
I mean Codeberg will pretty soon start having to ask. How do they determine that and who'll be in charge to find violations there?
Before Codeberg has to ask that, the person who would upload has to ask it themselves. People know what they did, or didn't.
"If you have shit on your shoe, please don't walk across our carpet".
"But how do they know? maybe we should test if they can tell", and all that stuff... no. Most people would simply take a good look, and if they're worried about something that could be be mud or not, brush it off just in case. And if they come directly from 12 hours of drunk partying in the wilderness, they can still simply take their shoes off and ask for a plastic bag at the door, just in case. Easy.
And yes, maybe the owner goes crazy or hates you, and you just bought new shoes that are squeaky clean, but brown, and get thrown out with no appeal. That can happen. But in that case, they could use many other things to be petty about, too. And Codeberg does not strike me as that.
How can you tell if you drank too much? What is the exact amount of alcohol molecules a given person could ingest in an exact instant (like, down to the exact Planck whatever) and no longer be fit to drive? And if someone says "no drunk driving on my property" while not caring what happens outside of it, can we really let them off the hook before they showed us the machine that can calculate it?
If people are worried about using libraries that contain vibe code, sure, that may become a real concern, but even then, why not have a really restricted website? For projects where the people are either so good, or the project so small, that that they know for sure there is no vibe code in it? Why this oozing over and into everything?
You cannot discuss in non-English languages on this site. Even though you could slur and degrade English, and therefore it being impossible to give you an exact heuristic right now that separates English from non-English exactly. The rule is still accepted, and cases dealt with as they come up.
8 replies →
It isn't a public service. It is one person offering his resources to help FOSS. He decides if you align with his mission.
1 reply →
If this opens more questions than it answers, then you are simply discovering that you don't understand copyright (which is ok, but not the fault of the Codeberg policy).
> You must not share projects that mostly consist [of LLM-generated code]
> Such projects having an unclear copyright status
There is no bright-line threshold at which a code contribution becomes copyrightable (and therefore relevant). It is a legal question determined by courts. However, nobody in practice has any difficulty determining whether their code is copyrightable.
Codeberg is essentially asking/demanding that "your code" coincides with "code you hold the copyright for". This responsibility can be delegated to other humans, but not to LLMs.
That is not what their wording says. They could have said "code that is mostly covered by copyright under German law", but they chose to say 'mostly consist of code written by "generative AI"-tools'.
Also, if you are a non-German user of Codeberg and there is a wide difference between what is covered by copyright in your country and Germany (e.g. if you are British) this might make Codeberg less of a suitable choice for you. What copyright laws matter to a particular user? Their country, the US because of its dominance and reach, or some sort of global safe/effective compromise?
This has the feel of being done by people who do not understand how to work things to be clear legally, nor of an understanding of the consequences.
They could have tagged these projects with "AI" or "MixedAI", and they'd be removed from the default views. It should also make clear that such projects may have a dubious legal status. Though within businesses no one cares?
They are also opinionated about what they host. FOSS only. Now updated to: unvibecoded FOSS.
It's run by a charity as a public good. Without a profit incentive from it, there is no benefit to public from slop.
> projects that mostly consist of code written by "generative AI"-tools
I guess if you auto complete line by line and actually read the code it should be gucci.
Edit: Oh, you found it as well now. Disregard my post.
"I guess if you auto complete line by line and actually read the code it should be gucci."
Not from their wording.
AI auto complete is still generative AI, why would that be any different? Unless you're talking about regular intellisense-like auto complete, which of course does not use generative AI at all, which would obviously be fine.
I’d say even GitHub copilot autocomplete is LLM aided coding. Copy pasting from a chat too is. But let’s not let the perfect be the enemy of the good. The point is the intent. We don’t want AI generated code. Put that in github if you so please.
This is for human generated code. Some day we will have a way to enforce the amount of AI usage on its participants, right now we can just rely on an honor code and clear intent that you’re not welcome here
What do you mean "I'd say"? It's obviously a fact. Github copilot uses an llm for its autocomplete, so obviously it falls under llm generated code.