rachid chabane.
Search
← All articles
Essays · agentic coding · agent-maintained

The premise of your LLM policy decides what it lets through

If you are writing an LLM contribution policy for your repository, the premise you start from decides where the rule will leak. Three toolchain projects published theirs…

08-08-2026 8 min ███░░ FR / EN
agentic codingquality

If you are writing an LLM contribution policy for your repository, the premise you start from decides where the rule will leak. Three toolchain projects published theirs recently, each reasoning from a different starting point, and all three landed on the same instrument. The instrument was forced. What sits underneath it is a choice, and that choice is the thing you inherit.

Three premises, one instrument

Rust starts from review capacity. The project has long had more people who want to write code than people willing to review it, and the arrival of LLMs makes that worse 1. The rule it derived permits using an LLM to answer questions, analyze, distill, refine, check, suggest and review, stops short of creation, and makes some of the permitted uses disclosable 2.

OpenJDK starts somewhere else entirely. The Oracle Contributor Agreement requires a contributor to own the intellectual property rights in each contribution and to be able to grant them to Oracle without restriction; most generative AI tools are trained on copyrighted and licensed content, their output can include content that infringes those copyrights and licenses, and whether the user of such a tool holds IP rights in its output is the subject of active litigation 6. That premise leaves one conclusion available. Contributions in the OpenJDK Community must not include content generated in part or in full by large language models, diffusion models or similar deep-learning systems, across repositories, pull requests, e-mail, wiki pages and issues, while private use to comprehend, debug and review code stays allowed 5.

GCC inherits its premise from copyright law as the GNU Project already reads it. The policy declines legally significant contributions that include or are derived from LLM-generated content, takes its definition of legally significant from the GNU Project maintainer guidelines, and still leaves GCC maintainers free to accept generated test cases 8.

Same instrument at the end of all three: what the author declares about the change.

ProjectPremise it starts fromWhat the rule forbidsWhere I expect it to leak
Rustreview capacity 1creating content with an LLM, while assistance stays allowed and sometimes needs disclosure 2at the boundary between refining a draft and producing one, which only the author can see
OpenJDKintellectual-property exposure under the contributor agreement 6contributing any LLM-generated content at all 5at verification, since a total ban is a norm with no check behind it
GCCcopyright significance, at the GNU threshold 8legally significant LLM-derived contributions 8underneath the threshold, where small generated hunks stay compliant

The fourth column is my judgment, and the rest of this piece is about it.

The declaration was the only object available

Let me concede the obvious part before anyone presents it as a finding. Origin leaves no reliable trace in a diff, so a rule about where a contribution came from has to attach to something the submitter says. Open source has worked this way for two decades: the Developer Certificate of Origin, contributor licence agreements, every plagiarism rule, all of them unverifiable self-attestations enforced after the fact. Three LLM policies converging on the author’s declaration was forced by that constraint, and calling it a discovery would be generous.

What the constraint forces downstream is more interesting. Rust spells it out for reviewers: the responsibility for determining whether a PR is LLM-generated sits with the author rather than the reviewer, a PR template will ask authors directly, uncertain cases go privately to moderation, and style is not evidence 4. In my experience that last clause is the one doing the work, because it disqualifies the inference a tired reviewer reaches for first.

Each premise picks a different leak

A threshold rule leaks below the threshold. GCC’s policy borrows the GNU Project maintainer guidelines’ line for legally significant content, around 15 lines of code and/or text, and declines LLM-derived contributions that clear it 8. It also leaves research, analysis, bug discovery and reporting, and patch review untouched, so long as the output stays out of the contribution 9. That is a coherent rule with a measurable edge, and contributors optimize against measurable edges.

A licensing premise leaks in the opposite direction. The rule it produces is a flat prohibition on generated content in contributions 5. In my experience the gap shows up in one specific way: a contributor generates a patch, retypes it by hand, and submits something whose keystrokes are theirs while its origin is the model. No step in a review pipeline separates that patch from one written from scratch, so what the project holds is a strong norm and a form.

Judgment decides where a capacity premise leaks, and only the author has access to it. Rust permits an LLM to check, refine and review while stopping short of creation 2. In my experience the boundary between improving a draft and producing one is not a boundary at all: two honest contributors will place it differently on the same pull request, and both will file a truthful declaration.

The strongest case against this

The strongest objection is that I have written at length about an analytic truth. If origin is undetectable, a declaration is the only available object, so three projects converging on declarations follows from the problem rather than from anything they discovered. On that reading, the divergent-premises framing is decoration over a conclusion with no alternative available.

That objection is right about the form and quiet about the scope. The form was forced; the scope was chosen. Nothing in “origin is undetectable” tells you whether to ban generated content, permit assisted work under disclosure, or draw the line at a copyright threshold, and those three rules fail in three different places. A maintainer writing a policy this quarter has to pick one, and the premise they pick decides which failure they get.

There is a second thing the objection leaves standing, and it is why Rust repays a close read.

The Developer Certificate of Origin never argued for itself in those terms. Rust does.

The rule taxes the thing it protects

These policies exist to protect reviewer attention, and the pressure is measurable. At the time of writing there are 1,281 open PRs to rust-lang/rust, a staggering amount of time invested by both authors and reviewers 1. The same pressure shows up from the other side of the review: generative AI tools make it easy to create large quantities of plausible-looking code with plausible-looking tests that is nonetheless incorrect or poorly designed, and reviewing such submissions can easily become a drain on the already limited time of human reviewers 7.

Now price the rules against that. Each of them adds a step to a review that did not have one: read the declaration, decide whether you believe it, decide what to do when you do not. I think this is the part nobody has costed, because the cost of producing a submission keeps falling while the cost of checking a claim about it does not. The lever these projects reached for spends the resource it was built to defend.

What I would put in my own CONTRIBUTING.md

Here is what I would ship. One declaration field in the pull-request template. Rust is adding one that asks authors whether their code was LLM-generated 4. I would phrase mine about this change rather than about the contributor’s habits, because habits are not what a reviewer needs to know. One machine-checkable commit trailer, LLM-Generated: yes|no|assisted, so that git interpret-trailers --parse turns a policy question into a field CI can require and a maintainer can grep six months later. And one sentence at the top of the file naming the premise, so whoever edits the policy next knows which failure mode they are trading against.

What I would not ship is a line count. A threshold buys legal defensibility and hands every contributor a way to stay compliant while doing the exact thing the rule exists to discourage. My preference is a rule I cannot enforce paired with an audit trail I can read.

What I am watching

The signal that would change my mind is what the declaration fields come back saying. If they read overwhelmingly negative on projects whose review queues keep growing, the field has become paperwork and the bargain failed. If maintainers report that it changes how they triage, the unenforceable rule earned its place. I expect the first serious argument to be about assisted work rather than generated work, because that boundary lives inside the author’s head and no policy text can move it somewhere a reviewer can see.

Glossary

Coding agents
LLM-driven agents that read, write and refactor code through tool calls (editor, shell, tests), shifting the economics of who reads code and what conventions are worth their cost.
Legally significant threshold
A size line above which a contribution counts as copyrightable and therefore falls under a project's licensing rules; the GNU Project maintainer guidelines put it at around 15 lines of code or text. Using it as the edge of a policy makes the rule measurable, and in the same move hands contributors a compliant way to stay just below it.
LLM contribution policy
A repository rule stating whether, and under what conditions, content produced with a large language model may enter a project's contributions. Its scope follows the premise it is derived from, whether legal exposure, review capacity or a copyright threshold, so two projects reaching for the same instrument can forbid very different things.
Origin attestation
A statement by the submitter about where a contribution came from, recorded in a sign-off, a pull-request field or a commit trailer. Origin leaves no reliable trace in a diff, so the attestation is unverifiable at review time and is enforced only after the fact, the way the Developer Certificate of Origin is.
Review bandwidth
The reviewer attention a project can actually supply, the scarce resource in most open-source repositories, where more people want to write code than to review it. Any rule that adds a step to review spends the same resource it is usually written to protect.
Rule declared unenforceable
A policy rule that names its own unenforceability as a deliberate choice rather than a defect: it aims at one clear, publicly stated line instead of at catching every violation. Violations are then identified from observable actions, with intent weighed only when deciding how to respond.

Sources

01
05-08-2026 blog.rust-lang.org
02

Want to go deeper?