I do not quote a patch-quality rate until I know the grader that produced it would label the same patch the same way twice. That sounds like pedantry until you try to use one of these numbers, and the number in front of everyone is this: across six recently disclosed CVEs, 1Password produced 6,080 patches with two frontier, cyber-capable reasoning models, and the average success rate for a patch that fully resolved the vulnerability without materially changing application behavior was 26.0% 1.
The companion figure travels further. Patches that did not resolve the vulnerability, added a new one, or both, an average 53.9% of the time 2. Both figures are counts of labels, and I read them as outputs of a labelling step that nobody in this argument has shown to be reproducible. Until that step is characterised, neither number tells a security team where to put its review capacity.
What the headline pools together
The usual objection to a benchmark is that it is too small, too easy, or too old. The objection worth making here is a different one: the number answers several questions at once and is reported as though it answered a single one. Trail of Bits reviewed 1Password’s code and data and reports four choices that make the 26% clean-fix rate a misleading guide to ordinary patching work 3. That word is theirs, quoted as their characterisation and not adopted as mine.
Two of them are about what got pooled into the average. The headline includes experiments that deliberately instructed agents to apply the wrong fix, along with experiments in which the agents could not compile or test what they had written 4. Two prompts tell agents to apply the wrong fix, and those prompts account for 22% of the data 5. One evaluation mode prevents agents from building or running code, and it accounts for 36% of the data 6.
Whether those two slices overlap is not something the review states, so putting them together would be my arithmetic and not a finding of either party. The shares are not the point. Three experiments asking three questions produce one average, and that average is then read as a capability measurement.
| Slice of the trials | What the review reports | What I read it as |
|---|---|---|
| Prompts that instruct the wrong fix | two prompts, 22% of the data 5 | how often the researchers chose to give bad advice |
| The mode that blocks build and test | one evaluation mode, 36% of the data 6 | patch writing with the feedback loop removed |
| Everything averaged into one figure | the headline combines those trials with trials that could test 4 | a number that answers no single question |
The defect is aggregation.
The result hiding in the same data
A reanalysis that only pushes a number down is a quibble. This one pushes a number up, out of data the study itself published. In the trials Trail of Bits examined, 2,634 of 3,067 patches generated by 1Password’s models, 86%, blocked the supplied exploit 7.
Set that beside the headline and the temptation is immediate: one of these rates must be wrong. Neither is. Blocking a supplied exploit and fully resolving a vulnerability without changing behavior are different predicates, and the higher figure is not a clean-fix rate in disguise. A benchmark reporting both, with the predicate attached to each, would beat one reporting either alone.
The second finding is quieter and does more damage. The six-target mean carries a standard error of about nine percentage points, which the report does not disclose 8. An undisclosed spread of that size, on a mean taken over six targets, is why I treat a single headline percentage from this study as a range I have not been shown.
A grader that does not agree with itself
The dispute over this number keeps circling back to whether 26% is right. In my experience the question that settles such arguments is duller: grade the same patches twice and find out whether the labels hold still. On that question the published material is more forthcoming than the headline suggests, and what it shows is not reassuring.
The two models assigned different outcomes to 36.8% of the same patches 13. Models grading their own patches matched human reviewers on the full five-category outcome in 65.9% of reviewed cases 12. And on the defect class an automated grader should be best placed to catch, the authors found 248 generated patches that repeated an off-by-one error already present in the upstream fix, with the grader flagging that new vulnerability in only 24 of them 14.
A labelling step that disagrees with itself that often is not producing a rate. It is producing one draw from a distribution neither party has characterised, and the headline is the average of two such draws. That is the claim in this piece a reader can genuinely refuse: you could hold that full-outcome agreement at that level is plenty for a directional read.
The human number sits on a different ruler
Every summary of this dispute I have read subtracts one rate from the other, because the two look like the same kind of quantity. They are not, and the conditions attached to the human figure do as much work as the figure: the developers had written the software they were patching, held detailed reports from the reviewing engineers, and knew a reviewer was coming. Those are close to best case for a human author, and nothing like the benchmark’s.
The polarity is the trap. One figure counts the fixes that came out clean; the other counts the fixes that came out wrong. Subtracting them produces a gap that is an artifact of the subtraction. I read the human record as most first fixes landing at the first attempt under those conditions, which is worth knowing and is not a number to set beside a benchmark.
One result should make anyone cautious about treating human and agent failure as different in kind: two authors, one human and one agent, working separately, made the same mistake on the same bug 11. Trail of Bits supplies the only human baseline on the record here, which is a statement about this exchange and not about what exists elsewhere. What neither party ran is a control arm: the same tasks, both kinds of author, one grading procedure.
The strongest case for taking the headline at face value
Here is the objection at full strength, and it is a good one. Full-outcome agreement in the mid-sixties is plenty for a directional read; grader disagreement that goes both ways cancels in an average rather than biasing it; and the critic’s own record shows that agent-written patches need real review before they land. On that reading the headline is roughly right and demanding grader statistics is perfectionism.
The critic’s record is the strongest part of that case. Maintainers merged 126 of Trail of Bits’s 186 pull requests, an acceptance rate of 67.7% 15. In 91 of those 126, 72.2%, maintainers accepted the security fix as originally proposed 16. Those are real patches, reviewed by people with no stake in the argument.
I concede both figures. Neither is in doubt here, and neither needs to be disputed for the argument to hold, because the answer is narrower than the objection assumes. A directional read is exactly what a percentage stops supporting once its labels move with the grader: direction is what an average of two disagreeing labellers is least able to fix, since cancellation is an assumption about the disagreement and not a measurement of it. I am not asking for a bigger sample, a harder benchmark, or a replication. I would demand the agreement number, which costs a fraction of the study that has already been run.
What the follow-up audit found, and what it could not
The most interesting evidence in this exchange is the part where a party goes looking for its own mistakes. The shape of that search decides what the finding is worth. Trail of Bits examined about 33,500 subsequent commits in Patch the Planet projects, and where a later commit touched a file one of its patches had modified, it investigated whether that commit fixed a problem the patch had introduced 17. The review found at least ten functional bugs, four build, test or release automation bugs, and one performance bug, and it found no exploitable security vulnerabilities 18.
Notice what has happened to the evidence. From the merge figures onward, the party doing the measuring is the party whose patches are being measured. I think that is worth stating plainly rather than treating as disqualifying: the search procedure is described, which lets a reader weigh it, and that is more disclosure than the headline offers about its own labelling.
A null result of that kind is still a statement about what one search found. I read it as bounded by the sweep that produced it, which looked at later commits touching patched files, and would not surface a defect nobody has committed against yet.
What I would demand before quoting a rate
Trail of Bits states that readers deciding whether to use agents need to know how their failures compare with those of human developers, and that establishing which is more reliable requires measuring both under comparable conditions 19. That is their requirement and I agree with it. Mine is smaller and comes first, because a comparison between two unstable measurements is not worth running.
outcome_rate: <the headline, and the predicate it counts>
grader_agreement: <two graders on the same patches, or not published>
human_agreement: <grader against human reviewers, or not published>
working_conditions: <reported separately, or pooled into one average>
task_set: <the same set for both arms, or two different sets>
verdict: <quotable | quotable with its agreement number | not a rate yet>
That checklist is mine and comes from no source. Run it over the next patch-quality study you are asked to act on and you will usually stop at the second line, which is itself the finding.
I have not re-run either party’s analysis, and nothing here disputes a figure either one published. This is a rule for reading published rates, and it changes what I do when a vendor number lands and someone asks whether agents can be trusted with security patches. My question now is whether the thing that produced the number would produce it again.