.
@1Password's FLAWED report says AI models produce a clean security fix only 26% of the time.
Defenders shouldn't take that number seriously.
• The six vulnerabilities were handpicked because their fixes were complex. Clean-fix rates ran from 3% to 60% depending on the bug, and the report averaged them together.
• Agents set up to fail were counted in the headline figure. Two of 1Password's prompts instructed the agent to apply the wrong fix. Those trials make up 22% of the data. One evaluation mode prevented the agent from compiling or running any code, and it accounts for 36% of the data.
• The report ran two models, GPT-5.5 at medium effort and Opus 4.8 at high. Neither was tested at its highest available setting, so the report says nothing about how more effort or stronger models change the results.
• Several instruction and grading errors further undercut the headline, and are elaborated upon in the attached blog.
We've spent four months submitting hundreds of AI-authored patches to widely adopted open-source projects as part of Patch the Planet.
Our experience didn't match 1Password's report, so we did a full analysis across 186 AI-authored pull requests and 33,500 subsequent commits, benchmarked against 2,265 human-authored patches we graded across years of security engagements.
blog.trailofbits.com/2026/09…