Home / Technology

Photo of artificial intelligence, network, robot
Image: via cdn.mos.cms.futurecdn.net
Technology

AI Security Patches Fail Majority of Tests, Researchers Find

WireByte Staff · August 7, 2026

Researchers from 1Password’s Off-by-1 Labs tested thousands of AI-generated security patches using frontier models and found that only a quarter fully resolved the targeted vulnerabilities. Many attempts failed to stop existing exploit paths, altered application behavior, or introduced new security flaws, raising concerns regarding the reliability of generative artificial intelligence in cybersecurity operations.

Key points

  • Researchers at 1Password’s Off-by-1 Labs analyzed security patches generated by two frontier AI reasoning models, ChatGPT 5.5 and Claude Opus 4.8.
  • The experiment tested 6,080 AI-generated patches across six recently disclosed Common Vulnerabilities and Exposures (CVEs).
  • Only 26% of the proposed patches successfully and fully resolved the underlying security issues.
  • Nearly half of the patches failed to fix existing exploit paths, while others altered application behavior or introduced new vulnerabilities.
  • Experts warn that relying on generative artificial intelligence for automated software patching remains highly unreliable and potentially risky.

A recent study by researchers at 1Password’s Off-by-1 Labs reveals significant limitations in the use of generative artificial intelligence for cybersecurity. Investigators tested thousands of AI-generated security patches to determine their effectiveness in addressing software vulnerabilities.

The experiment utilized two frontier reasoning models, ChatGPT 5.5 and Claude Opus 4.8, to generate 6,080 patches targeting six recently disclosed Common Vulnerabilities and Exposures. Out of all the proposed solutions, just 26% fully resolved the targeted security flaws.

Furthermore, the analysis showed that 49.3% of the patches failed to fix at least one existing exploit path. An additional 20.1% of the fixes resolved the original issue but altered application behavior, while 2.3% introduced entirely new security vulnerabilities. The findings highlight substantial risks for organizations considering automated LLM-based patching for critical infrastructure.

Sources

WireByte Staff — Editorial Team

The WireByte editorial team synthesises technology news from multiple primary sources, verifies the facts, and links every source. Articles are produced with AI assistance and reviewed under our editorial policy.