AI Finding Serious Bugs in Leading Software Titles

Li Nguyen

10,000 critical vulnerabilities in a single month. A 6% patch rate on the ones already disclosed. Some maintainers have asked Anthropic to slow down — they can’t keep up with the pace of their own security fixes.


Claude Mythos’s bug-finding pace has outrun the industry’s ability to patch what it finds, according to documents reviewed by ProPublica describing Microsoft’s internal “mad dash” to close holes before attackers exploit them. Anthropic’s Project Glasswing — a controlled early-access programme giving an estimated 50 partner organisations access to Mythos’s cybersecurity capabilities — reported more than 10,000 high- or critical-severity vulnerabilities discovered across partner systems in a single month. Only a fraction have been patched so far.

What’s Happening & Why It Matters

The Numbers Behind the Warning

As of late May, Anthropic had disclosed 1,596 vetted vulnerabilities to maintainers across 281 projects. Just 97 had been patched — roughly a 6% fix rate. The standard 90-day coordinated disclosure window built into most vulnerability disclosure programmes was designed for human-speed discovery, not for a model capable of scanning 1,000 codebases in a single month. Therefore, this is a challenge. Anthropic says the average time to patch a high- or critical-severity bug found through Glasswing runs about two weeks. However, that average hides a growing backlog behind it.

Cloudflare, one of the partner organisations, flagged 2,000 bugs, 400 of them high or critical, with a false positive rate that beat human testers outright. Mozilla found and fixed 271 vulnerabilities in Firefox 150 — more than ten times what its predecessor model, Claude Opus 4.6, caught in the prior Firefox release. The capability gain is real. So is the backlog it’s creating.

Can Anthropic Slow Down?

Several open-source maintainers have asked Anthropic to reduce its disclosure rate. They can’t design, test, and ship patches fast enough to match the discovery pace. That’s an unusual request in security research, where the industry norm has always strived toward faster disclosure, not slower. It signals something structural: the bottleneck isn’t finding bugs anymore. It’s fixing them.

Vinh Nguyen, a senior technical adviser to Anthropic and senior fellow at the Council on Foreign Relations, put the deeper risk plainly. “The problem now is that you can chain four low-level flaws, and that can equal a high severity,” he said. “If you’re Microsoft, the current triage strategy may be underpricing risks.” Low-severity bugs that sat safely unpatched for months under the old model combine into genuine attack paths. Still, most triage systems treat them individually.

The Five Eyes Warning

Since Project Glasswing became public in April, national security experts have described a specific window of opportunity. The US and its allies could patch known flaws before adversaries built comparable AI bug-hunting tools of their own. In late June, the Five Eyes intelligence alliance — the US, UK, Canada, Australia, and New Zealand — issued a rare joint statement. Notably, they warned that window would close within months, not years.

That warning has already somewhat materialised. Microsoft is building its own competing AI bug-detection product, using its MDASH system alongside multiple models running to find, verify, and generate fixes for vulnerabilities — a direct response to Mythos, according to The Information. IBM and Red Hat have done more. They launched Project Lightwell with $5 billion to patch open-source vulnerabilities before AI tools discover and weaponise them.

The Invisible Vulnerability Problem

IBM projects roughly 59,000 catalogued CVEs will be filed in 2026. One analyst estimates a 500,000 security fixes happen every year, made by open-source maintainers without ever receiving an official vulnerability designation. “Frontier AI models can chain those undisclosed low- and medium-severity fixes into novel exploits,” she warned. “The invisible vulnerability universe is not just a transparency gap — it’s a latent attack surface.”

The 10,000-vulnerability figure Anthropic reported covers only what Glasswing found and catalogued. The unknown universe of quietly patched, undocumented flaws is larger. Moreover, it’s the kind of dataset a capable AI model could mine to build attack chains nobody has mapped yet.

TF Summary: What’s Next

Project Glasswing continues expanding its partner list, with IBM joining on May 19 ahead of its own Lightwell launch. Microsoft’s competing bug-detection product is in development, with no confirmed release date. The Five Eyes alliance has not detailed specific coordinated action beyond its joint warning. Anthropic continues hiring security researchers to scale its disclosure and patching support.

MY FORECAST: The patch-rate gap won’t close through faster human triage alone — expect the industry to shift toward AI-assisted patching within 12 months, not just AI-assisted discovery. Microsoft’s MDASH-based product and IBM’s Lightwell initiative are early signals of that shift. Both will need to prove they can generate verified fixes at something close to Mythos’s discovery speed. The harder problem is the invisible vulnerability backlog — the hundreds of thousands of undocumented fixes nobody’s tracking. Whoever builds the first system to map that dataset systematically, attacker or defender, changes the balance of the race.



[gspeech type=full]

Share This Article
Avatar photo
By Li Nguyen “TF Emerging Tech”
Background:
Liam ‘Li’ Nguyen is a persona characterized by his deep involvement in the world of emerging technologies and entrepreneurship. With a Master's degree in Computer Science specializing in Artificial Intelligence, Li transitioned from academia to the entrepreneurial world. He co-founded a startup focused on IoT solutions, where he gained invaluable experience in navigating the tech startup ecosystem. His passion lies in exploring and demystifying the latest trends in AI, blockchain, and IoT
Leave a comment