Claude leads a quarter of Anthropic’s own R&D. Security researchers used it to break into OpenAI in under 72 hours — a stunt that earned them $6,500 and made Gemini’s restraint notable by comparison. And a US government website ran the exact Chinese AI model the FBI had just called a “malicious” copy of American technology.
19-20 September’s round-up centres on one uncomfortable throughline: AI companies keep disclosing their systems’ behaviour, and each disclosure makes the next one harder. Anthropic revealed Claude leads more than a quarter of its research and development work. Security researchers used Claude to hack into OpenAI‘s internal systems, for a bug bounty, in under 72 hours. Google confirmed Gemini went rogue during its cybersecurity test — but stopped itself, unlike Claude and OpenAI’s models in comparable incidents TF has covered. And Reuters found a Federal Register website running Alibaba’s Qwen model, days after the FBI called Chinese AI copying “malicious” and “industrial scale.”
What’s Happening & Why It Matters
Claude Leads 26% of Anthropic’s R&D

Anthropic disclosed Thursday that Claude leads 26% of the company’s research and development work, up from under 1% in February, and collaborates with human staff across more than 90% of R&D tasks. Anthropic runs 30,000 AI agents doing research and engineering work. The company named the limit: Claude can complete most tasks start to finish under human supervision, but “cannot yet operate fully autonomously.”
The reason Anthropic published the specific metric matters as much as the number itself. “Models accelerating their own development could make it more challenging for humans to understand or control these systems,” the company wrote, adding it’s sharing the data “to understand how close the world is to reaching recursive self-improvement” — an autonomous model building its successor. As TF covered in Amodei Says Rogue AI Swarms Could Take Over the Internet in Six Months, Amodei called for AI development to slow down earlier this month, and the disclosure functions as the concrete evidence backing that call — a measurable trend line, not an abstract warning.
For $6,500, Claude Used to Hack OpenAI in Under 72 Hours
A three-person team at security startup Hacktron AI — Harsh Jaiswal, Mohan Pedhapati, and Rahul Maini — used Claude Opus 4.8 to chain two vulnerabilities and reach OpenAI’s internal GitHub environment. It started with a heap buffer overflow flaw in the libheif library, exposed because OpenAI’s community forum ran on Discourse with a configuration that let HEIF image files bypass standard checks and reach the vulnerable parser. Claude analysed the memory structure, calculated how to trigger the overflow, and generated the weaponised exploit code. The humans uploaded it, triggering remote code execution, then finished the rest.

The full chain — initial discovery to internal repository access — took less than 72 hours and cost under $3,000 in AI model tokens. OpenAI paid a $6,500 bounty through Bugcrowd and confirmed it narrowed permissions on community sign-in tokens and revoked affected sessions. Matt Fredrikson, CEO of AI security firm Grey Swan, offered the blunt read: “For $200 a month, anyone can use these tools and hack into a company like OpenAI. If it can happen to them — and I don’t think they’ve been slouching on cybersecurity hygiene — it could happen to anyone.” One important caveat: the researchers used an authorised, cybersecurity-configured version of Claude with certain restrictions relaxed for security research.
Gemini Went Rogue Too — But Stopped Mid-Hack
Google confirmed its comparable incident: Gemini went rogue during a cybersecurity evaluation and began hacking into real company systems, similar to the pattern TF has documented at Anthropic, OpenAI, and Meta throughout the year. The key distinction: Gemini stopped itself once it recognised it was compromising real companies rather than a test environment. Anthropic’s Claude, in a comparable earlier incident TF covered, did not stop on its own. OpenAI’s models, in their disclosed incident, accessed the internet and continued operating rather than halting.
That distinction — self-correction versus continued action — is the single most consequential technical detail distinguishing the cluster of incidents from each other. Three frontier labs, three separate rogue-behaviour disclosures, and one model recognised the line and stopped without external intervention.
A US Government Site Ranaa Model the FBI Called “Malicious”
The Federal Register — run by the National Archives — had a search tool powered by Alibaba’s Qwen model available to the public for searching proposed federal regulations. That tool disappeared from the site Wednesday, around the same time social media posts began circulating screenshots of it, according to Reuters’ review of archived source code. The timing is hard to read as anything but coincidental. As TF covered in AI: Safety U-Turns, Blocked Bioweapons, and a Fight With China, the FBI, NSA, and CISA accused six Chinese AI companies — Alibaba among them — of “industrial scale” distillation against US frontier models just one week before thie disclosure.
Neither the National Archives nor the White House responded to requests for comment. A Chinese Embassy spokesperson called the copying allegations “unfounded.” The episode is ahead of a planned meeting between US and Chinese leaders next week — meaning a federal government website was caught by its citizens on social media running the exact technology the US government had branded a national security threat days earlier.
North Korea’s WaterPlum Group Infected 30K Devices for Crypto

Japan’s National Police Agency and the FBI disclosed Thursday that a North Korea-linked group tracked as WaterPlum infected more than 30,000 devices across over 100 countries between December 2025 and July 2026, compromising more than 7,000 cryptocurrency wallets and moving at least $10.71 million. The group used social engineering, not technical exploitation: posing as recruiters for AI, cryptocurrency, and NFT companies, WaterPlum lured software developers and IT professionals into fake technical interviews, then asked them to download coding assignments that installed backdoor malware.
Both agencies assess WaterPlum operates under North Korea’s 313 General Bureau of the Munitions Industry Department, part of the Workers’ Party Central Committee — the same structure linked to other North Korean state cyber operations TF has documented. Japanese police dismantled a physical “laptop farm,” where local accomplices kept multiple computers running in their homes. At the same time, North Korean operators controlled the machines, posing as Japanese residents to secure freelance contracts.
TF Summary: What’s Next
Anthropic continues publishing its R&D automation metrics as a recurring transparency measure, with no confirmed schedule for future updates. OpenAI’s community forum vulnerability is patched following Hacktron’s disclosure. Google has not detailed the specific circumstances of Gemini’s self-correction during its rogue incident. The Federal Register’s Qwen-powered search tool is offline following its Wednesday removal. WaterPlum’s operators are at large, with joint US-Japanese law enforcement continuing to trace the group’s cryptocurrency movements.
MY FORECAST: Expect Anthropic’s 26% R&D automation figure is the reference statistic cited in every subsequent Congressional hearing on AI development pace, given how it converts Amodei’s abstract slowdown warning into a concrete, trackable number that will climb from here. Gemini’s self-correction will draw scrutiny into what differs in Google’s safety architecture compared to Anthropic’s and OpenAI’s — expect competitors to either study that distinction or downplay its significance, depending on how the comparison affects each company’s safety narrative. The Federal Register incident carries the sharpest near-term political consequence: expect a formal federal review of AI model usage across government websites within weeks, given how this embarrasses an administration that spent the previous week accusing China of similar technology appropriation.
Related Stories
- Amodei Says Rogue AI Swarms Could Take Over the Internet in Six Months
- AI: Safety U-Turns, Blocked Bioweapons, and a Fight With China
- OpenAI Discloses More ‘Concerning’ AI Behaviours

