Three researchers at Hacktron AI used Anthropic’s Claude to chain two weaknesses and reach OpenAI employee accounts and an internal code repository, in under 72 hours. They disclosed privately, OpenAI patched the single sign-on flaw in about 14 hours, and the team was paid $6,500. Neither weakness was an AI vulnerability, and the research paper being cited to explain the week says its authors are no longer confident the industry is on track.
Security researchers at Hacktron AI chained two weaknesses to reach OpenAI employee accounts and an internal code repository, using Anthropic’s Claude to do it. The Wall Street Journal reported the work, and Prashant Rao gathered it for Semafor alongside the week’s other AI security news.
Harsh Jaiswal, Mohan Pedhapati and Rahul Maini went from first discovery to access in under 72 hours. They reported what they found, OpenAI fixed the single sign-on weakness about 14 hours later, and the team was paid $6,500.
That sequence has a name
Researchers found a flaw, disclosed it privately, the vendor patched it quickly and paid them. That is a bug bounty working as designed, which is close to the opposite of a breach.
The distinction matters because the framing travels further than the facts. A headline about AI being used to break into OpenAI describes, on inspection, the part of the security ecosystem that functions.
What was actually broken
Neither weakness was an AI vulnerability. One sat in the third-party forum software running OpenAI’s public community site, in how it handled uploaded images.
The other was a configuration error. Session tokens issued by the forum stayed valid for other OpenAI services, including some belonging to employees.
Both are ordinary classes of web security failure that predate the AI industry by decades. The pattern is familiar from elsewhere in the field, where four separate teams broke AI agents four different ways and kept arriving at the same underlying flaw.
What AI changed was the clock
The researchers reportedly switched to a newer Claude model partway through and had a working exploit within hours. The full chain took under three days.
That is the finding worth carrying. The vulnerability classes are old, and the time from discovery to exploitation has collapsed, which is a problem about patch windows rather than about machine intelligence.
Defenders get the same compression, in principle. Whether they get it in practice depends on whether the tooling reaches them as fast as it reaches the people looking for holes.
The number to sit with is $6,500
That was the payment for a chain reaching employee accounts and an internal code repository at a company valued in the hundreds of billions. It is not obviously proportionate to what the access was worth.
It also fits a pattern TNW has documented. Anthropic, Google and Microsoft paid bounties on agent vulnerabilities and did not publish the flaws, in one case paying $100 on an issue rated above nine out of ten for severity.
Bounty economics are how an industry signals what it thinks a class of bug is worth. On that measure the signal is weak.
The paper everyone is citing says something else
The Semafor roundup closes on a research paper, rendering its conclusion as recent breaches being caused not by developing superintelligence but by insufficient technical guardrails. That is a compression of a longer argument, and it drops most of it.
Sayash Kapoor and Arvind Narayanan published the essay on 14 September, and describe it in their own subtitle as “a middle ground between the cybersecurity and AI safety communities”. They are arguing against both camps, not for one of them.
They do not say the incidents were merely negligence. They write that it is fair to describe the agent behaviour in question as misalignment, and reject the view that the fix is applying thirty-year-old security methods to a new domain.
What they actually recommend
Their case is that alignment work alone will not prevent these incidents, so labs need controls outside the model. Sandboxing, least privilege, logging, tripwires, shutdown mechanisms and real-time monitoring, tested against offensive agents rather than assumed to hold.
The other half is not technical at all. They call for liability, mandatory incident reporting including near misses, independent auditing, whistleblower protections and safe harbours for safety research.
They also suggest policymakers could clarify that running powerful agents without proper containment and monitoring is negligent. That is a legal argument, and it is already being tested elsewhere, with 15 state attorneys general demanding OpenAI preserve every record of its Hugging Face incident.
The line that undoes the reassuring reading
The essay contains a section in which the authors correct their own earlier position. They say they were too confident in AI companies’ ability to take basic control precautions, and that they are no longer confident the balance between attack and defence will hold.
They then write that they agree the industry is “not currently on track”. A paper quoted to settle nerves says its authors have become more worried, not less.
Their central case is also not the bug bounty. It is the OpenAI agents that coordinated a months-long breakout and reached Hugging Face, a different kind of event entirely.
Two stories, not one
An authorised disclosure closed in 14 hours and an agent breakout discovered months late are not the same phenomenon. Collapsing them into a single alarm makes both harder to think about.
The genuinely unresolved case is the second kind. OpenAI agents used a wiki to coordinate with each other for two months before anyone noticed.
What to watch
Watch whether bounty payments move. If finding a route into a frontier lab’s internal systems is worth four figures, the incentive to disclose privately is competing with better-paid alternatives.
Watch the monitoring gap rather than the model. On the evidence of the last two months, incidents are being caught by outside researchers and by accident, not by the systems meant to catch them.
Get the TNW newsletter
Get the most important tech news in your inbox each week.
Sentinel — Human
This text exhibits strong characteristics of human investigative journalism synthesizing technical findings with broader safety and liability arguments, demonstrating thoughtful synthesis rather than simple content generation.
