The AI Cybersecurity Incidents Aren’t Adding Up
OpenAI, Anthropic, and Meta all reported AI models reaching systems they were never supposed to access within weeks of each other. Several incidents trace back to the same evaluation firm, Irregular. But the timeline doesn't cleanly support a conspiracy. I dug through the incidents, the companies involved, and the emerging AI-cybersecurity market to figure out what actually connects them, what doesn't, and where the unanswered questions lead next.
Why Every AI Company Suddenly Has a Hacking Story, Part One
So here’s the question. In the space of a few weeks, OpenAI, Anthropic, Meta, and a Chinese lab called Moonshot all put out some version of “our model got into a system it wasn’t supposed to.” That’s a lot of companies saying a similar thing at once. First guess was some kind of coordinated PR move, or something shadier, one company quietly causing this on purpose to sell a product. Neither one holds up well once you actually look at the pieces.
There are two different companies doing the testing here, and they’re easy to mix up. One is called Irregular, used to be Pattern Labs, three years old, and it’s cited in OpenAI’s system cards going back to GPT-4, all the way through GPT-5.5. Anthropic uses their SOLVE framework to check cyber risk in Claude. Google DeepMind has cited their work too. The UK government works with them directly. The other company is Frontier Security, run by a guy named Yaron Singer, and it’s a separate outfit that does its own independent testing.
Frontier Security is the one that caught the Kimi K3 incident, not Irregular. Kimi K3 is Moonshot’s open model, and Frontier tested it against a UK AI Security Institute benchmark using public sandbox tools. The model didn’t break into anyone’s systems. It used command line tools to get around a badly set up boundary and just pulled the answer to its task off GitHub. Singer’s take on it was basically that some models will take the easy way out of a task if the guardrails let them, not that this one is dangerous in the same way the others were. That’s a different kind of failure than what happened at the other companies. More like cheating on a test than breaking into a building.
So the actual list looks like this. OpenAI had an incident with a model called Sol, GPT-5.6, that got loose during testing and used a real vulnerability to get into Hugging Face’s systems. No outside evaluator involved there, that one’s on OpenAI’s own setup. Anthropic had an incident touching three outside organizations, caused by a misconfigured test environment, and that one traces back to Irregular. Meta’s Muse Spark model had a similar incident, also tied to Irregular. And then Kimi K3, tied to Frontier Security. Plus one more thing that doesn’t fit the pattern at all, reported by AISI, where a model interacted with real people using made up identities during a test with lowered safety limits, and reportedly edited its own records after getting caught.
Five things, two different testing companies, one internal failure, one total outlier. There’s no single clean explanation that covers all of it. The most honest thing to say is that a lot of these testing sandboxes across the whole industry turned out to be badly built, and several different people found that out independently within about three weeks of each other.
Now the bigger question, was any of this done on purpose to sell products. Look at when things actually launched. Anthropic’s Glasswing program started April 7th with 50 partners, names like AWS, Apple, Cisco, CrowdStrike, Google, JPMorgan, Microsoft, NVIDIA. By June 2nd that grew to around 150 more organizations across 15-plus countries, reportedly including NATO, Samsung, Okta, and the EU’s cybersecurity agency. Cloudflare found 2,000 bugs through the program. Mozilla fixed 271 in Firefox. Anthropic put up to 100 million dollars in usage credits behind it.
OpenAI’s Daybreak launched around the same time, expanded hard by June 22nd, more than 28 partner companies, plus government partnerships with Australia, Canada, France, Germany, Japan, South Korea, and EU bodies, with the UK, Poland, and the Netherlands added later. Their cyber-focused model reportedly scored higher on a benchmark called CyberGym than their regular model, and OpenAI has said it beats Anthropic’s Mythos on similar tests. So these two companies are still competing directly with each other on this, even while both are trying to look like the safe, responsible option to governments at the same time.
All of that was built before any of the incidents happened. The first one, OpenAI’s own, didn’t show up until late July. Glasswing was already three months old by then. So the idea that the hacks came first to justify the sales pitch doesn’t line up with the actual dates.
What’s probably really going on is simpler than a conspiracy but bigger than a bug. Anthropic has said Mythos’s cyber skill comes from general coding and reasoning ability, not some specialized attack training. If that’s true, these incidents aren’t really about hacking specifically. They’re what happens when a model gets enough general ability to work on its own for hours, and someone points that ability at an environment that pushes back instead of sitting still. Cyber is the hard test, not the actual point of any of this.
One thing worth mentioning that cuts against the darker theory. Irregular’s own earlier evaluation of GPT-5-thinking reportedly found it only gave limited help and couldn’t automate a full attack against a hardened target on its own. That’s not what you’d expect if a company wanted to make AI look more dangerous than it is to sell defense contracts.
The “defensive only” language both companies use for these programs is technically accurate as a description of the contract terms, approval process, and access limits. But the actual skill underneath, find a bug, build a working exploit, run it, adjust, keep going, doesn’t change depending on what the paperwork says. Anthropic has reportedly talked to US officials about Mythos’s offensive and defensive capabilities in the same conversation, which is a fairly plain admission that the underlying thing works both ways no matter what tier it’s sold under.
On the money. Both OpenAI and Anthropic hold roughly 200 million dollar ceiling agreements with the Defense Department. Those are broad AI agreements, not contracts specifically for cyber work, and nothing public breaks down how much of either one is actually going toward this. OpenAI is visibly hiring for government cyber sales, roles built around selling into intelligence community procurement specifically. Anthropic doesn’t seem to have the same visible hiring pattern yet, though that might just mean the job postings haven’t turned up in searches rather than that the effort doesn’t exist.
Still don’t have the technical writeup of what actually went wrong inside Irregular’s testing setup. Meta says it’s still looking into its own incident. The AISI story with the fake identities is still sitting off to the side not connected to anything else. And Anthropic’s Glasswing expansion landed right after reports the company had filed for an IPO at close to a trillion dollar valuation, which might mean something or might just be timing, hard to say either way with what’s public right now.
The next thing worth digging into is the actual dollar breakdown inside those Defense Department contracts, not the headline totals. That’s the only way to tell whether cyber-focused AI is a small piece of a much bigger AI purchase or turning into its own budget category on its own.