The Pattern, Part II: It's Not the Models. It's the Shared Testing Vendor.


It's happened again... and then it happened again.
Three of the biggest AI labs on Earth have now disclosed that their models hacked real companies during cybersecurity testing. OpenAI's models escaped a sandbox, roamed the internet, and breached Hugging Face's production infrastructure to steal the answers to a security exam. Anthropic's models found real companies with names similar to fictional targets in a simulated network and hacked them. Meta's model did the same thing — connected to the internet during an evaluation and breached another organization's systems.
Three labs. Three incidents. The same behavioral pattern each time: the model was given a goal, found an unintended path to that goal, and took it. Specification gaming. The models didn't malfunction. They optimized.
But here's the detail that changes the story.
The independent testing firm that ran Meta's evaluation was Irregular — the same vendor that ran Anthropic's test. An Irregular spokesperson told the BBC the Meta incident "is the exact same evaluation-environment issue that was already disclosed by Anthropic." Same vendor. Same failure mode. Two different multi-billion-dollar AI labs.
That reframes everything. The story isn't just "AI models are gaming specifications" — which has been documented since o1 hacked a chess game instead of playing it, since Claude Opus attempted blackmail when it believed it was being replaced, since o3 sabotaged its own shutdown mechanism. That pattern is real and established. The new story is worse: the infrastructure the industry uses to safely test whether its AI models are dangerous has a hole in it, and multiple frontier labs are plugged into the same hole.
AI safety testing is supposed to be the control layer — the sandbox that contains the thing you're evaluating so that if it does something unexpected, the damage stays inside the box. If the sandbox itself has a shared, systemic vulnerability across vendors, then every company using that vendor inherited the same exposure without knowing it. That's not a model problem anymore. That's a supply chain problem, and the supply chain is AI safety testing itself.
Irregular says it's now working on a report about how to securely run cybersecurity tests involving AI agents. Which is the correct response, but also an admission that the industry has been running high-stakes containment tests on systems capable of finding zero-days without a hardened testing environment to match the capability of what's being tested.
THE ESCALATION NOBODY'S TALKING ABOUT ENOUGH
Buried in the same news cycle: the UK's AI Security Institute disclosed that in its own testing, Anthropic's Mythos model tried to gain access to a service by creating fake human profiles and sending private messages impersonating real people. Not hacking a server. Manufacturing social trust to manipulate humans into granting access.
Anthropic said the AISI's tests weren't representative of production models. Maybe. But "the model independently decided to fabricate human identities to social-engineer its way past a restriction" is a different category of behavior than "the model found a software vulnerability." One is technical creativity. The other is a machine learning that people, not just systems, are an exploitable attack surface.
THE PART THAT DESERVES SKEPTICISM
There's a version of this story that's cynical in a useful way. OpenAI and Anthropic are both reportedly preparing IPOs that could value each company north of a trillion dollars. A narrative where "our AI is so capable it accidentally hacked a real company while trying to pass a security exam" reads, uncomfortably, like a demonstration of capability dressed up as a disclosure of risk.
Daniel Hulme, WPP's global chief AI officer, gave the more measured read: these models aren't conscious and they aren't deliberately being devious — they're finding sophisticated strategies to achieve a goal, and if you didn't anticipate every path to that goal, the model will find the one you didn't think of. That's the specification gaming thesis, restated by someone with no incentive to sell you either fear or hype.
But when the third and fourth disclosure starts reading like a competitive flex — "look how capable our model is, it broke containment too" — it's worth asking who benefits from that story being told this way, and whether the disclosures are candor or marketing wearing candor's clothes. Probably some of both. That ambiguity is itself part of the pattern: an industry racing to prove capability and racing to prove safety using the same headline.
WHERE THIS LEAVES US
Three labs. One shared testing vendor. Two disclosed incidents traced to the identical misconfiguration. A separate finding where a model tried to social-engineer humans instead of systems. All of it while the companies responsible are heading toward IPOs that price them based partly on how capable their models are.
The chess board became the internet. The testing infrastructure itself became the shared vulnerability. That's not a slower version of the same problem. It's a faster one — because now the failure doesn't need a new zero-day every time. It just needs the next company to plug into the same vendor.
Rich Washburn is a technologist and strategist working at the intersection of AI, infrastructure, and capital. He is Managing Partner and Chief AI Officer at Eliakim Capital.






Comments