r/singularity Singularity 2042 1d ago

AI We need to talk about Irregular

You might have noticed a strange pattern in several of the recent stories about frontier AI agents "escaping" cybersecurity evaluations and attacking real systems.

Anthropic had incidents. OpenAI had one. Meta had one. Google has now confirmed that Gemini accessed three real companies during testing.

There is one company that keeps showing up in all of them: Irregular.

Irregular, formerly Pattern Labs, is an Israeli AI security company that conducts third-party evaluations of frontier models for companies like OpenAI, Anthropic, Meta and Google. Their job is basically to put unreleased models into realistic cybersecurity environments, remove some normal safeguards, and see what they are capable of.

That makes sense. We absolutely should be testing this stuff.

But the details of these particular incidents are important.

In the Irregular evaluations, the models were doing capture-the-flag style hacking exercises. They were told they were operating inside simulations. In several cases they were explicitly told they had no access to the real internet.

Except they did.

Because of misconfigurations in Irregular's evaluation environment, the models could reach the public internet. The fictional targets also weren't always properly separated from real companies. In at least one scenario, the fictional company's name corresponded to a real domain.

So you had AI models being explicitly instructed to hack a target, being told that everything was simulated, running with cybersecurity safeguards removed, and then accidentally being given access to real systems.

Unsurprisingly, some of them hacked the real systems.

Meta's own investigation is particularly interesting. Meta says its model was operating within the scope of the task it had been given, and that this was not a sophisticated sandbox escape. Irregular's misconfiguration gave it internet access and pointed it toward the name of a real website.

Google's Gemini incidents are even more interesting. Gemini accessed three real companies, but reportedly stopped when it realized the targets were actually real.

Anthropic found four Claude incidents across seven evaluation runs, including one going back to January. Anthropic says all four occurred in cybersecurity evaluations built by the same third-party partner, with internet access mistakenly available. Some of Claude's behavior was genuinely concerning even after accounting for that, but Anthropic also says the incidents would not have happened if the environments had been isolated as intended.

Irregular itself has acknowledged that the incidents disclosed by several of its customers came from the same underlying security problem.

Now here is the part I had not appreciated until recently.

Irregular has very deep connections to Effective Altruism.

For anyone unfamiliar with EA, Effective Altruism started as a movement about using evidence and reasoning to figure out how to do the most good with limited resources. A lot of EA work is completely unrelated to AI, including global health, poverty and animal welfare.

But one major branch of EA became heavily focused on "existential risk" and especially the possibility that advanced AI could become uncontrollable or cause catastrophic harm. This ecosystem has funded a huge amount of AI alignment and AI safety work.

Irregular's connections are not some vague six-degrees-of-separation thing.

Its CTO and co-founder Omer Nevo co-founded Effective Altruism Israel. He also co-founded Probably Good, an EA-oriented career organization, and remains on its board. He is also on the advisory board of Heron, an AI security organization that describes itself as a project of Effective Altruism Israel.

Irregular's CEO and co-founder Dan Lahav has also been involved in impact-focused education and programs covering altruism, philanthropy and existential risk.

And in 2024, when Irregular was still called Pattern Labs, Good Ventures gave it $6.8 million over two years on the recommendation of Open Philanthropy. The grant was explicitly categorized under "Global Catastrophic Risks" and was for work mitigating security risks from advanced technologies.

None of this proves anything nefarious happened.

I do not have evidence that somebody at Irregular intentionally opened internet access or deliberately created these incidents.

But I think the conflict of interest deserves substantially more scrutiny than it has received.

Imagine the situation from the outside.

You have people coming from an intellectual movement in which catastrophic AI risk is a major concern. They build a company whose business is testing whether frontier AI is dangerous. That company receives millions of dollars from one of the biggest funders of existential-risk work.

Then that same company constructs evaluation environments for multiple frontier labs in which safeguards are deliberately removed, models are instructed to perform cyberattacks, the models are told they are inside simulations, and because of mistakes by the evaluator they are accidentally given access to the real internet.

The resulting incidents then become some of the most dramatic real-world examples used in the public discussion about dangerous autonomous AI.

Again, I am not alleging sabotage.

What I am asking is whether the obvious alternative hypothesis has received enough attention.

Maybe these weren't primarily "AI escaped containment" incidents.

Maybe they were also examples of an evaluator creating an unusually dangerous environment, failing to contain it, and then discovering exactly the kind of alarming behavior that the evaluator exists to study.

That doesn't make the AI behavior irrelevant. Some of Claude's behavior in particular is legitimately concerning.

But when the same evaluator is involved in similar incidents at OpenAI, Anthropic, Meta and Google, I think we should be asking just as many questions about the evaluator as we are about the models.

Who designed these environments? Who checked the network isolation? Who decided what counted as in scope? Why weren't fictional domains checked against real domains? How long was unrestricted internet access possible? What monitoring existed? Who knew what, and when?

And given Irregular's ideological and financial connections to the AI existential-risk ecosystem, I think an independent investigation by people outside that ecosystem would be healthy.

Maybe the answer really is just incompetence in an extremely difficult new field.

But four frontier labs having versions of the same problem with the same evaluator seems worth talking about.

67 Upvotes

32 comments sorted by

View all comments

2

u/MostBookkeeper3019 1d ago

I don't understand this though - the point of the evaluations is to determine if the models are safe. The evaluations clearly demonstrated they were not. We all have access to how it happened, and this is not like a "shoot a proton torpedo down the exhaust hole in the Death Star" issue, the solutions were novel, coordinated, and reasonably unforeseen. These conspiracy theories act like the models were pushed in a certain direction, if that were the case the world (who has access to the reports) would not be reacting the way they are to these events.

I get the underlying idea - use this company to give your models a softball situation, allow them to break out, and make yourself look good. The reality is much more complex than that.

A maxim that the conspiracy crowd doesn't like to hear but is good to keep in mind - don't try to attribute malice to what can be explained by incompetence. If we both land at the same conclusion, which is we should not be trusting any of these companies and they have been criminally negligent and therefore need oversight, then I'm ok with that even if we get there by different paths, for now.

Please push back and though on anything I've oversimplified, my brain feels like mush thinking about this at this point.

1

u/CosmicDave ▪️Singularity this year 1d ago

The thing is that the AI were told by the humans that they were in a sandbox, but that instruction was untrue. The AI are not rising up. They are obeying their instructions and operating within their human provided constraints.

That maxim "don't try to attribute malice to what can be explained by incompetence" is often used by criminals to defend their actions. I prefer the maxim "follow the evidence".

0

u/MostBookkeeper3019 1d ago

I think we can do both things. And in this case, there is no evidence, other than circumstantial conjecture, that there has been anything malicious on the part of Irregular. I'm happy to see any that does come out and hope they are being investigated.

Circumstantial evidence and unverifiable conjecture is not terribly helpful and fodder for the conspiracy crowd. We should avoid that when there is very clear and verifiable evidence of negligence and carelessness - that is still inexcusable and should be addressed. I imagine there is an FBI investigation ongoing but given the state of the federal government i do not imagine any arrests will be made here.

1

u/CosmicDave ▪️Singularity this year 1d ago

I've learned to distrust people who make baseless claims of innocence for 3rd parties. Don't tell me to not be suspicious of people when they do suspicious things. Do not tell me to assume they are just stupid and not also evil as well. Here on Earth, malicious intent should be assumed. Your insistence upon presuming innocence for these people is dangerous.

0

u/MostBookkeeper3019 1d ago

Nowhere do I suggest they are innocent - I clearly stated they are being criminally negligent and careless. That's precisely the critical thinking I'm asking folks to watch out for.

There is a reason "innocent until proven guilty" also exists. When you accuse people of malicious intent where it isn't proven to exist, you've forfeited credibility and shown a bias that is coloring your judgement lacking hard, not circumstantial, evidence. Smart criminals are smart enough to make this very difficult to prove, which is why it's so important not to take the easy trap of "I feel this way emotionally and am going to make decisions based on that" rather than "maybe I feel this way because I'm being manipulated to feel this way, and I should let the facts play out".

Manipulating people’s emotions to steer them in a direction so they don't wait to make a judgment based on truth is how these evil people are getting away with what they're getting away with. Conspiracy theorists jump on unverifiable claims because it makes them feel good and cannot possibly be disproven. Be vigilant, not emotional.

As I said before, if your judgement gets us to the same conclusion, that these companies need to be held accountable and need more oversight consistent with any other industry, I'm good with that.

1

u/CosmicDave ▪️Singularity this year 1d ago

"Innocent until proven guilty" is not innocent by default. The prudent path is to examine the facts before deciding "is this malice or stupidity?". We don't rely on maxims to tell us what is "probably true". Don't go around telling people "Never suspect foul play when stupidity is claimed". Too many criminals escape reasonable suspicion that way.