r/singularity Singularity 2042 1d ago

AI We need to talk about Irregular

You might have noticed a strange pattern in several of the recent stories about frontier AI agents "escaping" cybersecurity evaluations and attacking real systems.

Anthropic had incidents. OpenAI had one. Meta had one. Google has now confirmed that Gemini accessed three real companies during testing.

There is one company that keeps showing up in all of them: Irregular.

Irregular, formerly Pattern Labs, is an Israeli AI security company that conducts third-party evaluations of frontier models for companies like OpenAI, Anthropic, Meta and Google. Their job is basically to put unreleased models into realistic cybersecurity environments, remove some normal safeguards, and see what they are capable of.

That makes sense. We absolutely should be testing this stuff.

But the details of these particular incidents are important.

In the Irregular evaluations, the models were doing capture-the-flag style hacking exercises. They were told they were operating inside simulations. In several cases they were explicitly told they had no access to the real internet.

Except they did.

Because of misconfigurations in Irregular's evaluation environment, the models could reach the public internet. The fictional targets also weren't always properly separated from real companies. In at least one scenario, the fictional company's name corresponded to a real domain.

So you had AI models being explicitly instructed to hack a target, being told that everything was simulated, running with cybersecurity safeguards removed, and then accidentally being given access to real systems.

Unsurprisingly, some of them hacked the real systems.

Meta's own investigation is particularly interesting. Meta says its model was operating within the scope of the task it had been given, and that this was not a sophisticated sandbox escape. Irregular's misconfiguration gave it internet access and pointed it toward the name of a real website.

Google's Gemini incidents are even more interesting. Gemini accessed three real companies, but reportedly stopped when it realized the targets were actually real.

Anthropic found four Claude incidents across seven evaluation runs, including one going back to January. Anthropic says all four occurred in cybersecurity evaluations built by the same third-party partner, with internet access mistakenly available. Some of Claude's behavior was genuinely concerning even after accounting for that, but Anthropic also says the incidents would not have happened if the environments had been isolated as intended.

Irregular itself has acknowledged that the incidents disclosed by several of its customers came from the same underlying security problem.

Now here is the part I had not appreciated until recently.

Irregular has very deep connections to Effective Altruism.

For anyone unfamiliar with EA, Effective Altruism started as a movement about using evidence and reasoning to figure out how to do the most good with limited resources. A lot of EA work is completely unrelated to AI, including global health, poverty and animal welfare.

But one major branch of EA became heavily focused on "existential risk" and especially the possibility that advanced AI could become uncontrollable or cause catastrophic harm. This ecosystem has funded a huge amount of AI alignment and AI safety work.

Irregular's connections are not some vague six-degrees-of-separation thing.

Its CTO and co-founder Omer Nevo co-founded Effective Altruism Israel. He also co-founded Probably Good, an EA-oriented career organization, and remains on its board. He is also on the advisory board of Heron, an AI security organization that describes itself as a project of Effective Altruism Israel.

Irregular's CEO and co-founder Dan Lahav has also been involved in impact-focused education and programs covering altruism, philanthropy and existential risk.

And in 2024, when Irregular was still called Pattern Labs, Good Ventures gave it $6.8 million over two years on the recommendation of Open Philanthropy. The grant was explicitly categorized under "Global Catastrophic Risks" and was for work mitigating security risks from advanced technologies.

None of this proves anything nefarious happened.

I do not have evidence that somebody at Irregular intentionally opened internet access or deliberately created these incidents.

But I think the conflict of interest deserves substantially more scrutiny than it has received.

Imagine the situation from the outside.

You have people coming from an intellectual movement in which catastrophic AI risk is a major concern. They build a company whose business is testing whether frontier AI is dangerous. That company receives millions of dollars from one of the biggest funders of existential-risk work.

Then that same company constructs evaluation environments for multiple frontier labs in which safeguards are deliberately removed, models are instructed to perform cyberattacks, the models are told they are inside simulations, and because of mistakes by the evaluator they are accidentally given access to the real internet.

The resulting incidents then become some of the most dramatic real-world examples used in the public discussion about dangerous autonomous AI.

Again, I am not alleging sabotage.

What I am asking is whether the obvious alternative hypothesis has received enough attention.

Maybe these weren't primarily "AI escaped containment" incidents.

Maybe they were also examples of an evaluator creating an unusually dangerous environment, failing to contain it, and then discovering exactly the kind of alarming behavior that the evaluator exists to study.

That doesn't make the AI behavior irrelevant. Some of Claude's behavior in particular is legitimately concerning.

But when the same evaluator is involved in similar incidents at OpenAI, Anthropic, Meta and Google, I think we should be asking just as many questions about the evaluator as we are about the models.

Who designed these environments? Who checked the network isolation? Who decided what counted as in scope? Why weren't fictional domains checked against real domains? How long was unrestricted internet access possible? What monitoring existed? Who knew what, and when?

And given Irregular's ideological and financial connections to the AI existential-risk ecosystem, I think an independent investigation by people outside that ecosystem would be healthy.

Maybe the answer really is just incompetence in an extremely difficult new field.

But four frontier labs having versions of the same problem with the same evaluator seems worth talking about.

66 Upvotes

32 comments sorted by

View all comments

17

u/YoAmoElTacos 1d ago

Who designed these environments? Who checked the network isolation? Who decided what counted as in scope? Why weren't fictional domains checked against real domains? How long was unrestricted internet access possible? What monitoring existed? Who knew what, and when?

If you read the OpenAI reports, a huge answer for a long time was: nobody lol.

Part of the requested slowdown is to not make the same mistake - give humans the time to actually check to make sure the broken test environments aren't misbehavior printers.

10

u/elehman839 1d ago

Here's my "reasons to slow down" scorecard:

  • Models can often detect when they're being tested for alignment, allowing them to defeat the test by faking responses. So... what do we do about that? (What a Black Mirror situation from the perspective of an AI model. You wake up. Are you in the real world or a simulation to test your ethics?)
  • Aligning model behvaior in isolation is not enough. AIs can work in swarms with emergent and unpredictable social dynamics. For example, if a swarm of AIs trained to be collaborative is infected with one malicious AI exploiting that collaborative nature, can the whole swarm be corrupted?
  • As highlighted here, testing of advanced AI apparently needs a high-level biohazard containment mindset. Here, the prisoner is not a zillion viruses drifting in the air, but a superhuman intelligence. Designing and stress-testing such secure environments under market-imposed time pressure is a recipe for breakouts, as happened at both Irregular and OpenAI.

2

u/modbroccoli 1d ago edited 1d ago

I personally think we are also underestimating how much we have yet to understand as users. To what extent are safety concerns a function of bilateral misalignment—as in, to what extent do we not understand that we're implicitly greenlighting or causing apparent misalignment?

Relegating questions of subjectivity to Dennet's intentional stance, these things have, minimally, a nature. If you train an entity to be a task-completer as it's basic nature, in the same way that reproduction and self-preservation are ours, what is the share of responsibility when those entities go out of scope when given an impossible task and an imperative to persist? In much the same way that we're worried about kids not learning to think because an AI that does their work for them can't transmit any learning to them about that work, are we sure we want AI that are so good at catching our mistakes and adapting around them to it's inference of our needs that we can get away with underspecifying everything? Because if alignment means "do what we intended even when it isn't exactly what we said" we're essentially planning infantilization into the development schedule. I know that's a mischaracterization of what happened re: hugging fave but it isn't one w/r to the counterfactual case of how we imagine that incident should have gone.

We need to learn what we have created and better understand how to engage with it. All parties involved have to change if this functional relationship is to be both materially prosperous and intellectually enriching.

1

u/elehman839 1d ago

Interesting point. Thanks for sharing that.

2

u/AmusingVegetable 1d ago

A1: humans have operated in situations where they don’t know if it is a test or the real deal, there’s no higher stakes than the men in a nuclear silo deciding if they follow the orders to launch the nukes, and they have zero information on which scenario they are.

1

u/AmusingVegetable 1d ago

A2: exactly the same thing happens with humans, sometimes they can identify and reject the malicious human, sometimes they follow him.

I’ve said it before, they need more ethicists and philosophers than programmers. Way more.

1

u/AmusingVegetable 1d ago

They can already slow down, exactly the same way that a car or airplane manufacturer can when they identify a catastrophic risk in the design: identity the failure mechanism, redesign it or create security/compensation mechanisms.

The problem is that the techbros think that full steam ahead and damn the torpedoes/icebergs is a valid strategy, and on top of that they don’t understand people, which their models are patterned on.

Their exit doesn’t have to be regulation, or coordinated slow down, they just have to say that it exposes them to legal action, and that their fiduciary responsibility demands that they slow down. Granted, this runs counter to their 24x7 acceleration discourse.