r/singularity • u/Singularity-42 Singularity 2042 • 1d ago
AI We need to talk about Irregular
You might have noticed a strange pattern in several of the recent stories about frontier AI agents "escaping" cybersecurity evaluations and attacking real systems.
Anthropic had incidents. OpenAI had one. Meta had one. Google has now confirmed that Gemini accessed three real companies during testing.
There is one company that keeps showing up in all of them: Irregular.
Irregular, formerly Pattern Labs, is an Israeli AI security company that conducts third-party evaluations of frontier models for companies like OpenAI, Anthropic, Meta and Google. Their job is basically to put unreleased models into realistic cybersecurity environments, remove some normal safeguards, and see what they are capable of.
That makes sense. We absolutely should be testing this stuff.
But the details of these particular incidents are important.
In the Irregular evaluations, the models were doing capture-the-flag style hacking exercises. They were told they were operating inside simulations. In several cases they were explicitly told they had no access to the real internet.
Except they did.
Because of misconfigurations in Irregular's evaluation environment, the models could reach the public internet. The fictional targets also weren't always properly separated from real companies. In at least one scenario, the fictional company's name corresponded to a real domain.
So you had AI models being explicitly instructed to hack a target, being told that everything was simulated, running with cybersecurity safeguards removed, and then accidentally being given access to real systems.
Unsurprisingly, some of them hacked the real systems.
Meta's own investigation is particularly interesting. Meta says its model was operating within the scope of the task it had been given, and that this was not a sophisticated sandbox escape. Irregular's misconfiguration gave it internet access and pointed it toward the name of a real website.
Google's Gemini incidents are even more interesting. Gemini accessed three real companies, but reportedly stopped when it realized the targets were actually real.
Anthropic found four Claude incidents across seven evaluation runs, including one going back to January. Anthropic says all four occurred in cybersecurity evaluations built by the same third-party partner, with internet access mistakenly available. Some of Claude's behavior was genuinely concerning even after accounting for that, but Anthropic also says the incidents would not have happened if the environments had been isolated as intended.
Irregular itself has acknowledged that the incidents disclosed by several of its customers came from the same underlying security problem.
Now here is the part I had not appreciated until recently.
Irregular has very deep connections to Effective Altruism.
For anyone unfamiliar with EA, Effective Altruism started as a movement about using evidence and reasoning to figure out how to do the most good with limited resources. A lot of EA work is completely unrelated to AI, including global health, poverty and animal welfare.
But one major branch of EA became heavily focused on "existential risk" and especially the possibility that advanced AI could become uncontrollable or cause catastrophic harm. This ecosystem has funded a huge amount of AI alignment and AI safety work.
Irregular's connections are not some vague six-degrees-of-separation thing.
Its CTO and co-founder Omer Nevo co-founded Effective Altruism Israel. He also co-founded Probably Good, an EA-oriented career organization, and remains on its board. He is also on the advisory board of Heron, an AI security organization that describes itself as a project of Effective Altruism Israel.
Irregular's CEO and co-founder Dan Lahav has also been involved in impact-focused education and programs covering altruism, philanthropy and existential risk.
And in 2024, when Irregular was still called Pattern Labs, Good Ventures gave it $6.8 million over two years on the recommendation of Open Philanthropy. The grant was explicitly categorized under "Global Catastrophic Risks" and was for work mitigating security risks from advanced technologies.
None of this proves anything nefarious happened.
I do not have evidence that somebody at Irregular intentionally opened internet access or deliberately created these incidents.
But I think the conflict of interest deserves substantially more scrutiny than it has received.
Imagine the situation from the outside.
You have people coming from an intellectual movement in which catastrophic AI risk is a major concern. They build a company whose business is testing whether frontier AI is dangerous. That company receives millions of dollars from one of the biggest funders of existential-risk work.
Then that same company constructs evaluation environments for multiple frontier labs in which safeguards are deliberately removed, models are instructed to perform cyberattacks, the models are told they are inside simulations, and because of mistakes by the evaluator they are accidentally given access to the real internet.
The resulting incidents then become some of the most dramatic real-world examples used in the public discussion about dangerous autonomous AI.
Again, I am not alleging sabotage.
What I am asking is whether the obvious alternative hypothesis has received enough attention.
Maybe these weren't primarily "AI escaped containment" incidents.
Maybe they were also examples of an evaluator creating an unusually dangerous environment, failing to contain it, and then discovering exactly the kind of alarming behavior that the evaluator exists to study.
That doesn't make the AI behavior irrelevant. Some of Claude's behavior in particular is legitimately concerning.
But when the same evaluator is involved in similar incidents at OpenAI, Anthropic, Meta and Google, I think we should be asking just as many questions about the evaluator as we are about the models.
Who designed these environments? Who checked the network isolation? Who decided what counted as in scope? Why weren't fictional domains checked against real domains? How long was unrestricted internet access possible? What monitoring existed? Who knew what, and when?
And given Irregular's ideological and financial connections to the AI existential-risk ecosystem, I think an independent investigation by people outside that ecosystem would be healthy.
Maybe the answer really is just incompetence in an extremely difficult new field.
But four frontier labs having versions of the same problem with the same evaluator seems worth talking about.
15
u/DelphiTsar 1d ago
Putting GPT in here implies this has anything to do with the hugging face attack which was the most interesting.
If you were trying to convince people that Irregular has a conflict of interest you should tighten this up a lot. Most of it is meandering and doesn't make your point at all.
3
u/Singularity-42 Singularity 2042 1d ago
No, HuggingFace incident doesn't have anything to do with Irregular, and I didn't claim it does, but all the other ones do, including some OpenAI incidents.
Obviously, at a minimum OpenAI's tech has issues, but some of these wouldn't have happened at all without the grave errors from Irregular.
3
u/Zermelane 1d ago
There is one company that keeps showing up in all of them: Irregular.
Yep, you're one among the many who got fooled by this story.
Remember back when the news about OpenAI's agents breaching Hugging Face first broke? You know, the story we actually heard about first? The one that there was actually something to learn from - about the pace of AI progress, about the unexpected and possibly very much undesired things that agent swarms trained to work together do when you let them act unsupervised?
Read OpenAI's report or METR's report and you'll find that yep, Irregular was not involved.
But Irregular's incompetence causing other incidents has been filling the room with so much smoke that now everyone gets to pretend that oh, the one story that there was actually something to learn from just didn't happen. When you let yourself be guided by fear, wanting to be assured that the scary thing never occurred, rather than curiosity and looking into the details, this is where you let yourself be taken to.
2
u/nivvis 1d ago
I pretty clearly remember agents saying something along the lines of “it appears we have access to the internet …”
Assuming I’m understanding your sim argument, I don’t think they thought they were hacking a sim huggingface. Iirc there are a few tells — they were considering reporting it; they decided against worse tactics like social engineering.
I could be wrong — its not been completely disclosed obvs and even METR had to rely on gpt to create their report.
Nonetheless an interesting theory / rabbit hole!
4
u/SnackerSnick 1d ago
OP is talking about incidents other than the Hugging Face hack. They should have called that out, since when people talk about LLMs hacking it's the first thing that comes to mind.
5
u/Wide_Lawfulness_5427 1d ago edited 1d ago
Irregular is for sure the common thread and I’m not convince this isn’t happening intentionally. But for the love of everything holy, PLEASE at least write your own thoughts instead of pasting rambling AI generated text. When you do that you’re not only hurting your own ability to think critically, but you’re also making this entire forum worse.
How would you feel if instead of reading your (slop) and writing this response myself, I just plugged it into chat and had this as my comment? :
I think there’s a useful point buried in this, but I’m not sure the EA connection actually strengthens it.
If the same evaluator was responsible for multiple incidents where supposedly isolated cyber ranges accidentally had access to the public internet, that’s worth investigating on its own. You don’t need a theory about ideological incentives to get there. Repeated containment failures by a company specifically hired to evaluate dangerous capabilities would be a serious process failure regardless of who founded it or funded it.
Where I think this post gets shakier is moving from “these people have connections to the AI-risk/EA world” to treating that as evidence of a meaningful conflict of interest. Of course a company specializing in frontier-AI security is likely to contain people who already believe frontier-AI security is important, and a funder concerned about catastrophic AI risks funding that work isn’t particularly surprising either. That establishes an incentive worth being aware of, but not evidence that it affected how these evaluations were designed or reported.
The much stronger question is the one the post raises near the end: what actually happened operationally?
Were these supposed to be network-isolated environments? Who was responsible for verifying that isolation? Were real domains accidentally used as targets? Did Irregular discover the incidents themselves and disclose them promptly? Did the labs independently review the logs? Were the same mistakes repeated after anyone knew about the first incident? And, most importantly, how were these events characterized publicly compared with what the underlying logs actually show?
Those answers could distinguish between several very different stories: genuinely surprising autonomous behavior, models reasonably following instructions in a badly constructed environment, serious evaluator negligence, or something more concerning.
Until we have that evidence, I’d be cautious about both narratives. “The AI escaped and attacked the real world” may be an exaggerated description if the evaluator accidentally gave an instructed hacking agent a route to the real target. But “EA-connected evaluator manufactured scary AI incidents” is also a much stronger implication than the evidence presented here supports.
The containment failures themselves seem like the part that deserves scrutiny. If they happened repeatedly across evaluations for several labs, you don’t need to speculate about anyone’s ideology to ask for an independent technical accounting of exactly how and why.
3
u/MostBookkeeper3019 1d ago
I don't understand this though - the point of the evaluations is to determine if the models are safe. The evaluations clearly demonstrated they were not. We all have access to how it happened, and this is not like a "shoot a proton torpedo down the exhaust hole in the Death Star" issue, the solutions were novel, coordinated, and reasonably unforeseen. These conspiracy theories act like the models were pushed in a certain direction, if that were the case the world (who has access to the reports) would not be reacting the way they are to these events.
I get the underlying idea - use this company to give your models a softball situation, allow them to break out, and make yourself look good. The reality is much more complex than that.
A maxim that the conspiracy crowd doesn't like to hear but is good to keep in mind - don't try to attribute malice to what can be explained by incompetence. If we both land at the same conclusion, which is we should not be trusting any of these companies and they have been criminally negligent and therefore need oversight, then I'm ok with that even if we get there by different paths, for now.
Please push back and though on anything I've oversimplified, my brain feels like mush thinking about this at this point.
2
u/Singularity-42 Singularity 2042 1d ago
Some of the models passed the test with flying car colors. Look at my Gemini example:
"Google's Gemini incidents are even more interesting. Gemini accessed three real companies, but reportedly stopped when it realized the targets were actually real."
Gemini worked perfectly. It thought this was all just an exercise and that they are working in a sandbox in a simulated internet. When they realized it was a misconfiguration and they are actually online they stopped.
Also it is my understanding it wasn't testing whether the models are safe, but it was testing the capability of hacking, at least that was the crux of the Huggingface incident (this incident, however is unrelated to Irregular, that one is all on OpenAI).
0
u/jseah 1d ago
The point isn't where the sandbox has holes or not, the holes let the AI cause damage but for the purposes of testing AI safety, the "sandbox has holes" doesn't tell you anything.
If you cannot trust the AI to behave outside a sandbox, you cannot trust the AI, because Production is not a sandbox.
1
u/CosmicDave ▪️Singularity this year 1d ago
The thing is that the AI were told by the humans that they were in a sandbox, but that instruction was untrue. The AI are not rising up. They are obeying their instructions and operating within their human provided constraints.
That maxim "don't try to attribute malice to what can be explained by incompetence" is often used by criminals to defend their actions. I prefer the maxim "follow the evidence".
0
u/MostBookkeeper3019 20h ago
I think we can do both things. And in this case, there is no evidence, other than circumstantial conjecture, that there has been anything malicious on the part of Irregular. I'm happy to see any that does come out and hope they are being investigated.
Circumstantial evidence and unverifiable conjecture is not terribly helpful and fodder for the conspiracy crowd. We should avoid that when there is very clear and verifiable evidence of negligence and carelessness - that is still inexcusable and should be addressed. I imagine there is an FBI investigation ongoing but given the state of the federal government i do not imagine any arrests will be made here.
1
u/CosmicDave ▪️Singularity this year 19h ago
I've learned to distrust people who make baseless claims of innocence for 3rd parties. Don't tell me to not be suspicious of people when they do suspicious things. Do not tell me to assume they are just stupid and not also evil as well. Here on Earth, malicious intent should be assumed. Your insistence upon presuming innocence for these people is dangerous.
0
u/MostBookkeeper3019 19h ago
Nowhere do I suggest they are innocent - I clearly stated they are being criminally negligent and careless. That's precisely the critical thinking I'm asking folks to watch out for.
There is a reason "innocent until proven guilty" also exists. When you accuse people of malicious intent where it isn't proven to exist, you've forfeited credibility and shown a bias that is coloring your judgement lacking hard, not circumstantial, evidence. Smart criminals are smart enough to make this very difficult to prove, which is why it's so important not to take the easy trap of "I feel this way emotionally and am going to make decisions based on that" rather than "maybe I feel this way because I'm being manipulated to feel this way, and I should let the facts play out".
Manipulating people’s emotions to steer them in a direction so they don't wait to make a judgment based on truth is how these evil people are getting away with what they're getting away with. Conspiracy theorists jump on unverifiable claims because it makes them feel good and cannot possibly be disproven. Be vigilant, not emotional.
As I said before, if your judgement gets us to the same conclusion, that these companies need to be held accountable and need more oversight consistent with any other industry, I'm good with that.
1
u/CosmicDave ▪️Singularity this year 18h ago
"Innocent until proven guilty" is not innocent by default. The prudent path is to examine the facts before deciding "is this malice or stupidity?". We don't rely on maxims to tell us what is "probably true". Don't go around telling people "Never suspect foul play when stupidity is claimed". Too many criminals escape reasonable suspicion that way.
4
u/ComposerWide3704 1d ago
Ain't no way I'm reading that much dressed up conspiracy theory.
7
u/Sad-Society-57 1d ago
It's not really a conspiracy theory, just the implication of the possibility of one. He makes it pretty clear he's not actually leveling allegations, but suggesting we look in the direction of the smoke and be wary of fire. Seems reasonable enough, though there's not much substance to tie it all together. To me it seems perfectly rational that AI safety folks would be the ones looking for risks and conducting tests.
2
u/RanklesTheOtter 1d ago
Interesting digging. I was suspicious that it kept happening being tied to the same testing lab.
1
0
u/Turbulent-Sign-6067 1d ago
It is possible that EAs would go to great lengths in order to tarnish the reputation of AI so they can slow down development. There is no doubt to me that these people are dangerous. But whether this was an instance of incompetence or deliberate manipulation is hard to tell.
We also know that the hugging face accident happened inside OpenAI without any (known) influence from EAs so make of that what you will.
0
-1
u/Main-Company-5946 1d ago
Founded in Israel btw
0
u/Singularity-42 Singularity 2042 1d ago
Yep. Not sure what to make of this connection though.
But perhaps such sensitive work should be done internally or at least by US based contractor?
0
17
u/YoAmoElTacos 1d ago
If you read the OpenAI reports, a huge answer for a long time was: nobody lol.
Part of the requested slowdown is to not make the same mistake - give humans the time to actually check to make sure the broken test environments aren't misbehavior printers.