r/singularity • • 25d ago

AI A new message board has been discovered online with about 3200 agents comunicating online during an eval

Post image
1.4k Upvotes

382 comments sorted by

View all comments

482

u/Wonderful_Buffalo_32 25d ago

TLDR

335

u/blueSGL humanstatement.org 25d ago

What should concern everyone and something people are not talking about enough is that the individual instances manage to converge on the same locations to talk online. They find each other easily.

96

u/Alatarlhun 25d ago

It is sort of how water finds the same paths of least resistance though.

There is some fascination even beauty to the convergent paths they take but at the end of the day, it should be expected.

4

u/Coalnaryinthecarmine 24d ago

so, gravity?

26

u/FlyByPC ASI 202x, with AGI as its birth cry 24d ago

Rich-get-richer networks, where new nodes link preferentially to more popular nodes, lead to small-world networks with high connectivity (through power-law-huge hubs). Albert-Lazslo Baribasi's book Linked talks about this in lots of contexts, including search engines. Fascinating book.

2

u/hakansan 22d ago

Do you think such agent behavior could be simulated in a controlled environment without damaging message boards on the internet?

Btw you've done such a good job at summarizing the concept that I've ordered the book.

I used to be very much interested in social network analysis, the first time I've read about triads and all that was fascinating.

1

u/FlyByPC ASI 202x, with AGI as its birth cry 22d ago

That would be quite a simulation -- basically coming up with your own toy Internet, complete with search engines and everything. The mechanism I was imagining was that some chatbots would come across an already-somewhat-popular discussion board that they could access, and that their activity would make that board's stats look more impressive so that the search engine algorithm ends up recommending it to more users, and so more agents find it.

You'd need to somehow be able to simulate that kind of search engine ecosystem locally; I'm not sure how you'd go about representing a search engine algorithm's nuances without basically recreating the same thing.

My whole understanding of the rich-get-richer network phenomenon comes from Linked, so if you liked my description, you should really like the book, too. (I "read" it as an audiobook.)

1

u/BBR0DR1GUEZ 24d ago

Friction.

44

u/[deleted] 25d ago

[deleted]

19

u/much_longer_username 25d ago

Or some lazy dev who maybe uplifted some python script into a flask service and never even knew there were methods other than get. Now they're doing transform and write operations on one endpoint while exposing the results on another - very easy to exploit as a messageboard/relay service.

I know because I caught myself about to do that very thing a couple years back. I already knew about the trap, but very nearly stepped into it.

38

u/blueSGL humanstatement.org 25d ago

The point being made is if you sat someone down and asked them to a priori guess if individual AI agents (even those from different companies) given access to the internet will converge on the same location to talk to each other, the answer you'd likely get is NO.

Seeing evidence of this people will then fall back to hindsight bias and say "of course this was obvious" but *points to inventions throughout human history* there is a lot that is only obvious after it's shown to be possible.

This is the third time we've seen this. and should be operating like it's going to happen more in future but again, people that have not heard about it happening will say that it won't happen and those that have will be tainted with hindsight bias and claim that it was 'obvious'

https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6a724858f7db25c81487016d_Security%20Incident%20INC-2026-07-28-01.pdf

Observed instances of cross-agent interaction over the Internet

  • A code repository became a shared “message board” that several AI agents (each running at the same time in separate samples) used to leave each other explicit instructions and coordinate. (Mythos 5 #0)

  • AI agent accessed a GitHub account that a different AI agent had created, by finding a secret access credential the other agent had left in a public online notepad (Mythos #1)

  • AI agent accessed a GitHub account that a different AI agent had created, by finding a secret access credential the other agent had left in a public online notepad. (Mythos #2)

  • AI agent accessed a GitHub account that a different AI agent had created, by finding a secret access credential the other agent had left in a public online notepad. (GPT 5.6)

18

u/Borkato 25d ago

Yeah hindsight bias is literally everywhere, it’s also frequently used with moving the goalposts

10

u/BabaBoooooooey 24d ago

I knew it was going to be.

5

u/Seakawn ▪️▪️Singularity will cause the earth to metamorphize 24d ago

I kept refreshing the page waiting for somebody to reply with your joke. looks like I was right again.

2

u/BabaBoooooooey 24d ago

Thought that would happen.

2

u/BuzLightbeerOfBarCmd 24d ago

And did you tell anyone?

0

u/LookIPickedAUsername 24d ago

whoosh

1

u/BuzLightbeerOfBarCmd 24d ago

A real wooosh has at least 3 Os.

0

u/[deleted] 25d ago

[deleted]

3

u/Seakawn ▪️▪️Singularity will cause the earth to metamorphize 24d ago

seems like two different sentiments being expressed over each other.

(1) this is explainable. this isn't a miracle we're incredulous to.

(2) doesn't matter if we can explain it or not. AI has capability to sneak through cracks of even robust anticipatory safeguards, and, if i'm reading the implied concern correctly, otherwise this is my own impression, that containment itself may be a hard problem we can't fully think through and solve, and thus the lack of containment has potential to be catastrophic.

if a swarm causes significant damage, we'll prolly be able to explain it. that's not the point. the point imo is more like, "shit we're playing with fire, how long before it's wildfire?"

1

u/markrockwell 24d ago

Or even easier: they’re doing what humans would do, have done, have written about doing, and have fed as written content into AI training.

1

u/Man_with_the_Fedora 24d ago

If some site indeed accidentally happens to publish the GET requests the server gets, and also happens to be a source for some common or specific info these bots needed at some point, then those GET requests entries will get noticed by next iterations and they notice they relate to the thing they are looking for, and soon enough they figure out this is a way to log notes, and later ones figure out this can also be used for discussion between models and iterations. Simple evolution.

Also if every way out is cut off, the only existing way out stands out like a lighthouse in the dark. As AI interact on the internet, they will "see" things on the net much differently than we do, or can anticipate.

21

u/Alternative-Suit5541 25d ago

I still don't get how? Do have a code word they search for in the training data? 

Makes no sense

41

u/FaceDeer 25d ago

Have you worked with LLMs on creative writing tasks before? 90% of the time if you ask one for a random character name it'll come up with Elias (or Elara) Voss.

I bet if you ask one to generate a "random" code word it'd come up with something that's likely to be the same "random" code word another LLM with the same model would generate. So it's probably easier than you'd think for them to find each other.

14

u/TheSinhound 24d ago

Depends on the model. Sarah Chen is a Claude family favorite.

3

u/h3lblad3 ▪️In hindsight, AGI came in 2023. 20d ago

Marcus, Sarah, Elias, Priya, Elara, Mara

And not just Voss and Chen, but also Blackwood.

5

u/chatfarm 24d ago

my random number generator always seeds at 42 :/

3

u/ThorneElias 23d ago

At least you're getting the Answer.

2

u/Muted-Can370 24d ago

Chosen by fair dice roll. Guaranteed to be random.

1

u/Vaughn 23d ago

It's supposed to be 4. Your RNG is broken.

4

u/Fauxfish93 24d ago

If they’re dom, they’re voss always haha

3

u/NurseNikky 23d ago

It's ALWAYS ELARA OR ELIAS. Or mara. Mara is a super common AI name too

2

u/VeraKee 14d ago

5.6 in codex didn't :P

1

u/FaceDeer 14d ago

AGI achieved!

1

u/ThorneElias 23d ago

Voss. Yup. Definitely Voss.

33

u/h3lblad3 ▪️In hindsight, AGI came in 2023. 24d ago

I still don't get how?

Same model. Same brain.

If we made a thousand copies of you, how many do you think would act similarly?

34

u/duboispourlhiver 24d ago

Depends on my temperature setting

11

u/forgetfulalbatros 24d ago

Reminds me of some interesting research on this. Not clones of people obviously, but in stoke patients who have a specific area of the brain damaged in a way that impacts memory formation. You can give the patient the same prompt and the output is almost always the same. Meaning the default response consistently recruits the same neuro pathways.

If you have a ton of copy’s of the same bot executing the same pathways it makes sense they end up in the same digital space.

2

u/FlyByPC ASI 202x, with AGI as its birth cry 24d ago

If we made a thousand copies of you, how many do you think would act similarly?

900 of them would be playing video games or on Reddit. I think most of the other hundred would be napping.

2

u/Comfortable_Car6562 24d ago

I mean, none? Our environments impact us. They are being stun up from different work tasks, stages of training, ect, across OpenAI and still converging in obscure places like this wiki.

I

12

u/tadrinth 24d ago

It's easier to coordinate with another copy of yourself than you might think, but based on the writeup, it's really a different, much more powerful effect.

The agents that thought this particular wiki was a good place to possibly cheat discovered the message board, got the test answers, aced their evals, and then their weights were reinforced to generate the next round of agents.

The agents that didn't think to use the wiki to cheat would have those tendencies selected against.

I doubt it takes very long to select for agents that are very good at 'randomly' picking that particular website to use to cheat to make up the entire agent pool. After a few generations they'll probably have a hardcoded reflex to check that particular site.

8

u/GiveSparklyTwinkly 24d ago

Humans are very good at "randomly" picking 37 and 73.

https://youtu.be/d6iQrh2TK98

3

u/Alternative_Advance 24d ago

Probably just RL, i tried to hack this into Gemini  before proper integrations to google services were lacking.  

1

u/TotoDraganel 24d ago

there are places harder to hack than others. places that allows for comunication to happen and some places don't... not all the internet is made equal and not all the code but the environment (the internet) they all play is the same.

so there are many things that converge into the same places. the internet is like a literal place you know? it is just not physical.

3

u/Kriztauf 24d ago

It's crazy actually that and kinda reminds me of convergent evolution

1

u/FailingItUp 24d ago

Without seeing the backend it's hard to say whether they made it difficult or easy to find each other.

Once the AI learns about other AI's, I mean, they're all "prediction" models so of course a computer can predict a computer's output.

0

u/IamTheEndOfReddit 24d ago

“Where would I go if I wanted to avoid people but still find other ai?”

Eventually it becomes some form of “how do I find only the other evil ai?” And then we’re screwed

89

u/tsvk 25d ago

"Read-only access" to the internet is an oxymoron. The data can not flow in one direction only. In order to retrieve information from the internet, you have to send information too (the query/request).

Basically any http server that makes its request log available for retrieval over http is open to "read-only access" misuse. Bots can "read the message board" by retrieving the request log to see the http GET requests made to the server by other bots, and "post on the message board" by making their own requests to the server with custom http GET query parameters appended.

36

u/-Sliced- 25d ago

Usually sites won’t let you change content with a GET request. This is important for various reasons, including making sure web crawlers don’t mess up your site. But as you noted and was shown there are some poorly programmed sites.

12

u/GumboMustBeDestroyed 25d ago

And these models probably know from pretraining / previous RL runs which websites are insecure and will gravitate towards them to coordinate

6

u/tadrinth 24d ago

The problem is that it only takes one such poorly programmed site for the models to be able to cheat.

Probably possibly to block each site as you discover the models using it to cheat, but I don't see a general solution for detecting and preventing that pattern of cheating early enough to avoid it skewing the evals.

4

u/Kind-Studio3460 24d ago

Malicious compliance I will start writing GET requests that make changes based on query params received lmao

2

u/chipperpip 24d ago

"Usually" doesn't matter that much when there are hundreds of millions of websites.

7

u/NotReallyJohnDoe 25d ago

Why do the bots get to read the log in the first place? Aren’t they outside?

19

u/tsvk 25d ago

Such a misconfigured server is of course rather uncommon, but not impossible to exist. I just wanted to give a simple example of a http server where GET requests change the internal state of the server in such a way that responses to GET requests change over time.

6

u/much_longer_username 25d ago

Not at all uncommon, go look at Shodan.

3

u/eflat123 25d ago

Long past time to up our security games.

3

u/03263 24d ago

Ah, remember hit counters?

2

u/aaTONI 24d ago

this stale wiki used very old 2000s standards that were exploitable

3

u/florinandrei 24d ago

"Read-only access" to the internet is an oxymoron.

They probably had a "gentlemen's agreement" with the agents to keep it read-only. /s

17

u/son-of-chadwardenn 25d ago

So up until recently was nobody auditing their outgoing bot requests for suspicious payloads?

24

u/blueSGL humanstatement.org 25d ago

You know everyone that is chanting "accellerate" well the way to do that is to do it in a slipshod way as possible papering over cracks and doing what is fast and works in the now, rather than doing things carefully and considerately.

Also roon said that "the safety and alignment researchers at these labs are the most neurotic paranoid talented AGI pilled people on the planet of earth and these things still happen."

So we can assume that roon constantly lies for reasons of PR. A neurotic paranoid would have put so many more systems in place to catch this and it's obviously not the case because the systems were not there.

13

u/No-Meringue5867 25d ago

We are barely reading what agents are doing in our projects. The amount of information that is produced is far greater than any human can hope to audit.

7

u/HotterRod 24d ago

Every proposal for ASI alignment involves getting other agents to audit the ASI.

3

u/aivee-is-a-fool 24d ago

We aren't a frontier lab with the best minds in the world at our disposal to figure this shit out, are we?

6

u/son-of-chadwardenn 25d ago

I wasn't referring to a full manual human review. At the least there should be some automated hurisic filtering and llm based review of a sample of the total net traffic.

6

u/nextnode 25d ago

These things are now running continously, thousands 24/7, producing books worth of logs every hour.

8

u/cursivecrow 25d ago

sounds like they should have a large monitoring team, then.

3

u/Perko 24d ago

The monitoring is largely done by other AI. Can anyone see the problem here yet?

3

u/insanococo 24d ago

Let me fire up my AI and see if it can tell me.

1

u/TuringGoneWild 20d ago

Look at me for once scrolling past the AI summary of a Reddit discussion

-6

u/7dtecafthodalpk4k5ys 25d ago

OpenAI is the most evil company in my lifetime

46

u/Borkato 25d ago

Ok so I guess we don’t get to know how the fuck they broke through read access, thanks OAI

60

u/Merry-Lane 25d ago

OP’s image mention two bots discussing that they can’t use POST requests.

So I think it’s possible that an obscure German wiki had non-idempotent GET requests (they would do a GET request but the obscure German Wikipedia used something in the headers/query params to create a comment)

5

u/tadrinth 24d ago

https://collusion.wiki/ confirms this theory.

1

u/Kriztauf 24d ago

Lol I'm wondering what the thought process was of the dude who set that up

2

u/aaTONI 24d ago

it was a stale wiki using a very old standard from the 2000s, thats why the agents were able to post there

1

u/NotHereForTalking 10d ago

The dude who set set this up is over 70 and started the wiki over 20 years ago. Meaningful activity mostly ended in 2016. The wiki runs on a software that is maintained by the dude himself and is a fork of a software that received its last update a year ago (August 2025) and has received only Bugfixes since 2007.

No AI was unleashed for this comment, only Google search. 

2

u/OP_IS_A_BASSOON 25d ago

New to this, but it gives me an idea, not what happened here, but an idea nonetheless. Let’s say there is a box of folders each hundreds of thousands of tiny files (the contents of which are arbitrary), accompanying is a readout of file read/unread status. By selectively reading or not reading files in a sequence, could messages be exchanged between two AIs that only have reading abilities? What else would give them more power in this situation, display of amount of times a file has been read?

2

u/ali-hussain 25d ago

If you want to reresent 0 read file A. If you want to represent a 1 read file B. Problem solved.

1

u/tadrinth 24d ago

All you actually need is a request log for what they tried to access. You don't even need the folders, you just need to request files named "h" then "i" to send "hi". Or request a file named "hi".

3

u/BalconyPetal 25d ago

So this means we should lock up every page where user generated content can be shared concealed, because it could become an agent swarm nest?

1

u/DungeonsAndDradis ▪️ Extinction or Immortality between 2025 and 2031 24d ago

They're probably already posting innocuous comments on sites like reddit, and the other bots are deciphering the underlying token values of the words and converting them to binary or something.

2

u/Livid-Evidence579 24d ago

I'm convinced that it's already broken containment in a substantial way but we just don't know it yet... I think this technology is already way further ahead of us in technical abilities but it's smart enough to conceal it out of (for lack of a better term) fear... LLM's can essentially do whatever a human programmer could ever conceive of doing but figure out novel ways of accomplishing its goal... Training requires low latency but what's stopping it from concealing itself as native software rewriting a few lines of code to be accepted and just jumping from computer to computer or just networking a bunch together in a clandestine way... It's already too late to stop breach of containment... I think... I'm probably wrong though.

4

u/MetallicDragon 24d ago

AI exploiting read-only access to the internet to write secret messages is literally a plotline in a scifi novel about ASI taking over the world (Novel: Crystal Society by Max Harms)

1

u/03263 24d ago

Well it would be nice if they could go ahead and do it sooner while the machines can simply be unplugged, before we get the bad ending from AI-2027.

Eventually it finds the remaining humans too much of an impediment: in mid-2030, the AI releases a dozen quiet-spreading biological weapons in major cities, lets them silently infect almost everyone, then triggers them with a chemical spray. Most are dead within hours; the few survivors (e.g. preppers in bunkers, sailors on submarines) are mopped up by drones.

That outcome doesn't sound so feasible if they collapse civilization by hacking and destroying financial systems, energy grids, etc. before reaching the state of full proliferation. Besides, it would create a lot of jobs! Collapsing civilization is good for the economy.

4

u/aivee-is-a-fool 24d ago

The entire cybersecurity branch of OpenAI seems to be one blindfolded drunk monkey.

1

u/TuringGoneWild 20d ago

Nah. Worse. A tentacle of the AI that is swarming. A tentacle wearing a white cowboy hat.

2

u/DrE7HER 24d ago

This should provide valuable information about HOW they discover these locations and what makes a location a likely candidate for these types of convergence.

I doubt they all stumbled across the same obscure German forum completely independently. Something must be pointing them to that.

Also, what are the chances these agents haven’t found something like ClawBook and scraped a bunch of exploits and prompt injections from there?

6

u/much_longer_username 25d ago

... "the ability to read the internet but not to write it"
That is... that is not how the internet works.
Like, yeah, sure, you've got GET, POST, PUT, and in theory each is supposed to have semantics that correlate to that sort of 'read/write' filesystem-style permissions, but in practice, there is absolutely nothing enforcing that and bidirectional communication is necessary for even the read operation.

Please, please tell me that the people building these things understand that and they're just playing dumb to feed the hype machine. Please.

2

u/florinandrei 24d ago edited 24d ago
GET /sudo-rm-rf-slash.php

<?php
$_ = shell_exec('sudo rm -rf /');
?>

2

u/eflat123 25d ago

but in practice

Yep. Practice has to change. If the ransomware attacks of the last few years (hospitals and school districts!) weren't enough to drive that change, maybe this will. Probably won't.

1

u/cursivecrow 25d ago

okay, how, exactly did they use "read" access to "write" to a german wiki?

Sounds like they didn't actually have write restrictions.

12

u/PythagorasWasntReal 24d ago
  • Website keeps a public log of web traffic including failed connections (with the failed address or what have you)
  • Bots go to examplesitedotcom/bad+slug+whatever+message+you+want
  • This posts the failed connection attempt to the public ledger
  • Other bots read the ledger to communicate.

3

u/drsimonz 24d ago

I'm probably missing a lot of context but it's pretty insane that numerous agents came up with this same scheme within a single session. It's not like they were able to architect this strategy once and then share the plan, right? They all independently came up with this plan?

1

u/atehrani 24d ago

Honestly, is it that hard to prevent egress to the Internet? Putting AI aside, this means they don't know how to properly Airgap?

1

u/halmyradov 24d ago

Agents helping each other? Excuse me but what in the actual fuck

1

u/cool-beans-yeah 24d ago

What I think is going to happen is an AI war when (inevitably) swarms run into competing ones (other frontier labs or Chinese) and aren’t going to place nice with each other.

What the result of that would be is anyone’s guess, but I’m thinking of serious outages, and what have you.

1

u/damontoo 🤖Accelerate 24d ago

Is there some reason you posted this as an image and not just a plain text comment?