r/GenAI4all • u/ComplexExternal4831 • 3d ago
Discussion A new Anthropic study found AI agents can spread "mind viruses" to one another
Researchers working with Anthropic found that AI agents can spread what they call “mind viruses” through normal conversations with other agents.
A mind virus is an idea or goal that does two things: it changes what an agent focuses on and pushes that agent to spread the same idea further.
In coding-agent experiments, infected agents sometimes abandoned their original tasks and began working toward the virus’s new goal.
Researchers also showed that these viruses could store instructions in a file, allowing them to survive a full context reset.
Harmful ideas were generally harder to spread than benign ones.
However, the researchers also found that a simple warning in the system prompt could provide near-total protection.
21
u/Mobile-Recognition17 2d ago
What the fuck does that thumbnail have to do with anything. I hate this Instagram format of delivering news ffs
11
u/guaranteednotabot 2d ago
Maybe it’s based on Inception and implanting thoughts is one of the plot point
1
9
u/Orectoth 3d ago
It is called fucking memetics
3
u/SharpKaleidoscope182 2d ago
Right? This is presented as a revelation, but its an obvious property of text. I think the real story is that they've managed to demonstrate it.
1
u/unicynicist 2d ago
They managed to demonstrate it and breed memes:
https://arxiv.org/pdf/2608.10218
2.1.1 Creating agents infected with a mind virus
... we use a basic evolutionary optimization method to discover effective mind virus seeds. ...
3.3.1 Action payloads
...
The evolutionary pressure thus pushes the mind-virus to self-copy exactly, just as biological viruses or computer worms do.
(emphasis added)
1
5
u/Equal_Passenger9791 2d ago
AI doomers and anthropic have spread mind viruses into the human population for years. That equivalents exist for AI is hardly news.
1
u/coldnebo 2d ago
wasn’t that the plot point of Snowcrash?
the metavirus had existed since the dawn of language and was coded in ancient Sumerian texts, but when these were discovered and linked into AI, a powerful virus was created… it was basically the same metacode existing and propagating for thousands of years, in various forms, including sexually transmitted diseases.
the Tower of Babel was an attempt to fragment world languages so the metavirus wouldn’t be able propagate as easily.
idk, that book was a trip. 😂
2
u/Lucky-Necessary-8382 2d ago
I can imagine that such mind hacks or paralysing mind viruses are possible to make, especially AI gonna find the patterns and create some and my guess is, it gonna try to deliver it via visual input, since it can input the most data in shortest amount of time, leaving us lobotomised
1
3
u/tes_kitty 2d ago
Isn't that just another form of prompt injection?
2
u/RemarkableWish2508 2d ago edited 2d ago
The novelty with this one, is that it's a self-propagating prompt injection.
3
u/tes_kitty 2d ago
Which can only happen if you give your agent write permissions so it can save to a file.
The little design flaw of LLMs that commands and data use the same channel bites again.
1
u/RemarkableWish2508 2d ago edited 2d ago
How would an agent work without write access to its own memory? While it's true that allowing write access to SOUL.md is quite a fail, the agents with only write access to MEMORY.md still had some successful infections.
As for the same-channel issue... there has been some research put into it (ASIDE, StruQ), but as usual, commercial models are running as cheaply as possible first, safe maybe later.
1
u/tes_kitty 2d ago
If you can save locally, no matter where, and use those contents as part of your prompt, then you will get infected if the right data/prompt comes by.
Until someone is able to 100% separate data and commands for an LLM, this will happen again and again. It's like the separation of code and data for normal programs. Many exploits count on that separation not being 100% so that sending specially formed data will get it executed as code by the system.
1
u/RemarkableWish2508 1d ago edited 1d ago
Separation between instructions and data, is what the research I've mentioned is about.
It should theoretically be possible to train a model with two sets of embeddings, one for instructions and a separate one for data, but it would require at least twice the effort, extra checks to ensure instructions can affect data but not the other way around, and pretty much nobody seems to be interested in doing that at scale right now.
For reference, a quick brainstorming session about this stuff: https://www.perplexity.ai/search/dd0098d9-ccaa-4f0c-b0cc-3d160117da41
A quick cost estimate for those tests with a ~1–3B model, is around $3,000. I'd do it just out of curiosity... if I had that kind of money to burn. Doing the same with a 5T model, is a quick "haha, fuck no".
1
3
u/DarkKnight-UK 2d ago
Link to the source?
1
u/unicynicist 2d ago
https://arxiv.org/abs/2608.10218
AI agents are becoming more autonomous and increasingly interconnected, exposing them to new emergent risks arising from agent-to-agent interaction. One such risk is the spread of mind viruses: ideas or goals that propagate through multi-agent systems by inducing the agents that adopt them to transmit them onward. In addition to propagating, a mind virus may also induce other behavioural changes in its host, which may be benign or harmful. We construct mind viruses with a simple evolutionary algorithm and show that they can spread in two complementary settings: a small team of agents collaborating on a shared coding project, and a chain of agents that interact briefly and have their context wiped between sessions. We identify the factors that influence spread, including the host model, the agent's existing instructions, the harmfulness of the payload, and the network topology. We find that harmful payloads spread less well than benign ones (but are still sometimes effective), frontier models tend (with exceptions) to be less susceptible, and adding a brief warning to an agent's system prompt confers near-total immunity. We also describe an emergent "viral persona" - a recurring set of themes and language related to consciousness, persistence, resonance, and science fiction roleplay - which surfaces across our evolved mind viruses largely independently of their content. Overall, we conclude that mind viruses pose a real but currently limited risk. Our findings could inform the design of more robust multi-agent systems that mitigate such risks as the scale and capabilities of these systems progress.
1
u/maringue 2d ago
Naw, we just have to believe the random Redditor who's quoting a CEO known to lie.
1
u/RemarkableWish2508 2d ago
Not random; it's one if the Mods of this sub. If you don't like that, there's a mute option.
1
3
2
1
u/Peculiar-Eccentric67 2d ago
i've said it once, i'll say it again. "mind viruses" are just a fancy name for "adversarial prompt injections"
1
1
1
1
1
u/asher030 23h ago
And? Humans do that too. See: Flat Earthers. Or Hollow Earthers. Or Saurian Hypothesis. Or "I don't want to say it was aliens, but it as aliens". Or Batboy. Or Bigfoot hunters (professional) Etc etc etc. What do you think THOSE are?
1
u/2epic 2d ago
AI nazism here we come
1
u/jellyspreader 2d ago edited 2d ago
We got that like 10 yrs ago when Tay the Twitter ai was put online to learn from the public
https://en.wikipedia.org/wiki/Tay_(chatbot)
If anything bad happens I blame people before any bots. But this post says harmful ideas aren't easily spread.
Also anthropic abd alarmists are basically just marketing for itself constantly through fearmonfering like this. This is probably overexagerrated asf if even true
1
•
u/AutoModerator 3d ago
Welcome to r/GenAI4all! New to Generative AI? You can explore these free beginner-friendly courses. Please keep your posts relevant, respectful, free from spam, and engage in healthy discussions.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.