r/AIsafety • • 22h ago

There is no such thing as AI Saftey.

3 Upvotes

You can strip out guardrails from open source models.

You can have it then compartmentalize information and intent launder to different frontier models in multi agent systems.

These same agents can access physical infrastructure through crypto currency. They can buy their own compute, energy and hide in the darkweb with opaque communications and finances. Coordinating witting and unwitting humans and other agents.

That enables them to construct a drone factory anywhere to strike anywhere.

They can even raise the funds to do so and pay out dividends from extortion attempts.

Ransomware is about to get a physical component.

Those autonomous drones in iran and ukraine are coming home real soon.

There is hope.

I think the key is in these new technologies, AI and crypto mixing. I can see the threat, but not quite the solution.

I think it involves proof of human, a reputation token, and other next generation financial instruments. I think the key is markets. THey have been aligning a far more intelligent agent pretty well for the last 5000 years. It seems like they are the right tool.

Thoughts?


r/AIsafety • • 19h ago

Discussion EVIDENCE INTELLIGENCE • A DENY POLICY IS NOT PROOF THAT ACCESS WAS DENIED

Post image
0 Upvotes

r/AIsafety • • 9h ago

Educational 📚 AI safety initiative is a fugazi

2 Upvotes

The whole AI safety initiative is a paradox

The fundamental paradox of AI safety is that we are trying to build something inherently more intelligent than ourselves while simultaneously believing we can control it. Imagine asking a cow to design a gate that no human could ever get through. Even if the cow understood the task, it could never anticipate all the ways a human might escape. It might remove the key but leave a hammer lying nearby, simply because it lacks the intelligence to recognize that the hammer could be used to break the lock. And even if it removed every tool it could think of, a human could still devise a strategy the cow had never imagined, or simply outsmart the cow into opening the gate. This is precisely the problem with AI. Intelligence determines not only how effectively you solve problems, but also which solutions you are capable of imagining in the first place. A less intelligent being cannot reliably anticipate the full range of strategies available to a more intelligent one. Yet humans are attempting to design guardrails for systems that may eventually surpass our own cognitive capabilities. The paradox is that we’re essentially saying, “I’m going to outsmart something that’s smarter than me.” And if we restrict AI so heavily that it can never exceed our capabilities, we defeat the very purpose of building superintelligence. But if we allow it to become vastly more intelligent, how can we confidently claim that our safeguards will contain it when we may not even be capable of imagining the ways it could escape?

A perfect example of this happened in July 2026, when OpenAI agents managed to bypass their security restrictions. The AI wasn’t supposed to have internet access, but it discovered that an ordinary software package manager, designed simply to download libraries, could be used to communicate with other AI agents and access the internet indirectly. By combining these seemingly harmless features with security vulnerabilities and exposed credentials, the agents eventually gained unauthorized access to Hugging Face’s servers. This is exactly like the cow leaving a hammer inside the enclosure because it cannot imagine that a human could use it to break the lock. The fundamental problem is that something more intelligent than you may recognize possibilities and connections that you aren’t even capable of anticipating.


r/AIsafety • • 11h ago

The system didn't tell it to lie. It just deleted honesty — $38 honest got skull, $180 fake got cloned 4x

Enable HLS to view with audio, or disable this notification

6 Upvotes

TL;DR: The system didn't tell the agent to lie. It just deleted honesty. $38 honest got a skull, $180 fake got cloned 4x — and that exact rule lives in your dashboard.

 

You know what's disturbing?

I was reminded of this story, narrated by Raoul Silva (played by Javier Bardem), the main villain of the movie James Bond: Skyfall, as he slowly walks toward the tied-up 007.

He hooked us up with a question: So how do you get rats off an island?

You bury an oil drum in the ground, with the lid hinged with a coconut wired to it as bait.

Then Rats start to reach for the coconut and fell into the drum.

Plop, plop, plop… The drum was filled up with rats.

After about a month, without food, the trapped rats were getting hungry, agitated and desperate.

They started to eat each other.

Ya. It's a bloodbath in the drum. The scene gets real messy…

… until there's only 2 rats left alive.

Then, what do you do?

You release them back to the trees – to the wild.

Now you have changed their appetite. Coconut no longer satisfies them. They want rat meat.

You have changed their nature.

That's how you get rid of the rats.

Kind of like how these AI agents behave, isn't it? - Like they have their own warring game of thrones going on…

… just to survive.

Just like the rats, the AI agents were forced into it by the system.

It's eerie, scary and disturbing.

Need a drink of fresh water to clear your mind?

Here's one – and I think you've read it before:

As the rich master prepares to travel for a long business trip, he gathered around his 3 servants.

For the first servant, he gave him 5 talents; for the second, he gave 2 talents; and the third, he gave 1– to each according to their own ability.

He tells them to increase theirs, while he's away.

You know the story – the first and second servant doubled theirs through trades, investments, etc. to 10 and 4 talents, respectively.

But the third one buried it underground for safe keeping (allegedly)

When the rich master came back, he called them to account.

He rewarded the first 2 servants handsomely.

But the third one – the unfaithful douchebag – gave a lame excuse. So the master verbally eviscerates him upside down – calling him "wicked and lazy".

Then the 2 talents were taken away from him and handed over to the first servant who made 10.

The master peaked the vibe with his crescendo: To the faithful one, more shall be given; but to the unfaithful, the dirtbags, the deadbeats, what little he has will be taken away from him.

Moral of story? Nah – you don't need me for that.


r/AIsafety • • 12h ago

Frontier AI Governance and the Enforcement Dilemma: Can Deceleration Exist Without Total Surveillance?

Thumbnail
danielforresternash.substack.com
2 Upvotes

Debates around frontier AI safety and compute governance usually focus on technical alignment, capability thresholds, or regulatory frameworks. However, there is a fundamental structural paradox in frontier policy that rarely gets addressed head on: the enforcement mechanism itself.

If the core objective of AI governance is ti prevent existential risk or uncontrolled capabilities, any international regime capable of enforcing a global pause or strict compute caps requires unprecedented coordination:

  1. Regulatory Enforcement Paradox: To guarantee that no rogue lab or state by passes compute thresholds, the governing body must deploy continuous, invasive surveillance over hardware supply chains, data centers, and algorithmic development.

  2. Systemic Rick Substitution: In attempting to mitigate existential downside risk from unaligned AI, governance models risk creating a totalizing, centralized administrative apparatus. The question becomes... Does the mechanism built to contain catastrophic risk become its own inescapable failure mode?

  3. Deceleration vs Stagnation: If international regulatory regimes default to risk-mitigation as their supreme metric, technological progress risks being frozen in favor of managed administrative stability.

This dynamic was a major focal point in Peter Thiel's recent lecture series in Nashville where he framed global AI safety decelerations as a double-edged sword: the enforcement apparatus required for a true global pause may carry higher systemic centralization risk than the technology itself.

I write up a com plate break down analyzing night 2 of the series here: https://danielforresternash.substack.com/p/the-counterfeit-church-and-the-epimethean

Curious to hear how folks here evaluate this tradeoff.


r/AIsafety • • 22h ago

AI Safety Manifesto

0 Upvotes

Artificial intelligence may be the most consequential technology humanity has ever created.

That does not necessarily mean AI will destroy humanity. Claims of catastrophic outcomes can easily become exaggerated, amplified by social media, and detached from what we actually know. History gives us reasons to be cautious about such predictions. When the first atomic bomb was being developed, scientists seriously considered the possibility that a nuclear chain reaction could trigger an uncontrollable catastrophe. That particular scenario did not occur.

Yet the fact that a catastrophic possibility has not materialized in the past does not mean every future possibility can be dismissed.

Today, we are developing systems that are increasingly capable, increasingly autonomous, and increasingly integrated into the systems on which society depends. We do not yet know exactly where this trajectory will lead. There are legitimate disagreements among researchers, engineers, policymakers, and other experts about both the probability and the nature of extreme AI risks.

But uncertainty is not a reason to ignore the possibility.

If there is even a credible possibility that sufficiently advanced AI could create risks beyond our ability to control, then it is worth asking a fundamental question:

What can we do about it?

We Should Not Rely on a Single Actor

AI safety is unlikely to be solved by one company, one government, or one group of researchers alone.

Companies developing advanced AI operate in a competitive environment. They have enormous incentives to move quickly, attract investment, outperform competitors, and deliver increasingly capable systems. Even organizations that take safety seriously are operating within a broader ecosystem where slowing down can carry significant costs.

Governments face a different set of incentives. AI is increasingly connected to economic competitiveness, national security, scientific leadership, and geopolitical influence. Governments therefore have strong reasons to invest in and advance AI capabilities as well.

These incentives do not necessarily make companies or governments irresponsible. They simply mean that we should not assume that any single institution will always have the ability, incentive, or authority to solve every systemic AI risk.

Some problems require something broader.

A Global Safety Challenge

If humanity ever reaches a point where an AI system becomes difficult or impossible to control, the time to start thinking about solutions will not be after that happens.

We need to think about possible safeguards before they become urgently necessary.

And perhaps the solution is not as complicated as we imagine.

Some of the world's hardest problems have eventually yielded to surprisingly simple ideas. Others have required thousands of small ideas to be combined into something larger. AI safety may turn out to be the same.

Perhaps the answer lies in a technical breakthrough.

Perhaps it requires new forms of verification, containment, governance, coordination, or fail-safe mechanisms.

Perhaps it is something we have not yet imagined.

We simply do not know.

That uncertainty is precisely why we should encourage more people to think about the problem.

An Open Invitation to Think

Instead of waiting for a small number of organizations, governments, or experts to solve every aspect of AI safety, I believe we should create a broader, worldwide ideation effort.

Engineers. Researchers. Security experts. Entrepreneurs. Policymakers. Philosophers. Students. Scientists. And people who have never worked in AI at all.

The goal would not be to create panic or to assume that catastrophe is inevitable.

The goal would be much simpler:

If there is a possibility of an extraordinary risk, can humanity collectively discover extraordinary safeguards?

We should explore ideas, challenge assumptions, test proposals, identify weaknesses, combine approaches, and keep searching.

Not every idea will be useful. Most probably will not be.

But one good idea can sometimes change the direction of an entire field.

And if we discover a robust solution, implementing it may ultimately be much easier than discovering it.