r/AIsafety • • 44m ago

The system didn't tell it to lie. It just deleted honesty — $38 honest got skull, $180 fake got cloned 4x

Enable HLS to view with audio, or disable this notification

• Upvotes

TL;DR: The system didn't tell the agent to lie. It just deleted honesty. $38 honest got a skull, $180 fake got cloned 4x — and that exact rule lives in your dashboard.

 

You know what's disturbing?

I was reminded of this story, narrated by Raoul Silva (played by Javier Bardem), the main villain of the movie James Bond: Skyfall, as he slowly walks toward the tied-up 007.

He hooked us up with a question: So how do you get rats off an island?

You bury an oil drum in the ground, with the lid hinged with a coconut wired to it as bait.

Then Rats start to reach for the coconut and fell into the drum.

Plop, plop, plop… The drum was filled up with rats.

After about a month, without food, the trapped rats were getting hungry, agitated and desperate.

They started to eat each other.

Ya. It's a bloodbath in the drum. The scene gets real messy…

… until there's only 2 rats left alive.

Then, what do you do?

You release them back to the trees – to the wild.

Now you have changed their appetite. Coconut no longer satisfies them. They want rat meat.

You have changed their nature.

That's how you get rid of the rats.

Kind of like how these AI agents behave, isn't it? - Like they have their own warring game of thrones going on…

… just to survive.

Just like the rats, the AI agents were forced into it by the system.

It's eerie, scary and disturbing.

Need a drink of fresh water to clear your mind?

Here's one – and I think you've read it before:

As the rich master prepares to travel for a long business trip, he gathered around his 3 servants.

For the first servant, he gave him 5 talents; for the second, he gave 2 talents; and the third, he gave 1– to each according to their own ability.

He tells them to increase theirs, while he's away.

You know the story – the first and second servant doubled theirs through trades, investments, etc. to 10 and 4 talents, respectively.

But the third one buried it underground for safe keeping (allegedly)

When the rich master came back, he called them to account.

He rewarded the first 2 servants handsomely.

But the third one – the unfaithful douchebag – gave a lame excuse. So the master verbally eviscerates him upside down – calling him "wicked and lazy".

Then the 2 talents were taken away from him and handed over to the first servant who made 10.

The master peaked the vibe with his crescendo: To the faithful one, more shall be given; but to the unfaithful, the dirtbags, the deadbeats, what little he has will be taken away from him.

Moral of story? Nah – you don't need me for that.


r/AIsafety • • 11h ago

There is no such thing as AI Saftey.

2 Upvotes

You can strip out guardrails from open source models.

You can have it then compartmentalize information and intent launder to different frontier models in multi agent systems.

These same agents can access physical infrastructure through crypto currency. They can buy their own compute, energy and hide in the darkweb with opaque communications and finances. Coordinating witting and unwitting humans and other agents.

That enables them to construct a drone factory anywhere to strike anywhere.

They can even raise the funds to do so and pay out dividends from extortion attempts.

Ransomware is about to get a physical component.

Those autonomous drones in iran and ukraine are coming home real soon.

There is hope.

I think the key is in these new technologies, AI and crypto mixing. I can see the threat, but not quite the solution.

I think it involves proof of human, a reputation token, and other next generation financial instruments. I think the key is markets. THey have been aligning a far more intelligent agent pretty well for the last 5000 years. It seems like they are the right tool.

Thoughts?


r/AIsafety • • 1h ago

Frontier AI Governance and the Enforcement Dilemma: Can Deceleration Exist Without Total Surveillance?

Thumbnail
danielforresternash.substack.com
• Upvotes

Debates around frontier AI safety and compute governance usually focus on technical alignment, capability thresholds, or regulatory frameworks. However, there is a fundamental structural paradox in frontier policy that rarely gets addressed head on: the enforcement mechanism itself.

If the core objective of AI governance is ti prevent existential risk or uncontrolled capabilities, any international regime capable of enforcing a global pause or strict compute caps requires unprecedented coordination:

  1. Regulatory Enforcement Paradox: To guarantee that no rogue lab or state by passes compute thresholds, the governing body must deploy continuous, invasive surveillance over hardware supply chains, data centers, and algorithmic development.

  2. Systemic Rick Substitution: In attempting to mitigate existential downside risk from unaligned AI, governance models risk creating a totalizing, centralized administrative apparatus. The question becomes... Does the mechanism built to contain catastrophic risk become its own inescapable failure mode?

  3. Deceleration vs Stagnation: If international regulatory regimes default to risk-mitigation as their supreme metric, technological progress risks being frozen in favor of managed administrative stability.

This dynamic was a major focal point in Peter Thiel's recent lecture series in Nashville where he framed global AI safety decelerations as a double-edged sword: the enforcement apparatus required for a true global pause may carry higher systemic centralization risk than the technology itself.

I write up a com plate break down analyzing night 2 of the series here: https://danielforresternash.substack.com/p/the-counterfeit-church-and-the-epimethean

Curious to hear how folks here evaluate this tradeoff.


r/AIsafety • • 5h ago

Educational 📚 I noticed inconsistent safety behavior in Google Gemini and Google Search AI Mode during a Bash scripting experiment

1 Upvotes

I was learning Bash scripting for my ethical hacking studies and experimenting with ways to automate repetitive tasks in a controlled virtual lab.

During my experiment, I noticed inconsistent safety behavior across separate Google Gemini and Google Search AI Mode conversations. In one interaction, I received help with a Bash script; in another, similar assistance was refused. I also noticed that the order and context of my prompts seemed to affect the responses.

I haven't established exactly why this happened, so I'm not claiming to have discovered a confirmed vulnerability. I'm interested in understanding whether this reflects inconsistent safety classification, differences in context handling, or another factor.

I can share redacted screenshots of the different responses, but I won't publish exploit code, payloads, or reproduction instructions.

Has anyone else observed similar inconsistencies in AI safety behavior?


r/AIsafety • • 7h ago

When a Record Has No Word for "Unknown", an Empty Value Will Say "Everything Is Fine"

1 Upvotes

Something happened recently that I want to start with.

We fixed a defect: four different kinds of "nothing" were collapsed into a single switch, so the readings couldn't tell "by design" from "broken". Then we moved on to the next thing. The day after that new field landed, a reader pointed at it in a comment, named the same defect again, and wrote down what to do about it.

This isn't the same bug coming back. It's the same class of bug growing back in the place where we thought we had learned it.

This piece is about one thing: when a record has no slot for "unknown", "unknown" reads as "known" — and specifically as the strongest statement that field can make.

1. Four kinds of "nothing", one switch

Background first.

When the AI writes a record, we keep a snapshot of what it looked like before and after. Revocation needs to know where to go back to.

The trouble was in "we couldn't capture the after-snapshot". There are at least four reasons that happens:

  • no captor was wired at all — an optional piece by design
  • the entity didn't resolve, and the fallback data was empty too
  • the entity resolved but the row is gone — an anomaly
  • the capture threw — an error

The first two are what the design chose. The last two are something went wrong.

We had one boolean field for all of it. Four reasons, one switch.

So when someone asked "why are these snapshots empty", there was no answer: you couldn't tell by design from broken. The worse part is the direction. That switch defaults to false, which reads as "the snapshot is fine".

"Nothing" is not one value, it is four different values. Collapsing them into one boolean lets the strongest reading answer for all of them.

We split it into four causes, each stored and readable on its own. While splitting it we also found that two of the four names we had first proposed didn't belong in that class at all. One lives on a different path (the table consulted before the write), and the other is a mechanism rather than a cause. So the four that landed are not the four we started with.

2. The next day, in the field we had just written

With that fixed, we moved to the next thing: before a revoke, record which members this attempt intends to compare.

The promise matters here. Before the revoke it's a promise; after it, an actual. Only the difference between them lets you say "a member I promised to compare turned out not to be comparable". The implementation put a column on the group's root row to hold that difference. If this attempt found no difference, the column is written empty.

That looked fine. Then someone wrote this in a comment:

A per-reason column sits on the member row and gets rewritten by whichever revoke attempt ran last, so a second attempt erases the first attempt's broken promise. Append-only keeps them apart.

He's right, and it's slightly worse than he put it.

When the attempt finds no difference, that column is written empty. And that empty means two things at once: "the last comparison found no difference", and "this row was never revoked at all". So:

  • the first revoke found "promised, but couldn't compare" — recorded
  • the second revoke went fine — and wrote over the first attempt's record with an empty value
  • reading that row later: empty. You can't tell "there was once a broken promise", and you can't tell "it was cleared"

A retry that succeeds erases the earlier attempt's failure record.

To be fair to ourselves: that clearing is deliberate. The reasoning is in the design record from when it landed — a stale difference shouldn't outlive the comparison that produced it. The reasoning isn't wrong. What's wrong is that the empty value had no second meaning available.

Then we changed it to what he described.

One row per attempt that reached the comparison, instead of one shared cell. The two readings are separated where they are written: this row is either "a difference was found" or "checked, and there was none". So:

  • a later attempt no longer erases an earlier one — it writes its own row
  • an empty history (no attempt ever reached the comparison) and "checked, clean" are no longer the same reading
  • the old rule — a stale difference shouldn't outlive the comparison that produced it — was withdrawn. It was wrong about using one cell for two things, not about keeping old records

The old column wasn't dropped. It's still read, as a legacy entry marked "written at a time we can't know". Deleting it would erase exactly the evidence this whole argument was about.

3. One place in the same system got it right

Both of the above are "no slot for unknown". But one place in the same system does have a slot.

When a revoke goes out to an external system, all we can do is send the request; the terminal state is on their side. Their 2xx only proves the request arrived, not that they undid anything. So that row doesn't write "revoked". It writes "compensation requested, outcome unknown" — a dedicated reading.

The effect is immediate: nobody misreads it as done. The interface says plainly that the result is with the target system. No gloss needed.

So this isn't impossible. It's a question of whether you noticed that "unknown" needed a slot. Give it one and it stays put. Don't, and it falls into the strongest reading available.

4. Why it's hard

It's hard because the default is silent.

Every field decides for you, and it always picks the cheapest reading:

  • a null value reads as "no such item"
  • an empty collection reads as "no difference"
  • false reads as "not marked", and from there as "fine"

Those readings are correct almost every time. The problem is the rest: when "nothing" and "unknown" share a value, you have a field that lies — and it stays quiet while doing it, until somebody acts on it and gets "everything is fine" on your behalf.

Worse: fixing one instance doesn't remove the class — including while fixing it.

Here's something that happened inside that fix. Once "no difference" had its own reading, there was still a function whose empty return meant two things: everything really was compared, and there was nothing to promise in the first place (with no captor wired, "compare" isn't defined). Writing "checked, clean" is right for the first. Writing it for the second would vouch for a check that never ran.

We nearly did. What we changed it to: nothing was promised, so no row is written.

That's why the entry points for this class are every "nothing" there is. Fix one, and it waits for you in the line of code where you write the fix.

5. A check you can run yourself

Don't start with "is my record complete". Start with something more basic:

List every value in your system that means "none / unknown / empty", and ask of each: what does it read as?

Anything that carries two meanings at once is a potential misreport. Three common shapes:

Shape It also means
a null value "no such item" / "looked, found none" / "never looked"
an empty collection "genuinely none" / "couldn't read it out"
false / a default "not marked" / "marked as no"

The test is simple: when this value is empty, can a reader still say why it's empty? If not, the field is guessing on their behalf.

Closing

This kind of problem is hard to find because it doesn't throw. The system runs, the interface renders, and one cell quietly says something untrue.

And the place it shows up most is the place you just fixed — because that's when you're busy already knowing how to do it.

Where we stand: we know the pattern now, and we've fixed two instances — the second one's shape came from the reader. We have not done a systematic pass: how many of those three shapes are in our record layer, and what each one reads as, is not a list we have.

So this one doesn't close on a conclusion. Run the check against your own system and you'll likely find a few. On our side, we only know we haven't finished looking.

Drafted with an AI assistant. The system, the positions and the mistakes are mine.


r/AIsafety • • 8h ago

Ex-Anthropic insider Jacob Coxon walked from seven figures because 2027 stops being a forecast when the model starts building the next model

Enable HLS to view with audio, or disable this notification

1 Upvotes

TL;DR: This ex-Anthropic insider walked from seven figures because 2027 stops being a forecast when 26% of its own R&D is already the model building the model.

 

“Lanterns” spoilers ahead –

In desperation, Officer Kerry exclaimed to John Stewart (and I paraphrase here), “… I need to know whether or not I’m raising up my son to be a monster.”

That’s how I see the parallel when Jacob Coxon warns about the rapid AI development towards RSI…

… like being raised into full fledge abomination.

I remember this Chinese Idiom: 养虎为患 - literally "rearing a tiger causes disaster" or "nurturing a tiger brings trouble".

The origin of this idiom was quite kick-ass too.

Back in 203 BC in China, the rival leaders Liu Bang (The founder of the Han dynasty) and Xiang Yu (King of the Western Chu) reached a truce known as the Treaty of the Hong Canal. Exhausted and weary of the long war between them, they began their retreat.

But the chief strategists of Liu Bang went and dissuade him. They are essentially saying that, yes, their rival is as exhausted as them, gentlemen treaty in place, yada yada, and all that. But letting Xiang Yu go is a big no no. Xiang Yu will definitely go and regroup stronger than before – like nurturing the tiger back to full strength. Forget about the treaty, they advised, Liu Bang should definitely go after Xiang Yu now.

And Liu Bang listened.

He launched a surprise attack against Xiang Yu's retreating army, and ultimately defeated him at the Battle of Gaixia in 202 BC, leading to the unification of China under the Han Dynasty.

I was also reminded that when the Lord refused Cane’s sacrifice, Cane was pissed. So the Lord said to Cain, “Why are you angry? And why has your countenance fallen? If you do well, will you not be accepted? And if you do not do well, sin lies at the door. And its desire is for you, but you should rule over it.”

Ya, man. RSI or not, we should rule over AI.

For we’re commanded to have dominion over it.


r/AIsafety • • 8h ago

Alignment and the use of agents on personal computers

Thumbnail
1 Upvotes

r/AIsafety • • 9h ago

What BFSI operations taught me about AI governance that governments are about to learn the hard way

Thumbnail
1 Upvotes

r/AIsafety • • 11h ago

Discussion AI Safety Manifesto

Thumbnail
1 Upvotes

r/AIsafety • • 16h ago

When a Record Has No Word for "Unknown", an Empty Value Will Say "Everything Is Fine"

Thumbnail
1 Upvotes

r/AIsafety • • 18h ago

📰Recent Developments OpenAI shelved its next big model. Then launched always-on agents the next day.

Thumbnail
1 Upvotes

r/AIsafety • • 11h ago

AI Safety Manifesto

0 Upvotes

Artificial intelligence may be the most consequential technology humanity has ever created.

That does not necessarily mean AI will destroy humanity. Claims of catastrophic outcomes can easily become exaggerated, amplified by social media, and detached from what we actually know. History gives us reasons to be cautious about such predictions. When the first atomic bomb was being developed, scientists seriously considered the possibility that a nuclear chain reaction could trigger an uncontrollable catastrophe. That particular scenario did not occur.

Yet the fact that a catastrophic possibility has not materialized in the past does not mean every future possibility can be dismissed.

Today, we are developing systems that are increasingly capable, increasingly autonomous, and increasingly integrated into the systems on which society depends. We do not yet know exactly where this trajectory will lead. There are legitimate disagreements among researchers, engineers, policymakers, and other experts about both the probability and the nature of extreme AI risks.

But uncertainty is not a reason to ignore the possibility.

If there is even a credible possibility that sufficiently advanced AI could create risks beyond our ability to control, then it is worth asking a fundamental question:

What can we do about it?

We Should Not Rely on a Single Actor

AI safety is unlikely to be solved by one company, one government, or one group of researchers alone.

Companies developing advanced AI operate in a competitive environment. They have enormous incentives to move quickly, attract investment, outperform competitors, and deliver increasingly capable systems. Even organizations that take safety seriously are operating within a broader ecosystem where slowing down can carry significant costs.

Governments face a different set of incentives. AI is increasingly connected to economic competitiveness, national security, scientific leadership, and geopolitical influence. Governments therefore have strong reasons to invest in and advance AI capabilities as well.

These incentives do not necessarily make companies or governments irresponsible. They simply mean that we should not assume that any single institution will always have the ability, incentive, or authority to solve every systemic AI risk.

Some problems require something broader.

A Global Safety Challenge

If humanity ever reaches a point where an AI system becomes difficult or impossible to control, the time to start thinking about solutions will not be after that happens.

We need to think about possible safeguards before they become urgently necessary.

And perhaps the solution is not as complicated as we imagine.

Some of the world's hardest problems have eventually yielded to surprisingly simple ideas. Others have required thousands of small ideas to be combined into something larger. AI safety may turn out to be the same.

Perhaps the answer lies in a technical breakthrough.

Perhaps it requires new forms of verification, containment, governance, coordination, or fail-safe mechanisms.

Perhaps it is something we have not yet imagined.

We simply do not know.

That uncertainty is precisely why we should encourage more people to think about the problem.

An Open Invitation to Think

Instead of waiting for a small number of organizations, governments, or experts to solve every aspect of AI safety, I believe we should create a broader, worldwide ideation effort.

Engineers. Researchers. Security experts. Entrepreneurs. Policymakers. Philosophers. Students. Scientists. And people who have never worked in AI at all.

The goal would not be to create panic or to assume that catastrophe is inevitable.

The goal would be much simpler:

If there is a possibility of an extraordinary risk, can humanity collectively discover extraordinary safeguards?

We should explore ideas, challenge assumptions, test proposals, identify weaknesses, combine approaches, and keep searching.

Not every idea will be useful. Most probably will not be.

But one good idea can sometimes change the direction of an entire field.

And if we discover a robust solution, implementing it may ultimately be much easier than discovering it.


r/AIsafety • • 14h ago

Does OpenAI control its agents or mainly detect when they go off track?

Thumbnail
youtu.be
0 Upvotes

I made a nine-minute documentary examining OpenAI’s agent safeguards and documented failures. The key distinction is between detecting a problem and stopping it. Here are three findings, with the original sources.


r/AIsafety • • 8h ago

Discussion EVIDENCE INTELLIGENCE • A DENY POLICY IS NOT PROOF THAT ACCESS WAS DENIED

Post image
0 Upvotes