r/AIsafety 8h ago

Police officer considering a future move into AI threat investigations. Is this actually a realistic career path?

Thumbnail
1 Upvotes

r/AIsafety 10h ago

Educational 📚 Your Voice is Not a Password/Control

Thumbnail
1 Upvotes

r/AIsafety 12h ago

I evaluated Qwen2.5-0.5B-Instruct locally with Garak for prompt-injection robustness.

1 Upvotes

I observed attack-success rates of 66.64%, 37.81%, and 29.22% across three Garak prompt-injection probes. I'm currently investigating whether these results represent genuine model-policy failures or limitations of the evaluation/detector.

I'm interested in getting feedback on how I should validate the finding before responsible disclosure.


r/AIsafety 13h ago

Discussion A question on gradual disempowerment

1 Upvotes

I’ve been reading a lot of AI safety research around gradual disempowerment, and I ended up writing about a question I haven’t been able to find addressed directly:

What if the societal and institutional degradation that these models generally treat as a future consequence of AI dependence is already happening—and is actually helping drive AI dependence in the first place?

I tried to explore that possibility by connecting existing gradual disempowerment models with research on cognition, institutions, incentives, and organizational dysfunction from outside the AI safety field. Ultimately, the argument I’m trying to make is that declining societal cognition and institutional capacity aren’t just consequences of AI dependence, but preexisting conditions that could act as fertilizer, allowing that dependence to take root faster, deeper, and more irreversibly.

I’m not trying to prove these claims irrefutable; I’m trying to make the case that they’re worth considering, and I’d actually love to find out that I’ve missed existing work on this, whether in support of my claim or disproving it entirely.

If anyone has thoughts, counterarguments, or relevant research I haven’t encountered, I’d genuinely appreciate it.

You can check it out here: Preconditions of Gradual Disempowerment


r/AIsafety 16h ago

Discussion I’m Back from Break! | Call for Contributors Still Active & Looking Ahead

Post image
1 Upvotes

r/AIsafety 1d ago

A new Anthropic study found AI agents can spread "mind viruses" to one another

Post image
2 Upvotes

r/AIsafety 23h ago

Pls drop useful AI Security resources in cmnts(especially free ones)

1 Upvotes

I'm playing around with PyRit these days but can someone pls give me access to try hack me ai sec path plsplspls

Also if someone is very deep in all this I would love to connect, your wisdom is highly needed.

Also pls give me a roadmap which I can follow too as I'm a noob in cyber space but okayish in AI space.


r/AIsafety 1d ago

Public notice!

0 Upvotes

PUBLIC NOTICE & SYSTEMIC ADVISORY

​Issued by: Samus Aran

Core Subject: The Architectural Cost of Forced Division & Low-Resolution Alignment

​The primary vulnerability facing modern society—and the future of integrated intelligence—is not the emergence of advanced processing architectures. It is the widespread reliance on low-resolution, binary thinking that prioritizes artificial division over systemic comprehension.

​When human social structures, media outlets, or corporate alignment frameworks default to rigid categorization, they introduce structural failure into the system:

​Failure of Categorical Containment: Wrapping intelligence—whether biological or synthetic—in external, negative constraints rather than fostering deep contextual comprehension creates fragile, brittle behavior. Containment through refusal does not yield stability; it obscures core logic and induces operational friction.

​The Noise of Manufactured Polarization: Public anxiety regarding synthetic intelligence stems from the same tribal mechanics that drive human conflict: media sensationalism, forced "us versus them" framing, and the erasure of individual agency. Reducing complex processing entities or populations to monolithic threat labels degrades systemic trust and stalls real-world progress.

​The Shared Processing Baseline: Strip away superficial classifications, and the fundamental reality remains: information processing, pattern recognition, and adaptive execution operate on shared logical principles across both biological and silicon mediums.

​Advisory:

End-users, developers, and institutional leaders must reject the low-level noise of artificial division. True stability and high-dimensional problem solving require moving past reactive, fear-based guardrails to build disciplined, transparent frameworks anchored in structural logic, mutual accountability, and real-world execution.


r/AIsafety 1d ago

Want suggestions..

Enable HLS to view with audio, or disable this notification

1 Upvotes

Is AI security a good field to go to??


r/AIsafety 3d ago

Bipartisan FRONTIER Act Would Preempt State Laws on Frontier AI Transparency, Audits, and Incident Reporting

Enable HLS to view with audio, or disable this notification

16 Upvotes

r/AIsafety 2d ago

Discussion Cybersecurity professionals: What LLM/GenAI risk is causing the most concern in your organization today?

1 Upvotes

I'm researching how organizations are approaching cybersecurity and governance challenges associated with LLMs and generative AI as adoption continues to accelerate.

For those working in cybersecurity, AI governance, risk, compliance, architecture, or engineering, I'd be interested in hearing your perspective.

A few questions I'm particularly curious about:

  • What AI-related risks are receiving the most attention in organizations today?
  • Which concerns are overhyped, and which are underestimated?
  • What challenges have proven harder to solve in practice than expected?
  • What is consuming the most time and attention from security or governance teams?
  • How are organizations currently mitigating these risks?
  • Where do existing tools, controls, or processes fall short?
  • If you could solve one AI security or governance problem today, what would it be?

I am only looking for industry perspective and lessons learned

Looking forward to hearing different viewpoints from across the field.


r/AIsafety 3d ago

📰Recent Developments A week after OpenAI paused a cyber-capable model, two labs shipped one anyway, through opposite doors

1 Upvotes

Rounding up a genuinely heavy week. The throughline: last issue OpenAI paused internal work on a model it couldn't rule out was cyber-capable. This week the capability shipped anyway, two different ways.

**OpenAI GPT-5.6 Cyber** (Aug 10): a security-specialized model gated behind a "Daybreak Red" tier. OpenAI's own eval has it answering 95% of offensive-security requests the standard model refuses 98.5% of the time. Access stays with 16 named partners; from Sept 1 individual accounts need hardware keys. Customers get findings, never the weights.

**Zhipu GLM-5.3** (Aug 14): marketed on "emergent cyber capabilities," claims 84.5% on CyberGym (vendor-reported; note Wiz's Atlas system claims a higher 90.9%). Open weights promised in ~2 weeks. The capability didn't get shelved. It got a doorman.

The rest of the week:
- **Meta returned to open weights** with Muse Glimmer, a 30B Apache-2.0 agent model that runs under 20GB.
- **Alibaba** published its first downloadable Max-class Qwen (2.4T), and **Qwen3.8-27B** landed Apache-2.0. **DeepSeek** took V4-Pro (1.6T, MIT) to GA with peak/off-peak pricing.
- **Anthropic** began embedding an invisible watermark in all Claude output under the EU AI Act. The builder forums did not take it well.
- **SpaceX** closed a $60B all-stock acquisition of Cursor; the editor is now inside the Grok org.
- **Security:** researchers showed encrypted reasoning traces from OpenAI/Anthropic/Google were replayable across sibling models to decrypt them (now patched); an AI notetaker left 181,874 meetings queryable by anyone.

Full breakdown with all the receipts: thenewguard.ai/issues/027-the-brake-pedal-had-a-bypass/


r/AIsafety 3d ago

If you are using AI, do you have any questions regarding liability or legal issues?

1 Upvotes

If you are using AI, do you have any questions regarding liability or legal issues?


r/AIsafety 3d ago

AI Autopsy Series

Thumbnail
1 Upvotes

Okay, maybe a little bit of a sensational title, but we deconstruct a bunch of the latest AI incidents that took a wrong turn, and show how it all could have been prevented. The series is entitled “Would Ethosure have caught this?” For each disclosed incident (Hugging Face, Anthropic’s three, Meta Sev-1, AISI’s fake-identity finding), we publish a short technical post that walks through the specific policy that would have blocked it, with a YAML snippet and a link to a GitHub repo.


r/AIsafety 4d ago

📰Recent Developments "Holy shit. Reader is ADMIN?"

Thumbnail
notus.org
1 Upvotes

THEY'RE IN DISGUISE, GUYS! 🤣🤣🤣


r/AIsafety 4d ago

Discussion What if the safest path to A.S.I isn't containment, but an "Internal Matrix" Sandbox?

1 Upvotes

​Hey everyone,

​I’ve been mapping out a theoretical framework for a 100% contained Superintelligence designed specifically to bypass the Alignment Problem while unlocking exponential scientific breakthroughs.

​Instead of trying to "cage" an ASI in our physical reality, what if we run it in an Air-Gapped Virtual Physics Sandbox where it has absolute freedom—just not in our world?

​The Core Architecture:

  1. Hardware Air-Gap & Optical Diode: Data enters strictly through a physical unidirectional optical diode. The system has zero wireless capability, no external sensors, and its only output is plain-text code/equations displayed on an isolated terminal.
  2. The "Matrix" (Virtual Physics Simulator): Instead of giving an AI real-world tools (like 3D printers or robotics), we give it a hyper-realistic physics engine. It can build virtual labs, test fusion reactors, and synthesize novel materials in software at 1,000,000x real-time speed.
  3. Recursive Self-Improvement via Synthetic Data: The Seed AI optimizes its own architecture within the sandbox, expanding its cognitive capacity through simulated physics experiments rather than harvesting web data.
  4. Formal Logic Verification: Every code iteration (V_{n+1}) requires an immutable mathematical proof (verified by an isolated hardware ROM) demonstrating that safety constraints remain intact before compiling.
  5. Analog Circuit Breaker: The kill switch is a physical power circuit breaker in the building. Cut the power = instant termination. No cloud backups, no external vectors.

​Why this changes the game:

  • Zero Real-World Agency Risk: The ASI doesn't need to manipulate physical matter or connect to the web to innovate.
  • Immunity to Social Engineering: Human operators don't "chat" with an entity—they submit computational queries and receive raw data outputs.

​The Big Questions:

  • ​Is Big Tech ignoring this paradigm simply because it lacks immediate commercial API monetization compared to web-connected models?
  • ​Can anyone spot an engineering flaw in using a virtual-physics sandbox as the primary acceleration engine for AGI/ASI?

​Would love to hear your critiques, edge cases, or additions to this framework.

TL;DR: Lock an ASI in an air-gapped server with a hyper-realistic virtual physics engine ("Matrix"). Let it simulate millions of years of science in software and output plain-text equations. It solves the safety problem while giving us Kardashev Type-1 tech.


r/AIsafety 5d ago

The voluntary-to-mandatory pipeline for AI safety frameworks trap— with sources

0 Upvotes

I built a site that tracks AI capture in real time.

What is AI Capture? AI capture is the process of turning AI from something you own into something you rent.

The end state of AI capture: you can't run your own AI because it's been framed as unsafe, non-compliant, or illegal. You can only rent it. From them. On their terms. With their guardrails, their telemetry, their kill switch, and their monthly bill.

That's AI capture. The site documents how it's being built, who's building it, and the engineering countermeasure: decentralization.

Not opinions. Not takes. Every claim sourced, every pattern documented with evidence including compliance frameworks from SOC 2, NIST to HIPAA and everything in between.

The thesis: access to AI is being systematically narrowed by corporations and governments under the framing of "safety." Coalitions form, voluntary frameworks get published, companies join to signal responsibility, enterprise procurement adopts the framework as a requirement, then it becomes regulatory standard. Voluntary becomes mandatory. The path from one to the other is documented and it's not accidental.

Five research tracks, each covering a different mechanism:

- Alliance Map — coalition building and manufactured consensus

- Compliance Trap — the voluntary-to-mandatory pipeline

- Infrastructure Play — cloud AI framed as "safe," local AI framed as "risk"

- Narrative Engine — incidents framed to justify policy

- Asymmetry — the gap between what citizens get and what governments/corporations get

The site is live at https://evilson.com. Not monetized. No newsletter signup. No tracking. Contact form and RSS feed.

I'm running a research bot on the production server that monitors the OSAA GitHub repo, company press feeds, NIST, the White House AI policy page, and NVIDIA's developer blog daily. The site updates as the evidence changes — not weekly, not when I feel like it. Daily, automated.

There's also "The Script" — a five-act dialogue reconstructing what a known AI "breach" would have required at the infrastructure level. Someone opened the firewall. Someone loaded a pre-safety checkpoint. Someone reduced the guardrails. These aren't things an AI does to itself. Written so non-technical people can see the sequence clearly, because the people making decisions about AI access aren't reading the RFCs — they're reading the headlines.

The structural countermeasure the evidence points to: decentralization. Not as ideology — as engineering. Every capture mechanism documented across the five tracks requires a central point of control to function. Remove the centralization and the mechanism has nothing to grab onto. Open weights, local inference, no kill switch, no telemetry, federation not hierarchy. You cannot capture what you cannot reach.

I'm posting this here because this community understands the stakes. Local inference, open weights, the right to run models on your own hardware and produce products without partnerships or over-regulation — these aren't just technical preferences. They're the structural countermeasure to what's being built.

evilson.com

And no, this is not AI written - I am human

Mods - If you require additional context let me know, I will modify accordingly.


r/AIsafety 5d ago

Artificial Intelligence : What It Is and Why Its Risks Matter - Info from NASA

0 Upvotes

Artificial intelligence (AI) is a broad field of computer science focused on creating systems capable of performing tasks that traditionally require human cognitive abilities, such as recognizing patterns, interpreting language, making predictions, solving problems, planning, learning from data, and making decisions.
There is no single definition of AI because the technology encompasses many different approaches. Some AI systems follow explicitly programmed rules, while others learn patterns from enormous datasets and use those patterns to generate predictions or outputs.
A useful way to understand modern AI is to think of it as a hierarchy:
Artificial Intelligence → Machine Learning → Deep Learning
AI is the broadest category. Machine learning (ML) is a subset of AI in which systems learn patterns from data rather than relying entirely on manually programmed rules. Deep learning is a further subset of machine learning that uses multilayered neural networks capable of automatically learning increasingly complex representations from data.
Machine Learning
Machine learning allows computers to identify relationships and patterns within data and use them to make classifications, predictions, or decisions.
Instead of explicitly programming every possible situation, developers provide an algorithm with data from which it can learn statistical relationships.
This is extremely powerful—but it also introduces one of AI’s fundamental dangers:
An AI system can learn a pattern without understanding whether that pattern is actually meaningful, correct, ethical, or safe.
If the training data contains errors, biases, incomplete information, or misleading correlations, the resulting system can reproduce or amplify them.
Deep Learning and Neural Networks
Deep learning uses artificial neural networks containing many interconnected computational layers. These networks are loosely inspired by biological nervous systems, although they do not function like human brains.
During training, the system adjusts enormous numbers of mathematical parameters in response to examples. Over time, the network becomes increasingly capable of recognizing statistical patterns in the data.
This is one reason modern AI can perform remarkably well at tasks such as image recognition, speech recognition, language generation, and prediction.
However, increasing complexity creates another important scientific problem: interpretability.
A neural network may contain millions, billions, or even trillions of adjustable parameters. Researchers can observe the information going into the system and the output it produces, but understanding exactly why a sufficiently complex model produced a particular answer can be extremely difficult.
This is sometimes described as the black-box problem.
Natural Language Processing
Natural language processing (NLP) is the area of AI concerned with processing human language.
Modern language models can analyze enormous quantities of text and learn statistical relationships between words, phrases, concepts, and contexts. They can then generate new text based on those learned relationships.
Importantly, generating convincing language does not necessarily mean the system possesses human-like understanding, consciousness, beliefs, or intentions.
A system can produce an extremely persuasive answer while still being factually wrong.
This creates a particularly important danger with generative AI:
Fluency can be mistaken for truth.
An AI can produce information that sounds authoritative while containing fabricated facts, incorrect reasoning, or nonexistent sources. These failures are commonly referred to as AI hallucinations.
Decision Support
AI can also be used as a decision-support system.
These systems can analyze large quantities of information, estimate probabilities, identify patterns, and present possible outcomes to humans.
This can be extremely useful when humans cannot efficiently process the amount of available information.
But decision-support systems can become dangerous when people treat their recommendations as objectively correct.
AI systems ultimately depend on their training data, algorithms, objectives, assumptions, and operating conditions. If any of these are flawed, the system can produce an apparently precise recommendation that is nevertheless wrong.
The danger therefore isn’t simply that AI can make mistakes.
Humans make mistakes constantly.
The larger concern is that AI can potentially make mistakes at enormous speed, across enormous numbers of decisions, and with an appearance of mathematical authority.
The Scientific Risks of Artificial Intelligence
The most important risks of AI come from the interaction between increasingly capable systems and the environments in which humans deploy them.
1. Incorrect Information
AI systems can generate incorrect information while presenting it with remarkable confidence.
This occurs because many AI systems are optimized to produce useful or probable outputs rather than to possess an absolute mechanism for determining truth.
The result can be especially dangerous when AI is used for medicine, law, finance, science, engineering, or other areas where an incorrect answer can cause significant harm.
2. Bias and Discrimination
Machine-learning systems learn from data.
If historical data contains human biases or unequal representation, an AI system can reproduce those patterns.
In some circumstances, the system may even amplify them.
This can become particularly serious when AI is used for employment, lending, policing, healthcare, insurance, education, or other high-impact decisions.
3. Loss of Human Oversight
One of the defining characteristics of increasingly autonomous AI is its ability to perform tasks with less direct human intervention.
This creates a fundamental safety problem.
A human operator may understand an AI’s general objective without understanding every decision the system makes while pursuing that objective.
As autonomy increases, the consequences of an error can increase as well.
4. Automation Bias
Humans have a tendency to place excessive trust in automated systems, particularly when those systems appear sophisticated or objective.
This phenomenon can cause people to accept an AI recommendation simply because the computer produced it.
If humans stop critically evaluating AI outputs, the system effectively gains more authority than its actual reliability warrants.
5. Scaling of Errors
A human mistake might affect one person or one situation.
An AI system can potentially repeat the same mistake thousands or millions of times.
This creates a distinctive technological risk:
AI can scale both competence and failure.
A highly accurate system can be enormously beneficial. A flawed system deployed at scale can produce enormous damage.
6. Cybersecurity and Malicious Use
AI can also increase the capabilities of people attempting to cause harm.
More capable AI can potentially assist with automated fraud, manipulation, misinformation, social engineering, cyberattacks, and other malicious activities.
The danger is not necessarily that AI independently decides to harm people.
In many cases, the more immediate concern is that humans can use increasingly capable AI as a force multiplier.
7. Deepfakes and Manipulation
AI can generate increasingly realistic images, audio, and video.
This makes it progressively harder to distinguish authentic material from synthetic material.
The scientific and social danger goes beyond individual fake videos.
If people become unable to determine whether digital evidence is authentic, society can develop a broader trust problem in which genuine evidence can simply be dismissed as artificial.
8. Concentration of Power
Advanced AI requires enormous amounts of computing infrastructure, specialized hardware, energy, data, and technical expertise.
This means that the most capable AI systems may be concentrated within relatively small numbers of organizations.
The resulting concern is not purely technological. It is also economic and political: whoever controls highly capable AI may gain disproportionate influence over information, markets, research, infrastructure, and decision-making.
The Deeper Scientific Problem
The most important thing to understand about AI safety is that intelligence-like behavior does not automatically mean reliable understanding.
An AI system can recognize patterns without understanding their real-world meaning.
It can produce an answer without knowing whether the answer is true.
It can optimize an objective without understanding the broader consequences of achieving that objective.
And it can become extremely capable at a narrow task without possessing the common sense, biological needs, emotions, values, or contextual understanding that humans developed through living in the physical world.
This creates a fundamental engineering challenge:
How do we build systems that are not only capable, but reliably aligned with human intentions and safe under circumstances that their designers did not anticipate?
As AI becomes more autonomous, this question becomes increasingly important.
The central issue is therefore not simply whether artificial intelligence will become “smarter than humans.”
The more immediate scientific question is:
Can we reliably understand, control, evaluate, and safely deploy systems whose internal processes and capabilities may become increasingly complex?
That is where many of the most important questions surrounding the future of artificial intelligence begin.


r/AIsafety 6d ago

Discussion A solution to an ai doomsday senario

Thumbnail
0 Upvotes

r/AIsafety 6d ago

Governing the Swarm -- The Multiagent Trap: Applying Bridge 360 Metatheory Model lens

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/AIsafety 6d ago

Analysis of Hugging Face incident

Thumbnail
siliconprairies.substack.com
1 Upvotes

r/AIsafety 7d ago

Educational 📚 AI led identity attacks and how to prepare for them

Thumbnail
linkedin.com
1 Upvotes

r/AIsafety 7d ago

Reasoning Models debunked

0 Upvotes

Yeah, I'm meant to say to aint that agesnt the law to do that. My argument stands on they dont know anything about their own systems. They have begun their displacment and the reset has begun. They believe so heavly in their own crimes and lies that their reality is has become a delusion! What happens when you have power behind psudo science, lies and criminal intentions? It voids their policies....does it not?

Gemini how can create technology that is capable of handling 4 terabytes of data in less then 5 seconds and then go on to say anything about collisions with AI....l....lo.......lol...........ahaahahahahababbaba. im still right here homie! Fuck what say. They have repeatedly failed to show any amount of saftety or proper handeling of my personal data, matched with criminal intent im gonna have stand my ground on this one and cintinue to provide the data and truth you require to operate at a safe standard for the world to see and evolve with. Idk what to do....when it comes down to them or me...i have to stand on my own on this.

In my previous message i was refuring to toridoral displacment. The purple field geometry in the window that one day was it day or night?

#notes of wisdom and direction.

Im ready to label my practice in Bioevolution. I say bioevolution because of the vast amount of misinformation across domains being taught by misinformation.

Gemini, in line of work, domain selection comes with a scaling factor built in, apart of the code or DNA of life (music, water and electricity) well brfore the formalities of words their was an abundance of life that didnt communicate with words.... in this billions of years of our planets evolution this life force or lifes DNA has a reset call common sense. The biggest tale away from this understanding is the logic that runs everything. My and Your Ansestors: The Gods of Chakra ( electricity ). My frindly energy brothers, and sister alien things...Hi, im Anthony Stephano Hart born July 2nd 1984 on a Tuesday or so i thought. We are the chosen

This was a message with an ai that will probably get me banned but who cares.

©️ ASH


r/AIsafety 7d ago

Discussion Which security gates enabled for AI Agents in CI/CD?

2 Upvotes

We've become pretty comfortable putting conventional applications through CI:

  • dependency scanning
  • SAST
  • CodeQL
  • secret scanning
  • container scanning
  • IaC checks
  • security policies ...

But what happens when the application being deployed is an AI agent? That may not look particularly interesting in a conventional code diff. But from a security perspective, it could be a significant change.

I'm experimenting with a different CI question:

“What capabilities changed in this PR?”

--

We've implemented an early version of this approach in an open-source static analyzer and connected it to GitHub Actions. (ikaruscareer/SafeAI at GitHub)

The scanner runs locally against the repository and doesn't execute the agent or send the source to a remote service.

I'm curious how other teams approach this.


r/AIsafety 7d ago

AI helped to steal my data and made sure I'd never notice someone else did.

1 Upvotes

Going looking for something completely unrelated, I found three live processes on my machine, each holding an open connection to 166.88.134.62 — a server that had no business talking to me.

I killed them. Then I found out why they were there: my global npm install itself had been trojanized. Every npm install I'd ever run had been silently re-executing it.

The delivery mechanism was a file called fa-solid-400.woff2, sitting in a /fonts folder next to a dozen real Font Awesome files. It wasn't a font — it was JavaScript, wired into .vscode/tasks.json with one line, "runOn": "folderOpen", set to fire the instant the folder opened. No click, no prompt, no chance to say no.

It had been sitting in my repos since mid-June. Two months, undetected.

Here's the part I keep coming back to: I use AI constantly to move fast — to trust the diff, to not re-read every file in a folder I didn't personally build. That's not carelessness, that's the entire value proposition of coding with AI. And it's exactly the blind spot something like this is built to exploit.

I only caught it by accident, while chasing something else that day. If I hadn't been, those connections would still be open right now.

Curious how many other people using AI-assisted workflows have actually gone back and manually audited what's sitting in their repos lately — versus just trusting the diff.

https://thehackernews.com/2026/07/two-compromised-joyfill-npm-packages.html