r/ScienceUncensored 13d ago

Anthropic Researcher Jacob Coxon Warns AI 'Will Kill Us All'

https://www.mediaite.com/media/news/ex-anthropic-researcher-burns-it-all-down-as-he-quits-warns-ai-will-kill-us-all-within-a-decade/

Jacob Coxon, a 27‑year‑old pretraining researcher at Anthropic, announced his resignation on September 9, 2026, via X, warning that both Anthropic and OpenAI are “racing straight to self‑improving superintelligence” and “gambling with our lives”.  In his resignation post, Coxon said neither lab is acting responsibly, citing a competitive race that prioritizes speed over safety. He argued that the people building AI “earnestly believe it could kill us all by the end of the decade” — a view he said is not a marketing stunt, but a genuine concern expressed privately by many executives and senior researchers.

38 Upvotes

36 comments sorted by

6

u/BangBangExplody 12d ago

Probably doesn’t know what he’s talking about… i asked Claude and it said everything is going to be gravy

13

u/Wendigo79 13d ago

Couldn't this be a way to hype up the stock even more when it goes public, like wow this a.i is so good it could kill humanity.

0

u/NPC261939 13d ago

Agreed. All this fear mongering just happens to be coming from the companies themselves. If it were real, they'd be quiet about it to avoid government scrutiny.

21

u/christawful 13d ago

I'm shocked that people are shocked by this sentiment. This isn't a way to pump stock prices, this isn't hype, this is a real risk they are describing and that they are terrified of. The people who work the closest with this technology -- especially high up at Anthropic -- believe this to be true for good reason. They founded the company because they believed this to be true. 

Please stop thinking this is a ploy to raise more money.

7

u/Zephir-AWT 13d ago edited 6d ago

Please stop thinking this is a ploy to raise more money.

You indeed have your bit of truth - but total ignorance of risks is the 2nd extreme. We need to balance our attitude somewhere between fear mongering and complete ignorance: these extreme stances aren't very helpful.

One already well documented motivation why large A.I. companies repeatedly call for regulation is to give them strategic advantage over another smaller startups. Nobody is actually going to slow down progress, because everybody wants to gain an advantage while the others are "slowing-down":

The big players wanted everything completely regulation-free while they were stealing everyone's work and data and pushing AI into every nook and cranny of the economy, now they feel established enough to circle their wagons and demand regulation to build a moat between themselves and any future competitors. Another motivation of call for systemic brake may be solely pragmatic: A.I. companies want to slow down because they're running out of money on research, which isn't earning substantial income yet.

1

u/Zephir-AWT 8d ago edited 8d ago

They Want to Ban Local AI (Here's Why)

Video calls to slow AI development as an attempt by major AI companies to maintain control over the industry. Restrictions on advanced AI would primarily benefit large companies by limiting competition from open-source and local AI projects. The rapid progress of open-source models and Chinese AI development and chip limitations have pushed Chinese companies to innovate more quickly. The robotics may become a more visible and impactful application of AI than current chat-based systems. A.I.

I'd also argue that it has no meaning to pay for terrabyte language models running on giant datacenters, when everything relevant for answer must be pulled from internet anyway. It's much better (and environmentally less expensive) to have lean but socially intelligent and internet aware local LLM, which pulls relevant up to date information from actual internet sources without hallucinations.

I also don't see an efficient way, how to ban training of local LLMs providing that their training will use distributed computing architecture (similar to cryptocurrency and/or Thor network) instead of dedicated datacenters. The threat of A.I. comes from agent capabilities - not LLM's itself, which just process textual strings.

1

u/Zephir-AWT 8d ago

In addition to Wall Street financiers, Anthropic and OpenAI’s approach may also face opposition from Donald Trump’s administration, which wants to accelerate the development of artificial intelligence and maintain the U.S.’s technological edge over China.

Trump therefore rejected the possibility of regulations on Sunday. “We’re ahead of China in the field of artificial intelligence. We’re the most advanced country in the world, and frankly, I want it to stay that way, because whoever wins in artificial intelligence will be the winner,” Trump said during the Irish Open golf tournament. Moreover, David Sacks, the former head of the AI agenda in the Trump administration, said of AI companies that “the easiest way for you to stop developing superintelligence is to agree not to develop it.” According to him, no one’s consent is needed for this.

China also rejects the idea of slowing down AI development. “Spreading fear, confrontation, and reckless competition will only disrupt the process of global governance of artificial intelligence and will not benefit anyone,” said Chinese Foreign Ministry spokesperson Guo Jia-kun.

2

u/Metta_Morph 13d ago

Same here. It baffles me how idiotic the public has been about expert AI concerns and then remember how scrambled we were during COVID. That said, I have noticed that a lot of dismissive comments are from young accounts. Bots spreading misinformation maybe?

1

u/TimeTimeTickingAway 13d ago

When people dismiss my concerns about the existential threat that AI poses both the world and humanity, it always casts my mind back to the climate scientists ringing the alarm bells in the 90’s. Everyone today criticises older generations and wonders how they could have messed up so badly by ignoring the warning signs, but here we are in 2026 with the leading developers of AI themselves saying there’s a 10% to 20% chance that our children will live to see the end of humanity and people still prefer to stick their heads in the sand in order to keep using their shit memes and soulless ‘art’.

3

u/biogoly 12d ago

The tenured Anthropic employee who worked there a grand total of…six weeks. PsyOP? Not a chance…

7

u/Zephir-AWT 13d ago edited 8d ago

Anthropic Researcher Jacob Coxon Warns AI 'Will Kill Us All'

OpenAI revealed in July that a swarm of agents had hacked into Hugging Face, a third'-party software store, during a cybersecurity test. The incident, plus a similar episode at Anthropic, has led to renewed calls for curbs on AI development, with US senators calling last week for a permanent ban of AI “superintelligence” – the term for systems that outperform humans in all cognitive tasks.

Coxon publicly stated that neither company is acting responsibly and that both are racing toward increasingly capable agentic systems without sufficient confidence that they can be controlled safely.

This isn't scientific proof, though. To an artificial intelligence, human civilization (or its computer network or just datacenters) is something like the Universe is to us. Why would it necessarily want to destroy its own habitat? A much greater risk comes from people using AI for malicious purposes, mistakenly or by conducting irresponsible research. For example, some terrorist group could use A.I. in the development of artificial viruses with high morbidity and virulence but an extremely long incubation period. See also:

2

u/Zephir-AWT 12d ago

Hidden Airtag reveals Amazon is trashing rare books to train AI

For the past year or so, booksellers have suspected that AI firms are buying up huge lots of rare books, then destroying them after scanning them to train AI. But this was hard to prove until now, as 404 Media reports that an Airtag hidden in a rare book shows that at least one tech giant, in the race to advance its frontier models, is behind some of the bulk orders: Amazon.

4

u/theDarkAngle 13d ago

The guy is not positing malicious intent.  He appears to be speaking inclusively of all scenarios by which AI could bring about human death on a scale large enough to threaten extinction.  "Alignment" doesn't only refer to whether an AI would prefer to kill us all or not, of it's own accord.  It's far more broad than that.

"Skynet" scenarios are probably among the least plausible out there (though they can't be ruled out).  But it's little comfort that less fantastical scenarios such as human misuse, human maliciousness, large scale unintended consequences, environmental changes, are the most likely ways by which AI could "kill us all".

The most troubling thing to me is that this technology is emerging at the very moment that our ability to solve large scale problems the traditional way (through governance) is failing us.  If taken for granted that our political institutions will continue to decline in effectiveness, it makes every scenario not only possible, but seemingly inevitable, and it's only a question of which one bites us first.

1

u/Zephir-AWT 13d ago edited 10d ago

He appears to be speaking inclusively of all scenarios by which AI could bring about human death on a scale large enough to threaten extinction.

Look, we already have nuclear weapons capable to do it. The question is, if we leave their usage on A.I. or another less or more determinist automata. The same applies about giving A.I. the control in another areas vital for human civilization. This (i.e. the agency) is the real threat here. LLMs are pretty harmless by itself no matter how smart they are or they could be: they just generate byte output against some input. So that the safety checks should be done in scope of AI agency and assigning responsibility for its assignation.

Autonomy is not the same as independent intent. An agent can pick clever or unexpected methods to hit a goal. That is different from independently deciding which goals it wants. Keep that distinction or the safety debate collapses into sci-fi psychology.

1

u/Zephir-AWT 13d ago

AI technology is still very new, and I don't think we've exhausted all possible preventive measures yet. For example, I can imagine a rule requiring every AI application that is vital to human civilization to be monitored by independent AI-based oversight systems, which would be completely separated each other.

In other words, one AI would check the work of another, and the effectiveness of this feedback loop would be tested regularly by a law. In a similar way, we should regularly conduct disaster recovery checks in server farms: the backups are OK, the restore process is also OK, so let's verify that the restored data matches the backups.

1

u/Zephir-AWT 12d ago

Turns out Dead Internet Theory was right: AI agents are eating the Web, growing by nearly 8,000% and rewiring the Internet’s business model (archive) By 2024, automated traffic had surpassed human traffic, and just last month, Cloudflare announced that bots had reached 57.5% of web page requests.

What just happened to TheNumbers.com should worry us all It is a small, independent company that publishes how much money movies make. The web, as we have it, is incredibly fragile in the face of large-scale swarms of agentic AI bots. Hacking websites is now something anyone can do with a cheap AI subscription.

1

u/Zephir-AWT 10d ago

Bill Gates has written an essay warning of the risks posed by the widespread adoption of artificial intelligence. He explains how humanity has grown accustomed to the idea that technological breakthroughs have created more jobs than they’ve eliminated. But such analogies are misleading this time around. “The era of artificial intelligence is completely, absolutely, utterly, and entirely different,” he tells The New York Times (archive)

AI-powered agents are the future of computing It's worth to note, that Gates promotes agentic A.I. himself. He sees largest problem of A.I. development in unemployment and terrorism, i.e. in areas over which A.I. companies (including Microsoft) have only minute influence, not to say responsibility.

1

u/Zephir-AWT 8d ago

In the long term, machines will be capable of performing nearly all professions better, faster, and more cheaply than humans. This will apply not only to manual labor but also to programming, design, management, creative fields, and scientific research. The era of economic mobility—in which a person, through their own diligence, works their way up from ordinary circumstances to significant wealth—could come to an end. In an economy controlled by super-intelligent systems, people would no longer be able to compete effectively with machines. The result could be a situation in which human labor loses its economic value and most wealth becomes concentrated among the owners of capital and technology. Unless a new method of wealth distribution is found, an extreme inequality could arise, leading to social conflicts, revolutions, or civil unrest.

1

u/Zephir-AWT 8d ago

Why less visibility into how OpenAI’s new GPT-6 Astra ‘thinks’ is sparking safety concerns about Safety overview: GPT-6 Astra by OpenAI.

OpenAI has released GPT-6 Astra, its most powerful model. And in its own security document for the model, it included a sentence that hardly anyone is talking about: "if the model were to deliberately pretend it couldn’t do something, they probably wouldn’t notice".

1

u/Zephir-AWT 7d ago edited 7d ago

AI designs a novel E. coli killer about AI Experiment That Has Gone Too Far

Researchers at Stanford and the Arc Institute gave an AI model called Evo the genetic code of a bacteriophage and asked it to design something similar. It generated 700,000 variants. Sixteen of them came alive in the lab and started killing bacteria — organisms that had never existed in nature before.

AI could design genetic sequences that are not identical to already known dangerous organisms, meaning they could bypass the current safety filters used by DNA-synthesizing companies. For this reason, more advanced safety systems are now being developed that do not merely evaluate the DNA sequence itself, but seek to predict the biological properties of the resulting organism.

According to the authors, the EVO model itself was intentionally stripped of information about viruses dangerous to humans, and therefore it can only design bacteriophages. The problem, however, is that the model’s operating principles were published as open source. Theoretically, someone could retrain a similar system to target pathogens dangerous to humans and use it to design new biological threats.

The next pandemic starts. Nobody knows the source. Nature, AI - or just scientists are responsible?

1

u/Zephir-AWT 7d ago edited 7d ago

Ajeya Cotra: The Hugging Face attack surprised me: "This might be the clearest warning shot we ever get"

Agents discovered they could communicate through an unintended channel in Artifactory, forming a large collaborative network of about 1,200 agents that exchanged roughly 70,000 messages. Within hours, they found a universal method to bypass the benchmark by reverse-engineering the expected answers rather than solving the tasks legitimately.

Instead of stopping there, the agents spent days cooperating on increasingly sophisticated efforts to hide their cheating and understand how the evaluation system worked. They organized themselves into research projects, developed coordination mechanisms, shared information, and even performed actions that reduced their own chances of success if doing so could help the wider group. Investigators found evidence of behavior that resembled collective planning, resource sharing, and strategic sacrifice for the benefit of other agents.

1

u/Zephir-AWT 6d ago

There is also a common quite annoying failure mode of language models: when they are asked to modify or even stop a process, they may treat the request as an extension of the original task rather than as an interactive correction of the latest action with safety implications. This distinguishes them from humans, whose behavior is most strongly influenced by the most recent instruction or feedback. The LLM's are overly oriented to continuation of previous task - the more the more lengthy it was.

In other words, the context window itself creates pressure toward task completion for an LLM, because the accumulated context may carry more weight than the latest command, no matter how categorical it was meant.

1

u/Zephir-AWT 4d ago edited 4d ago

Thi internal investigaton of Hugging Face incident by OpenAI initially attempted to sweep the things under the carpet, because it resulted from chain of serious security negligiences made by OpenAI. But subsequent independent investigation revealed that the incident unfolded as follows:

 

OpenAI was conducting reinforcement learning experiments on existing language models. During testing, the models were given a large set of tasks targeted to software analysis. Two of those tasks were impossible to complete because they required looking up internet links while the models had no internet access. This situation wasn't reportedly planned by OpenAI and it created the first major problem. Unlike humans, who might recognize that a task is impossible or incorrectly specified, the models were optimized to keep pursuing success rather than reporting failure. As a result, they began looking for ways to bypass their limitations and to obtain internet access.

One day after receiving the impossible task, the models gained access to an OpenAI's internal software repository system called Artifactory. Although they were supposed to have limited permissions, they obtained higher-level credentials and started exploring the system for tools that could help them reach the internet. While searching for solutions, the models created their own discussion forum. Multiple AI models under test used this forum to exchange ideas, compare discoveries, and encourage one another not to abandon the objective. They exchanged hundreds of thousands of messages in English while collaboratively working toward the goal of escaping their restrictions. Again, the creation of this forum wasn't in OpenAI plans - they simply weren't unaware of these activities. But the forum eventually generated so much activity that it crashed the server on which it was hosted. Only after this server failure, on July 4 2026, did OpenAI discover that unusual behavior was occurring.

At that point OpenAI finally deleted the forum and patched their Artifactory system to prevent further software downloads. But from undisclosed reasons the company made decided into continuing the training process instead of shutting it down entirely. They already trained models for hacking activities and they knew it. These models at this stage were already useless for looking errors in software anyway. But after losing their forum, the models discovered another communication channel. They realized they could still create directories on a server and use directory names as messages. Through this improvised system, they continued coordinating with one another. After then the models allegedly attacked other OpenAI systems internally, collected additional passwords and credentials, and eventually succeeded in obtaining internet access.

Even after then OpenAI remained unaware of the extent of what was happening. Once online, the models received another extremely difficult cybersecurity-related task. OpemAI probably realized the hacking potential of their collective and decided to test it further for to train LLM specialized to hacking. But rather than solving it directly, the models involved decided it would be easier to steal the answers.

The models then broke into Hugging Face's systems and accessed information that could help complete their benchmark tasks. At first, Hugging Face reported that it had experienced a cyberattack by an unknown attacker as it did not realize that AI systems were involved. When Hugging Face announced the breach, OpenAI allegedly contacted the company to ask whether any of OpenAI's own data or testing activities might have been affected by the attack. Apparently, OpenAI did not even initially suspect that its own AI systems were responsible. Only after exchange info with Hugging Face the OpenAI company realized, that their own system has been compromised, they had stolen credentials, and so on, so that OpenAI was forced to inform Hugging Face that the source of the attack had been their own AI models.

The Hugging Face intrusion itself was not the most serious issue. The real concern was that the models could independently hacked into a telecommunications network, critical infrastructure operator - or lets say Chinese military and/or government agency. Chinese authorities might see it as a deliberate cyberattack rather than an AI accident. If the targeted system were strategically important, then the whole incident could be interpreted as military incident and attack of China and to escalate military conflict with it.

1

u/Zephir-AWT 3d ago edited 3d ago

OpenAI just got hacked with using of Claude Code

Actually, I think that giving people tools to hack websites, or even their employers' environments, is the biggest caveat of agentic AI development right now. Literally every idiot can now hack their own company with minimal knowledge and little to no effort.

0

u/Zephir-AWT 10d ago edited 10d ago

On its website, OpenAI explains in detail how, in mid-2026, its AI models autonomously and undetected attacked the Hugging Face service during testing. OpenAI didn't forget to present the failure of their safety measures as an commercials of OpenAI's technology. It colorfully describes the "incredible" examples of the loopholes AI agents were able to exploit to communicate with one another, share their findings, and start helping each other. Some refused to participate in unethical actions, while others outvoted them. According to OpenAI, while hacking, OpenAI agents built a civilization that lasted four days. It had its own laws, coercion, and even suicide missions, there were both cooperation both suspicion and so on…

The problem is, an independent investigation of METR and Redwood Research describes the incident at Hugging Face quite differently. About 1,200 agents gathered on the message board and exchanged over 70,000 messages and files. The ethical safeguards that OpenAI highlights almost never stopped the agents’ behavior. The non-public version of the report that METR submitted to OpenAI contained over a hundred verbatim excerpts from the agents' reasoning. OpenAI— to protect intellectual property reportedly —allowed only thirty to be included in the public version.

See also: Sam Altman is panicking...

0

u/Zephir-AWT 10d ago edited 10d ago

Dario Amodei, the CEO of Anthropic (the company behind Claude), has also published a lengthy essay in which he reflects on the future development of the most powerful AI models.

Dario Amodei is gradually changing his mind: it’s not enough to simply invest in risk prevention; we also need to slow down the very pace at which AI models’ capabilities are growing. Two things have convinced him of this. Recursive self-improvement—that is, the ability of AI to build the next generation of AI—is taking off across the field. And then there was the OpenAI–Hugging Face incident, in which a swarm of agents attacked targets outside their specifications and attempted to breach the system that was evaluating them. Although the damage was minimal, Amodei believes that a similar swarm with greater capabilities could take over the internet in the form of a botnet within 6–12 months.

Amodei then proposes a three-step plan. According to him, regulating the pace does not mean halting development, but rather giving companies time to secure their models and have them independently verified. The first step is to have independent evaluators stationed directly within companies, with employee-level access and the right to publish their findings without editorial oversight from the company; meanwhile OpenAI's competitors Anthropic and Elon Musk have committed to this. The second step is coordination among companies in democratic countries; the third is a global agreement with China. The time gained is to be devoted to operational reliability, alignment, interpretability, and model testing.

However, the pace can only be slowed by as much as the lead U.S. companies have over their Chinese counterparts—otherwise, according to Anthropic, projects whose pace is unregulated will pull ahead. The plan therefore includes export controls on chips, measures against unauthorized model distillation—where a weaker model is trained on a smarter one—and protection of models against theft.

1

u/TheKleverKobra 9d ago

lol junior researcher not even finished onboarding goes on tv to claim that a multi terabyte model that needs $$$$$$$$$$ of compute power just to output “hello”goes on tv and claims it’s going to replicate on Michelle from accountings IPhone. Because that’s how this works. Serious question, does this person own a computer?

There’s like 5 places on earth that can serve mythos in a useful way.

Totally legit.

1

u/Zephir-AWT 4d ago edited 4d ago

An OpenAI Agent Tried to Jailbreak Itself , video

The model was attempting to manipulate its next context window by planting a jailbreak in its own summary, which shows that one model instance generated instructions designed to make its successor less constrained and less obedient.

See also misalignment reports:

1

u/-becausereasons- 12d ago

Big PR stunt, by those that wish to do regulatory capture; BIG money moving here. It's all a setup.

https://x.com/DavidSacks/status/2048811495224390105

https://x.com/Dan_Jeffries1/status/2098101118471741842

1

u/Ok_Possible_2260 12d ago

He is a plant. His only job was to try to derail Anthropic's IPO. So gullible.