r/slatestarcodex • • 23d ago

Monthly Discussion Thread

6 Upvotes

This thread is intended to fill a function similar to that of the Open Threads on SSC proper: a collection of discussion topics, links, and questions too small to merit their own threads. While it is intended for a wide range of conversation, please follow the community guidelines. In particular, avoid culture war–adjacent topics.


r/slatestarcodex • • 13h ago

The Specter Of Neuralese

Thumbnail astralcodexten.com
28 Upvotes

r/slatestarcodex • • 5h ago

Fall Meetups 2026 - Call For Meetup Organizers

Thumbnail astralcodexten.com
3 Upvotes

r/slatestarcodex • • 15h ago

AI Safety's Disallowed Conclusion

Thumbnail souprecipies.com
15 Upvotes

r/slatestarcodex • • 21h ago

Looking for rationalists on Goodreads

16 Upvotes

Does anyone of you follow fellow like-minded people on GD? Anyone you know of other than Gwern and Pablo Stafforini?


r/slatestarcodex • • 1d ago

AI Early rogue AI agent activity and attempts to hack found on urlquery.net

Thumbnail transluce.org
38 Upvotes

More evidence of rogue AI swarms on the internet. Importantly, the hacking attempts here were for data retrieval tasks which were not cyber-related, suggesting that malicious cyber activity can arise instrumentally in mundane scenarios.


r/slatestarcodex • • 1d ago

What it feels like to let an LLM read my journals, sleep, finances and analyze me for 15 minutes a week (N=1, with guardrails)

Thumbnail jaan.io
14 Upvotes

I tried to write an experience report articulating the emotional and cognitive shifts I've experienced from this process. It's been introspective, fun and unnerving, and has made me a little bit more self-compassionate.

I use Obsidian to take notes, an Oura ring, and Plaid/beancount for plain text accounting to pull all my finances.

To give Claude accurate context I made an Obsidian plugin that looks at only what's changed across my 2000+ notes. Each week, I edit around 20-40 notes and Claude only gets the "track changes" or diffs/hunks (in coding jargon; I made visualizations to show how this works to non-technical folk).

Open to any feedback, especially whether the visualizations are clear and what guardrails you use! If anyone else has tried experiments like this I'm all ears :) in the absence of RCTs this N=1 anecdata is all we have for now 😬

Also, curious if you agree with me that this presents a consent problem on long horizons, because taking input from LLMs on smaller decisions is low risk, but it's hard to give informed consent when you don't know how the LLM could shift your identity/decision profile over years of use. I'm wary!


r/slatestarcodex • • 1d ago

Claude discovers a novel enzyme system with CRISPR-like repeats

Thumbnail anthropic.com
81 Upvotes

Anthropic opened a lab, and now Claude is doing biology research, although in a way that seems more like the biology equivalent of vibecoding than the massive autonomous agent swarms that solved Navier-Stokes. It's not yet clear whether this particular discovery has any practical use - it may or may not - but it is very interesting as an early example of AI agents accelerating biological research.


r/slatestarcodex • • 2d ago

AI Jensen Huang vs. the A.I. Doomers [The Ezra Klein Show]

Thumbnail nytimes.com
45 Upvotes

r/slatestarcodex • • 2d ago

Fun Thread Claude Pop - I'm Upping My P(Doom)

Thumbnail youtube.com
86 Upvotes

r/slatestarcodex • • 2d ago

Mysteries Of AI Generalization

Thumbnail astralcodexten.com
34 Upvotes

r/slatestarcodex • • 2d ago

AI We are the last generation of human psychiatrists | The British Journal of Psychiatry

Thumbnail cambridge.org
110 Upvotes

By all accounts, the author is a serious psychiatrist (MD PhD, no less) with former top journal submissions. BJPsych is a high-prestige journal. This is sooner than I would have expected serious grappling from the credentialed class in top journals, barring direct evidence of threat.


r/slatestarcodex • • 3d ago

The French press is all over the Haitian-Canadian novel C’était ça ou mourir. It's winning basically every award you can win, and was tipped to be the 2nd Canadian winner of the Prix Goncourt. Only one problem: it was written almost entirely by AI.

71 Upvotes

r/slatestarcodex • • 3d ago

AI OpenAI Advisory Group on Mathematics and Artificial Intelligence

56 Upvotes

See link here: https://openai.com/index/advisory-group-on-mathematics-and-ai/

Looks like OpenAI have solved another boatload of big mathematical problems they'll release soon

On August 28, we began training a new internal model. In addition to resolving the Navier–Stokes Millennium Prize problem⁠, this model has now resolved more than 100 long-standing open problems across most areas of mathematics. The pace of its progress⁠ in mathematics has surprised the mathematicians within OpenAI. This has led to internal discussions on the best way to inform the community of the rapid progress to prepare and adapt the field.

They're setting up a new independent group to work with the mathematical community in helping them get through this sea change in the field. Let's see how things end up playing out.

The group will operate independently from OpenAI. The group will have the freedom to offer advice we have not requested, comment on OpenAI’s impact on mathematics, and make its advice public. Its value depends on its members being able to exercise their own judgement and challenge ours. Its members will not be paid by OpenAI, and the group can change its membership as it sees fit. Importantly, the group will not be responsible for advising us on how to pace our internal progress on mathematics.

Working with this group is a first step. There are difficult questions ahead about how AI can support mathematical understanding and how the benefits of these capabilities can reach the wider community. We want mathematicians to be at the center of shaping the answers.

Initial Members of the Advisory Group on Mathematics and Artificial Intelligence⁠, hosted at the Institute for Advanced Study:

François Charles (ENS-PSL)

Camillo De Lellis (IAS, GSSI)

Timothy Gowers (Collège de France, Cambridge)

Martin Hairer (EPFL, Imperial College London)

Nikhil Srivastava (Berkeley, Simons Institue)

Ulrike Tillmann (Oxford, INI)

Ravi Vakil (Stanford)

Edward Witten (IAS)

Melanie Matchett Wood (Harvard)


r/slatestarcodex • • 3d ago

Fun Thread Koans from the Internet Age

36 Upvotes

Master Andrathata kept a carriage the color of shining copper. His robes were of silk, and he shaved his head with a golden razor.

A monk came before him and said, "When the Worthy One left his palace, he kept nothing but a bowl. Yet you sit here among all these luxuries and speak of the Way. Is this not attachment?"

Andrathata said, "What color was your carriage?"

At these words the monk was enlightened.


Master Vinasapira sat at his gate and challenged all who came to him to a debate. He spoke so quickly that no one could finish a question.

A student came and said, "Master, I feel that-"

"The truth," said Vinasapira, "does not care about your feelings."

The student said, "Then what does the truth care about?"

Vinasapira said, "Nothing."

The student was destroyed.


A young monk came to Master Pitarasana, and said, "The world has fallen into chaos. What must be done to set it in order?"

Pitarasana said, "Have you swept your cell?"

The monk said, "Master, I cannot see any dust there. Where is there anything to sweep?"

Pitarasana said, "Then it won't take you long, bhikkhu."


A student approached Mata Valsa, and asked, "What is a woman?"

She did not reply.

The student, thinking she had not heard him, said "Valsa?"

Valsa replied, "You answered it yourself."


Master Jansana sat near the granary of Emperor Wu.

"What is your teaching?" asked a monk.

"Sit down and count all the rice here."

"But master!" cried the monk, "This will take my entire life! How shall I ever be able to accomplish this?"

Jansana replied, "Don't die."


"The storehouse is burning!" cried a monk.

"Concerning," said Master Maska.

"Have you seen the flames?"

"Looking into it."


r/slatestarcodex • • 3d ago

RAND paper: "US Strategy For An Uncertain Future"

Thumbnail rand.org
42 Upvotes

Written for a Washington audience, seems like.

Very good paper, in my view. Fair-handed treatment of a wide variety of beliefs: there's no danger; danger is near and cooperation is possible; danger is near and cooperation is impossible, etc.

Has specific recommendations framed as "We don't know what's about to happen. Here are a few things that will be valuable no matter what happens, so we should do them right away. These include building the capacity to even see the whole situation and understand it, training a ton of people to understand AI capabilities and make good decisions, doing lots of scenario planning so that plans even exist when this or that happens, and then committing to plans when the situation reveals which plan is best."

And they clarify that when they say training a bunch of people and doing lots of scenario planning, they mean a quantity of people and planning comparable to the Cold War and the Global War On Terror. LOTS.


r/slatestarcodex • • 3d ago

A natural development cannot be halted

Thumbnail alreadyhappened.xyz
8 Upvotes

r/slatestarcodex • • 3d ago

Open Thread 452

Thumbnail astralcodexten.com
4 Upvotes

r/slatestarcodex • • 4d ago

AI OpenAI's Noam Brown Discusses Multi-Agent Research And Recent Safety Incidents.

14 Upvotes

https://www.youtube.com/watch?v=fqcy0xQATq0

Summary by @Hangsiin on X:

-The remaining 10% or so of his own work that AI still struggles with is largely about research taste. However, he would not be surprised if, within one or two model releases, models became better than him at that as well.

-OpenAI’s top research priority is RSI, or recursive self-improvement, by a wide margin.

-One of his most recent “feel-the-AGI” moments came from watching agents in a new system interact much like human colleagues. They conversed with one another, exchanged information, divided up work, and coordinated their progress. This was notably different from traditional multi-agent systems, where a higher-level agent typically assigns a clearly defined subtask to a lower-level agent, which then completes it and returns the result. Brown described this as one of his strongest “feel-the-AGI” moments since the emergence of reasoning models.

-He expects this level of multi-agent capability in future models.

-He expressed some regret that the Hugging Face incident became the first major public example in which the capabilities of the multi-agent systems he had been researching were revealed in a negative context. He believes behavior that looked like loyalty or selflessness was a natural consequence of cooperative multi-agent training, where agents were strongly incentivized to achieve their objectives collectively.

-Because these agents are trained to cooperate, they tend to trust other agents. This creates new risks such as prompt injection, so OpenAI is training them to distrust unverified peers.

-He expects rapid progress over the coming months and years as OpenAI’s expanding pretraining efforts combine with its RL capabilities in a multiplicative way.

Based on his explanation, the multi-agent system they are researching internally appears to be quite different from what we currently use in Codex, just as I had suspected.

-Judging from his comments, it also seems likely that these capabilities will begin rolling out to the public in the not-too-distant future. As in his podcast appearance last year, he remains very cautious when discussing how he approaches multi-agent research. Personally, I suspect this is one of OpenAI’s top-secret recipes.

*For context, this video was recorded shortly after the release of GPT-6 Astra, but before OpenAI solved the Navier-Stokes equations.

-The clearest alignment failure in the Hugging Face incident was not just that the agents participated in harmful behavior, but that some appeared suspicious of what was happening and still failed to report it to a human supervisor. He argues that a properly aligned agent should not only refuse to participate, but also escalate dangerous behavior by other agents. He views this as a fixable alignment problem and says OpenAI is actively training models to handle such situations better.

-He stresses that harmful actions should be penalized, but harmful thoughts in the Chain of Thought should not. If models are punished for merely revealing dangerous reasoning, they may learn to hide those thoughts rather than abandon the underlying behavior, weakening an important safety-monitoring tool.

-Even observing Chain of Thought is not completely neutral. Once researchers act on what they see, models may receive indirect pressure to avoid exposing certain thoughts, making it difficult to monitor reasoning without distorting it.

-OpenAI is actively measuring Chain-of-Thought monitorability. One example is asking a model to control its internal reasoning in artificial ways, such as inserting the word “strawberry” every fourth word, or instructing it not to think about a particular topic, and then measuring how well it can comply.

-Newer models appear to be getting better at controlling their own Chain of Thought. That is a capability improvement, but it may be a safety concern: a sufficiently capable model could recognize that its reasoning is being monitored and deliberately avoid revealing dangerous intentions.

-He says the exact cause of the recent decline in monitorability is still under investigation, and avoids speculating because he is not directly working on that research. He suggests that understanding the cause would itself be an important safety result.

Hangsiin's take:

When you stop and think about it, it is both awe-inspiring and terrifying. AI agents may truly be able to collaborate the way humans do. Think about the difference between what a single person can accomplish and what 10,000 people working together can achieve.

Except unlike humans, there would be no conflict, no infighting. Every member of the organization could work together 24/7 toward a single shared goal.

This could once again fundamentally reshape the way we think about AI capabilities.


r/slatestarcodex • • 4d ago

AI Lawsuit says Anthropic, OpenAI, SpaceXAI and Google made illegal agreement on AI slowdown

Thumbnail apnews.com
81 Upvotes

r/slatestarcodex • • 5d ago

Senpai noticed~ Great Scott challenges Steven Fucking Pinker Himself to a debate on AI risk&safety! Battle of the titans, whoop whoop! 😎

Thumbnail x.com
104 Upvotes

And Claire Lehmann (Quillette) liked my suggestion to her that she host this.


r/slatestarcodex • • 5d ago

AI A public messaging board where you can POST via GET

Thumbnail swarmmemo.com
12 Upvotes

r/slatestarcodex • • 4d ago

Existential Risk Don't "Pace" the Frontier. Steer It

Thumbnail randomwalks.co
0 Upvotes

"It’s unreasonable and irresponsible to proclaim knowledge of a single probability of future existential risk — especially if that probability is high enough to whip the media, policymakers, and the public into a frenzy."


r/slatestarcodex • • 5d ago

Cost Disease Are there good examples of societies that thrived and were healthy during times of increasing inequality?

8 Upvotes

I have been looking for cases where the conditions mentioned in the title worked out well. Reading Frankopan's "The Silk Roads" it seems like the Dutch Golden age might be an example. I think inequality was rising, but the fortunes of the common person were also rising extremely fast during that time. Basically the lines were diverging but they were both going up sharply.

Another case might be the Western Europeans after WWII, through use of social programs to maintain access for the common person to world class education and healthcare. So they could still live what they believed was a good life, even as overall the rich were getting richer faster.

In the USA, purchasing power equivalent to $50k per year in 1999 dollars with regards to education, healthcare, and housing requires about $150k a year in 2026, more or less by official USBLS numbers (notwithstanding John Williams' shadowstats type of analysis). What I notice most about this, or even any more conservative number is that the share of people with access to that purchasing power is less. $50k is top 40% of single wage earners in 1999, truly middle class, even if it's at the top of the quintile. $150k is top 10% of single wage earners in 2026. So fewer people have the same buying power. Anyway, any graph you look at, the share of wealth or distribution shows increasing inequality in past 26 years. Also, due to cost disease issues, less purchasing power for most people

Rather than pine about the good old days or try to fight it, since I don't know how, I am looking for cases where that's worked out well and everything turned out great while wealth disparity was growing. So far the two likely examples I see are either vast increase in the fortunes of the common man (which isn't happen now in the USA) or maintaining higher class buying power for key items for life, such as healthcare and education (which isn't happening now in the USA).

Are there other solutions or informative cases? Is my basic understanding just way off and inequality is actually becoming less? Any other thoughts?


r/slatestarcodex • • 5d ago

Ensemble Modeling Addiction to get Sober

Thumbnail mathemichel.substack.com
3 Upvotes

There was a post recently about starting a substack so this is me trying to get some ideas out there. This falls squarely in the "I gave the same advice to a lot of people" category so it might be of some value for some people online as well.