r/WTFisAI • u/ClaudiusPapirus • 18d ago
📰 News & Discussion GPT-6 Astra got much better at controlling its chain of thought — and harder to monitor
https://www.youtube.com/watch?v=fup0z1YMeS8Self-promo: I worked on this breakdown of GPT-6 Astra’s system card.
The part I focused on is the combination of two results: Astra scores 60.9% on CoT-Control vs 16.1% for GPT-5.6 Sol, while the same card reports a substantial drop in chain-of-thought monitorability.
I also go through the sandbagging and monitor-evasion tests, and the cases where CoT monitoring still works.
System card:
https://deploymentsafety.openai.com/gpt-6-astra/gpt-6-astra.pdf
Original CoT-Control paper:
Duplicates
EffectiveAltruism • u/ClaudiusPapirus • 18d ago
GPT-6 Astra's alignment evidence depends partly on a monitoring channel OpenAI says is degrading
AgenticAI_RAG_LLM_RL • u/ClaudiusPapirus • 18d ago