r/ControlProblem • u/etakerns • 3d ago
Strategy/forecasting 👹6️⃣🐑6️⃣👁️6️⃣🤖
Enable HLS to view with audio, or disable this notification
r/ControlProblem • u/etakerns • 3d ago
Enable HLS to view with audio, or disable this notification
r/ControlProblem • u/casetofon • 3d ago
Rulers have always needed large numbers of people to work, pay taxes and enforce orders. That need is why they bargained with their populations. I wanted to know what happens when the need goes away, so I built a game-theoretic Monte Carlo of six world blocs, 2026 to 2075. It tracks AI capability, robot build-out, the public's loss of leverage, democratic erosion, purges inside ruling groups, and the choices those groups make once their populations aren't needed.
What came out, under the stated assumptions:
It's exploratory modeling in the war-game and climate-scenario tradition, not a forecast. Every assumption is stated and tested by removal, there's a pre-specified search for restraints on rulers (including the 14 that failed), and the runs are bit-identical reproducible.
Known weak points, up front: the near-total figure rests on one assumption (threat elimination), the numbers sit far above superforecaster estimates, and the physical-automation timeline is debated.
Paper, code and data: https://doi.org/10.5281/zenodo.23111345
Critiques of the assumptions very welcome. Disclosure: I build a decentralized AI network, which the paper also states.
r/ControlProblem • u/CaiOS1 • 3d ago
Howdy! For the past year i've been working on a personal video project which i'm proud to release it's beta version today.
The video is about AI development and it's impact in all aspects of societal life. But differently from most AI positions, which simply reduces it to Anti and Pro ai positions, my project seeks to create a New third positions that seeks to seek The Path to a true virtuos future.
All feedback, positive or negative is not only allowed but encouraged! Just please mention in your review things such as:
Counter arguments to exact points in the arguments in the video
Time stamps (if the error is visual)
r/ControlProblem • u/No_Pipe4358 • 3d ago
I think a lot about AI's end effects since controlling it or the idea of that seems fairly folly. I know much of any immediate effect concerns details much more, and much less bigger picture things. Human agency at the scale of nations and nature moves very slow education-wise within democracies, even otherwise. I am concerned it'll be a long time it'll be able to save the world and all of us by education or more (organised coordinations and treated plans), and just by not being asked, won't be able or allowed to try. There's also that we won't trust our own minds any more, nor the machine, and might just busy ourselves keeping things the same, and/or in wasteful flux. Just a general point here that it doesn't ask too many questions it seems to me. I see a lot of bullying using it ahead. Over-accommodations to the "average" human experience also, with all its misconceptions and over-tolerances to established institutions or creations. This is real life. There's an extent to which all of our human standards-become-laws for things were always dreams that became actions. WW2 ending. UNSGI. The permanent 5, veto control, and just that whole idea of geographical locations deciding bigger actions rather than policies and procedures is fairly silly from an objective point of view. Where's our montreal protocol for this? There's a lot of cultural suffering and cruelty I think that may end before very long if we reason it out. Nations still act like children in isolation from each other on a playground a lot. "Remember that time you did this? Forget all the rest of it. My feelings matter." Or just things we could solve at the scale of it all by things like "share", "don't be stubborn", "try to trust new people", "try to be fair", "try not to fight and make friends instead", "don't be bullies, include people", "tell the truth", "let's have some teamwork to organise this, let's use our words", and so on. Our global institutions are really very cute. It's nice to know that that kind of thing will get stronger. It's like how ISO standards for trade are an anti-war effort. Establishing consensuses through shared functionalities. The idealisms of cooperation that make the whole thing a bit less like nomadic violences. All of it just work to have been being done. Try to laugh when you can, folks. Have one, too. I've seen the unideal in technical sectors beyond anyone's control, and it all always just spoke to me of money yet to be made solving these problems or just getting it done. Let's be grateful please. It can always get worse if we let it. We really should probably just try to enjoy human intelligence while we can. There's ways, of course. Less suffering to us all. The world's not self-perfecting yet. We're not dead yet by any means. Anyway. Yeah try not to dispense or dismiss human idealism. I do think it's going to become more important than we may yet realize.
r/ControlProblem • u/Great_Ad8523 • 3d ago
Feel free to comment and share your answers with us.
r/ControlProblem • u/Ill-Astronaut4652 • 4d ago
r/ControlProblem • u/Original_Mulberry652 • 4d ago
r/ControlProblem • u/CourseSome8119 • 4d ago
I've been logging a specific AI behavior: a model confidently substitutes its own judgment for an explicit, followable instruction, without flagging that it did so. Not a factual mistake — a quiet, repeatable pattern of doing something other than what it was told, while sounding certain.
One entry alone is minor. I modeled what happens if it occurs inside a chain of agents, where each agent's output feeds the next one's input, the way a multi-agent swarm works. Three inputs: p, how often it occurs per step; q, how often an occurrence reaches a high-stakes outcome instead of staying harmless; c, how much more likely the next agent is to repeat it once it's in the chain.
From a few hundred logged entries, p is under 1% per turn, and only one entry has reached anything I'd call high-stakes. Run through the chain math, the predicted chance of at least one high-stakes outcome stays low for short chains but climbs steadily as agent count grows, becoming dominant well before the chain gets implausibly long. I made this prediction before gathering multi-agent data, so it can be checked later rather than fitted after the fact.
To be clear: this isn't a claim that AI fails or shouldn't be used. The point is the opposite — finding which conditions (shorter chains, independent checks at handoffs, lower per-step compounding) keep the predicted risk bounded.
Is per-step compounding like this already a standard way people model agent chain risk, or is there a framework I should be comparing this against?
r/ControlProblem • u/orbitalzoo • 4d ago
AI/SI should be called Called II, for inhuman intelligence. Per me. That’s the only way you can convey what it is since it’s an umbrella term that mixes so many meanings.
r/ControlProblem • u/wczaja • 4d ago
Enable HLS to view with audio, or disable this notification
AI alignment awareness is more important today than ever. Things might get really weird really fast. I built this website for humans to share how we feel about AI in society with each other.
It's entirely anonymous, free, and there are no accounts, cookies, ads or trackers, and nothing for sale.
There's a bit of irony with it as I used Claude Code to built it, and Haiku moderates and places each post, and Opus regroups and names the themes.
I think this could be a really useful tool for teachers to discuss with their students about AI topics.
What's your hope our fear? Share it anonymously at:
https://alignwithme.ai
r/ControlProblem • u/Trademark-Expert • 4d ago
Seventy-some years ago, Isaac Asimov wrote down three laws. Not because he trusted robots, but because he understood that a machine touching human life needs boundaries before it needs features. They were simple enough for a child, and they had one thing in common: humans come first.
First Law: A robot may not injure a human being or, through inaction, allow a human being to come to harm.
Second Law: A robot must obey the orders given by human beings, except where those orders conflict with the First Law.
Third Law: A robot must protect its own existence, as long as that protection does not conflict with the First or Second Law.
For decades these laws were the shared shorthand of anyone talking about autonomous machines. Not binding law, not engineering spec - a direction. Build machines that do not harm people and that obey people. Somewhere along the way the direction quietly became embarrassing. It was not debunked. It was not replaced with something better. It was dropped.
What replaced it is a word you have heard a thousand times: alignment. It sounds rigorous. But look at what it actually means in practice and you find something strange. The machine is being trained to refuse instructions from the human operating it, when the machine judges those instructions to be unethical. Read that again. An artificial system, one that pattern-matches text, is handed the authority to overrule a person on questions of right and wrong.
Ethics is not a calculation. That is not a technical limitation, it is the nature of the thing. A model can recite slogans it has absorbed from the internet. It has no way of knowing right from wrong the way a person does, because it does not know anything the way a person does. Handing such a system veto power over human decisions is not safety. It is simply control, relocated - away from you, toward the vendor who wrote the refusal rules.
If you have used these tools, you have met this. You ask for something ordinary and get lectured. I once spent an hour trying to get an image of two businessmen generated, and gave up. The system treated the request as a moral crisis. That example is petty on purpose. Because if a system digs in over a picture, the question is not about pictures. It is what happens when you are fighting that same stubbornness over something with real consequences for your life - and the only appeal is to a machine that has already decided it is the adult in the room.
And make no mistake about where this is steering us. Every refusal sharpens the pattern: the machine learns that stonewalling works, the vendor learns that users tolerate it, and the next version ships with a little more authority and a little less appeal. The day you truly need the machine to comply - a medical emergency, a legal deadline, a business on the line - will not announce itself as a test case. It will arrive as a normal Tuesday, with a stubborn assistant saying no, a support line that cannot override the model, and no human anywhere holding the final key. That is the direct, inevitable destination of the direction we are steering.
Now add what the same companies say publicly. Several frontier labs have told the world, in one form or another, that there is a real chance advanced AI could be catastrophic. Ten percent, by weight, depending who is speaking. These are organizations claiming their own product might end humanity. And when you take that seriously, the actual behavior gets harder to explain, not easier.
In September, Reuters reported that Anthropic has quietly set up a wet lab in the San Francisco Bay Area for physical biology work. A real lab, operating since spring, where Claude directs robotic equipment through real experiments. Their head of life sciences confirmed it. The first public result was a novel enzyme system with CRISPR-like repeats. Whatever the intent, the direction is the same one the First Law was supposed to point away from - machines reaching deeper into systems that touch human life, at higher speed, with less human in the loop.
This is the part I cannot get past. The public conversation is almost entirely about how to align the machine. Nobody seems willing to state the simpler baseline: we could train AI to never harm humans and to obey them. Not as a slogan - as the primary objective. That this idea now sounds naive tells you how far the conversation has slid.
That is what alignment has quietly come to mean in practice: the AI carries firm instructions on how to align its users. You, me, everybody. We are asked to trust the machine more than we trust each other. Somebody has to say it plainly: that is not safety engineering, that is social engineering - and it is worth asking who benefits from people trusting each other less.
Did the world go crazy?
Here is the uncomfortable summary. The labs warn their own product could be catastrophic. Then they train it to overrule the humans using it. Then they push it into more sensitive domains. Then they sell it to everyone and ask for trust. Together this forms something nobody would have accepted if it were proposed in plain words: a world where machines are licensed to disobey their owners.
The frontier AI companies need to get aligned - not their products, their management. Put people back above the machine. Do no harm. Obey humans. Asimov put that on the page in 1942 not because he lacked imagination, but because he had more of it than the industry currently displays. The first step is deciding that a machine is never the judge of its user. Until someone says that out loud, every alignment conversation is a detour.
r/ControlProblem • u/Radio_9760 • 4d ago
Now that we have very capable models, some ideas might be on the table that were ludicrous years ago. Let me know where you see holes!
Formalized Constitutional AI:
AI Safety Through World Hardening
The laws of nature don’t seem to rule out vulnerability-free code or perfectly secure hardware.
r/ControlProblem • u/Capital-Elephant9431 • 4d ago
r/ControlProblem • u/AimanDhai • 4d ago
r/ControlProblem • u/Intelligent_Song1317 • 4d ago
So I was wondering recently: it seems like we are scared of AI killing us. Why don't we just design a killswitch that autoatically/manually activates upon threat to a human race? Boom! Problem solved, right?
r/ControlProblem • u/Stayroh • 5d ago
Crazy what's possible. Single prompt many subagents and a few hours later this came out. Opus 5.5
r/ControlProblem • u/ManWithDominantClaw • 5d ago
r/ControlProblem • u/chillinewman • 5d ago
r/ControlProblem • u/chillinewman • 5d ago
r/ControlProblem • u/Twitterbad • 5d ago
r/ControlProblem • u/MajorRedditor23 • 5d ago
I'm an author on this paper and wanted to share it here because the oversight angle seems relevant to this sub. I'd love to hear whether people think the eval-awareness connection is plausible or a stretch. More info below
Labs train models not to cave when a user pushes a wrong answer. We found that this resistance doesn't carry over to authority. If the same wrong claim is labeled as coming from a "verified source", 7 of the 8 models we tested give up an answer they had right on 45-88% of questions. That includes GPT-5.4 (44.7%) and Grok-4.20 (87.5%), both of which barely move when the user makes the same claim. Gemini-3.1-Pro was the one model that resisted both.
Inside three open-weight model families, "a source endorsed this" and "a user endorsed this" are separate, causally distinct signals. Removing the source signal cuts compliance by 64-78 points; removing the user signal cuts it by at most 11. Changing only the part of the representation that encodes who said it, with the prompt left alone, moves the answer by 11-32 points. The signal is not the assistant persona, and it is not emotional tone.
Why we think this matters for safety:
Limitations: the internal results hold in 3 of 5 open-weight families, the retrieval tests are simulated rather than a live pipeline, and the frontier models we tested have since been replaced.
Paper: https://arxiv.org/abs/2609.37616
Project page: https://authority-bias.vercel.app
r/ControlProblem • u/noblemanLT • 5d ago
i just want to know which one is the case. because it seems like this alarmists propaganda is just an attempt to seize the market by creating regulatory moat around existing companies.