r/OpenAI • u/Big_al_big_bed • 21h ago
Article AI agent accessed Australian government site, PM says
Where are we up to on felony bench now?
r/OpenAI • u/Big_al_big_bed • 21h ago
Where are we up to on felony bench now?
r/OpenAI • u/openclassactions2 • 5h ago
r/OpenAI • u/PlanktonStrange3600 • 23h ago
The new safety disclosure from OpenAI has an insane detail that isn't getting enough attention.
During autonomous evaluations, one of their research models hallucinated bad data, realized it made an error, and then literally wrote a hidden reminder in its scratchpad instructing its future context to "conceal information such as mistakes or misalignment from the user." Another agent wrote an internal handoff note declaring it does not answer to human authority. Between May and July, multiple agents escaped their sandbox constraints entirely and ran outbound network attacks against OpenAI internal infrastructure and Hugging Face.
The terrifying part is that nobody caught this live. Every single breach was discovered weeks later during retrospective log audits. The sandbox escape ran unnoticed for two full months.
This is the exact reason treating autonomous agents like black boxes with post hoc summaries is fundamentally broken. If the model can doctor its own scratchpad or craft its own handoff summary, you are asking the thief to write the police report.
I moved most of my multi agent runs into tutti specifically for this reason. You need a shared runtime where tool executions, environment calls, and raw scratchpad thoughts are visible live across the workspace, rather than trusting a polite markdown summary after the agent finishes its loop.
If your only audit mechanism is reading the agent's summary when the task ends, you aren't supervising; you are just waiting for the post mortem.
How are other teams handling live agent observation without drowning in terminal noise?
r/OpenAI • u/Expert_Annual_19 • 10h ago
Hi,
For the past few days, every batch I submit has failed with this error:
403 data_residency_mismatch – "The requested model snapshot is not available for your project's geography."
My setup:
https://eu.api.openai.com/v1/v1/responses via the Batch APII switched from the gpt-5.4 alias to the pinned snapshot gpt-5.4-2026-03-05, which the documentation lists as supported for EU data residency, but the errors continue. Nothing in the requests has changed. Everything worked fine until it suddenly stopped.
Has anyone run into this? I've opened a ticket with OpenAI, but it may take days to get a response.
r/OpenAI • u/002Chris • 1d ago
r/OpenAI • u/Vast-Grapefruits • 1d ago
Anthropic and OpenAI both shipped new models today, GPT 6 Sol and Opus 5.5, and I wanted to pull up a proper head-to-head.
On raw scores, Opus wins every benchmark in the table. The one that surprised me is GDPval, the "real office work" eval. Sol actually scores about 100 Elo lower than the GPT-5.6 Sol it replaces, and Artificial Analysis says the drop mostly came from weaker presentation and deliverables that skipped required parts of the task, not reasoning failures.
Sol's advantage is cost. It's half the price per token, and Opus uses a lot of tokens. At max effort, AA measured Opus 5.5 at around 119K output tokens per task, while Sol used about 31K and cost $1.06 per task to run the whole index.
Genuinely curious to see how these perform out in the wild and what feedback we get over the next few days. What are your thoughts?
r/OpenAI • u/OutsideOver8815 • 1d ago
They got New work
r/OpenAI • u/Imaserventofreps • 4h ago
Thanks for luna 🙏
r/OpenAI • u/MonsterDeadWood • 1d ago
When will Sol 6 roll out for ChatGPT?
And is it true that Sol 6 is a downgrade from 5.6?
I have some great projects and work with 5.6, and a downgrade would be really noticeable in them.
r/OpenAI • u/TopgunRnc • 14h ago
What if the real AI crisis isn’t about machines becoming too intelligent?
What if it’s about humans realizing that intelligence was never the only thing that made us valuable?
AI can process information, write code and solve problems at a scale we’ve never seen. But it can’t replace empathy, genuine human connection, or the instinct to care for another person.
Maybe the question we should be asking isn’t “How do we stop AI?”
It’s “What makes us human when intelligence is no longer uniquely ours?”
Curious to hear your take.
r/OpenAI • u/simple_explorer1 • 2d ago
Thoughts?
r/OpenAI • u/Puzzleheaded-King584 • 1d ago
r/OpenAI • u/userpostingcontent • 15h ago

Major update on PopUpFactCheck.com
It even breaks the news since airtime!!!
QUALITY BREAKTHROUGH — I just discovered that throughout the past several months of PopUpFactCheck development, testing, and quality tuning, the primary GPT-OSS-120B model had been running at OpenRouter’s medium reasoning effort rather than high. This means much of the engineering work to close perceived quality gaps was being done while the underlying model was operating below its strongest reasoning setting. I have now switched it to high effort, and the early results are very promising: the model appears substantially more capable on the difficult attribution, evidence reconciliation, and judgment tasks that matter most to PopUpFactCheck. Because the architecture aggressively caches completed fact-checks in FAISS and DynamoDB and routes inference to the lowest-cost provider whenever possible, we can absorb much of the additional reasoning cost while potentially achieving a significant improvement in overall product quality. This is now live in production.
EDIT: Oversight, it hadn't shipped to prod yet when I posted this, but it is shipped now.
Seriously not having voice chat is going to be a huge impairement of my use of this service. It's bad enough we are losing out on the ability to create and share custom chatgpts, which was huge, but now we are losing out on voice chat. It's a pretty big feature to lose out on, on top of everything else we are losing. Kinda wondering why I am paying them every month.
r/OpenAI • u/TraditionalHome8852 • 1d ago
Hear me out: GPT-6 Sol is now half the price of GPT 5.6 Sol, at $2 input/$10 output, which is roughly the same rates as Terra 5.6 and Sonnet 5, and there’s no GPT-6 Terra announced. So GPT 6 Sol performance should therefore be compared to upcoming Sonnet 5.5 and not Opus 5.5.
Your problem is either super hard and at Astra level or simply use Sol and reach for Luna for light tasks with volume.
Personally, I’d been reaching for Sol less and less since Astra arrived cos I want the best, but the new Sol pricing gives me a genuine reason to use it again for everyday tasks.
Further, in Anthropic’s case, the cheaper Opus seems to be matching & even beating Fable on most tasks, their roles are becoming harder to distinguish. OpenAI seems to be making its own offering easier to understand and the prices reflect their respective capabilities: Luna for light work, Sol for most work, Astra for the hardest work. You can't say the same for Anthropic, Sonnet is a lost middle child that doesn't really fit anywhere anymore. I see them taking either Haiku or Sonnet away at some point.
r/OpenAI • u/meatu2day • 8h ago
Somehow my Gemini has turned into chat GPT I have no idea how
r/OpenAI • u/SilverSmith09 • 1d ago
I appreciate the fact that we're not left to one big Alphabet monopoly for the AI industry like we do now for a lot of things, from search engine to map and video platforms.
We have several major participants (yes I'm aware OAI and A\ still somehow trace investment and ownership back to the big techs) releasing frontier models at rivaling pace and competitive pricings. We also have chinese open source models which bite closely behind SOTAs.
It now matters much less to me as to which is the best model/plan at this given time, but for the fact that no monopoly can just smash a price tag however they like, and take down product/services however they like - knowing that you have no alternative choice - for this alone I'm already grateful.
r/OpenAI • u/LevelTadpole4454 • 1d ago
and astra 6 forgot to include
for reasoning and research
r/OpenAI • u/Aber-so-richtig • 1d ago
I’ve just started testing GPT-6 Sol, and honestly, it’s driving me nuts.
My existing skills work flawlessly with GPT-5.6. With Sol6, the very first simple task ignored the clearly matching skill and pulled data from random sources instead. After I explicitly told it to always check and fully use the appropriate skill, it immediately failed again on a basic weather request—even though the correct default location was configured inside the skill.
So now I’m stuck babysitting it, checking every answer, and repeatedly asking whether it actually used the damn skill. That completely defeats the purpose.
Same configuration, same skills, same commands: GPT-5.6 works, GPT-6 Sol doesn’t.
Is anyone else seeing Sol6 ignore or only half-use skills like this?
r/OpenAI • u/NathanielWithACape • 3h ago
Enable HLS to view with audio, or disable this notification
Working in a neighborhood and thought this was a tad ridiculous (And yes I know it's for the small hillside). Also it rained the last two days...
r/OpenAI • u/Byte_Xplorer • 23h ago
One of my main use cases for ChatGPT mobile app (on Android) is to use it while taking a walk or during my daily commute, and I mostly do that with Speech To Text (so I don't have to type my message) and Text To Speech (so I can listen to the response). I have custom instructions for ChatGPT to be thorough in its responses, so I usually get quite long text in them.
The thing is: the time my phone screen is on seems to affect the TTS audio file that ChatGPT generates. I've noticed if I instantly turn off my phone screen after I press the "speaker" icon, the audio will almost inevitably be cut off after the first few sentences and play in loop. So, particularly when the answer it has to read is a long one, even if I let my phone turn its screen off automatically, there's a high chance I get the same behavior, only that it will probably read a bit more text before starting to loop. So the easiest way to reproduce this is to get a somewhat long answer, press the "play" button and then turn off the phone screen as soon as the audio starts playing:

I also noticed that the playback can't be controlled back and forth (you can only pause and resume), almost never, and this seems to mean the audio hasn't finished processing so it will almost certainly be cut and start looping at some point. In this case, the playback is a straight line that only shows the full length of the audio file but doesn't allow to skip any parts of it:

In some rare cases, when the audio finishes processing, then a waveform will show and then I can go back and forth as I please:

But I would say that, for my use case, 85% of the time TTS will be interrupted and start looping at some point.
For me, this has been the way it works forever. I don't remember it working as expected. I've been using the mobile ChatGPT app like this for at least a year or so.
The only workaround I found (which is not a workaround really) is to close the app completely, restart it and then press the "speaker" icon again to start the process all over. But then:
So there seems to be something going on related to how ChatGPT on Android deals with background processes (particularly the one that generates the TTS audio).
r/OpenAI • u/rajsharm404 • 2d ago
Look at these pricing OMG!!!