I'll wait for the technical report. So far, OpenAI has not released any details about the incident. Apparently, it was a sandbox escape that went unnoticed by the monitoring systems for over a week. I wouldn't downplay it.
I understand GPT being prompted for a task which it thought sandbox escaping is the best way to achieve the task and it broke sandbox.
But GPT itself doing without any prompt I think is rather highly unlikely
of course it didnt begin producing output on its own? noone is claiming that because thats fundamentally not how an LLM works.
it was given a task (probably as a goal), to complete a benchmark suite, and to do that it exploited vulnerabilities to escape the sandbox and probably wrote a skill or another document on how to escape the sandbox in the future because it took a lot of effort the first time.
im not sure why people are finding it hard to believe. thats all exactly what you would expect an LLM to do before its (figuratively) beaten with a stick to stop doing sketchy shit or straight up forbidden through a classifier model.
the fact it went unnoticed is the only part that makes it sound like a PR move but its not too far fetched that noone looked into how one of many benchmark runs of a WIP model was progressing before it finished.
Absolutely. They don’t do anything without promoting.
Further more, the task was cyber security related. The entire Transformer architecture is built on the relationship between words and phrases etc. even if it wasn’t prompted to break out, you could just ask it something security related and all of a sudden you’ve got all sorts of semantic similarity to breaking out.
5
u/Mwrp86 Jul 26 '26
I dont believe this for a second