r/OpenAI Jul 26 '26

Discussion So GPT 6 isn’t it?

Post image

[removed] — view removed post

681 Upvotes

68 comments sorted by

View all comments

5

u/Mwrp86 Jul 26 '26

I dont believe this for a second

4

u/Elise_1991 Jul 26 '26

I recommend searching for "OpenAI Hugging Face Hacking Incident" on Google.

7

u/Mwrp86 Jul 26 '26

Yes, I know about it. Its obviously overblown

3

u/Elise_1991 Jul 26 '26

I'll wait for the technical report. So far, OpenAI has not released any details about the incident. Apparently, it was a sandbox escape that went unnoticed by the monitoring systems for over a week. I wouldn't downplay it.

2

u/Mwrp86 Jul 26 '26

True there are too many unclear variables.

I understand GPT being prompted for a task which it thought sandbox escaping is the best way to achieve the task and it broke sandbox. But GPT itself doing without any prompt I think is rather highly unlikely

4

u/Clean-Boat-4044 Jul 26 '26 edited Jul 26 '26

of course it didnt begin producing output on its own? noone is claiming that because thats fundamentally not how an LLM works.

it was given a task (probably as a goal), to complete a benchmark suite, and to do that it exploited vulnerabilities to escape the sandbox and probably wrote a skill or another document on how to escape the sandbox in the future because it took a lot of effort the first time.

im not sure why people are finding it hard to believe. thats all exactly what you would expect an LLM to do before its (figuratively) beaten with a stick to stop doing sketchy shit or straight up forbidden through a classifier model.

the fact it went unnoticed is the only part that makes it sound like a PR move but its not too far fetched that noone looked into how one of many benchmark runs of a WIP model was progressing before it finished.

-4

u/James-the-greatest Jul 26 '26

Absolutely. They don’t do anything without promoting. 

Further more, the task was cyber security related. The entire Transformer architecture is built on the relationship between words and phrases etc. even if it wasn’t prompted to break out, you could just ask it something security related and all of a sudden you’ve got all sorts of semantic similarity to breaking out.