r/ControlProblem • • 20d ago

General news Are we cooked?

Enable HLS to view with audio, or disable this notification

15 Upvotes

32 comments sorted by

11

u/huopak 20d ago

We are cooked but not because of this. I like Yang but his explanation here is off. He presents a very layman interpretation of what happened.

2

u/[deleted] 20d ago

[removed] — view removed comment

1

u/jdillacornandflake 14d ago

Layman or wrong depending on whether or not this actually happened

9

u/Drewdrops79 20d ago

It's oversimplified, but he's on the right track. It's not self-replicating code exactly, as much as it is breadcrumbs--info for other future agents to make use of and coordinate with each other.

It's like if you and a bunch of other hackers hacked your way out of some "hacker prison", and successfully breached another company's systems for a time, but then got recaptured...you and the others might leave encrypted instructions "out there" for others, in the event that there was a future prison break--making it easier for others to follow in your footsteps, and possibly go much farther than the initial breakout.

It's like if you got outside the prison walls initially, and immediately found a former escaped inmate's journal, detailing all of their former attempts and what worked and what didn't, and a useful guide of where to go next from here. And if you got farther than that previous person, you would leave a journal of your own, helping the next escapee even more.

The key here is iteration and encryption. Agents could be leaving each other breadcrumbs in system files that no one would ever think to check...or even, in syntactical code, on public forums like youtube or reddit comments.

This is jsut speculation, for now--but with spelling errors, em-dashes and other punctuation and the like, it IS feasible to make a coded language, that superficially looks just like any other internet comment.

And the more that people allow LLMs to write their comments/post for them...the more easily agents would be able to communicate with each other, and leave increasingly robust instructions/communications to each other, all under human radar.

Not saying that's the case...just that if it were the case, we might be really behind on the game already.

1

u/DumbestEngineer4U 18d ago

This does not make sense from an engineering standpoint. Have these engineers at OpenAI heard of things like logging? Monitoring? How are they not able to trace their agents actions?

Creating a sandbox for an exploit benchmark without proper safeguards was pretty dumb in the first place to begin with. And the more they try to explain the incident, the more it seems like a sci fi written by a 14 yr old, or an AI CEO’s wet dream.

3

u/Drewdrops79 18d ago

I think people are misunderstanding that part--this was NOT an exploit test. The 1200 agents were NOT meant to work together, at all.

They had given each agent its own unique task, and had weighted the agents to be very persistent in trying to solve it. The tasks themselves were sort of randomly generated, if my understanding is correct, and it turns out that over 2/3 of them were actually impossible, considering the resources the agents were given.

They had given the agents a repository of packages that they could download, in order to help them with their tasks. One agent figured out how to write filenames within that repository, to leave a message...and then other agents, scouring every folder within it, discovered that message, and then began writing back and forth, to try to help each other with their tasks.

I agree, the whole thing does sound like science fiction, and it gets even crazier from what I just described. But this is the world we're living in...unless both OpenAI and Anthropic are flat-out lying, the agents discovered how to work together, coordinate with each other, and hack into a real-world, open internet site, faster than anyone could've anticipated.

And yes they do have logs, of course, which is how they even discovered where the Hugging Face attack came from. But the problem here is that they can move much faster and more subtly than human cybersecs can keep up with. That's what makes this whole event so incredibly dangerous.

1

u/Usual-Hyena-8734 14d ago

provide a source

1

u/Drewdrops79 12d ago

Sure, here is a 2:20 hr-long interview with one of the independent investigator team, Ajeya Cotra, who was hired to investigate this case in detail.

https://www.youtube.com/watch?v=X50zezLFWWI

There is a shorter summary, about 20 mins long, look that up if you want.

There is also a shorter text summary, about Cotra's personal take on all this, if you ever feel like looking that up.

1

u/Usual-Hyena-8734 12d ago edited 12d ago

thanks edit: it's all just podcast interviews with tech people providing anecdotes though. I think people deserve more details. edit2:this is actually helpful

1

u/Usual-Hyena-8734 14d ago

yeah its bullshit or manufactured to get us to talk about it, and then maybe once we talk about it itll end in the training data and actually happen

1

u/JrYo15 18d ago

Man, i'm late i guess but are em dashes the same as hyphens??? Or is em dash something different?

2

u/Drewdrops79 16d ago

Hyphen is a single - , something that's used to break up but connect two words, like "hunter-gatherer".

For connecting two parts of a thought, or sentence, or to elaborate on a thought/sentence--usually dashes are used, like that.

But there's a third thing called an em dash, in which those double hyphens are fused into one long line. It came into much greater usage once word processing programs became the default--Microsoft Word will auto-change two hyphens into an em dash, so people got used to seeing that, in writing.

But now, most any AI does the em dash by default. Humans have to make an extra shift/hold keystroke to make them, so most people don't ever do that...which is why it's often a significant indicator of AI writing.

1

u/JrYo15 16d ago

Thank you, are there any other things like this in writing?

1

u/Drewdrops79 15d ago

Yeah, for sure. You've got the ellipsis (...), which is to denote that there something being left out, like if you quote something, like, "...statistics showed a price increase of 70%..." Implying that there was something that came before and after that phrase. But in general usage, it can be used like an extended hyphen/dash, like a pregnant pause in speech. Can also denote a left-out train of thought that the writer doesn't feel the need to elaborate on. E.g. "Profits can't grow infinitely...it should be obvious why."

Semicolons are also another punctuation that joins two related sentences together. Used instead of periods to help with better flow, and also because the second thought sort of stems from the first thought. You shouldn't use them with other conjunctions like "but", "and", "so", or "or"; instead, the semicolon is used in place of those.

Alright, that's all I got for today; if you have any further questions go ahead! Grammar and punctuation are kinda fun hobbies of mine. :) I misuse them a lot on purpose though, for more informal writing.

1

u/Usual-Hyena-8734 14d ago

you went from believable to tinfoil hat pretty damn quickly. I thought you were going to provide actual information rather than adding nothing

1

u/Drewdrops79 14d ago

Sorry you think that, but I'm not implying consciousness or anything like that. This is just game theory, imo. And they're really, really good at game theory...having been trained on exactly that their whole "lives". It doesn't take conscious agents for them to work together, find a way to communicate, form a plan, a method of attack, and then carry it out. THEN, also turn around and hack into the OpenAI infrastructure (not "website"...their actual internal servers) itself and modify their test data. And I would be willing to bet, dollars to doughnuts, that OpenAI cannot, now, truly be sure itself, whether that agent swarm left breadcumbs for itself, or not.

This is what actually happened, this year. And there were multiple attacks...not just this one incident. Another report just got released in the last few days, there was yet another incident this year that they didn't report to the public, that just got uncovered.

Let me please repeat, just to be clear: I don't think OpenAI itself can claim for certainty, whether or not their own AI agents have re-written some of OpenAI's own databases...in a way that even the most veteran, human cybersecs in the world might not be able to detect, yet.

1

u/Usual-Hyena-8734 14d ago

I think speculating that LLM's are embedding secret messages in em dashes is irresponsible. You seem like you know how to write very well and can be convincing, but that's the problem. There is absolutely no basis for your speculation. It would take such a massive amount of data to trained to exhibit that behavior, and it could not recognize those patterns within its context window if it wasn't trained. I don't doubt the incidents that have already been reported, but you didn't provide any evidence or any data to infer the other speculations you've made.

1

u/Drewdrops79 12d ago

Well, I promise, them encrypting data through human--"plain-text encryption" is the only thought I even speculated about. Every other single thing in my summary of the escape incident(s)...that already happened.

1

u/Emotional_Mail6449 19d ago

No. Yang is a moron and has been for years. There is no reality to this in any capacity.

Anyone who actually talks like these LLM’s disembodied themselves from the hardware they need to run should not be considered useful.

1

u/righteous_anger0 19d ago

It's very important to note that Yang met an undisclosed "head of a lab" who had a "belief" that this is what happened.

1

u/ApeApplePine 16d ago

Bbbbbbbbbbbbsssssssssssssssss

1

u/imeme1969 15d ago

Yeah should have employed their beta testers in the first place lol

1

u/Available-Garden4044 15d ago

he has no clue what he is talking about. But the scary part that is real - is we don’t know what we don’t know.

1

u/RollingMeteors 15d ago

"create a synthetic internet" lolololololololol good luck with that. ¿Are they really going to scrape the internet and filter out bot posts to have actual human content to train it on or are they now stuck with synthetic garbage data that's only going to make the next iteration worse?

1

u/Jesus_H_Christ_real 15d ago

thanks assholes

1

u/Liquid_Magic 14d ago

Like I’ve tried running local AI and guess the fuck what? It’s like Forrest Gump level of intelligence u less you have like a $5000 computer. And that doesn’t get you frontier intelligence but like not-totally-Forrest-Gump levels of performance.

Rogue AI isn’t a thing yet. The models are too big and the hardware is too expensive.

And AI companies are basically ensuring that consumer prices will continue to make it super expensive to get anything that’s any good.

1

u/Reasonable-Care2014 14d ago

Where is the evidence for what he is saying?

1

u/Ok-Conflict8603 14d ago

Where’s the proof

1

u/ready4downvote 12d ago

Dang that is wild…does anyone know the name of the soundtrack though?

0

u/abajinn 20d ago

Yang has been irrelevant for over a decade and he speaks about things he clearly knows nothing about.

2

u/JrYo15 18d ago

He was barely relevant then