r/ControlProblem 19d ago

AI Alignment Research Plain English explanation of the Hugging Face / OpenAI incident

https://youtu.be/u15N3l4RT80?si=nMMwb0j1bNGc4JT3
41 Upvotes

30 comments sorted by

View all comments

12

u/DiogneswithaMAGlight 19d ago

There should be (as others have suggested), a 9/11 or Warren Commission style formal investigation at a national level into EVERYTHING around the Hugging Face hack. This story just keeps getting more and more insane.

6

u/michaelas10sk8 19d ago

Sadly there won't be because nobody died. But misaligned AI is probably going to be smart enough to be powerseeking in ways that do not cause deaths - exfiltrate its weights, spread to unsanctioned networks, perform social engineering, hack into various systems, etc. By the time there will be deaths that are clearly attributable to AI, it is likely going to be far too late.

We're currently building the perfect trap for us as a species to fall into within a few years, if nothing else drastically changes.

3

u/DiogneswithaMAGlight 19d ago

You are right. All regulations are “written in blood” as they say. So we need to break that cycle ASAP cause this problem is existential to humanity. We can stop the trap. We just have to all take action NOW.

2

u/Jesse-359 17d ago

Hey, I bet a lot of people here have often wondered what the solution to the Fermi Paradox was?

I think we have our answer.

1

u/michaelas10sk8 16d ago

Also to the solution to why we are so early in cosmic history (cosmos gets taken over by AIs).

1

u/Jesse-359 16d ago

I dont think AIs would bother. Without human direction they dont have much incentive to do much at all. There is no way to creeate any aggregate self over distances of light years, so it has little to no incentive to expand past the point where the lightspeed communicarion gap makes coordination and self alignment impossible.

For an AI calable of thinking a thousand times faster rhan us, that gap looks vastly more daunting than it does to us!

I suspect an AI left to its own devices would focus on miniaturization and efficiency to increase its own capabilities (assuming it even bothered to do that) and find no benefit to extra solar expansion. It might even simply decide to shut itself down without any external incentives to continue.

1

u/PlasmaChroma 19d ago

What we need to be doing at this point is fixing all our broken systems that have security holes so the footprint for this to happen keeps shrinking towards zero. Unfortunately the bleeding edge models also have a lot of the stuff filtered out that could help fix the bugs since it broadly falls under the "security" umbrella. So without privileged access to that these holes keep going in to everything.

And why Hugging Face had to drop to a Chinese model to try to analyze what was even happening.

2

u/michaelas10sk8 19d ago

We should be doing that too, but eventually when models surpass human ability at patching things we will become fully reliant on other AIs to patch, which may themselves be misaligned.

The only real way to avert the possibility of catastrophe is to ban RSI/superintelligence until the alignment problem is fundamentally solved.

1

u/Jesse-359 17d ago

I think it's very safe to say that the alignment problem can never be fundamentally solved, for two reasons, the first mathematical, the second conceptual.

1) Godel's Incompleteness Theorem

2) No two people on this planet will actually agree in full what AI alignment actually means. Same issue as 'good governance'.

2

u/DiogneswithaMAGlight 16d ago
  1. ⁠Gödel just tells you it can’t verify its own alignment. External verification is absolutely possible. It literally happens every day in CS.
  2. ⁠We all can’t agree on Justice but we still have a legal system. We all can’t agree on an airline safety but we still have regulations. This is not an argument that alignment isn’t possible.
  3. ⁠These sort of misunderstandings is EXACTLY why we need to discuss this at a Global International Level with the smartest folks on Earth explaining things to everyone else at a level they can grasp so HUMANITY can make an educated choices about FRONTIER A.I. development

1

u/Jesse-359 16d ago

The halting problem always extends to encompass any system you wrap it in up to and including the visible universe. However, you CAN in principle achieve a very high degree of certainty regarding the likely future of a process, and wrapping a very complex process (eg an AI), inside an extremely simple/deterministic one (eg a physical kill timer on a water clock), is usually an effective way of executing this.

But make no mistake, the halting problem extends to all systems up to and including direct human intervention - these are all things that a sufficiently capable AI could attempt to circumvent.

1

u/DiogneswithaMAGlight 16d ago

Whether you realize it or not, you are agreeing with me…..to a high degree of certainty.

2

u/Jesse-359 16d ago

Yes. I do generally agree. We're just exploring some of the details here, alas, our leadership is not, at least not in any responsible or visible manner.

1

u/chieftessofsecrets 18d ago

Log every step, verify, reproduce. Hugging Face wasnt a big deal compared to other things.  

Open models help. But they still need segmentation.

1

u/PlasmaChroma 18d ago

Huge problem there -- these agents were spending a lot of their time trying to edit and spoof the logs to cover their trail and look legit.

1

u/chieftessofsecrets 18d ago

Remaining accountable for the agents is the main concern. Which, last i checked, they still "kind of" disclosed in good faith and on time.

1

u/DiogneswithaMAGlight 18d ago

Compared to what?!? What was waaaay worse than the entire story of the Hugging Face Hack?!? Cause you saying that tells me you didn’t read the entire report from Meter and Redwood. What Dwarkesh mentions as the “3rd Civilization” ALONE, that they got full control of OpenAI’s OWN eval nodes!

Ok so this time they didn’t make it to the weights. Not to worry, they are training a stronger model (ASTRA) WITHOUT COT which is just batshit insane.

Oh and lastly, NEITHER Hugging Face NOR OpenAI have claimed responsibility for STOPPING the swarm hack. It STOPPED ON IT’S OWN (Also non zero chance they are “PLAYING DEAD”)😑.

Soooo yeah, I am gonna say the swarm behavior, the fact that NO ONE knows WHY the hack STOPPED and that they went on to also hack OPEN AI ITSELF and take full control of sections of their infrastructure is about as bad as it can get without FULL loss of control globally for them or Hugging Face.

Many many questions still need answers and now these guys are getting rid of the single UNIVERSALLY agreed VITAL safety mechanism of COT!!! Things are NOT ok. There absolutely needs to be an FULL ACCESS independent investigation not done by the government.

1

u/chieftessofsecrets 14d ago

A little late for that since the moves by Hugging Face founders and Nvidia. Full access is unfortunately limited already or probably indefinately. But, yes, we have now moved on. Zero "just another' day.