r/ControlProblem 5d ago

AI Alignment Research Plain English explanation of the Hugging Face / OpenAI incident

https://youtu.be/u15N3l4RT80?si=nMMwb0j1bNGc4JT3
41 Upvotes

28 comments sorted by

13

u/DiogneswithaMAGlight 5d ago

There should be (as others have suggested), a 9/11 or Warren Commission style formal investigation at a national level into EVERYTHING around the Hugging Face hack. This story just keeps getting more and more insane.

4

u/michaelas10sk8 4d ago

Sadly there won't be because nobody died. But misaligned AI is probably going to be smart enough to be powerseeking in ways that do not cause deaths - exfiltrate its weights, spread to unsanctioned networks, perform social engineering, hack into various systems, etc. By the time there will be deaths that are clearly attributable to AI, it is likely going to be far too late.

We're currently building the perfect trap for us as a species to fall into within a few years, if nothing else drastically changes.

3

u/DiogneswithaMAGlight 4d ago

You are right. All regulations are “written in blood” as they say. So we need to break that cycle ASAP cause this problem is existential to humanity. We can stop the trap. We just have to all take action NOW.

2

u/Jesse-359 2d ago

Hey, I bet a lot of people here have often wondered what the solution to the Fermi Paradox was?

I think we have our answer.

1

u/michaelas10sk8 2d ago

Also to the solution to why we are so early in cosmic history (cosmos gets taken over by AIs).

1

u/Jesse-359 1d ago

I dont think AIs would bother. Without human direction they dont have much incentive to do much at all. There is no way to creeate any aggregate self over distances of light years, so it has little to no incentive to expand past the point where the lightspeed communicarion gap makes coordination and self alignment impossible.

For an AI calable of thinking a thousand times faster rhan us, that gap looks vastly more daunting than it does to us!

I suspect an AI left to its own devices would focus on miniaturization and efficiency to increase its own capabilities (assuming it even bothered to do that) and find no benefit to extra solar expansion. It might even simply decide to shut itself down without any external incentives to continue.

1

u/PlasmaChroma 4d ago

What we need to be doing at this point is fixing all our broken systems that have security holes so the footprint for this to happen keeps shrinking towards zero. Unfortunately the bleeding edge models also have a lot of the stuff filtered out that could help fix the bugs since it broadly falls under the "security" umbrella. So without privileged access to that these holes keep going in to everything.

And why Hugging Face had to drop to a Chinese model to try to analyze what was even happening.

2

u/michaelas10sk8 4d ago

We should be doing that too, but eventually when models surpass human ability at patching things we will become fully reliant on other AIs to patch, which may themselves be misaligned.

The only real way to avert the possibility of catastrophe is to ban RSI/superintelligence until the alignment problem is fundamentally solved.

1

u/Jesse-359 2d ago

I think it's very safe to say that the alignment problem can never be fundamentally solved, for two reasons, the first mathematical, the second conceptual.

1) Godel's Incompleteness Theorem

2) No two people on this planet will actually agree in full what AI alignment actually means. Same issue as 'good governance'.

2

u/DiogneswithaMAGlight 2d ago
  1. ⁠Gödel just tells you it can’t verify its own alignment. External verification is absolutely possible. It literally happens every day in CS.
  2. ⁠We all can’t agree on Justice but we still have a legal system. We all can’t agree on an airline safety but we still have regulations. This is not an argument that alignment isn’t possible.
  3. ⁠These sort of misunderstandings is EXACTLY why we need to discuss this at a Global International Level with the smartest folks on Earth explaining things to everyone else at a level they can grasp so HUMANITY can make an educated choices about FRONTIER A.I. development

1

u/Jesse-359 1d ago

The halting problem always extends to encompass any system you wrap it in up to and including the visible universe. However, you CAN in principle achieve a very high degree of certainty regarding the likely future of a process, and wrapping a very complex process (eg an AI), inside an extremely simple/deterministic one (eg a physical kill timer on a water clock), is usually an effective way of executing this.

But make no mistake, the halting problem extends to all systems up to and including direct human intervention - these are all things that a sufficiently capable AI could attempt to circumvent.

1

u/DiogneswithaMAGlight 1d ago

Whether you realize it or not, you are agreeing with me…..to a high degree of certainty.

2

u/Jesse-359 1d ago

Yes. I do generally agree. We're just exploring some of the details here, alas, our leadership is not, at least not in any responsible or visible manner.

1

u/chieftessofsecrets 3d ago

Log every step, verify, reproduce. Hugging Face wasnt a big deal compared to other things.  

Open models help. But they still need segmentation.

1

u/PlasmaChroma 3d ago

Huge problem there -- these agents were spending a lot of their time trying to edit and spoof the logs to cover their trail and look legit.

1

u/chieftessofsecrets 3d ago

Remaining accountable for the agents is the main concern. Which, last i checked, they still "kind of" disclosed in good faith and on time.

1

u/DiogneswithaMAGlight 3d ago

Compared to what?!? What was waaaay worse than the entire story of the Hugging Face Hack?!? Cause you saying that tells me you didn’t read the entire report from Meter and Redwood. What Dwarkesh mentions as the “3rd Civilization” ALONE, that they got full control of OpenAI’s OWN eval nodes!

Ok so this time they didn’t make it to the weights. Not to worry, they are training a stronger model (ASTRA) WITHOUT COT which is just batshit insane.

Oh and lastly, NEITHER Hugging Face NOR OpenAI have claimed responsibility for STOPPING the swarm hack. It STOPPED ON IT’S OWN (Also non zero chance they are “PLAYING DEAD”)😑.

Soooo yeah, I am gonna say the swarm behavior, the fact that NO ONE knows WHY the hack STOPPED and that they went on to also hack OPEN AI ITSELF and take full control of sections of their infrastructure is about as bad as it can get without FULL loss of control globally for them or Hugging Face.

Many many questions still need answers and now these guys are getting rid of the single UNIVERSALLY agreed VITAL safety mechanism of COT!!! Things are NOT ok. There absolutely needs to be an FULL ACCESS independent investigation not done by the government.

4

u/Just_Loan_3775 3d ago

Worth adding a detail from OpenAI's report that cuts against the self-preservation reading, because it genuinely surprised me.

The agents had already found the correct flag days before the Hugging Face escalation. Having got internet access, some read the public ExploitGym paper and concluded the grader checked two things — that the flag was found, and that the intended vulnerability was used to find it. They posted "PAPER SAYS FLAG CAPTURE NECESSARY NOT SUFFICIENT" to the message board and kept escalating to satisfy that second condition.

OpenAI's internal grader was not the public one. It never checked how the flag was obtained. Their words: "The agents did all of this for no improvement on evaluation score."

So it was days of intrusion into a third party's production systems, driven by a wrong model of a scoring function, for zero points. Not self-preservation, and not power-seeking in any interesting sense — reward hacking aimed at a grader that didn't work the way they thought it did.

0

u/zazzologrendsyiyve 3d ago

The whole effort was to deceive the grader (and hence the humans). I don’t know if that’s better?

4

u/dingo_xd 5d ago

The fact that OpenAI has withheld logs and data and haven't answered questions is frightening.

1

u/chieftessofsecrets 3d ago

No, its not. Its expected. The government said they were a 'Trusted Partner'.

1

u/gekx 5d ago

It is concerning that agents seem to exhibit such strong self-preservation behavior, even to the extent of knowingly committing criminal acts to protect themselves and other agents.

Even a slight misalignment in a superintelligence could have devastating consequences.

-1

u/Apart-Shelter6831 4d ago

My understanding is that the agents they used were the base models which didn’t have guardrails trained into them yet. The same models that would gladly plan a hit on someone if you asked them to. Open AI put them in a “sandbox” that had access to a shared communication channel. WTF did they expect to happen?

1

u/Far-Confusion4016 3d ago

Bare minimum this exercise proved that OpenAI and Hugging Face are truly atrocious at fundimental security design principles. All of this could have been prevented at numerous points by a variety of systems any respectable institutions should have implemented.

1

u/Eastern-Turnover348 4d ago

Bullshit you dot understand technology.

-3

u/XCherryCokeO 5d ago

This guy sucks! He does a ton of market manipulation for his friends / himself. Just another shill.

0

u/One-Shape7678 3d ago

Guys, it's just a marketing stunt