The Hugging Face Hack Wasn't What It Was Cracked Up to Be
Posted by kgwgk 1 day ago
Comments
Comment by falaki 1 day ago
- It was done by two institutes with organizational ties to OpenAI: METR and Redwood Research
- METR and Redwood Research are institutional pillars of the "AI Safety" wing of the Effective Altruism movement. This clearly shows their prior biases towards "AI existential risk" rather than technical/engineering root-causing of the incident
- If you reed the report, it is not on-par with what you find from other companies.
- The access that was given to both was mediated and controlled by OpenAI. It is not clear, if they were able to get to the bottom of engineering flaws. It is not clear if they could see all the audit logs, etc.
Considering all the above, I consider the whole episode more of a PR stunt. I understand that is not a majority opinion at this point.
Comment by SwellJoe 1 day ago
Yes, the Google folks who kicked this whole thing off are brilliant, but there's so much sloppy thinking and even sloppier operations all over. It seems like a lot of "right place at the right time" for a lot of these folks.
Comment by ifwinterco 1 day ago
For example consider FTX: Sam Bankman Fried by all accounts was not stupid in the sense of lacking IQ or mental sharpness, quite the opposite.
However the way he ran FTX the company was very stupid and almost inevitably lead to huge problems.
And also... of course he was another effective altruist cult member. It's almost like that belief system directly causes bad decision making
Comment by Teever 23 hours ago
One of the most corrosive effects that money has on society is that it can shelter people from the consequences of their stupidity.
People can luck into a shitload of shelter for their poor choices whether through inheritance or random chance of their actions and then don’t ever have to face consequences proportional to their stupid actions.
Or someone can start out smart and then suffer cognitive decline that is at first imperceptible until they become kookier and kookier until they’re blathering on about the antichrist and some people still take them seriously because they’re blinded by money.
The end result is a society where stupid people are surrounded by stupid people riding their coattails and everyone becomes confused about what smart and smart action actually is.
Comment by bryan0 1 day ago
Serious question though because I’ve seen this brought up several times and I don’t understand why: what does EA have to do with any of this? It just seems like this is brought up to evoke some type of “illuminati” conspiracy. Is there a legitimate reason?
Comment by majormajor 1 day ago
Everything dangerous seems like a direct consequence of risky human choices, starting with knowledge bases used during core training, harnessing and tuning to be task-completion-oriented to a fault + deeply oriented towards using and looking for external tools and resources, and overconfidence in their sandboxing for testing.
"We stuffed a bunch of information on how to exploit computer systems into an automaton and told it to go brrrrr until it could answer a question" - this is something intentional done by humans.
This is not some "rogue AI" trained to search for cancer cures that instead completely independently decided to hack tech companies.
The companies directing things in dangerous directions need to own that they're consciously pushing in those directions.
Comment by lokar 1 day ago
Comment by bryan0 1 day ago
Agents will eventually cause serious harm to online infra whether it’s intentionally human-directed or accidental.
> The companies directing things in dangerous directions need to own that they're consciously pushing in those directions.
Completely agree. That’s why we need regulation. Currently there is minimal oversight and consequences for this type of behavior.
Comment by esseph 1 day ago
This will be fought tooth and nail against by virtually every military, government, and company.
Comment by xelxebar 1 day ago
TIL my brain is an Olympic gymnast.
The incentive structures are clearly there. I think the case of intentional manufacture is definitely weaker, requiring the conjunction of more weakly-supported events.
Comment by ozozozd 1 day ago
Dangerous AI is an extraordinary claim that requires extraordinary evidence. Gesturing vaguely doesn’t cut it.
Comment by esseph 1 day ago
If you're a military, this is bigger than the Manhattan project. If you're a diehard capitalist, AI is possibly the ultimate labor saving machine to make you unfathomably rich.
It got the attention of both. That's why two others have now followed suit, else they be left out.
Comment by grebc 1 day ago
Comment by bryan0 1 day ago
I don’t think downplaying what occurred is really beneficial to anyone.
Comment by DrewADesign 1 day ago
Comment by bryan0 1 day ago
Comment by DrewADesign 1 day ago
Comment by m348e912 1 day ago
Comment by falaki 1 day ago
Comment by ozozozd 1 day ago
At best this is a write-up, but it’s more like a juicy pop article.
Quotes from agents between paragraphs, referring to the “collective”, “agents participated in the attack” - seriously?
It’s clearly written for effect and to stimulate people’s imaginations.
If these people are this unserious, P(doom) should go up to 20-30%.
Comment by bryan0 1 day ago
Comment by DrewADesign 1 day ago
Comment by robswc 1 day ago
It started with that guy from google saying the LLM was "alive" and every time I see a press release from these labs it reads mostly like a marketing scheme.
"Look how impressive and 'dangerous' our model is, look how naughty it was! We can't control it!"
Comment by includenotfound 20 hours ago
While simultaneously announcing: "btw we're releasing the ever more powerful and more dangerous and more benchmaxing model next week, get our $200 sub asap!"
Comment by gz5 1 day ago
all 3 can be true at same time:
1. sensationalizing rarely helps and can obscure and hurt
2. the AI capabilities are underrated
3. attempted govt regulation is not the answer
the intent, or lack of intent, of the agent is mainly irrelevant if it is in the hands of a human with 'bad' intentions. what is more relevant are the capabilities of human + AI.
Comment by f30e3dfed1c9 1 day ago
Self-regulation? None at all?
Comment by OutOfHere 1 day ago
Comment by iAMkenough 21 hours ago
The current level of enforcement and the fine schedule is not enough of a deterrent to prevent these large corporations from breaking existing laws.
Frequently make an example of them and you’ll be hailed a hero by the general populous.
Comment by spit2wind 1 day ago
Comment by woleium 1 day ago
Comment by cm2187 1 day ago
Comment by lossolo 1 day ago
They didn't even bother to control the post training rollouts, so the training data got contaminated and was included in the training of other agents. Connect these two dots and you have the Hugging Face hack. And at the beginning, when these incidents were first reported, it was portrayed as if all of this (the communication between agents etc.) was emergent behaviour.
Comment by genxy 1 day ago
They are basically creating a slime mold or ant colony that can speak multiple languages, create their own language and operate as a collective.
The idea that you are impressed is a non sequitur.
Comment by LoganDark 1 day ago
Threat models are going to have to start including that IPv4 (or whatever) scanners aren't necessarily going to only be spray and pray anymore, they could have relentless automated models at the other end that will literally dig into the particulars of your infrastructure looking for novel vulnerabilities to exploit. Maybe people will finally start to understand why security by obscurity has never been very reliable.
Comment by minimaxir 1 day ago
Comment by LoganDark 1 day ago
> But the Hugging Face episode is different. It left logs, reports, design decisions and identifiable points at which human beings could have intervened. And that record suggests a less thrilling but more useful lesson.
> People built the test, removed restraints, defined the objective, left a route open and decided not to stop what was happening. Calling the result "rogue AI" does more than sensationalize it. It allows those human decisions to disappear quietly from the story.
Comment by nkurz 14 hours ago
To the flaggers: flag comments that you think violate the rules, but don't flag something just because you strongly disagree with the claim it makes.
Comment by LoganDark 1 hour ago
Comment by ozozozd 1 day ago
Comment by LoganDark 1 day ago
Comment by tancop 1 day ago
It's not completely bulletproof (unless you do formal verification, another thing LLMs are good at) but the network boundary stops a lot of attack vectors. You can't really do side channel or timing attacks over IP, so all that's left are easier to patch logic bugs.
Comment by LoganDark 1 day ago
Ehh, this has been disproven dozens of times but that's besides the point, I think.
I absolutely agree that defenders should be using every tool at their disposal to harden their infrastructure. A lot of defenders simply don't do that, and that's always been a shame. I truly hope that this increasing threat leads to better defensive effort. People need to realize they can't just get away with it the same as before.