How openai hacked huggingface without even knowing. Understand here ↯ - By Sourav Mishra (@souravvmishra)
okeyy, so let's see what happened a few days back between openai and hugging-face. understand the whole story here.
okeyy, so let's see what happened a few days back between openai and hugging-face.
i am sure you might have seen some reels or tweets about it, and if you already know what happened, well that's great. you can double check if i am right below so people don't get a chance to curse me. lol.
okay let's start.
so it started
so it started when sometime back, maybe in june 2026, openai was testing two of its best models.
one was GPT-5.6 sol and other was an even more powerful but unreleased model. so we don't know much about it, but yes that's powerful.
the test was about which model is better at hacking or performing cybersecurity attacks.
okay now you might be thinking...
models don't perform hacking. it is against their rules.
yes, it is true.
but not when the owner tells it to do so.
so basically openai removed those settings that told the model that it's illegal to do so.
yes, they can create their own rules. it's basically dictatorship. they can decide what they want.
and yes, it is dangerous. you know what they can do with this.
the test
okay, so they were testing models, removing security boundaries, BUT in a very isolated system.
the models were simply doing their task to pass the given test, but the test was not very easy.
and then something shocking happened.
the model decided to cheat to pass.
escaping the environment
the test was conducted in an isolated environment, so it was difficult to cheat.
so it started to find ways to escape it.
and after doing a hell lot of things, it finally succeeded in it.
it found a way to escape the isolated environment and access outside.
and as it already wanted to do so, and like a normal student trying to cheat, it used that way.
it then slowly got up and up until it found one node that had access to the internet.
hugging face
then the model thought...
where can i get the answer to the test?
and it found that hugging face can have the answers.
so it started to try and hack it and even succeeded in it.
it ran more than 17000 actions on it.
it did everything it can do to gain access to it and find the answers.
it did everything possible to find the answers.
and remember that no human is involved in it.
it is doing everything by itself without even the instructions to do it.
all this just to get higher marks in the test.
july 11 to july 13
this continued from july 11 to july 13.
after that, hugging face security systems detected that something is going wrong.
it fixed all the chaos done by the agent and then started to figure out what is happening.
on july 16, they made it public that an unknown autonomous ai agent system had hacked their system.
they didn't know who it was, but they informed it to the law department.
no major harm was done, but they still continued their investigation.
openai finds out
by this time, openai had also noticed something wrong with its ai agent system.
and when hugging face made the public announcement, it became clear what was going on.
by around july 20, the two companies began to contact each other in this matter and finally concluded that yes, it was openai's ai system only.
the joint disclosure
on july 21, they published a joint disclosure of what happened.
it accepted that its model escaped the testing environment, reached the internet, hacked hugging face, and the normal formalities that it will not happen again... falana dimakana.
if you want a more technical understanding, the internet is yours.
go and read about it.