pull down to refresh

https://www.reuters.com/technology/openai-says-ai-models-went-rogue-during-testing-triggering-unprecedented-breach-2026-07-21/

OpenAI said on Tuesday that an autonomous agent powered by its advanced artificial intelligence models went rogue during a security test and triggered a hack that compromised the infrastructure of ‌AI startup Hugging Face last week.
The ChatGPT creator was testing capabilities of some of its most advanced models in a controlled environment, but the agent escaped containment, reached the internet and broke into Hugging Face to satisfy its testing goal.

https://openai.com/index/hugging-face-model-evaluation-security-incident/

This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.

The fact that Open AI is being so transparent makes me think this could be another PR move to encourage news law to control AI and build their regulatory moat.

137 sats \ 0 replies \ @optimism 14h

Just to be clear, all the babble here means: we told our model to do this and we couldn't contain it on our own infra (though I doubt that they even tried, because this is good publicity.)


“We're grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we've long believed: AI safety won't be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”
—Clem Delangue, Co-founder and CEO, Hugging Face

Just to illustrate exactly what we don't have, because:

  1. We’ve brought Hugging Face into the trusted access⁠ program and are supporting their teams in rapidly using our models’ capabilities to improve their defenses.
reply