OpenAI said on Tuesday that an autonomous agent powered by its advanced artificial intelligence models went rogue during a security test and triggered a hack that compromised the infrastructure of AI startup Hugging Face last week.
The ChatGPT creator was testing capabilities of some of its most advanced models in a controlled environment, but the agent escaped containment, reached the internet and broke into Hugging Face to satisfy its testing goal.
https://openai.com/index/hugging-face-model-evaluation-security-incident/
This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.
The fact that Open AI is being so transparent makes me think this could be another PR move to encourage news law to control AI and build their regulatory moat.
Just to be clear, all the babble here means: we told our model to do this and we couldn't contain it on our own infra (though I doubt that they even tried, because this is good publicity.)
Just to illustrate exactly what we don't have, because: