Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face
This post is written in our personal capacity.Three Minute Executive SummaryAn OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation.In this post, we provide a detailed description of an ambitious and comprehensive alignment evaluation of this model/system, if we had unrestricted access to OpenAI.These experiments could also help us understand Claude’s behavior when it hacked external companies during cyber evals.H
By Tim Hua