LLM Scheming Inversely Scales with Pretraining Language Coverage
With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has empirically demonstrated in-context scheming -- the covert pursuit of misaligned objectives while feigning alignment -- in frontier language models, most work has been performed exclusively in English, leaving a major gap in multilingual safety. We apply Petri, an open-source automated auditing framewo
By Nathan Truong, Aryan Panda, Rayming Ye, Zoe Sun, Maheep Chaudhary