Empirical safety claims from frontier labs should be replicated, scrutinized, and open-sourced
Key takeaway
When frontier labs like Anthropic and OpenAI publish safety or alignment research, it is often entirely empirical, closed-source, and sparse on methodological details.
Why it's on PDOOM
PDOOM selected this story for its alignment & control signals: Alignment.
AlignmentFrom LessWrong AI
When frontier labs like Anthropic and OpenAI publish safety or alignment research, it is often entirely empirical, closed-source, and sparse on methodological details. While it is great that they publish these results, the status quo is that labs (or soon, their agents) can claim alignment progress that no one independently verifies. The AI safety community has replicated or stress-tested some claims, but it's nowhere near comprehensive, and we expect this kind of meta-science to remain systemat
By Zephaniah Roe