Tracking AI existential risk. Source-backed context. Original reporting always linked.
MONITORING CORE AI-RISK FEEDS · UPDATED HOURLY
Research

Debate Training Reduces Reward Hacking in RLAIF

Zac Boring August 19, 2026 1 min read
Read original source →

Paper: Debate Training Reduces Reward Hacking in RLAIFLinkpost for GDM Alignment blogpostWork done by the GDM Amplified Oversight team (we're hiring).TL;DR: When you RL against an LLM judge, the judge gets hacked i.e. fooled into incorrectly giving high reward; adding a debate opponent reduces this.Many of the most impressive capabilities of current AI systems are produced by training on crisp tasks, like math and coding, where task success can be automatically verified. However, much of AI beha

By zac_kenton

Read the full article at Alignment Forum →