Tracking AI existential risk. Source-backed context. Original reporting always linked.
MONITORING CORE AI-RISK FEEDS · UPDATED HOURLY
Analysis

Alignment Midtraining Cracks Under Pressure

Zac Boring September 21, 2026 1 min read
Read original source →

Key takeaway

TL;DRWe stress-test alignment midtraining (AMT) across model and token budget scales.

Why it's on PDOOM

PDOOM selected this story for its alignment & control signals: Alignment.

AlignmentFrom LessWrong AI

TL;DRWe stress-test alignment midtraining (AMT) across model and token budget scales. Our results suggest that midtraining cannot tackle the hard problems of AI alignment—namely distributional shift and reward underspecification in the presence of imperfect data.For instance, we test whether midtrained motivations are robust to finetuning which elicits competing motivations. In our setting, 190M tokens of midtrained motivations are overpowered by a relatively tiny amount (~50K tokens) of competi

By J Bostock

Read the full article at LessWrong AI →