Tracking AI existential risk. Source-backed context. Original reporting always linked.
MONITORING CORE AI-RISK FEEDS · UPDATED HOURLY
Research

Fixed-weight models are adversarially vulnerable: hence misaligned

Zac Boring September 28, 2026 1 min read
Read original source →

Key takeaway

This post argues that fixed-weight models (at least as we understand them today) will a) always be vulnerable to adversarial examples in their concept-spaces, and b) hence will be misaligned, under sufficient optimisation pressure.Boundaries in concept…

Why it's on PDOOM

PDOOM selected this story as relevant to alignment & control.

From Alignment Forum

This post argues that fixed-weight models (at least as we understand them today) will a) always be vulnerable to adversarial examples in their concept-spaces, and b) hence will be misaligned, under sufficient optimisation pressure.Boundaries in concept spaceTo serve any purpose whatsoever, an AI will have to draw boundaries inside its world-model - to distinguish world A from world B, and reach some comparison between them.If we want the AI to follow our goals and values, we want it to be able t

By Stuart_Armstrong

Read the full article at Alignment Forum →