Fixed-weight models are adversarially vulnerable: hence misaligned
Key takeaway
This post argues that fixed-weight models (at least as we understand them today) will a) always be vulnerable to adversarial examples in their concept-spaces, and b) hence will be misaligned, under sufficient optimisation pressure.Boundaries in concept…
Why it's on PDOOM
PDOOM selected this story as relevant to alignment & control.
From Alignment Forum
This post argues that fixed-weight models (at least as we understand them today) will a) always be vulnerable to adversarial examples in their concept-spaces, and b) hence will be misaligned, under sufficient optimisation pressure.Boundaries in concept spaceTo serve any purpose whatsoever, an AI will have to draw boundaries inside its world-model - to distinguish world A from world B, and reach some comparison between them.If we want the AI to follow our goals and values, we want it to be able t
By Stuart_Armstrong