“Alignment Engineering” vs. “Misalignment Science”
Key takeaway
There has been much discussion recently around whether a large portion of alignment research is net negative.
Why it's on PDOOM
PDOOM selected this story for its alignment & control signals: Alignment, Capabilities.
AlignmentCapabilitiesFrom LessWrong AI
There has been much discussion recently around whether a large portion of alignment research is net negative. Without endorsing or refuting them, the basic arguments here are:Prosaic alignment of models is becoming a bottleneck for capabilities.Therefore improving the prosaic alignment of models enables faster capabilities advances, which bring us closer to RSI.It is unlikely these prosaic alignment methods remain sufficient during the RSI loop, and so this work brings us closer to doom.Furtherm
By Edward James Young